Astr research: what survives of your CV text inside a PDF?

Astr publishes small, reproducible experiments on how software reads CVs, using synthetic CVs that belong to no real person. Every number on this page is computed directly from the published results files, alongside the method, its limits and the original files to download. The results apply only to the stated samples, tools and settings; they are not ATS scores or hiring odds. You may reuse them under CC BY 4.0 with attribution.

Key findings

  • English text PDFs: every reference word was recovered in 16/16 files with pypdf and 16/16 with PyMuPDF.
  • Arabic text PDFs: word recovery ranged from 74.79% to 93.02% with pypdf and from 88.19% to 91.60% with PyMuPDF; all seven fields were complete in 0/16 and 0/16 files respectively.
  • Image-only copies: both tools returned empty text in 16/16 measurements; OCR was not run.

Source: version 2 of the PDF reading experiment, measured on October 1, 2026. Synthetic CVs: 8; files: 40; measurements: 80.

Reading Arabic and English CV PDFs — version 2

Independent synthetic CVs for a graduate, a teacher, an engineer and a lawyer, in Arabic and English, each in several text layouts plus an image-only copy, measuring what two open-source tools extract directly from the file's text.

Sample language: Arabic, English · Measured: October 1, 2026 · Files: 40; measurements: 80

Results

GroupToolWord recoveryAll wordsAll seven fieldsComplete employer nameSix headings in order
Arabic text PDFspypdf 6.10.074.79%–93.02%0/160/160/1616/16
Arabic text PDFsPyMuPDF 1.26.788.19%–91.60%0/160/160/168/16(8 unmeasurable)
English text PDFspypdf 6.10.0100%16/1616/1616/1616/16
English text PDFsPyMuPDF 1.26.7100%16/1616/1616/1616/16
Image-only copiesBoth tools0%0/160/160/160/16(16 unmeasurable)

Method

Files were generated in Tajawal at 16 px on A4 paper with 18 mm margins, then text was extracted directly with pypdf 6.10.0 and PyMuPDF 1.26.7, without OCR, spelling repair or reordering. The experiment measures literal word recovery, seven complete fields (name, email, phone, two date ranges, degree and employer) and the order of six headings.

Limits

Purposively chosen synthetic samples with one font and one generator, and different content in each language; they do not show that English always works better, and they do not measure Astr's exports, grader scores, commercial ATS products or hiring outcomes.

Cite as: Astr Team, “Reading Arabic and English CV PDFs: a reproducible experiment,” version 2, October 1, 2026.

Read the full study with examples

Astr grader example on a synthetic teacher CV

A local run of Astr's grader rule engine on the reference text of a synthetic teacher CV, with an editorial review of its match against a fictional job. The grader page uses this example to explain how a result reads.

Sample language: Arabic · Measured: September 28, 2026

Results

  • Grader score in this example: 83/100.
  • Job match: an editorial review, with no model-generated percentage.

Method

The CV fields were entered as manually structured verbatim text, with no PDF parser or AI extraction, then scored by the rule engine the grader uses; the code revision and the hashes of input and output are recorded.

Limits

Not a PDF upload test and not a customer result; an uploaded file can score differently because of extraction and field assignment. The match table is an editorial comparison, not model output.

Cite as: Astr Team, “Astr grader example on a synthetic teacher CV,” September 28, 2026.

See the example on the CV grader

Reading Arabic CV PDFs — version 1

One synthetic Arabic CV of a mathematics teacher in PDF variants that differ in font, size and layout, plus image-only copies for comparison, measuring what two open-source tools extract directly from the text.

Sample language: Arabic · Measured: September 25, 2026 · Files: 20; measurements: 40

Results

GroupToolWord recoveryAll wordsAll seven fieldsComplete employer nameSix headings in order
Arabic text PDFspypdf 6.10.081.62%–91.18%0/160/160/1616/16
Arabic text PDFsPyMuPDF 1.26.779.41%–90.44%0/168/1616/160/16(16 unmeasurable)
Image-only copiesBoth tools0%0/80/80/80/8(8 unmeasurable)

Method

Tajawal and Cairo at 12 and 16 px, in four layouts: one column, two columns, a table and a forced page break. Text was extracted directly with pypdf 6.10.0 and PyMuPDF 1.26.7 without OCR, measuring word recovery, seven fields and the order of six headings.

Limits

A small, correlated set from one CV and one generator; Word files, other fonts, commercial ATS products and hiring outcomes were not tested.

Cite as: Astr Team, “Reading Arabic CV PDFs: a reproducible 20-file experiment,” September 25, 2026, version 1.

Read the full study with examples

License and reuse

You may copy, reuse and adapt the results and synthetic specimens of these experiments under the Creative Commons Attribution 4.0 license (CC BY 4.0), crediting Astr with a link to the study. Fonts and tools included in the archives keep their own licenses, listed inside them.

The CC BY 4.0 license

Try it on your own CV

Paste your PDF's text into a plain-text editor and check your name, email, employers and dates, then check your CV with Astr's grader.

Check your CV with Astr's grader