Reading Arabic and English CV PDFs — version 2
Independent synthetic CVs for a graduate, a teacher, an engineer and a lawyer, in Arabic and English, each in several text layouts plus an image-only copy, measuring what two open-source tools extract directly from the file's text.
Sample language: Arabic, English · Measured: October 1, 2026 · Files: 40; measurements: 80
Results
| Group | Tool | Word recovery | All words | All seven fields | Complete employer name | Six headings in order |
|---|---|---|---|---|---|---|
| Arabic text PDFs | pypdf 6.10.0 | 74.79%–93.02% | 0/16 | 0/16 | 0/16 | 16/16 |
| Arabic text PDFs | PyMuPDF 1.26.7 | 88.19%–91.60% | 0/16 | 0/16 | 0/16 | 8/16(8 unmeasurable) |
| English text PDFs | pypdf 6.10.0 | 100% | 16/16 | 16/16 | 16/16 | 16/16 |
| English text PDFs | PyMuPDF 1.26.7 | 100% | 16/16 | 16/16 | 16/16 | 16/16 |
| Image-only copies | Both tools | 0% | 0/16 | 0/16 | 0/16 | 0/16(16 unmeasurable) |
Method
Files were generated in Tajawal at 16 px on A4 paper with 18 mm margins, then text was extracted directly with pypdf 6.10.0 and PyMuPDF 1.26.7, without OCR, spelling repair or reordering. The experiment measures literal word recovery, seven complete fields (name, email, phone, two date ranges, degree and employer) and the order of six headings.
Limits
Purposively chosen synthetic samples with one font and one generator, and different content in each language; they do not show that English always works better, and they do not measure Astr's exports, grader scores, commercial ATS products or hiring outcomes.
Downloads
Cite as: Astr Team, “Reading Arabic and English CV PDFs: a reproducible experiment,” version 2, October 1, 2026.
Read the full study with examples