# Astr bilingual PDF extraction cohort, version 2.0

Run date: 2026-10-01. Eight independently authored synthetic CVs: graduate,
teacher, engineer and lawyer, with one Arabic and one English record per group.
The two language records are DIFFERENT careers, not translations or matched
language controls. No customers, personal data, real employers, real licences,
AI API calls, employer ATS or Astr production export/scoring engine are used.
The fictional reference content was AI-assisted and reviewed for this test.

## Planned matrix and comparability

Each record: four text layouts (single column, two columns, table, page break
after the first three sections) and one image-only control of ALL single-column
pages. Total: 32 text PDFs + 8 raster controls = 40 files, 80 measurements.
These are eight content samples, not 40 independent people. All planned cases
and all raw outputs, including empty and damaged text, are retained.

Fixed font Tajawal, body 16 CSS px (12pt), A4, 18mm margins, tagged Chrome PDF,
144dpi raster controls. Compare this with the corresponding Tajawal/16px v1
conditions, not the entire two-font/two-size v1 matrix. The generator is a
research harness, NOT Astr's CV builder. Actual browser, OS, Node, Puppeteer and
font hashes are in manifest.json; exact parser versions are in results.json.

## Measurements unchanged from v1

pypdf 6.10.0 `page.extract_text()`; PyMuPDF 1.26.7
`page.get_text('text', sort=False)`. No OCR, AI repair or geometric reordering.
PDF.js only rasterizes controls; it is not a measured extractor.

Unicode NFKC, remove Arabic diacritics/tatweel/bidi controls, collapse spaces.
Do not reverse characters, join broken words, or correct spelling. Count
Unicode letter/number tokens with multiplicity; no punctuation, stemming or
case-folding. Each CV has its OWN denominator, published in results.json.
Word recall is not semantic accuracy, precision, an ATS score or a hiring rate.

Seven exact normalized fields per record: name, email, deliberately invalid
phone, two date ranges, degree and employer/internship organization. Field
presence does not measure validity, entity recognition or date association.
Six section headings: all must be present for sequence to be measurable.
`headingOrder: null` means unmeasurable, not an observed order failure. Reference
order for columns is all of the first column, then the second (right then left
in Arabic, left then right in English). No claim about every line's order.

## Repeat extraction from the original bytes

Unzip pdf-cohort-study.zip and work inside it. Node 22.15+ and Python 3.11+.
Use an isolated Python environment:

```sh
python -m venv .venv
# Activate .venv for your operating system.
python -m pip install pypdf==6.10.0 PyMuPDF==1.26.7
npm install --no-save puppeteer@24.4.0 pdf-lib@1.17.1 jszip@3.10.1 pdfjs-dist@5.7.284 @napi-rs/canvas
STUDY_OUTPUT=. STUDY_PYTHON=.venv/bin/python node pdf-cohort-study.mjs --remeasure
```

PowerShell last command:

```powershell
$env:STUDY_OUTPUT='.'
$env:STUDY_PYTHON=(Resolve-Path '.venv/Scripts/python.exe').Path
node pdf-cohort-study.mjs --remeasure
```

Compare results-rerun.json with results.json: timestamps differ, original PDF
hashes, raw-text hashes and measurement rows should match on the recorded tools.
The command never overwrites original raw text or results. Generation requires
a NEW `STUDY_OUTPUT` and omitting `--remeasure`; set `STUDY_CHROME` if necessary.
Font bytes can be read from the included HTML if repository fonts are absent.
Regeneration is not byte-identical and is not a substitute for the original PDFs.

## Evidence and limits

- manifest.json: every planned specimen, PDF/HTML hashes, generation environment.
- references.json: all eight complete inputs, fields, headings, reference text.
- results.json: every per-file, per-engine measurement and raw-text hash.
- pdfs/, html/, raw/: complete original evidence, including all raster pages.
- Included scripts: cohort inputs, generator, unchanged v1 metrics and parsers.

This is a small purposive synthetic sample from one generator and one font,
with correlated layout variants. It cannot estimate population failure rates,
rank languages/professions/fonts, measure Astr's export quality or predict
employer decisions. English and Arabic differences are confounded with different
content. Empty raster extraction says nothing about OCR quality; none was run.
Human review of the PDF and checking imported application fields still matter.

Reuse the synthetic records, scripts and measurements with attribution to Astr
and https://astrsa.com/blog/arabic-cv-pdf-text-extraction-study-2026.
Font licences are in FONT-LICENSES.txt; third-party tools retain their licences.
Version 1 remains separately available, unchanged, under
https://astrsa.com/research/arabic-pdf-2026-09-25/arabic-pdf-study.zip.
