Benchmarks¶
Phase 2 measurement report for ParseCraft: real backends over the pinned
local corpus, produced by the in-repo harness
(parsecraft.benchmark) on 2026-09-27.
What was measured¶
Documents (all local, never fetched at run time) from the corpus staged in
tests/downloads/:
| document | format | pages |
|---|---|---|
…-pg1342.txt (Gutenberg, Pride & Prejudice, 754 KiB) |
text/plain | 1 |
…-rfc5322.txt (119 KiB) |
text/plain | 1 |
…-commonmark-spec.md (201 KiB) |
text/markdown | 1 |
whatwg-html.html (151 KiB) |
text/html | 1 |
wikipedia-calculus.html (792 KiB) |
text/html | 1 |
itu-t-p863.pdf (licensed, 2.3 MiB) |
application/pdf | 80 |
etsi-ts-103558.pdf (licensed, 2.7 MiB) |
application/pdf | 68 |
…-nist-sp-800-53r5.pdf (5.8 MiB) |
application/pdf | 492 |
Backends: native-text, native-markdown, native-html, native-pdf
(CPU), plus liteparse and the GPU OCR backends where noted below. The
harness measures every backend that is eligible for each document
(is_hard_eligible + the document's media type), one conversion per cell.
Environment: Windows laptop, CPython 3.13, CPU only for the table below;
light extras (download, web, pdf-lite, pdf, liteparse) installed.
No GPU work is included in this snapshot — see Gaps.
Reproduce¶
uv run --extra pdf --extra liteparse --extra download --extra web --extra pdf-lite \
parsecraft benchmark tests/downloads/*.html tests/downloads/*.md \
tests/downloads/*.pdf tests/downloads/*.txt \
-o docs/benchmarks/
Notes:
--extra pdfpulls PyMuPDF (AGPL) explicitly — the copyleft choice stays an operator decision (ADR-0003);native-pdfrecords a typeddependency_missingrow instead of crashing when it is absent.- The
--extraflags install transiently: plainuv rundoes not prune them afterwards, and the canonical dev environment (whichty's optional-import ignore comments assume) is restored withmise dev. Runmise devbefore the quality gate if you benchmarked first. - The report writers are deterministic: stable key and row ordering, fixed rounding, LF newlines, no wall-clock fields. Re-running on the same inputs yields the same schema and row set — the measured values (elapsed, pages/s, peak bytes) are run data and will vary.
tests/test_benchmark_report.pyguards the committed artifacts: byte round-trip through the current writers, pinned schema, sorted rows.
Results¶
The full harness output is committed alongside this page:
benchmark.md and benchmark.json.
Headlines (CPU pass, 2026-09-27; values as committed in
benchmark.json):
| document | backend | elapsed | pages/s | coverage |
|---|---|---|---|---|
| itu-t-p863.pdf (80 pp) | native-pdf | 0.612 s | 130.8 | 0.999 |
| nist-sp-800-53r5.pdf (492 pp) | native-pdf | 4.002 s | 122.9 | 0.995 |
| etsi-ts-103558.pdf (68 pp, image-heavy) | native-pdf | 6.468 s | 10.5 | 0.963 |
| itu-t-p863.pdf (80 pp) | liteparse | 5.672 s | 14.1 | 1.017 |
| nist-sp-800-53r5.pdf (492 pp) | liteparse | 115.647 s | 4.3 | 0.904 |
| etsi-ts-103558.pdf (68 pp) | liteparse | 6.011 s | 11.3 | 1.012 |
| pg1342.txt (754 KiB) | native-text | 0.002 s | 409.9 | 0.988 |
| rfc5322.txt | native-text | 0.009 s | 116.6 | 0.976 |
| commonmark-spec.md (201 KiB) | native-markdown | 0.739 s | 1.4* | 0.717 |
| whatwg-html.html | native-html | 0.873 s | 1.1* | 0.271 |
| wikipedia-calculus.html | native-html | 2.460 s | 0.4* | 0.106 |
* Single-page formats: "pages/s" is one page over elapsed time — read the elapsed figure instead. Coverage may exceed 1.0 (liteparse emits slightly more characters than the analyzer counted — different segmentation, not a bug).
GPU OCR legs (bounded fixtures)¶
The four OCR models ran in the isolated GPU environment (.venv-gpu, RTX
A2000 8 GB) over 2-page fixtures of the three corpus PDFs
(.tmp/ocr-bench/*-p1-2.pdf) — full-corpus OCR would take hours. At run time
the four models pinned two different transformers majors, so the snapshot is
two passes over the same fixtures. That split was a pinning fact,
not a model requirement: pc-4u7.36 unified the extras on
transformers>=5.17,<6 and the TeleOCR/Unlimited models now load from local
vendored modeling code, so a re-run needs one environment and one pass
(see Deferred runs).
| fixture | pass A: ovis (5.17) | pass A: qianfan (5.17) | pass B: tele (4.57.1) | pass B: unlimited (4.57.1) |
|---|---|---|---|---|
| etsi-ts-103558-p1-2 | 197.6 s | 713.7 s | 220.2 s | 751.7 s |
| nist-…-p1-2 | 212.6 s | backend_error |
81.7 s | 128.0 s |
| itu-t-p863-p1-2 | 413.8 s | 1024.0 s | 84.0 s | 133.6 s |
Cross-version rows behave exactly as typed: the wrong-major pair records
backend_error rows whose detail names the cause (e.g. qwen3_5/
qianfan_ocr unknown to transformers 4.57.1), never a silent pass.
liteparse rows (0.07–0.16 s for the fixtures) appear in pass B after an
entry-point refresh.
Reading rules (also pc-4's run notes):
elapsed_sincludes per-document model load — the runner constructs the backend per document; these are 2-page fixtures, not per-page rates you can multiply to the full corpus.- The two passes are different virtualenvs: merge rows per backend, never sum across passes.
- Full-corpus
qianfanis a dated deferral, not a gap — see Deferred runs. selectedis recorded before load success/failure: aselected: true+failures: [backend_error]row means “planner picked it, then it could not load in this transformers” — not a planner bug.
Deferred runs¶
The legs below were deliberately deferred on 2026-09-28, with the reason recorded so a later run is a scheduling decision rather than a missing result. None of them failed: the measured fixtures already prove the adapters work, and the full-corpus runs only cost machine time.
| Run | Deferred on | Reason | Command |
|---|---|---|---|
| Full-corpus OCR, all four models | 2026-09-28 | Cost/time only: the 2-page fixtures already take 82–1024 s per document, so the corpus is many hours of exclusive GPU time. Not an install conflict — the extras are jointly installable on one transformers window |
Full-corpus command |
Full-corpus ocr-qianfan |
2026-09-28 | Cost/time: ≈ 145 s/page measured → ≈ 20 h for NIST SP 800-53r5 alone. Weights are staged and sha-verified (20 GB cache), so the run is offline-ready | same command (qianfan is one of the measured backends) |
| Full-corpus docling | 2026-09-28 | Runtime budget: docling is ≈ 8–15 s/page warm on this laptop → ≈ 1.8 h for NIST SP 800-53r5 alone, and the incident shape (>1 h on 113 pp) says hours. Not a CI-shaped job | Full-corpus docling |
| Quantised variants (int8 / 4-bit) | 2026-09-28 | Out of scope by design: quantisation is a backend-installation choice, not a harness dimension | — |
Full-corpus OCR command¶
The GPU environment is .venv-gpu (RTX A2000 8 GB). All four OCR extras share
one transformers>=5.17,<6 window (pyproject.toml), so the full run is a
single pass in a single environment. The command prints JSON on stdout and never
overwrites the committed CPU report; --extra pdf is required for PDF
rasterisation (PyMuPDF, AGPL — ADR-0003).
uv run --project .venv-gpu --extra ocr-ovis --extra ocr-qianfan \
--extra ocr-tele --extra ocr-unlimited --extra pdf --extra download \
parsecraft benchmark tests/downloads/itu-t-p863.pdf \
tests/downloads/etsi-ts-103558.pdf tests/downloads/nist-sp-800-53r5.pdf \
--json > docs/benchmarks/benchmark-ocr-full.json
Full-corpus docling command¶
Resumable: each (document, backend) pair is persisted as it finishes.
uv run --extra docling --extra pdf --extra pdf-lite \
python scripts/bench_docling.py tests/downloads/*.pdf --backends docling
The committed bounded runs stay the evidence for adapter correctness:
benchmark-ocr-pass-a.json and
benchmark-ocr-pass-b.json.
Cross-check: the motivating incident¶
ADR-0001 records the incident this package exists to fix: Docling ran for >1 hour on a 113-page ITU-T P.863 PDF at full CPU/RAM and was cancelled (≈ >31.9 s/page, never finished).
Measured here on the staged ITU-T P.863 artifact (80 pages, same document family — the staged file is not byte-identical to the incident's 113-page bundle): native-pdf extracts 80 pages in 0.612 s (≈ 7.7 ms/page), and the 492-page NIST SP 800-53r5 in 4.00 s (122.9 pages/s). That is more than three orders of magnitude below the incident's per-page cost — with honest caveats: this is text-layer extraction into typed chunks, not Docling's full layout/table pipeline. The routing design leans exactly on this gap: run the cheap native pass first, escalate to OCR/layout backends only for pages whose signals say native extraction cannot do the job (ADR-0001 §6 keeps the 15-minute cold-cache budget for the whole document, escalations included).
Docling vs native-pdf (bounded head-to-head)¶
The gap the incident cross-check only inferred is now measured on 2-page
fixtures of the three corpus PDFs (pc-4u7.39, docling 2.130.0, 2026-09-28).
Environment: Windows laptop, CPython 3.13, CPU only. Docling ran its default
pipeline untuned (layout, tables, and its bundled RapidOCR); native-pdf used
PyMuPDF. Both were measured by the in-repo harness.
| fixture (2 pp) | docling | native-pdf | ratio |
|---|---|---|---|
itu-t-p863 (of 80) |
20.6 s | 0.014 s | ≈ 1470× |
etsi-ts-103558 (of 68) |
30.5 s | 2.98 s | ≈ 10× |
nist-sp-800-53r5 (of 492) |
15.5 s | 0.024 s | ≈ 650× |
- Docling's first conversion in a process is far slower — 191.9 s for the
p863 fixture — because it loads its layout/OCR models. The table is warm; its
peak Python memory is 158–405 MB against 0.03–15 MB for
native-pdf. - Docling emits more structure (headings, tables) at comparable text coverage; the cost of that is runtime and memory, not correctness.
- Ratios vary per document (the ETSI fixture is the most docling-favourable of the set) — read the per-row numbers, not a single ratio.
This is an on-demand measurement, never part of mise test/CI:
# 1. two-page fixtures of the pinned corpus PDFs (pymupdf, transient)
uv run --extra pdf python - <<'PY'
from pathlib import Path
import pymupdf
out = Path(".tmp/docling-bench")
out.mkdir(parents=True, exist_ok=True)
for src in Path("tests/downloads").glob("*.pdf"):
doc = pymupdf.open(src)
two = pymupdf.open()
try:
two.insert_pdf(doc, from_page=0, to_page=min(1, doc.page_count - 1))
two.save(out / f"{src.stem}-p1-2.pdf")
finally:
two.close()
doc.close()
PY
# 2. one (document, backend) pair at a time, resumable
mise run bench-docling
The equivalent explicit command:
uv run --extra docling --extra pdf --extra pdf-lite \
python scripts/bench_docling.py .tmp/docling-bench/*.pdf
scripts/bench_docling.py writes each measured (document, backend) pair to
.tmp/bench-docling-partial.json immediately and skips pairs already present on
re-invocation, so a long run resumes instead of restarting. It never uses the
conversion cache — a cache hit would report ~0 s and falsify the measurement.
Honest gaps¶
- Full-corpus OCR was deferred, not skipped. The GPU legs cover 2-page
fixtures only (see GPU OCR legs); scaling
any OCR number here to a whole document would be fabricated. The dated
deferral and exact commands are in Deferred runs. The pinned
corpus fetch tier (
mise run test-corpus) stays opt-in; the licensed PDFs are hand-staged and never downloaded by the harness. - Full-corpus
qianfanis a dated deferral (≈ 145 s/page measured; ≈ 20 h for NIST alone) — see Deferred runs. All four models are staged and sha-verified (20 GB cache), so the deferred run passes offline asset checks. - VRAM is not measured here. Peak bytes are
tracemalloc(host Python allocations). GPU VRAM is covered by the separategputest tier (mise run test-gpu, RTX A2000 8 GB pins inpyproject.toml). - No quantisation variants were benchmarked — model quantisation stays a backend-installation concern, not a harness dimension.
exoplanets.csvandwikidata-q42.jsonwere skipped (unsupported source suffix in the CLI media-type map). The skips are recorded in the report, not hidden.- HTML coverage is a structure proxy.
native-htmlemits headings, paragraphs, lists, tables, code, and quotes; raw inline markup dominates the analyzer's character count, so 0.11–0.27 coverage is expected for markup-heavy pages, not data loss. itu-t-p863.pdfshowsselected: -: its page signals routed some pages to OCR intents while no OCR backend was installed in this environment, so auto-mode planning correctly declined (selectedis null) — the direct per-backend measurement still ran.
Artifacts¶
| file | contents |
|---|---|
benchmark.md |
full Markdown report (harness output) |
benchmark.json |
full JSON report, schema-guarded by tests |
benchmark-ocr-pass-a.json |
GPU pass A (ovis + qianfan, transformers 5.17.0), 12 rows — schema-guarded |
benchmark-ocr-pass-b.json |
GPU pass B (tele + unlimited, transformers 4.57.1), 15 rows — schema-guarded |
benchmark-docling-vs-native.json |
docling vs native-pdf, 2-page fixtures, 6 rows |