Skip to content

Benchmarks

Phase 2 measurement report for ParseCraft: real backends over the pinned local corpus, produced by the in-repo harness (parsecraft.benchmark) on 2026-09-27.

What was measured

Documents (all local, never fetched at run time) from the corpus staged in tests/downloads/:

document format pages
…-pg1342.txt (Gutenberg, Pride & Prejudice, 754 KiB) text/plain 1
…-rfc5322.txt (119 KiB) text/plain 1
…-commonmark-spec.md (201 KiB) text/markdown 1
whatwg-html.html (151 KiB) text/html 1
wikipedia-calculus.html (792 KiB) text/html 1
itu-t-p863.pdf (licensed, 2.3 MiB) application/pdf 80
etsi-ts-103558.pdf (licensed, 2.7 MiB) application/pdf 68
…-nist-sp-800-53r5.pdf (5.8 MiB) application/pdf 492

Backends: native-text, native-markdown, native-html, native-pdf (CPU), plus liteparse and the GPU OCR backends where noted below. The harness measures every backend that is eligible for each document (is_hard_eligible + the document's media type), one conversion per cell.

Environment: Windows laptop, CPython 3.13, CPU only for the table below; light extras (download, web, pdf-lite, pdf, liteparse) installed. No GPU work is included in this snapshot — see Gaps.

Reproduce

uv run --extra pdf --extra liteparse --extra download --extra web --extra pdf-lite \
  parsecraft benchmark tests/downloads/*.html tests/downloads/*.md \
    tests/downloads/*.pdf tests/downloads/*.txt \
  -o docs/benchmarks/

Notes:

  • --extra pdf pulls PyMuPDF (AGPL) explicitly — the copyleft choice stays an operator decision (ADR-0003); native-pdf records a typed dependency_missing row instead of crashing when it is absent.
  • The --extra flags install transiently: plain uv run does not prune them afterwards, and the canonical dev environment (which ty's optional-import ignore comments assume) is restored with mise dev. Run mise dev before the quality gate if you benchmarked first.
  • The report writers are deterministic: stable key and row ordering, fixed rounding, LF newlines, no wall-clock fields. Re-running on the same inputs yields the same schema and row set — the measured values (elapsed, pages/s, peak bytes) are run data and will vary.
  • tests/test_benchmark_report.py guards the committed artifacts: byte round-trip through the current writers, pinned schema, sorted rows.

Results

The full harness output is committed alongside this page: benchmark.md and benchmark.json.

Headlines (CPU pass, 2026-09-27; values as committed in benchmark.json):

document backend elapsed pages/s coverage
itu-t-p863.pdf (80 pp) native-pdf 0.612 s 130.8 0.999
nist-sp-800-53r5.pdf (492 pp) native-pdf 4.002 s 122.9 0.995
etsi-ts-103558.pdf (68 pp, image-heavy) native-pdf 6.468 s 10.5 0.963
itu-t-p863.pdf (80 pp) liteparse 5.672 s 14.1 1.017
nist-sp-800-53r5.pdf (492 pp) liteparse 115.647 s 4.3 0.904
etsi-ts-103558.pdf (68 pp) liteparse 6.011 s 11.3 1.012
pg1342.txt (754 KiB) native-text 0.002 s 409.9 0.988
rfc5322.txt native-text 0.009 s 116.6 0.976
commonmark-spec.md (201 KiB) native-markdown 0.739 s 1.4* 0.717
whatwg-html.html native-html 0.873 s 1.1* 0.271
wikipedia-calculus.html native-html 2.460 s 0.4* 0.106

* Single-page formats: "pages/s" is one page over elapsed time — read the elapsed figure instead. Coverage may exceed 1.0 (liteparse emits slightly more characters than the analyzer counted — different segmentation, not a bug).

GPU OCR legs (bounded fixtures)

The four OCR models ran in the isolated GPU environment (.venv-gpu, RTX A2000 8 GB) over 2-page fixtures of the three corpus PDFs (.tmp/ocr-bench/*-p1-2.pdf) — full-corpus OCR would take hours. At run time the four models pinned two different transformers majors, so the snapshot is two passes over the same fixtures. That split was a pinning fact, not a model requirement: pc-4u7.36 unified the extras on transformers>=5.17,<6 and the TeleOCR/Unlimited models now load from local vendored modeling code, so a re-run needs one environment and one pass (see Deferred runs).

fixture pass A: ovis (5.17) pass A: qianfan (5.17) pass B: tele (4.57.1) pass B: unlimited (4.57.1)
etsi-ts-103558-p1-2 197.6 s 713.7 s 220.2 s 751.7 s
nist-…-p1-2 212.6 s backend_error 81.7 s 128.0 s
itu-t-p863-p1-2 413.8 s 1024.0 s 84.0 s 133.6 s

Cross-version rows behave exactly as typed: the wrong-major pair records backend_error rows whose detail names the cause (e.g. qwen3_5/ qianfan_ocr unknown to transformers 4.57.1), never a silent pass. liteparse rows (0.07–0.16 s for the fixtures) appear in pass B after an entry-point refresh.

Reading rules (also pc-4's run notes):

  • elapsed_s includes per-document model load — the runner constructs the backend per document; these are 2-page fixtures, not per-page rates you can multiply to the full corpus.
  • The two passes are different virtualenvs: merge rows per backend, never sum across passes.
  • Full-corpus qianfan is a dated deferral, not a gap — see Deferred runs.
  • selected is recorded before load success/failure: a selected: true + failures: [backend_error] row means “planner picked it, then it could not load in this transformers” — not a planner bug.

Deferred runs

The legs below were deliberately deferred on 2026-09-28, with the reason recorded so a later run is a scheduling decision rather than a missing result. None of them failed: the measured fixtures already prove the adapters work, and the full-corpus runs only cost machine time.

Run Deferred on Reason Command
Full-corpus OCR, all four models 2026-09-28 Cost/time only: the 2-page fixtures already take 82–1024 s per document, so the corpus is many hours of exclusive GPU time. Not an install conflict — the extras are jointly installable on one transformers window Full-corpus command
Full-corpus ocr-qianfan 2026-09-28 Cost/time: ≈ 145 s/page measured → ≈ 20 h for NIST SP 800-53r5 alone. Weights are staged and sha-verified (20 GB cache), so the run is offline-ready same command (qianfan is one of the measured backends)
Full-corpus docling 2026-09-28 Runtime budget: docling is ≈ 8–15 s/page warm on this laptop → ≈ 1.8 h for NIST SP 800-53r5 alone, and the incident shape (>1 h on 113 pp) says hours. Not a CI-shaped job Full-corpus docling
Quantised variants (int8 / 4-bit) 2026-09-28 Out of scope by design: quantisation is a backend-installation choice, not a harness dimension —

Full-corpus OCR command

The GPU environment is .venv-gpu (RTX A2000 8 GB). All four OCR extras share one transformers>=5.17,<6 window (pyproject.toml), so the full run is a single pass in a single environment. The command prints JSON on stdout and never overwrites the committed CPU report; --extra pdf is required for PDF rasterisation (PyMuPDF, AGPL — ADR-0003).

uv run --project .venv-gpu --extra ocr-ovis --extra ocr-qianfan \
  --extra ocr-tele --extra ocr-unlimited --extra pdf --extra download \
  parsecraft benchmark tests/downloads/itu-t-p863.pdf \
    tests/downloads/etsi-ts-103558.pdf tests/downloads/nist-sp-800-53r5.pdf \
  --json > docs/benchmarks/benchmark-ocr-full.json

Full-corpus docling command

Resumable: each (document, backend) pair is persisted as it finishes.

uv run --extra docling --extra pdf --extra pdf-lite \
  python scripts/bench_docling.py tests/downloads/*.pdf --backends docling

The committed bounded runs stay the evidence for adapter correctness: benchmark-ocr-pass-a.json and benchmark-ocr-pass-b.json.

Cross-check: the motivating incident

ADR-0001 records the incident this package exists to fix: Docling ran for >1 hour on a 113-page ITU-T P.863 PDF at full CPU/RAM and was cancelled (≈ >31.9 s/page, never finished).

Measured here on the staged ITU-T P.863 artifact (80 pages, same document family — the staged file is not byte-identical to the incident's 113-page bundle): native-pdf extracts 80 pages in 0.612 s (≈ 7.7 ms/page), and the 492-page NIST SP 800-53r5 in 4.00 s (122.9 pages/s). That is more than three orders of magnitude below the incident's per-page cost — with honest caveats: this is text-layer extraction into typed chunks, not Docling's full layout/table pipeline. The routing design leans exactly on this gap: run the cheap native pass first, escalate to OCR/layout backends only for pages whose signals say native extraction cannot do the job (ADR-0001 §6 keeps the 15-minute cold-cache budget for the whole document, escalations included).

Docling vs native-pdf (bounded head-to-head)

The gap the incident cross-check only inferred is now measured on 2-page fixtures of the three corpus PDFs (pc-4u7.39, docling 2.130.0, 2026-09-28). Environment: Windows laptop, CPython 3.13, CPU only. Docling ran its default pipeline untuned (layout, tables, and its bundled RapidOCR); native-pdf used PyMuPDF. Both were measured by the in-repo harness.

fixture (2 pp) docling native-pdf ratio
itu-t-p863 (of 80) 20.6 s 0.014 s ≈ 1470×
etsi-ts-103558 (of 68) 30.5 s 2.98 s ≈ 10×
nist-sp-800-53r5 (of 492) 15.5 s 0.024 s ≈ 650×
  • Docling's first conversion in a process is far slower — 191.9 s for the p863 fixture — because it loads its layout/OCR models. The table is warm; its peak Python memory is 158–405 MB against 0.03–15 MB for native-pdf.
  • Docling emits more structure (headings, tables) at comparable text coverage; the cost of that is runtime and memory, not correctness.
  • Ratios vary per document (the ETSI fixture is the most docling-favourable of the set) — read the per-row numbers, not a single ratio.

This is an on-demand measurement, never part of mise test/CI:

# 1. two-page fixtures of the pinned corpus PDFs (pymupdf, transient)
uv run --extra pdf python - <<'PY'
from pathlib import Path

import pymupdf

out = Path(".tmp/docling-bench")
out.mkdir(parents=True, exist_ok=True)
for src in Path("tests/downloads").glob("*.pdf"):
    doc = pymupdf.open(src)
    two = pymupdf.open()
    try:
        two.insert_pdf(doc, from_page=0, to_page=min(1, doc.page_count - 1))
        two.save(out / f"{src.stem}-p1-2.pdf")
    finally:
        two.close()
        doc.close()
PY

# 2. one (document, backend) pair at a time, resumable
mise run bench-docling

The equivalent explicit command:

uv run --extra docling --extra pdf --extra pdf-lite \
  python scripts/bench_docling.py .tmp/docling-bench/*.pdf

scripts/bench_docling.py writes each measured (document, backend) pair to .tmp/bench-docling-partial.json immediately and skips pairs already present on re-invocation, so a long run resumes instead of restarting. It never uses the conversion cache — a cache hit would report ~0 s and falsify the measurement.

Honest gaps

  • Full-corpus OCR was deferred, not skipped. The GPU legs cover 2-page fixtures only (see GPU OCR legs); scaling any OCR number here to a whole document would be fabricated. The dated deferral and exact commands are in Deferred runs. The pinned corpus fetch tier (mise run test-corpus) stays opt-in; the licensed PDFs are hand-staged and never downloaded by the harness.
  • Full-corpus qianfan is a dated deferral (≈ 145 s/page measured; ≈ 20 h for NIST alone) — see Deferred runs. All four models are staged and sha-verified (20 GB cache), so the deferred run passes offline asset checks.
  • VRAM is not measured here. Peak bytes are tracemalloc (host Python allocations). GPU VRAM is covered by the separate gpu test tier (mise run test-gpu, RTX A2000 8 GB pins in pyproject.toml).
  • No quantisation variants were benchmarked — model quantisation stays a backend-installation concern, not a harness dimension.
  • exoplanets.csv and wikidata-q42.json were skipped (unsupported source suffix in the CLI media-type map). The skips are recorded in the report, not hidden.
  • HTML coverage is a structure proxy. native-html emits headings, paragraphs, lists, tables, code, and quotes; raw inline markup dominates the analyzer's character count, so 0.11–0.27 coverage is expected for markup-heavy pages, not data loss.
  • itu-t-p863.pdf shows selected: -: its page signals routed some pages to OCR intents while no OCR backend was installed in this environment, so auto-mode planning correctly declined (selected is null) — the direct per-backend measurement still ran.

Artifacts

file contents
benchmark.md full Markdown report (harness output)
benchmark.json full JSON report, schema-guarded by tests
benchmark-ocr-pass-a.json GPU pass A (ovis + qianfan, transformers 5.17.0), 12 rows — schema-guarded
benchmark-ocr-pass-b.json GPU pass B (tele + unlimited, transformers 4.57.1), 15 rows — schema-guarded
benchmark-docling-vs-native.json docling vs native-pdf, 2-page fixtures, 6 rows