Skip to content

Backends

A backend turns one source format into IR pages. Each tool stays isolated behind a DocumentBackend, and ParseCraft translates its result into the typed IR — the protocol and registry are described in Backends.

This page catalogs the in-package and external backends. Extra is the optional install that carries the backend; Availability separates what ships today from what is planned. Each third-party tool keeps its own licence.

Copyleft extras are opt-in

native-pdf extraction and OCR PDF input need PyMuPDF (AGPL-3.0-or-commercial, pdf extra), and pandoc needs the Pandoc binary (GPL-2.0-or-later). These are never core or dev dependencies. The OCR extras deliberately exclude PyMuPDF so the choice stays explicit. Installing one is your own licence decision; see ADR-0003.

Native (in-package)

Backend Extra Formats Upstream / licence Strengths Weaknesses Availability
native-text — .txt (text/plain) in-package — MIT Zero dependencies, fully offline, deterministic Paragraphs only; no headings, tables, or layout Available
native-markdown — .md (text/markdown) markdown-it-py — MIT CommonMark token structure: headings, lists, tables, code, quotes CommonMark only; no non-CommonMark dialects Available
native-html — .html (text/html) in-package (stdlib html.parser) — MIT Zero dependencies; headings, lists, tables, code, quotes Simplistic parser; scripts and styles skipped; not a full HTML5 engine Available
native-pdf pdf-lite (analysis) + pdf (extraction) .pdf (application/pdf) pypdf — BSD-3-Clause; pymupdf — AGPL-3.0-or-commercial Fast native text and layout extraction; analysis works without the AGPL extra Extraction requires the AGPL pdf extra; scanned PDFs need OCR Available

The native backends are registered through [project.entry-points."parsecraft.backends"] and appear in parsecraft backends.

External converters

Backend Extra Formats Upstream / licence Strengths Weaknesses Availability
pandoc parsecraft[pandoc] (pypandoc 1.17, MIT) docx, pptx, xlsx (read-only), odt, rtf, epub — the verified MIME set pandoc.org — GPL-2.0-or-later binary, MIT wrapper Broad office and e-book coverage through one wrapper Needs the external Pandoc binary; .ods/.odp are unsupported by pandoc 3.11; copyleft gate Available
liteparse parsecraft[liteparse] (liteparse 2.14.7, Apache-2.0) application/pdf, image/jpeg, image/png, image/tiff run-llama/liteparse — Apache-2.0 Broad and permissive; no copyleft gate Office/ODF formats require a system LibreOffice; .html is not supported; larger extra dependency surface Available
docling parsecraft[docling] (docling 2.130.0, MIT) application/pdf, text/html, text/markdown, text/plain docling — MIT Layout, tables, and reading order Heavy dependency graph; office/image formats are registered upstream but undeclared pending conversion verification; the motivating incident ran >1 h on a 113-page PDF Available

OCR / document-VLM (GPU)

Model pins and licences are recorded once in src/parsecraft/backends/ocr/_models.py. The adapters are implemented but not yet benchmarked with weights — GPU runs are pending.

Each backend has its own extra (pip install "parsecraft[ocr-ovis]"), and all four share one transformers window — transformers>=5.17,<6, declared once in pyproject.toml (the authoritative source) — plus torch>=2.5, torchvision, pillow, and accelerate. The per-model card pins (transformers==4.57.1 and similar) are advisory origin only: the project range supersedes them, and the TeleOCR/Unlimited models ship as local vendored modeling code so they load on the unified major.

The OCR extras are not mutually exclusive. Every extra installs jointly, and uv sync -U --all-extras --all-groups --all-packages is expected to succeed (root AGENTS.md rule 10). On a GPU host, install torch/torchvision from the PyTorch CUDA index first — the PyPI Windows wheels are CPU builds.

The optional vllm extra is Linux/WSL2-only: its marker (vllm>=0.11 ; sys_platform != 'win32') keeps Windows installs resolvable, but the vLLM runtime does not run there, so those backends use the default transformers runtime on Windows.

A host whose installed transformers falls outside the shared window gets a typed UnsupportedDependencyVersionError before any model load — the detail names the package, the actual version, and the required range; it is never a crash inside weight loading, and it is distinct from a missing extra.

The OCR extras exclude PyMuPDF, so PDF input additionally needs parsecraft[pdf] (AGPL — see the warning above). All OCR backends accept application/pdf, image/jpeg, and image/png.

Every backend declares supported_formats as MIME types — one vocabulary shared with RoutingConstraints.formats, which is built from the source's media type. The OCR backends declare application/pdf, image/png, and image/jpeg (the payloads their page-access layer accepts).

Backend Extra Input Model / licence Strengths Weaknesses Availability
ocr-ovis ocr-ovis page images ATH-MaaS/OvisOCR2 — Apache-2.0 ~0.9 B; formulas and tables; vLLM wrapper (optional vllm extra) GPU required (~1 GB VRAM); not yet benchmarked Available (adapter implemented; not yet benchmarked — GPU/weights pending)
ocr-tele ocr-tele images, scans StarDoc-AI/TeleOCR — Apache-2.0 (code) Geometry-aware; camera captures GPU required (~1.2 GB VRAM); not yet benchmarked Available (adapter implemented; not yet benchmarked — GPU/weights pending)
ocr-unlimited ocr-unlimited multi-page baidu/Unlimited-OCR — MIT Long-horizon multi-page documents 3 B; quantization required under an 8 GB budget; not yet benchmarked Available (adapter implemented; not yet benchmarked — GPU/weights pending)
ocr-qianfan ocr-qianfan images baidu/Qianfan-OCR — Apache-2.0 (Layout-as-Thought) Strong element, box, and reading-order control 4 B; quantization required under an 8 GB budget; not yet benchmarked Available (adapter implemented; not yet benchmarked — GPU/weights pending)

Model licences apply to the weights; the ParseCraft adapter code is MIT.

Languages

Backends declare BCP-47 tags in capabilities.languages; an empty tuple means no claim. A RoutingConstraints.language request only narrows a declaration — it never excludes a language-agnostic backend (see Language).

Backend Declared languages
native-text, native-markdown, native-html, native-pdf agnostic (no claim)
liteparse agnostic (no claim)
ocr-tele zh, en
ocr-ovis, ocr-unlimited, ocr-qianfan agnostic (multilingual)
pandoc, docling agnostic (no claim)

Choosing a backend

Document trait Route Why
Digital text PDF native-pdf Fast native text extraction, no GPU
Scanned pages or camera photos ocr-tele, ocr-ovis No native text; geometry-aware, formula-capable
Multi-page dense tables ocr-unlimited Long-horizon multi-page handling
Formulas and equations ocr-ovis Formula-aware extraction
Broad office formats (docx, pptx, xlsx) pandoc Convert to a readable format first
Layout-heavy, reading order matters docling Layout, tables, and reading order

Native extraction runs before OCR. OCR is selective and expensive: it is used per page range, under a time budget, only where analysis shows native text is missing or unusable.

Auto mode

auto mode is decided by parsecraft.routing. analyze() produces per-page signals, one rule table classifies each page into an Intent, eligibility rules filter the catalog, and plan_route() returns a RoutingPlan with ordered fallbacks. Planning is deterministic, pure, and offline. parsecraft.pipeline.execute() runs that plan: it groups contiguous pages, converts each group with the chosen backend, retries fallback candidates on a typed failure, and aggregates a DocumentResult.

Hard constraints stay code-owned: installed extras, VRAM budget, format coverage, max_passes, allow_ocr, and offline are enforced before any judge sees a candidate, and a NATIVE page requires at least one eligible non-OCR backend. An optional RoutingJudge may re-rank eligible candidates only; DeterministicJudge is the default, and a judge cannot override a constraint. A Jev / System One-backed judge is a planned optional extra.

See Routing and auto mode and ADR-0004. Tracking bead: pc-5ub.