Skip to content

ParseCraft

Document intelligence: align existing converters, parsers, and OCR/VLM models behind one workflow and one typed output format

Python 3.13 | 3.14+ License CI PyPI

ParseCraft is a thin aggregation and alignment layer for document intelligence. Capable converters, parsers, and OCR/VLM models already exist, but each emits a different shape — and many ship demo-grade wrapper code pinned to an outdated Python version and dependency set. ParseCraft keeps every tool isolated behind a backend, translates its output into a single typed intermediate representation (IR), and projects that IR to deterministic Markdown. You get one workflow and one output format, while every heavy runtime stays optional.

The IR is the source of truth. Backends read a document and produce typed pages and chunks; every rendering — Markdown today, HTML/JSON/consumer trees later — is a projection of that same IR, never parsed back into state.

Requirements

  • Python 3.13 (GIL only — the free-threaded 3.13t build is not supported)
  • Python 3.14 or newer, including free-threaded builds (3.14t)
  • uv (recommended) or pip to install the package

Install and Quick Start

Install the released package and list the available backends:

uv add parsecraft      # or: pip install parsecraft
uv run parsecraft backends

First release: 2026.9.2. The current version is on PyPI.

Building from a source checkout, running the test suite, and working on the codebase are covered in the development setup.

CLI

Command Description
parsecraft backends List registered backends; load warnings go to stderr
parsecraft backends --json Emit backend descriptors as JSON
parsecraft config check Validate the effective configuration
parsecraft config show Show the resolved configuration with provenance
parsecraft --version (-v) Print the package version
uv run parsecraft backends
# No backends registered.        (no third-party backend installed yet)

See the CLI reference for every command, flag, exit code, and environment variable.

How it works

source document → backend (analyze / convert) → typed IR → Markdown projection
  • Backends are the extension point. A third party ships a BackendFactory under the parsecraft.backends entry-point group; the core package does not change to add one.
  • One typed IR. DocumentResult, PageResult, and StructuredChunk move between layers, and Markdown is a deterministic, dependency-free projection.
  • Heavy runtimes stay optional. Model weights, CUDA, and OCR/VLM stacks load only inside a backend factory, so importing ParseCraft stays offline-clean.
  • Failures are typed. A failed pass records a PassFailure rather than silently dropping a page.

Continue with the architecture overview, the IR model, the backend catalog, and the backend authoring guide.

Documentation

License

MIT — see LICENSE for details.


Documentation