liteparse vs pdf-inspector
LlamaIndex's local document parser in Rust: PDFium text with bounding boxes at ~2-5 ms/page, selective Tesseract or HTTP OCR, Markdown/JSON output, screenshots; Python, Node, WASM. — versus — Firecrawl's Rust PDF triage: classifies text-based vs scanned in ~10-50ms, extracts positioned text and clean Markdown without OCR — routing the ~54% of PDFs that never needed a model.
Both are Rust fast paths that extract positioned PDF text and skip OCR when a page does not need it; pdf-inspector triages text versus scanned, LiteParse also runs selective OCR, converts Office files and renders screenshots.
| liteparse | pdf-inspector | |
|---|---|---|
| Stars | 13k | 19k |
| Forks | 870 | 1.3k |
| Language | Rust | Rust |
| License | Apache-2.0 | MIT |
| Last activity | today | 3 days ago |
| Topics | ocr, rag | ocr, rag |
| Curated connections | 4 | 8 |
liteparse — the curator's take
The fast first pass of a document pipeline: PDFium spatial text with bounding boxes in a few milliseconds per page, OCR only where a page needs it, a cheap complexity check that tells you before parsing whether a file needs heavier treatment, and worker pools that kill runaway documents instead of stalling. Runs locally from Rust, Python, Node or the browser. Its limits are stated up front: dense tables, multi-column layouts, charts, handwriting and scans need a layout model (marker, mineru or chandra locally, or LlamaParse, the cloud product this README points to). Office files and images are converted through LibreOffice, which is one more dependency on a server.
pdf-inspector — the curator's take
The router your document pipeline is missing: most stacks OCR everything, but ~54% of PDFs are text-based — this classifies in tens of milliseconds (with confidence and per-page routing), extracts locally in under 200ms, and only the genuinely scanned pages go to an expensive OCR model. Pure Rust, one dependency, bindings for Python, Node and browser WASM. NOT an OCR engine and NOT a layout-analysis heavyweight: scanned documents still need a model downstream, and complex-layout fidelity trails ML parsers — its job is knowing when you don't need them.