StackMap
Subscribe
Explore / liteparse
run-llama

liteparse

LlamaIndex's local document parser in Rust: PDFium text with bounding boxes at ~2-5 ms/page, selective Tesseract or HTTP OCR, Markdown/JSON output, screenshots; Python, Node, WASM.

12,756 870 Rust Apache-2.0updated today
View on GitHubDispute this mapping →
Curator's take

The fast first pass of a document pipeline: PDFium spatial text with bounding boxes in a few milliseconds per page, OCR only where a page needs it, a cheap complexity check that tells you before parsing whether a file needs heavier treatment, and worker pools that kill runaway documents instead of stalling. Runs locally from Rust, Python, Node or the browser. Its limits are stated up front: dense tables, multi-column layouts, charts, handwriting and scans need a layout model (marker, mineru or chandra locally, or LlamaParse, the cloud product this README points to). Office files and images are converted through LibreOffice, which is one more dependency on a server.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside liteparse. Ranked by curator confidence.

pairs wellpairs wellalternativealternativeLlamaIndexmarkerpdf-inspectoropendataloader-pdfliteparse
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md2 min read

LiteParse

CI | Crates.io version | npm version | wasm version | PyPI version | License | Docs

English | 简体中文

out

Looking for LiteParse V1? Follow this link to the old code

LiteParse is a standalone OSS PDF parsing tool focused exclusively on fast and light parsing. It provides high-quality spatial text parsing with bounding boxes, without proprietary LLM features or cloud dependencies. Everything runs locally on your machine.

Hitting the limits of local parsing? For complex documents (dense tables, multi-column layouts, charts, handwritten text, or scanned PDFs), you'll get significantly better results with LlamaParse, our cloud-based document parser built for production document pipelines. LlamaParse handles the hard stuff so your models see clean, structured data and markdown.

Sign up for LlamaParse free

Overview

  • Fast Text Parsing: Spatial text parsing using PDFium, ~2-5ms per page
  • Flexible OCR System:
    • Built-in: Tesseract (zero setup, bundled with the library)
    • HTTP Servers: Plug in any OCR server (EasyOCR, PaddleOCR, custom)
    • Standard API: Simple, well-defined OCR API specification
  • Complexity Detection: Cheaply check whether a document needs OCR or heavier parsing — route, reject, or estimate cost before a full parse
  • Screenshot Generation: Generate high-quality page screenshots for LLM agents
  • Multiple Output Formats: Markdown, JSON, and Text
  • Markdown Output: Structured Markdown with headings, tables, lists, images, and links — great for feeding LLMs and RAG pipelines
  • Bounding Boxes: Precise text positioning information
  • Multi-language: Use from Rust, Node.js/TypeScript, Python, or the browser (WASM)
  • Worker Pool Mode (Python & Node.js): Parse in persistent worker processes for true parallelism (PDFium otherwise serializes concurrent parses) and hard per-parse timeouts — rogue documents are killed, identified by name, and never stall the pipeline
  • Multi-platform: Linux, macOS (Intel/ARM), Windows
flowchart LR
      subgraph Input["Input Formats"]
          direction TB
          PDF["PDF"]
          DOCX["DOCX"]
          XLSX["XLSX"]
          PPTX["PPTX"]
          IMG["Images"]
      end

      subgraph Core["Rust Core"]
          direction TB
          CONV["Format Conversion\nLibreOffice / Rust image + resvg + usvg crates"]
          EXTRACT["Text Extraction\nPDFium C library"]
          OCR["Selective OCR\nTesseract / HTTP / Custom"]
          MERGE["OCR Merge\nNative text + OCR results"]
          PROJ["Grid Projection\nSpatial layout reconstruction"]
          CONV --> EXTRACT
          EXTRACT --> OCR --> MERGE --> PROJ
          EXTRACT --> MERGE
      end

      subgraph Output[" Output "]
          direction TB
          JSON["Structured JSON\ntext + bounding boxes"]
          TEXT["Plain Text\nlayout-preserved"]
          SCREEN["Screensh