[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:liteparse":3},"\u003Ch1>LiteParse\u003C\u002Fh1>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Frun-llama\u002Fliteparse\u002Factions\u002Fworkflows\u002Fci.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Frun-llama\u002Fliteparse\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"CI\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fcrates.io\u002Fcrates\u002Fliteparse\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fcrates\u002Fv\u002Fliteparse.svg\" alt=\"Crates.io version\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fwww.npmjs.com\u002Fpackage\u002F@llamaindex\u002Fliteparse\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fnpm\u002Fv\u002F@llamaindex\u002Fliteparse.svg\" alt=\"npm version\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fwww.npmjs.com\u002Fpackage\u002F@llamaindex\u002Fliteparse-wasm\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fnpm\u002Fv\u002F@llamaindex\u002Fliteparse-wasm.svg\" alt=\"wasm version\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fliteparse\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fliteparse.svg\" alt=\"PyPI version\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fopensource.org\u002Flicenses\u002FApache-2.0\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache%202.0-blue.svg\" alt=\"License\" \u002F>\u003C\u002Fa>\n|\n\u003Ca href=\"https:\u002F\u002Fdevelopers.llamaindex.ai\u002Fliteparse\u002F\" rel=\"nofollow ugc noopener\">Docs\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>English | \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Frun-llama\u002Fliteparse\u002Fblob\u002FHEAD\u002FREADME.zh-CN.md\" rel=\"nofollow ugc noopener\">简体中文\u003C\u002Fa>\u003C\u002Fp>\n\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F07ba6a82-6bb1-4dea-b0ef-cad7df7d1622\" alt=\"out\" width=\"600\" \u002F>\u003Cblockquote>\n\u003Cp>Looking for LiteParse V1? Follow this link to \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Frun-llama\u002Fliteparse\u002Ftree\u002Flogan\u002Fliteparse-v1\" rel=\"nofollow ugc noopener\">the old code\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>LiteParse is a standalone OSS PDF parsing tool focused exclusively on \u003Cstrong>fast and light\u003C\u002Fstrong> parsing. It provides high-quality spatial text parsing with bounding boxes, without proprietary LLM features or cloud dependencies. Everything runs locally on your machine.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Hitting the limits of local parsing?\u003C\u002Fstrong>\nFor complex documents (dense tables, multi-column layouts, charts, handwritten text, or\nscanned PDFs), you'll get significantly better results with \u003Ca href=\"https:\u002F\u002Fdevelopers.llamaindex.ai\u002Fpython\u002Fcloud\u002Fllamaparse\u002F?utm_source=github&amp;utm_medium=liteparse\" rel=\"nofollow ugc noopener\">LlamaParse\u003C\u002Fa>,\nour cloud-based document parser built for production document pipelines. LlamaParse handles the\nhard stuff so your models see clean, structured data and markdown.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp> \u003Ca href=\"https:\u002F\u002Fcloud.llamaindex.ai?utm_source=github&amp;utm_medium=liteparse\" rel=\"nofollow ugc noopener\">Sign up for LlamaParse free\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Ch2>Overview\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Fast Text Parsing\u003C\u002Fstrong>: Spatial text parsing using PDFium, ~2-5ms per page\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Flexible OCR System\u003C\u002Fstrong>:\u003Cul>\n\u003Cli>\u003Cstrong>Built-in\u003C\u002Fstrong>: Tesseract (zero setup, bundled with the library)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>HTTP Servers\u003C\u002Fstrong>: Plug in any OCR server (EasyOCR, PaddleOCR, custom)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Standard API\u003C\u002Fstrong>: Simple, well-defined OCR API specification\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Complexity Detection\u003C\u002Fstrong>: Cheaply check whether a document needs OCR or heavier parsing — route, reject, or estimate cost before a full parse\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Screenshot Generation\u003C\u002Fstrong>: Generate high-quality page screenshots for LLM agents\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Multiple Output Formats\u003C\u002Fstrong>: Markdown, JSON, and Text\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Markdown Output\u003C\u002Fstrong>: Structured Markdown with headings, tables, lists, images, and links — great for feeding LLMs and RAG pipelines\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Bounding Boxes\u003C\u002Fstrong>: Precise text positioning information\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Multi-language\u003C\u002Fstrong>: Use from Rust, Node.js\u002FTypeScript, Python, or the browser (WASM)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Worker Pool Mode\u003C\u002Fstrong> (Python &amp; Node.js): Parse in persistent worker processes for true parallelism (PDFium otherwise serializes concurrent parses) and hard per-parse timeouts — rogue documents are killed, identified by name, and never stall the pipeline\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Multi-platform\u003C\u002Fstrong>: Linux, macOS (Intel\u002FARM), Windows\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cpre>\u003Ccode class=\"language-mermaid\">flowchart LR\n      subgraph Input[\"Input Formats\"]\n          direction TB\n          PDF[\"PDF\"]\n          DOCX[\"DOCX\"]\n          XLSX[\"XLSX\"]\n          PPTX[\"PPTX\"]\n          IMG[\"Images\"]\n      end\n\n      subgraph Core[\"Rust Core\"]\n          direction TB\n          CONV[\"Format Conversion\\nLibreOffice \u002F Rust image + resvg + usvg crates\"]\n          EXTRACT[\"Text Extraction\\nPDFium C library\"]\n          OCR[\"Selective OCR\\nTesseract \u002F HTTP \u002F Custom\"]\n          MERGE[\"OCR Merge\\nNative text + OCR results\"]\n          PROJ[\"Grid Projection\\nSpatial layout reconstruction\"]\n          CONV --&gt; EXTRACT\n          EXTRACT --&gt; OCR --&gt; MERGE --&gt; PROJ\n          EXTRACT --&gt; MERGE\n      end\n\n      subgraph Output[\" Output \"]\n          direction TB\n          JSON[\"Structured JSON\\ntext + bounding boxes\"]\n          TEXT[\"Plain Text\\nlayout-preserved\"]\n          SCREEN[\"Screensh\n\u003C\u002Fcode>\u003C\u002Fpre>\n",1790887801925]