[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:pdf-inspector":3},"\u003Ch1>pdf-inspector\u003C\u002Fh1>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fcrates.io\u002Fcrates\u002Fpdf-inspector\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fcrates\u002Fv\u002Fpdf-inspector.svg\" alt=\"Crates.io\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fwww.npmjs.com\u002Fpackage\u002F@firecrawl\u002Fpdf-inspector\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fnpm\u002Fv\u002F@firecrawl\u002Fpdf-inspector.svg\" alt=\"npm\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fpdf-inspector\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fpdf-inspector.svg\" alt=\"PyPI\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Ffirecrawl\u002Fpdf-inspector\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-MIT-blue.svg\" alt=\"License: MIT\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Ffirecrawl\u002Fpdf-inspector\u002Fblob\u002FHEAD\u002Fdocs\u002Fpython.md\" rel=\"nofollow ugc noopener\">Python\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Ffirecrawl\u002Fpdf-inspector\u002Fblob\u002FHEAD\u002Fnapi\u002FREADME.md\" rel=\"nofollow ugc noopener\">Node.js\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Ffirecrawl\u002Fpdf-inspector\u002Fblob\u002FHEAD\u002Fwasm\u002FREADME.md\" rel=\"nofollow ugc noopener\">browser WebAssembly\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Built by \u003Ca href=\"https:\u002F\u002Ffirecrawl.dev\" rel=\"nofollow ugc noopener\">Firecrawl\u003C\u002Fa> to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.\u003C\u002Fp>\n\u003Ch2>Features\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Smart classification\u003C\u002Fstrong> — Detect TextBased, Scanned, ImageBased, or Mixed PDFs in ~10-50ms by sampling content streams. Returns a confidence score (0.0-1.0) and per-page OCR routing.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Text extraction\u003C\u002Fstrong> — Position-aware extraction with font info, X\u002FY coordinates, and automatic multi-column reading order.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Markdown conversion\u003C\u002Fstrong> — Headings (H1-H4 via font size ratios), bullet\u002Fnumbered\u002Fletter lists, code blocks (monospace font detection), tables (rectangle-based and heuristic), bold\u002Fitalic formatting, URL linking, and page breaks.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Table detection\u003C\u002Fstrong> — Dual-mode: rectangle-based detection from PDF drawing ops, plus heuristic detection from text alignment. Handles financial tables, footnotes, and continuation tables across pages.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>CID font support\u003C\u002Fstrong> — ToUnicode CMap decoding for Type0\u002FIdentity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Multi-column layout\u003C\u002Fstrong> — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Encoding issue detection\u003C\u002Fstrong> — Automatically flags broken font encodings so callers can fall back to OCR.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Single document load\u003C\u002Fstrong> — The document is parsed once and shared between detection and extraction, avoiding redundant I\u002FO.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Browser WebAssembly\u003C\u002Fstrong> — Run the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Lightweight\u003C\u002Fstrong> — Pure Rust, no ML models, no external services. Single dependency on \u003Ccode>lopdf\u003C\u002Fcode> for PDF parsing.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Benchmark\u003C\u002Fh2>\n\u003Cp>Evaluated on the \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fopendataloader-project\u002Fopendataloader-bench\" rel=\"nofollow ugc noopener\">opendataloader-bench\u003C\u002Fa> corpus (200 PDFs). Only local engines without model-based PDF parsing are shown; OCR was disabled. Scores are 0-1, higher is better.\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Engine\u003C\u002Fth>\n\u003Cth>Overall\u003C\u002Fth>\n\u003Cth>Reading Order (NID)\u003C\u002Fth>\n\u003Cth>Tables (TEDS)\u003C\u002Fth>\n\u003Cth>Headings (MHS)\u003C\u002Fth>\n\u003Cth>Speed (200 docs)\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>pdf-inspector\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>0.875\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>0.915\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>0.814\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>0.788\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>2.8s\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>liteparse\u003C\u002Ftd>\n\u003Ctd>0.870\u003C\u002Ftd>\n\u003Ctd>0.908\u003C\u002Ftd>\n\u003Ctd>0.693\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>0.811\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>13.9s\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>opendataloader\u003C\u002Ftd>\n\u003Ctd>0.843\u003C\u002Ftd>\n\u003Ctd>0.912\u003C\u002Ftd>\n\u003Ctd>0.489\u003C\u002Ftd>\n\u003Ctd>0.760\u003C\u002Ftd>\n\u003Ctd>9.8s\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>pymupdf4llm\u003C\u002Ftd>\n\u003Ctd>0.735\u003C\u002Ftd>\n\u003Ctd>0.886\u003C\u002Ftd>\n\u003Ctd>0.401\u003C\u002Ftd>\n\u003Ctd>0.424\u003C\u002Ftd>\n\u003Ctd>15.5s\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>markitdown\u003C\u002Ftd>\n\u003Ctd>0.583\u003C\u002Ftd>\n\u003Ctd>0.879\u003C\u002Ftd>\n\u003Ctd>0.000\u003C\u002Ftd>\n\u003Ctd>0.000\u003C\u002Ftd>\n\u003Ctd>6.7s\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\n\u003Cp>Results were refreshed on July 16, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Speed is the median of three complete corpus runs.\u003C\u002Fp>\n\u003Cp>For context, engines that use OCR or model-based document parsing (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the top of that range without either, in 2.8 seconds.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Best fit:\u003C\u002Fstrong> Native-text PDFs where speed, reading order, and table structure matter. pdf-inspector delivered the highest overall, reading-order, and table scores, along with the fastest complete run in this benchmark. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.\u003C\u002Fp>\n\u003Cp>Use the [paired benchmark harness](docs\u002F\u003C\u002Fp>\n",1784844295980]