[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:ontocast":3},"\u003Ch1>OntoCast \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fgrowgraph\u002Fontocast\u002Frefs\u002Fheads\u002Fmain\u002Fdocs\u002Fassets\u002Ffavicon.ico\" alt=\"OntoCast logo\" \u002F>\u003C\u002Fh1>\n\u003Cp>\u003Cstrong>Agentic ontology-assisted extraction of RDF knowledge graphs from documents.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.12%2B-blue.svg\" alt=\"Python\" \u002F>\n\u003Ca href=\"https:\u002F\u002Fbadge.fury.io\u002Fpy\u002Fontocast\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fbadge.fury.io\u002Fpy\u002Fontocast.svg\" alt=\"PyPI version\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpepy.tech\u002Fprojects\u002Fontocast\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fstatic.pepy.tech\u002Fbadge\u002Fontocast\" alt=\"PyPI Downloads\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgrowgraph.github.io\u002Fontocast\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fdocs-growgraph.github.io-orange.svg\" alt=\"Docs\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fopensource.org\u002Flicenses\u002FApache-2.0\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache_2.0-blue.svg\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fgrowgraph\u002Fontocast\u002Factions\u002Fworkflows\u002Fpre-commit.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fgrowgraph\u002Fontocast\u002Factions\u002Fworkflows\u002Fpre-commit.yml\u002Fbadge.svg\" alt=\"pre-commit\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.17796467\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fzenodo.org\u002Fbadge\u002FDOI\u002F10.5281\u002Fzenodo.17796467.svg\" alt=\"DOI\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>OntoCast turns unstructured text into queryable RDF: it \u003Cstrong>co-evolves\u003C\u002Fstrong> domain ontologies and fact graphs in a parallel map\u002Freduce pipeline, with RDF 1.2 provenance, entity disambiguation across chunks, and optional vector-backed ontology retrieval. Run it as a REST service, a batch CLI, or embed the pipeline in your own LangChain \u002F LangGraph agent.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Documentation:\u003C\u002Fstrong> \u003Ca href=\"https:\u002F\u002Fgrowgraph.github.io\u002Fontocast\u002F\" rel=\"nofollow ugc noopener\">growgraph.github.io\u002Fontocast\u003C\u002Fa>\u003C\u002Fp>\n\u003Chr \u002F>\n\u003Ch2>Why OntoCast\u003C\u002Fh2>\n\u003Cp>Most extractors dump triples and leave ontology drift to you. OntoCast treats schema and instance data as one loop: per-chunk render → critic → merge, with GraphUpdate patches (insert\u002Fdelete) instead of regenerating whole graphs, SHACL validation with LLM-free autofix, and a light install so you can embed the core without pulling Docling, gRPC, or ONNX.\u003C\u002Fp>\n\u003Chr \u002F>\n\u003Ch2>Features\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Parallel ontology + facts loops\u003C\u002Fstrong> — concurrent per-unit render\u002Fcritic with configurable workers\u003C\u002Fli>\n\u003Cli>\u003Cstrong>GraphUpdate patches\u003C\u002Fstrong> — token-efficient insert\u002Fdelete ops, not full-graph regeneration\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Entity disambiguation\u003C\u002Fstrong> — embedding + symbolic alignment across chunks\u003C\u002Fli>\n\u003Cli>\u003Cstrong>RDF 1.2 provenance\u003C\u002Fstrong> — quoted triples \u002F provenance artifacts; optional \u003Ccode>strip_provenance\u003C\u002Fcode>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Ontology context\u003C\u002Fstrong> — catalog selection, vector retrieval (LanceDB or Qdrant), or a fixed ontology\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Facts validation\u003C\u002Fstrong> — invariants, SHACL, and machine repairs without an extra LLM pass\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Stores\u003C\u002Fstrong> — in-memory pyoxigraph by default; Fuseki for persistence; tenancy by tenant\u002Fproject\u003C\u002Fli>\n\u003Cli>\u003Cstrong>LLM caching\u003C\u002Fstrong> — disk cache, in-flight limits, optional read-only \u002F batch pre-warm\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Embeddable\u003C\u002Fstrong> — \u003Ccode>ontocast_tools\u003C\u002Fcode>, \u003Ccode>run_unit_pipeline\u003C\u002Fcode>, or a LangGraph node\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Chr \u002F>\n\u003Ch2>Install\u003C\u002Fh2>\n\u003Cp>Pick at least one LLM provider extra. Add \u003Ccode>server\u003C\u002Fcode> for the CLI and HTTP API:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-sh\">uv add \"ontocast[server,openai]\"\n# or: pip install \"ontocast[server,openai]\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Common add-ons: \u003Ccode>doc-processing\u003C\u002Fcode> (PDF\u002FDOCX), \u003Ccode>lancedb\u003C\u002Fcode> or \u003Ccode>qdrant\u003C\u002Fcode> (ontology retrieval), \u003Ccode>shacl\u003C\u002Fcode> (shape validation).\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-sh\">uv add \"ontocast[server,openai,doc-processing,lancedb,shacl]\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Full extras table: \u003Ca href=\"https:\u002F\u002Fgrowgraph.github.io\u002Fontocast\u002Fgetting_started\u002Finstallation\u002F\" rel=\"nofollow ugc noopener\">Installation\u003C\u002Fa>.\u003C\u002Fp>\n\u003Chr \u002F>\n\u003Ch2>Quick start\u003C\u002Fh2>\n\u003Cpre>\u003Ccode class=\"language-bash\">cp .env.example .env\n# Set LLM_API_KEY (and LLM_PROVIDER \u002F LLM_MODEL_NAME as needed)\n\nontocast serve\ncurl -X POST http:\u002F\u002Flocalhost:8999\u002Fprocess -F \"file=@document.pdf\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Batch without a server:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">ontocast process --input-path .\u002Fdocument.pdf --head-chunks 5 --output-dir .\u002Fout\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Omit \u003Ccode>FUSEKI_URI\u003C\u002Fcode> for in-memory pyoxigraph. Details: \u003Ca href=\"https:\u002F\u002Fgrowgraph.github.io\u002Fontocast\u002Fgetting_started\u002Fquickstart\u002F\" rel=\"nofollow ugc noopener\">Quick Start\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch3>Supplying Your Ontologies\u003C\u002Fh3>\n\u003Cp>OntoCast uses seed ontologies (in Turtle \u003Ccode>.ttl\u003C\u002Fcode> format) to guide extraction. Provide yours in two ways:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Directory Seed:\u003C\u002Fstrong> Set \u003Ccode>ONTOCAST_ONTOLOGY_DIRECTORY=\u002Fpath\u002Fto\u002Fyour\u002Fontologies\u003C\u002Fcode> in your \u003Ccode>.env\u003C\u002Fcode>. All \u003Ccode>.ttl\u003C\u002Fcode> files in that folder sync automatically on startup.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>API Upload:\u003C\u002Fstrong> Register schemas dynamically with the running server:\u003Cpre>\u003Ccode class=\"language-bash\">curl -X POST\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003C\u002Fli>\n\u003C\u002Fol>\n",1787530085307]