StackMap
Subscribe
Explore / topics

RAG & Retrieval

Connect LLMs to your data — ingestion, indexing, retrieval.

arkon logo
3
nduckmink avatarnduckmink 1.5k · 4 months ago
arkon

Self-hosted enterprise knowledge hub + MCP server: an LLM pipeline compiles SOPs and docs into a traceable, human-reviewed wiki, then serves it to AI clients scoped by department and role.

rag
→ alternative to llm_wiki, MaxKB
chroma-core avatar
16
chroma-core avatarchroma-core 29.4k · 5 days ago
Chroma

Open-source embedding database for building AI apps with retrieval.

ragmemory
→ pairs well with xberg, chunkr
chunkr logo
8
lumina-ai-inc avatarlumina-ai-inc 4.2k · 22 days ago
chunkr

Document intelligence API in Rust: layout analysis, OCR with bounding boxes, and semantic chunking that turn PDFs, PPTs and Word docs into RAG-ready structured chunks.

ocrrag
→ pairs well with LlamaIndex, Chroma
cocoindex logo
7
cocoindex-io avatarcocoindex-io 11.6k · 3 days ago
cocoindex

Rust-core incremental indexing engine: declare Target = F(Source) in Python and it keeps vector/graph/relational targets fresh forever, reprocessing only the delta — with per-row lineage.

rag
→ pairs well with latticedb, LlamaIndex
code-graph-rag logo
5
vitali87 avatarvitali87 5.2k · 3 days ago
code-graph-rag

Parses a polyglot monorepo with Tree-sitter into a Memgraph knowledge graph: query it in plain English (NL→Cypher), trace data flow, find dead code, edit via AST-surgical patches.

code-intelrag
→ alternative to codegraph-mcp, gortex
codegraph-mcp logo
12
cognis-digital avatarcognis-digital 7 · 2 months ago
codegraph-mcp

No-train, on-prem code knowledge graph served to AI agents over MCP — symbols, call edges, cross-language links and blast-radius queries, with a hash-chained audit log of every read.

code-intelrag
→ pairs well with omnigraph, tokensave
crawl4ai logo
6
unclecode avatarunclecode 84.4k · 6 days ago
crawl4ai

The 74k-star LLM-native crawler: turns any site into clean, RAG-ready Markdown — adaptive crawling, JS rendering, extraction strategies, Docker deploy. Python, Apache-2.0.

webrag
→ pairs well with browser, Agent-Reach
DSPy logo
3
stanfordnlp avatarstanfordnlp 38.4k · 3 days ago
DSPy

Program — don't prompt — your language models. Compile declarative pipelines into optimized prompts.

orchestrationrag
→ alternative to LangGraph, labs-OO-Agents
Graft logo
6
trailhq avatartrailhq 9.3k · 3 days ago
Graft

Context layer for large codebases: a graph of plain-English markdown nodes — no embeddings, no index — agents read like any repo file. Claude Code hooks + MCP; 42% fewer tokens in its bench.

code-intelrag
→ alternative to mex, ontology-atlas
HelixDB avatar
3
HelixDB avatarHelixDB 6.1k · 3 days ago
helix-db

Rust OLTP graph database on object storage with native vector and BM25 search: property graph, traversal-prefiltered ANN and full-text in one transactional engine; Rust, TS, Go, Python SDKs.

knowledge-graphsrag
→ alternative to hydradb, latticedb
Hyper-Extract logo
10
yifanfeng97 avataryifanfeng97 4k · 3 days ago
Hyper-Extract

Knowledge-extraction CLI: LLMs turn documents into structured graphs, hypergraphs and spatio-temporal knowledge — with an MCP server for agents and Obsidian vault export.

knowledge-graphsrag
→ pairs well with vault-ld, open-ontologies
hyperresearch logo
6
jordan-gibbs avatarjordan-gibbs 3.7k · 6 days ago
hyperresearch

Deep-research harness for Claude Code: a 16-step, tier-adaptive pipeline with adversarial critics, cite-checking and 250+ sources per run, every source kept in a persistent markdown+SQLite vault.

agentsragskills
→ pairs well with open-notebook, SurfSense
knowledge_graph logo
3
rahulnyk avatarrahulnyk 4.1k · 1 months ago
knowledge_graph

Notebook recipe that turns any text corpus into a concept graph with a local Mistral 7B via Ollama — chunk, extract concepts and relations, add proximity edges — for Graph RAG and KG QA.

knowledge-graphsraglocal
→ alternative to Hyper-Extract, docling-graph
ktx logo
8
Kaelio avatarKaelio 1.6k · 20 days ago
ktx

Self-improving context layer for data agents — ingests dbt/Looker/wikis, maps your warehouse, builds a semantic layer with approved metrics, and serves Claude Code/Codex via CLI and MCP.

ragagents
→ pairs well with data-formulator, duckle
jeffhajewski avatar
6
jeffhajewski avatarjeffhajewski 694 · 8 days ago
latticedb

Embedded single-file graph database in Zig: graph traversal, HNSW vector search and BM25 full-text in one query language, plus a durable event log — built for Graph RAG and local agent memory.

knowledge-graphsraglocal
→ pairs well with cocoindex, Chroma
liteparse logo
4
run-llama avatarrun-llama 12.8k · today
liteparse

LlamaIndex's local document parser in Rust: PDFium text with bounding boxes at ~2-5 ms/page, selective Tesseract or HTTP OCR, Markdown/JSON output, screenshots; Python, Node, WASM.

ocrrag
→ pairs well with LlamaIndex, marker
run-llama avatar
14
run-llama avatarrun-llama 52.3k · 4 days ago
LlamaIndex

Data framework for connecting custom data sources to LLMs — ingestion, indexing, retrieval.

ragagents
→ pairs well with Ragas, MemMachine
llm_wiki logo
5
nashsu avatarnashsu 20k · 3 days ago
llm_wiki

Desktop app implementing Karpathy's LLM Wiki pattern: an LLM ingests your documents into a persistent, interlinked wiki with a knowledge graph, hybrid search, deep research and an MCP server.

ragknowledge-graphs
→ pairs well with MinerU, arkon
marker logo
8
datalab-to avatardatalab-to 40k · 18 days ago
marker

Datalab's 39k-star PDF-to-Markdown/JSON converter: a layout pipeline plus an optional LLM pass for tables, forms and equations, with chunk output and form-value extraction built in.

ocrrag
→ pairs well with pdf-inspector, contextgem
1Panel-dev avatar
6
1Panel-dev avatar1Panel-dev 22.9k · 3 days ago
MaxKB

Open-source enterprise agent platform: RAG pipelines (upload or crawl docs), a visual workflow engine with MCP tool-use, and zero-code embedding into existing business systems.

agentsrag
→ pairs well with Ollama, HugAgentOS
metronix-memory logo
8
mtrnix avatarmtrnix 103 · 4 days ago
metronix-memory

Self-hosted agent memory stack in Docker: Postgres + Qdrant + Neo4j hybrid retrieval, a temporal knowledge graph, ontology layer and freshness checks behind one MCP-native API.

memoryragknowledge-graphs
→ pairs well with STALE, MemMachine
MinerU logo
14
opendatalab avataropendatalab 80.8k · 3 days ago
MinerU

Heavyweight document-to-markdown/JSON parser — PDFs plus Office (docx/pptx/xlsx) through layout analysis and OCR into LLM-ready output for RAG and agentic pipelines. 73k stars, self-hostable.

ocrrag
→ pairs well with llm_wiki, marker
morphik-core logo
3
morphik-org avatarmorphik-org 3.7k · 6 days ago
morphik-core

Multimodal retrieval engine for visually rich documents: ingestion, visual-first search over charts, tables and diagrams, knowledge graphs and cache-augmented generation — one engine, not a pipeline.

rag
→ pairs well with WeMM-Embedding, PixelRAG
olmocr logo
12
allenai avatarallenai 19.7k · 6 months ago
olmocr

Open toolkit that linearizes messy PDFs — scans, tables, equations, handwriting — into clean ordered Markdown with a self-hosted vision-language model. Built for LLM training data and RAG ingestion.

ragocr
→ pairs well with LlamaIndex, Chroma
open-notebook logo
4
lfnovo avatarlfnovo 39.6k · 4 days ago
open-notebook

Self-hosted NotebookLM alternative: multi-modal sources, vector plus full-text search, context-aware chat and multi-speaker podcast generation — 18+ model providers incl. Ollama, full REST API.

ragstorage
→ pairs well with Ollama, hyperresearch
opendataloader-pdf logo
5
opendataloader-project avataropendataloader-project 29.4k · 3 days ago
opendataloader-pdf

Deterministic PDF parser for AI pipelines: #1 extraction accuracy (0.907) on its public bench, bounding boxes on every element, 0.015s/page — plus the first open PDF auto-tagging for accessibility.

ocrrag
→ pairs well with PageIndex, MinerU
OpenViking logo
8
volcengine avatarvolcengine 38.9k · 3 days ago
OpenViking

Volcengine's context database: memories, resources and skills as one `viking://` filesystem agents ls, tree and grep — L0/L1/L2 tiers, traceable retrieval, sessions distilled into memory.

memoryragskills
→ pairs well with deepseek-harness, LangGraph
PageIndex logo
7
VectifyAI avatarVectifyAI 35.9k · 7 days ago
PageIndex

Vectorless, reasoning-based RAG — builds a hierarchical tree index from long documents so an LLM retrieves by relevance instead of similarity. No chunking, no embeddings, no vector DB.

rag
→ pairs well with opendataloader-pdf, Ragas
firecrawl avatar
8
firecrawl avatarfirecrawl 19.4k · 3 days ago
pdf-inspector

Firecrawl's Rust PDF triage: classifies text-based vs scanned in ~10-50ms, extracts positioned text and clean Markdown without OCR — routing the ~54% of PDFs that never needed a model.

ocrrag
→ pairs well with marker, anydoc
pgGraph logo
4
Evokoa avatarEvokoa 1.1k · 8 days ago
pgGraph

PostgreSQL extension adding graph search, traversal and shortest-path over your existing tables — a derived graph index queried from plain SQL, no separate graph DB or query language. Rust.

knowledge-graphsrag
→ alternative to ontobricks, omnigraph
PixelRAG logo
4
StarTrail-org avatarStarTrail-org 10.1k · 4 days ago
PixelRAG

Berkeley's visual RAG: render pages and PDFs to screenshot tiles and retrieve with a VLM embedder — tables, charts and layout survive. pixelshot CLI plus a hosted 8.28M-page Wikipedia index.

ragvision
→ pairs well with WeMM-Embedding, MinerU
Ragas logo
8
vibrantlabsai avatarvibrantlabsai 15.9k · 7 months ago
Ragas

Evaluation toolkit for your RAG and agent pipelines — faithfulness, relevance, and more.

evalsrag
→ pairs well with LlamaIndex, AutoGen
searchbox logo
3
hanxiao avatarhanxiao 53 · 3 months ago
searchbox

Airgapped closed-corpus QA testbed: a local Qwen agent in a Pi harness explores a .zip dataroom with grep/embeddings/rerankers under a token budget — a bed to study search as test-time compute.

localragagents
→ pairs well with Ollama, Ragas
sie logo
5
superlinked avatarsuperlinked 3.3k · 3 days ago
sie

Self-hosted inference cluster for everything agents call besides the big LLM: embeddings, rerankers, OCR, NER, guardrails and small LLMs — 100+ models, one OpenAI-compatible API, K8s stack included.

localrag
→ pairs well with FreeToken, kvcached
supavec logo
3
supavec avatarsupavec 1.2k · 9 months ago
supavec

Open-source RAG-as-a-service (the Carbon.ai alternative): upload any data source, get vector search and a chat API in minutes — Supabase-based, multi-tenant with RLS, streaming responses.

rag
→ alternative to PageIndex, MaxKB
supermemory logo
4
supermemoryai avatarsupermemoryai 31k · 3 days ago
supermemory

Memory and context engine for AI: fact extraction, user profiles, contradiction handling and forgetting, hybrid RAG + memory search, connectors, agent plugins and a one-binary local mode.

memoryrag
→ alternative to MemOS, hindsight
SurfSense logo
6
MODSetter avatarMODSetter 16.3k · 5 days ago
SurfSense

Open-source competitive-intelligence platform for agents: live Reddit/YouTube/TikTok/Maps/search connectors; scheduled agents produce briefs and alerts into a cited knowledge base. REST + MCP.

webrag
→ pairs well with flowsint, hyperresearch
swiftide logo
4
bosun-ai avatarbosun-ai 787 · 3 days ago
swiftide

Rust framework for LLM apps: an agent harness, compile-time-typed task graphs, and streaming RAG pipelines — MCP toolboxes, human-in-the-loop approval, tracing with Langfuse support.

ragagents
→ pairs well with turbovec, LangGraph
turbovec logo
5
RyanCodrai avatarRyanCodrai 17.3k · 18 days ago
turbovec

Rust vector index with Python bindings built on Google's TurboQuant: no training step, online ingest, hand-written SIMD kernels — 10M x 1536 vectors in ~4 GB, with allowlist-filtered search.

raglocalstorage
→ pairs well with LlamaIndex, swiftide
txtai logo
2
neuml avatarneuml 13k · 4 days ago
txtai

All-in-one AI framework around an embeddings database — dense, sparse, graph and relational fused — with pipelines, workflows, agents and MCP/web APIs. Python, bindings for JS/Java/Rust/Go.

rag
→ alternative to LlamaIndex, Chroma
Unlimited-OCR logo
5
baidu avatarbaidu 26.5k · 2 months ago
Unlimited-OCR

Baidu's open OCR VLM that parses entire multi-page documents in one shot — 'unlimited' long-horizon parsing pushing DeepSeek-OCR further. MIT weights on HF; serve via transformers, vLLM or SGLang.

ocrrag
→ alternative to olmocr, MinerU
unstract logo
6
Zipstack avatarZipstack 7.3k · 3 days ago
unstract

LLM-driven platform turning unstructured documents into structured data: a no-code Prompt Studio to define extractions, then deploy as APIs or ETL pipelines. Self-hosted, AGPL + enterprise.

ocrrag
→ alternative to chunkr, contextgem
utopia logo
6
deeplethe avatardeeplethe 8k · 3 days ago
utopia

DeepLethe's open 'enterprise world model': one Rust binary plus Postgres running a bitemporal knowledge graph with ontology packs, cited hybrid search, an agent harness and MCP. Air-gap ready.

knowledge-graphsrag
→ pairs well with open-ontologies, orionbelt-ontology-builder
WeMM-Embedding logo
3
Tencent avatarTencent 1.7k · 11 days ago
WeMM-Embedding

Tencent WeChat's universal multimodal embedding family (2B/4B/9B): one vector space for text, images, video, visual documents and interleaved inputs, with Matryoshka dimensions from 64 to 4096.

ragvision
→ pairs well with PixelRAG, morphik-core
xberg logo
7
xberg-io avatarxberg-io 9.3k · 3 days ago
xberg

Rust-core extraction orchestrator with 15 language bindings: 96 formats — PDF, Office, images, audio, code, web — to clean text, tables and RAG-ready chunks. OCR and structured extraction built in.

ocrraglocal
→ pairs well with Chroma, LlamaIndex
zvec-grep logo
5
zvec-ai avatarzvec-ai 3.8k · 3 days ago
zvec-grep

zg: ripgrep, BM25 and vector search behind one local-first CLI for humans and agents — index a workspace once, search code, docs and data by meaning, then verify with exact text or regex.

code-intelraglocal
→ pairs well with serena, cocoindex-code
zvec logo
2
alibaba avataralibaba 16k · 3 days ago
zvec

Alibaba's open-source in-process vector database: billion-scale similarity search embedded in your app, with DiskANN on-disk indexing, native full-text search and hybrid retrieval.

rag
→ alternative to Chroma, zvec-grep