StackMap
Subscribe
Explore / topics

Local / Inference

Run and serve open models on your own hardware.

agentacct logo
8
mikehasa avatarmikehasa 755 · 3 days ago
agentacct

Local-first work receipts for coding agents: reads the session logs Claude Code, Codex, OpenCode and Hermes already write and reports what each task did, what it cost, and whether a check proved it.

codinglocal
→ pairs well with token-optimizer, zeroshot
agentic-api logo
5
vllm-project avatarvllm-project 301 · 3 days ago
agentic-api

Rust gateway that gives vLLM a stateful, OpenAI-compatible Responses API: server-side conversation state, server-side tool loops, SSE and WebSocket streaming, background runs — Codex-ready.

gatewaylocal
→ pairs well with openinterpreter, llm-d
agentsight logo
7
eunomia-bpf avatareunomia-bpf 710 · 18 days ago
agentsight

System-level observability for AI agents via eBPF and TLS tracing: correlates prompts and model calls with the processes, files and network the agent actually touched — no SDK, no proxy.

securitycodinglocal
→ pairs well with geiger, abtop
airllm logo
8
lyogavin avatarlyogavin 35.1k · 3 days ago
airllm

Layer-by-layer inference that runs 70B models on a 4GB GPU — no quantization required; 405B on 8GB, DeepSeek-V3 671B on ~12GB. One AutoModel line for most open model families.

local
→ alternative to colibri, BigMoeOnEdge
AnyJev logo
4
nokia-applied-research avatarnokia-applied-research 998 · 3 days ago
AnyJev

Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM.

decision-modelslocal
→ alternative to kev, laya
arcboxlabs avatar
5
arcboxlabs avatararcboxlabs 4.3k · 3 days ago
arcbox

Open-source container and VM runtime for macOS in pure Rust: drop-in Docker engine, sub-100ms agent sandboxes (`abctl claude`), full Linux VMs and throwaway macOS guests on one daemon.

sandboxeslocalagents
→ alternative to CubeSandbox, forkd
BigMoeOnEdge logo
7
Helldez avatarHelldez 595 · 3 days ago
BigMoeOnEdge

Run MoE models bigger than your RAM: keep the always-needed weights resident and stream each token's experts from flash — a 284B model on a 12 GB phone, CPU only, byte-identical output.

local
→ alternative to colibri, FreeToken
ClaraVerse logo
6
claraverse-space avatarclaraverse-space 3.9k · 1 months ago
ClaraVerse

Self-hosted private AI workspace: chat, multi-agent crews with human review, a visual workflow builder and Telegram integration in one app — local models via Ollama/llama.cpp or your own keys.

localagents
→ pairs well with cccc, lobehub
cocoindex-code logo
8
cocoindex-io avatarcocoindex-io 2.7k · 9 days ago
cocoindex-code

AST-based semantic code search for coding agents: pipx install, zero config, local embeddings out of the box — a CLI/skill/MCP that cuts agent context ~70% vs grepping. Built on CocoIndex.

code-intellocal
→ pairs well with tgrep, tokensave
codebase-memory-mcp logo
10
DeusData avatarDeusData 45.3k · 3 days ago
codebase-memory-mcp

Code intelligence MCP in pure C: tree-sitter knowledge graph over 158 languages, average repo indexed in milliseconds, sub-ms queries, 10x fewer tokens. Single static binary, zero deps.

code-intellocal
→ alternative to tokensave, gortex
codeburn logo
4
getagentseal avatargetagentseal 11.3k · 4 days ago
codeburn

Local-first cost ledger for AI coding: reads the session files 36 tools already write and breaks every token and dollar down by task, model, project. TUI, web, desktop, menubar — no proxy, no keys.

codinglocal
→ alternative to tare, agentacct
colibri logo
6
JustVugg avatarJustVugg 38.1k · 3 days ago
colibri

Pure-C, zero-dep MoE runtime that runs GLM-5.2 (744B) on a 25GB-RAM consumer box by streaming experts from disk — VRAM/RAM/NVMe as one tiered hierarchy, never touching precision.

local
→ alternative to airllm, FreeToken
apple avatar
3
apple avatarapple 2.2k · 5 days ago
coreai-models

Apple's official Core AI toolkit: recipes exporting Hugging Face models to .aimodel, PyTorch primitives for authoring, Swift runtime for macOS/iOS apps — plus skills for coding agents.

local
→ pairs well with MLX-LoRA-Studio, SiliconScope
CubeSandbox logo
17
TencentCloud avatarTencentCloud 12.7k · 3 days ago
CubeSandbox

Hardware-isolated microVM sandboxes for AI agents — sub-60ms boot, <5MB overhead, E2B-compatible API, self-hosted on your own KVM nodes.

sandboxesagentslocal
→ pairs well with LangGraph, AutoGen
exxperts logo
4
EXXETA avatarEXXETA 360 · 3 days ago
exxperts

Local-first agentic runtime with persistent AI rooms and approval-gated memory: every memory write needs your OK; rooms, KB and artifacts are plain files on disk.

memoryagentslocal
→ pairs well with litellm, vLLM
forkd logo
5
deeplethe avatardeeplethe 2.9k · 17 days ago
forkd

fork() for agent microVMs: children fork copy-on-write from a warm Firecracker parent — 100 KVM-isolated VMs in ~100ms, live-VM branching in ~56ms, portable snapshots from a hub.

sandboxesagentslocal
→ alternative to CubeSandbox, superserve
FreeToken logo
8
FlashML-org avatarFlashML-org 13.9k · 5 days ago
FreeToken

Edge-native MoE serving engine: bandwidth-adaptive CPU-GPU co-execution, global LRU expert caching and elastic VRAM run 290B+ frontier MoE models on a gaming PC at interactive speed.

local
→ pairs well with sie, marker
gigatoken logo
2
marcelroed avatarmarcelroed 4.1k · 29 days ago
gigatoken

Tokenization at GB/s: ~1000x faster than HuggingFace tokenizers with drop-in compatibility modes for HF and tiktoken — Rust reading your files directly. pip install gigatoken.

traininglocal
→ pairs well with DataFlow, train-llm-from-scratch
gortex logo
7
zzet avatarzzet 1.8k · 3 days ago
gortex

Code-intelligence engine in one static Go binary: tree-sitter graph over 257 languages, compiler-grade resolution for 17, multi-repo, 175 configurable MCP tools — up to 50x fewer tokens. 100% local.

code-intellocal
→ pairs well with mex, ripwire
gridex logo
4
gridex avatargridex 1.5k · 24 days ago
gridex

Native database IDE for Postgres, MySQL, SQLite, Redis, MongoDB, SQL Server and ClickHouse, with a built-in MCP server — 13 tools, 3-tier permissions, audit trail — and schema-aware AI chat.

datalocalcoding
→ pairs well with WrenAI, duckle
antirez avatar
3
antirez avatarantirez 2.8k · 1 months ago
h3.c

antirez's C/Metal engine running MiniMax-H3 video and audio generation natively on Apple Silicon — interactive prompting, 4-step schedules, SSD block streaming down to ~2 GB DiT residency.

videolocal
→ pairs well with SiliconScope, sdnext
Handy logo
2
cjpais avatarcjpais 32.4k · 3 days ago
Handy

Push-to-talk offline dictation: hotkey, speak, text lands in whatever field has focus. Whisper or Parakeet fully on-device; cross-platform Rust/Tauri, built to be forked.

voicelocal
→ alternative to voicebox, meetily
kev logo
10
jaredpalmer avatarjaredpalmer 7.6k · 3 days ago
kev

Open reproduction of TypeSafe's Jev: a LoRA + readout head on Qwen that answers many typed questions about one document in a single prefill pass, returning calibrated probabilities.

decision-modelstraininglocal
→ pairs well with Adala, jev-ultrafast
kimi-k3-in-c logo
3
FareedKhan-dev avatarFareedKhan-dev 8.8k · 9 days ago
kimi-k3-in-c

Kimi K3 (2.78T-parameter MoE) inference in portable C99 on one CPU: the dense trunk stays in RAM, 4-bit experts stream from disk, and output is byte-identical from 8 GB to 224 GB.

local
→ alternative to colibri, BigMoeOnEdge
knowledge_graph logo
3
rahulnyk avatarrahulnyk 4.1k · 1 months ago
knowledge_graph

Notebook recipe that turns any text corpus into a concept graph with a local Mistral 7B via Ollama — chunk, extract concepts and relations, add proximity edges — for Graph RAG and KG QA.

knowledge-graphsraglocal
→ alternative to Hyper-Extract, docling-graph
kvcached logo
5
ovg-project avatarovg-project 1.5k · 6 days ago
kvcached

Virtualized, elastic KV cache for LLM serving on shared GPUs: reserve virtual memory, back it with physical GPU memory only when used — vLLM and SGLang, with a memory-limit CLI, router and sleep mode.

local
→ pairs well with LMCache, llm-d
jeffhajewski avatar
6
jeffhajewski avatarjeffhajewski 694 · 8 days ago
latticedb

Embedded single-file graph database in Zig: graph traversal, HNSW vector search and BM25 full-text in one query language, plus a durable event log — built for Graph RAG and local agent memory.

knowledge-graphsraglocal
→ pairs well with cocoindex, Chroma
laya logo
5
NandhaKishorM avatarNandhaKishorM 29.8k · today
laya

Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint.

decision-modelslocal
→ pairs well with LangGraph, LlamaIndex
linecast logo
2
ashuttl avatarashuttl 563 · 3 days ago
linecast

Six dependency-free terminal apps for weather, sun, moon, tides, radar and maps, drawn from free public data with no accounts or API keys — mouse-friendly TUIs that try to match your terminal theme.

weblocal
→ alternative to Crucix, worldmonitor
llm-d logo
9
llm-d avatarllm-d 4.7k · 3 days ago
llm-d

Distributed inference stack for Kubernetes from Red Hat, Google and IBM (CNCF) — prefix-cache-aware routing, tiered KV-cache, prefill/decode disaggregation and SLO autoscaling above vLLM/SGLang.

local
→ pairs well with LMCache, Model-Optimizer
LMCache logo
4
LMCache avatarLMCache 11.9k · 3 days ago
LMCache

KV-cache layer for scalable LLM serving: offload and reuse KV across GPU/CPU/disk/remote tiers to cut TTFT and prefill cost. vLLM-first; used by NVIDIA Dynamo and llm-d.

localstorage
→ pairs well with llm-d, Model-Optimizer
githubnext avatar
7
githubnext avatargithubnext 792 · 13 days ago
localjev

GitHub Next's local Jev bridge: a Bun/TypeScript POST /v1/systemone that translates typed decision questions into prompts for DiffusionGemma behind any OpenAI-compatible endpoint.

decision-modelsgatewaylocal
→ pairs well with Ollama, vLLM
ysharma3501 avatar
3
ysharma3501 avatarysharma3501 5.4k · 3 months ago
LuxTTS

Lightweight voice-cloning TTS — 48kHz speech at 150x realtime, fits in 1GB VRAM and runs on CPU or MPS. SOTA cloning from a ~3s reference sample, rivaling models 10x larger.

voicelocal
→ pairs well with ai-avatar-system, pocket-tts
Lynkr logo
9
Fast-Editor avatarFast-Editor 598 · 8 days ago
Lynkr

Self-hosted LLM gateway wrapping Claude Code, Cursor or Codex with zero code changes — strips unused tools, compresses JSON tool results ~88%, semantic-caches, tier-routes easy work to local models.

codinglocalgateway
→ pairs well with Ollama, prompt-cache-skills
magnitude logo
2
magnitudedev avatarmagnitudedev 5.4k · 3 days ago
magnitude

Local inference desktop app + CLI: profiles your hardware, estimates tok/s per model before download, tunes the one you pick, and connects Pi, OpenCode, Hermes, Codex or Claude Code in a click.

local
→ alternative to Ollama, ODS
meetily logo
4
Zackriya-Solutions avatarZackriya-Solutions 31.2k · 16 days ago
meetily

Privacy-first meeting note-taker that runs 100% on-device: live Whisper/Parakeet transcription, speaker diarization, local Ollama summaries. Desktop app for macOS & Windows — no cloud, no call bots.

localvoice
→ alternative to call.md, Handy
memanto logo
5
moorcheh-ai avatarmoorcheh-ai 2.3k · 3 days ago
memanto

Companion memory agent for 20+ coding agents, built on Moorcheh — its own information-theoretic engine, no third-party vector DB to manage. Runs local (Docker + Ollama, keyless) or on their cloud.

memorylocal
→ alternative to ai-memory, memsearch
memvid logo
3
memvid avatarmemvid 16.6k · 2 months ago
memvid

Single-file memory layer for agents: data, embeddings, index and metadata in one portable .mv2 — append-only Smart Frames, time-travel queries, sub-5ms recall, no server. Rust core, Node/Python SDKs.

memorylocal
→ alternative to Memoria, Chroma
mesh-llm logo
7
Mesh-LLM avatarMesh-LLM 3.5k · 3 days ago
mesh-llm

Distributed LLM inference in Rust: pool GPUs across machines into one OpenAI-compatible endpoint — local fit first, mesh routing, and stage splits for models too large for any single box.

local
→ alternative to Personal-AI-Router, colibri
MLX-LoRA-Studio logo
5
Goekdeniz-Guelmez avatarGoekdeniz-Guelmez 271 · 1 months ago
MLX-LoRA-Studio

Native Mac app for on-device LLM fine-tuning via mlx-lm-lora: pick a model, choose SFT/LoRA/DPO-family algorithms, watch loss fall live, push to Hugging Face. No cloud, no code.

traininglocal
→ pairs well with SiliconScope, coreai-models
Model-Optimizer logo
7
NVIDIA avatarNVIDIA 5k · 3 days ago
Model-Optimizer

NVIDIA's model-compression library: quantization (PTQ/QAT, FP8/NVFP4), pruning, distillation, NAS and speculative decoding over HF/PyTorch/ONNX, exported to TensorRT-LLM, vLLM and SGLang.

traininglocal
→ pairs well with vLLM, trl
needle logo
4
cactus-compute avatarcactus-compute 12.8k · 3 days ago
needle

A 45M-parameter tool-calling model shipped as one 14MB binary that runs a full session in ~28MB RAM — grammar-constrained JSON, calibrated confidence, tool retrieval, LoRA fine-tuning.

localagentstraining
→ pairs well with pocket-tts, xybrid
ODS logo
3
Osmantic avatarOsmantic 6.9k · 3 days ago
ODS

One installer that turns a PC, Mac or Linux box into a private AI server: Ollama, Open WebUI, n8n, ComfyUI wired together — inference, chat, voice, agents, RAG and image gen, no cloud.

local
→ alternative to ClaraVerse, magnitude
ollama avatar
29
ollama avatarollama 181.8k · 4 days ago
Ollama

Run Llama, Mistral and other open models locally with a single command and a clean API.

local
→ pairs well with Lynkr, pocket-tts
OpenSandbox logo
8
opensandbox-group avataropensandbox-group 15.6k · 3 days ago
OpenSandbox

CNCF-landscape sandbox platform for AI agents: multi-language SDKs, unified API, CLI and MCP over Docker/Kubernetes runtimes — coding agents, GUI agents, evals and RL training.

sandboxesagentslocal
→ pairs well with gym-anything, agent-sandbox
Personal-AI-Router logo
4
NVIDIA avatarNVIDIA 1.5k · 8 days ago
Personal-AI-Router

NVIDIA's local inference router: pairs home machines running Ollama or LM Studio behind OpenAI-, Anthropic- and Ollama-compatible endpoints, sending each request to the best available node.

localgateway
→ pairs well with litellm, Lynkr
pocket-tts logo
7
kyutai-labs avatarkyutai-labs 9.7k · 3 days ago
pocket-tts

Kyutai's 100M-parameter CPU-only TTS — streaming audio in ~200ms, ~6× real-time on two laptop cores, voice cloning, six languages. pip install and it talks.

localvoice
→ pairs well with Ollama, ai-avatar-system
rtk-ai avatar
8
rtk-ai avatarrtk-ai 81.9k · 3 days ago
rtk

Rust CLI proxy compressing dev-command output 60-90% before your agent reads it — git, tests, linters, docker, 100+ commands; hooks auto-rewrite bash calls. Single binary, <10ms overhead.

codinglocal
→ pairs well with headlong, tokensave
searchbox logo
3
hanxiao avatarhanxiao 53 · 3 months ago
searchbox

Airgapped closed-corpus QA testbed: a local Qwen agent in a Pi harness explores a .zip dataroom with grep/embeddings/rerankers under a token budget — a bed to study search as test-time compute.

localragagents
→ pairs well with Ollama, Ragas
sie logo
5
superlinked avatarsuperlinked 3.3k · 3 days ago
sie

Self-hosted inference cluster for everything agents call besides the big LLM: embeddings, rerankers, OCR, NER, guardrails and small LLMs — 100+ models, one OpenAI-compatible API, K8s stack included.

localrag
→ pairs well with FreeToken, kvcached
SiliconScope logo
5
kennss avatarkennss 966 · 5 days ago
SiliconScope

Sudoless Apple Silicon monitor: SwiftUI dashboard plus menu-bar suite tracking ANE, Media Engine and memory bandwidth Activity Monitor won't show — with DVR-style record & replay.

local
→ pairs well with MLX-LoRA-Studio, Ollama
tare logo
7
kelviq avatarkelviq 292 · 1 months ago
tare

Claude Code skill that reads your local session logs and answers 'where did my tokens go' in plain English: deduplicated totals, cost charged to the tool that caused it, 5-hour window state.

codingskillslocal
→ pairs well with token-optimizer, agent-orchestrator
tokensave logo
16
aovestdipaperino avataraovestdipaperino 650 · 4 days ago
tokensave

Code-intelligence MCP server for coding agents — a pre-indexed semantic graph (libSQL + FTS5) they query instead of grepping: symbols, callers, impact radius in one call. 100% local, 50+ languages.

code-intellocal
→ pairs well with Lynkr, headroom
turbovec logo
5
RyanCodrai avatarRyanCodrai 17.3k · 18 days ago
turbovec

Rust vector index with Python bindings built on Google's TurboQuant: no training step, online ingest, hand-written SIMD kernels — 10M x 1536 vectors in ~4 GB, with allowlist-filtered search.

raglocalstorage
→ pairs well with LlamaIndex, swiftide
vLLM logo
25
vllm-project avatarvllm-project 92.9k · 3 days ago
vLLM

High-throughput, memory-efficient inference and serving engine for LLMs.

local
→ pairs well with Model-Optimizer, litellm
xberg logo
7
xberg-io avatarxberg-io 9.3k · 3 days ago
xberg

Rust-core extraction orchestrator with 15 language bindings: 96 formats — PDF, Office, images, audio, code, web — to clean text, tables and RAG-ready chunks. OCR and structured extraction built in.

ocrraglocal
→ pairs well with Chroma, LlamaIndex
xybrid logo
4
xybrid-ai avatarxybrid-ai 460 · 3 days ago
xybrid

Cross-platform on-device AI toolkit: run LLMs, ASR and TTS natively from Flutter, Unity, Kotlin, Swift or Rust on a llama.cpp and ONNX Runtime core. Private, offline, no cloud.

localvoice
→ pairs well with pocket-tts, needle
zvec-grep logo
5
zvec-ai avatarzvec-ai 3.8k · 3 days ago
zvec-grep

zg: ripgrep, BM25 and vector search behind one local-first CLI for humans and agents — index a workspace once, search code, docs and data by meaning, then verify with exact text or regex.

code-intelraglocal
→ pairs well with serena, cocoindex-code