{"site":"https://stackmap.shipwithai.xyz","version":"2026.07.24","build":"698e6e6","schema_version":1,"repos":{"9router":{"name":"9router","owner":"decolua","slug":"9router","stars":23183,"image":"https://raw.githubusercontent.com/decolua/9router/master/images/9router.png?1","avatar":"https://avatars.githubusercontent.com/u/8282593?v=4&s=96","forks":3879,"language":"JavaScript","license":"MIT","updated":"4 days ago","topics":["gateway","coding"],"summary":"Self-hosted gateway pointing any AI coding CLI (Claude Code, Codex, Cursor, Cline) at 40+ providers, with subscription→cheap→free auto-fallback and tool_result compression for 20-40% token savings.","curator_note":"The 'never hit a rate limit again' gateway: run it on localhost, point Claude Code / Codex / Cursor / Cline / Copilot at its OpenAI-compatible endpoint, and it round-robins your accounts and cascades Subscription → cheap ($0.2-0.6/1M) → free (Kiro, OpenCode) providers so a job never dies mid-flight, while RTK compresses tool_result payloads (git diff, grep, ls) to shave 20-40% of tokens. Best when you juggle several provider subscriptions/keys and keep exhausting quota. NOT for you if you want one stable premium model (free-tier fallbacks vary in quality/availability), if routing traffic through third-party free providers raises data-privacy concerns, or if you want savings from smarter code retrieval rather than payload compression. Overlaps heavily with lynkr — pick 9router for free/multi-account fallback breadth, lynkr for tool-stripping, semantic caching and local-model routing.","edges":[{"to":"lynkr","type":"alternative","why":"Both are self-hosted LLM gateways that wrap coding agents (Claude Code/Codex/Cursor/Cline) with zero code changes and cut tokens by compressing tool_result payloads, then route across backends. 9router optimizes for cost/uptime — 40+ providers, multi-account round-robin, subscription→cheap→free auto-fallback. lynkr optimizes for efficiency — strips unused tools, semantic-caches, and tier-routes easy work to local models. Pick by whether you need free-provider breadth or local-model routing.","confidence":0.85,"status":"approved"},{"to":"tokensave","type":"complements","why":"Both cut a coding agent's token bill but at different layers, so they stack: 9router compresses tool_result payloads at the gateway and routes to cheap/free models, while tokensave gives the agent a pre-indexed code graph to query instead of burning tokens on grep/read. Front your agent with 9router and hand it tokensave as an MCP tool.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-08T11:58:30.312Z","linkCount":8,"related":{"complements":[{"slug":"tokensave","why":"Both cut a coding agent's token bill but at different layers, so they stack: 9router compresses tool_result payloads at the gateway and routes to cheap/free models, while tokensave gives the agent a pre-indexed code graph to query instead of burning tokens on grep/read. Front your agent with 9router and hand it tokensave as an MCP tool.","dir":"out","confidence":0.5,"name":"tokensave"}],"alternative":[{"slug":"lynkr","why":"Both are self-hosted LLM gateways that wrap coding agents (Claude Code/Codex/Cursor/Cline) with zero code changes and cut tokens by compressing tool_result payloads, then route across backends. 9router optimizes for cost/uptime — 40+ providers, multi-account round-robin, subscription→cheap→free auto-fallback. lynkr optimizes for efficiency — strips unused tools, semantic-caches, and tier-routes easy work to local models. Pick by whether you need free-provider breadth or local-model routing.","dir":"out","confidence":0.85,"name":"Lynkr"},{"slug":"freellmapi","why":"Both are self-hosted OpenAI-compatible gateways that multiplex providers with automatic fallback for coding CLIs. 9router optimizes paid usage (subscription→cheap→free routing, token compression); FreeLLMAPI exists purely to stack and stay under 18 providers' free-tier caps.","dir":"in","confidence":0.75,"name":"freellmapi"},{"slug":"free-claude-code","why":"Both are self-hosted gateways pointing Claude Code/Codex-class CLIs at many providers. 9router leans on routing policy (subscription→cheap→free fallback, token compression); FCC leans on UX — launchers, validation UI, native model-picker integration, IDE configs.","dir":"in","confidence":0.65,"name":"free-claude-code"},{"slug":"cliproxyapi","why":"Two directions of the same arbitrage: 9router points coding CLIs at 40+ providers; CLIProxyAPI exposes your CLI subscriptions as an OpenAI/Claude/Gemini-compatible endpoint for any client.","dir":"in","confidence":0.6,"name":"CLIProxyAPI"},{"slug":"headroom","why":"Same wire, same goal — cut coding-agent token spend at a local proxy. 9router is a multi-provider routing gateway with tool_result compression as a feature; Headroom is compression-first with content-aware routers, reversibility and no routing opinions.","dir":"in","confidence":0.6,"name":"headroom"},{"slug":"codexmate","why":"Overlapping job of pointing coding CLIs at arbitrary providers: 9router is a dedicated self-hosted gateway with auto-fallback across 40+ providers; Codex Mate does it via built-in Codex/Claude protocol bridges as one feature of its control plane.","dir":"in","confidence":0.55,"name":"codexmate"},{"slug":"litellm","why":"Both are self-hosted gateways routing to many providers with fallback; 9router aims at coding CLIs, LiteLLM at app/SDK developers.","dir":"in","confidence":0.55,"name":"litellm"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/decolua/9router"},"aaron-marketing-skills":{"name":"aaron-marketing-skills","owner":"aaron-he-zhu","slug":"aaron-marketing-skills","stars":2447,"avatar":"https://avatars.githubusercontent.com/u/139607425?v=4&s=96","forks":333,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["skills"],"summary":"120 marketing skills + 8 commands for Claude Code across 7 disciplines — SEO/GEO, influencer, paid ads, email, launch, social, brand — on one contract with 8 benchmark-driven auditor gates.","curator_note":"The industrialized option among marketing skill packs: every skill shares one contract and passes discipline-specific auditor gates (CORE-EEAT for content, CITE for GEO citations, ROAS for ads, SEND for email…), with keyless data connectors so nothing needs API keys to start. That structure makes it closer to a framework than a prompt dump — and it's why we ingested this repo instead of seo-geo-claude-skills, which is now just a signpost pointing here (its standalone line is frozen at v9.9.12). NOT for a one-off copy tweak — the machinery earns its weight when marketing IS your recurring workflow. Actively developed, Apache-2.0.","edges":[{"to":"ai-marketing-claude","type":"alternative","why":"Both turn Claude Code into a marketing practitioner: ai-marketing-claude is 15 skills with parallel audit agents and agency-shaped PDF deliverables; aaron-marketing-skills is 120 skills across 7 disciplines with benchmark-driven auditor gates and one shared contract.","confidence":0.65,"status":"approved"},{"to":"claude-ads","type":"alternative","why":"Overlapping on the paid-ads slice: claude-ads goes deep — 12 platforms, account reads and capability-gated writes; aaron's ads discipline stays at the strategy/creative level with a ROAS auditor gate and no account plumbing.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-16T16:07:34.505Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"ai-marketing-claude","why":"Both turn Claude Code into a marketing practitioner: ai-marketing-claude is 15 skills with parallel audit agents and agency-shaped PDF deliverables; aaron-marketing-skills is 120 skills across 7 disciplines with benchmark-driven auditor gates and one shared contract.","dir":"out","confidence":0.65,"name":"ai-marketing-claude"},{"slug":"pm-skills","why":"Same shape, adjacent discipline: large curated skill marketplaces for Claude Code that encode a profession's frameworks — 120 marketing skills there, 68 PM skills with chained workflows here. Run both if you own both functions.","dir":"in","confidence":0.65,"name":"pm-skills"},{"slug":"claude-ads","why":"Overlapping on the paid-ads slice: claude-ads goes deep — 12 platforms, account reads and capability-gated writes; aaron's ads discipline stays at the strategy/creative level with a ROAS auditor gate and no account plumbing.","dir":"out","confidence":0.5,"name":"claude-ads"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/aaron-he-zhu/aaron-marketing-skills"},"abtop":{"name":"abtop","owner":"graykode","slug":"abtop","stars":3381,"image":"https://raw.githubusercontent.com/graykode/abtop/main/assets/demo.gif","avatar":"https://avatars.githubusercontent.com/u/10525011?v=4&s=96","forks":302,"language":"Rust","license":"MIT","updated":"17 days ago","topics":["coding"],"summary":"htop for AI coding agents: every Claude Code, Codex and OpenCode session in one TUI — tokens, context-window %, rate limits, child processes, orphan ports. Read-only, no API keys. Rust.","curator_note":"The moment you run 3+ agents in parallel this earns its place: live per-session context-window bars, real-time rate-limit quota, and the sleeper feature — orphan port detection for the dev server an agent forgot to kill. Read-only from local process/file state, cross-platform, Enter jumps to the agent's terminal (tmux/cmux aware). NOT an analytics tool: it's a live dashboard, not history — for cost breakdowns, replays and budgets over time you want a session-log analyzer, not a top. OpenCode monitoring needs the sqlite3 CLI installed.","edges":[{"to":"ai-token-monitor","type":"alternative","why":"Same watch-your-agents job, different surface: ai-token-monitor is a menu-bar app focused on token spend and plan limits; abtop is a terminal TUI adding context-window bars, child processes and port detection per live session.","confidence":0.75,"status":"approved"},{"to":"cc-lens","type":"alternative","why":"Live vs retrospective: abtop shows what your agents are doing right now; cc-lens analyzes what they did — replays, cost/cache breakdowns, budgets. Many desks end up running both.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.869Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"ai-token-monitor","why":"Same watch-your-agents job, different surface: ai-token-monitor is a menu-bar app focused on token spend and plan limits; abtop is a terminal TUI adding context-window bars, child processes and port detection per live session.","dir":"out","confidence":0.75,"name":"ai-token-monitor"},{"slug":"cc-lens","why":"Live vs retrospective: abtop shows what your agents are doing right now; cc-lens analyzes what they did — replays, cost/cache breakdowns, budgets. Many desks end up running both.","dir":"out","confidence":0.6,"name":"cc-lens"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/graykode/abtop"},"adala":{"name":"Adala","owner":"HumanSignal","slug":"adala","stars":1615,"image":"https://raw.githubusercontent.com/HumanSignal/adala/master/docs/src/img/logo.png","avatar":"https://avatars.githubusercontent.com/u/48309720?v=4&s=96","forks":155,"language":"Python","license":"Apache-2.0","updated":"3 days ago","topics":["agents","training"],"summary":"HumanSignal's autonomous data-labeling agent framework: define a skill, give it ground truth, and the agent iterates — learn, apply, reflect — until it hits your accuracy threshold.","curator_note":"From the Label Studio company, and the ground-truth-first design is the differentiator: instead of prompt-tuning a classifier by hand, you hand the agent labeled examples and `agent.learn()` iterates against them (student/teacher runtimes) until accuracy clears your bar — then you run it on the unlabeled pile. Skills cover classification, summarization, QA, translation, and compose into sequences; any OpenAI-compatible endpoint works (OpenRouter for Claude/Gemini). Use it for scaled labeling and dataset bootstrapping where you already have some ground truth. NOT a general agent framework despite the name — it's specialized for data processing, and note the trailing Python 3.8–3.11 support window: check activity before adopting for something new.","edges":[{"to":"dataflow","type":"alternative","why":"Both build LLM-powered training data at scale: DataFlow is operator pipelines you compose for generation/cleaning/filtering; Adala is agents that LEARN the labeling skill from ground truth and self-improve to a target accuracy.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.577Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"dataflow","why":"Both build LLM-powered training data at scale: DataFlow is operator pipelines you compose for generation/cleaning/filtering; Adala is agents that LEARN the labeling skill from ground truth and self-improve to a target accuracy.","dir":"out","confidence":0.55,"name":"DataFlow"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/HumanSignal/adala"},"agent-flow":{"name":"agent-flow","owner":"patoles","slug":"agent-flow","stars":1341,"image":"https://res.cloudinary.com/dxlvclh9c/image/upload/v1773924941/screenshot_e7yox3.png","avatar":"https://avatars.githubusercontent.com/u/4022964?v=4&s=96","forks":150,"language":"TypeScript","license":"Apache-2.0","updated":"13 days ago","topics":["coding"],"summary":"Live node-graph visualization of Claude Code and Codex sessions — watch agents think, branch into subagents and call tools in real time. VS Code extension or npx web app. Apache-2.0.","curator_note":"The black-box opener: hooks stream Claude Code events (Codex via rollout-file tailing) onto an interactive canvas where you watch a session branch into subagents, spot the slow tool call, trace the decision chain that went wrong, and replay any JSONL log. Zero-friction start — `npx agent-flow-app` and go — and it earns its keep two ways: debugging orchestration, and building prompt intuition by literally watching how the agent interprets you. Born inside a real product (CraftMyGame), not a demo. NOT analytics — no cost aggregation or history (that's claude-code-karma's job). Note the npx binary ships opt-out telemetry: aggregate-only, documented field-by-field, honors DO_NOT_TRACK; dev and extension builds send nothing.","edges":[{"to":"claude-code-karma","type":"complements","why":"Two halves of Claude Code observability over the same local data: Agent Flow is the live execution graph you watch while a session runs; Karma is the historical dashboard you consult afterwards — costs, timelines, tool and skill analytics.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.453Z","linkCount":1,"related":{"complements":[{"slug":"claude-code-karma","why":"Two halves of Claude Code observability over the same local data: Agent Flow is the live execution graph you watch while a session runs; Karma is the historical dashboard you consult afterwards — costs, timelines, tool and skill analytics.","dir":"out","confidence":0.6,"name":"claude-code-karma"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/patoles/agent-flow"},"agentfield":{"name":"agentfield","owner":"Agent-Field","slug":"agentfield","stars":2424,"image":"https://raw.githubusercontent.com/Agent-Field/agentfield/main/assets/harness-banner.png","avatar":"https://avatars.githubusercontent.com/u/204899035?v=4&s=96","forks":384,"language":"Go","license":"Apache-2.0","updated":"2 days ago","topics":["agents","orchestration"],"summary":"Open-source control plane that runs AI agents as microservices: write plain Python/Go/TS functions, get REST endpoints with routing, queues, retries, memory and tracing — one laptop to 10k agents.","curator_note":"The 'agents as a backend' play: write plain functions (no DSL, no graph wiring), and the Go control plane turns each into a REST endpoint any service can call — with fan-out to thousands of parallel agents, queues, retries, versioned deploys, observability and identity/audit built in. Reach for it when agents must be production infrastructure callable by frontends, cron jobs and other services — not a chat window. NOT for notebook experiments or a single local agent (a control plane + SDK is real operational commitment), and it won't give you reasoning-pattern libraries — you still design the agent logic it hosts. Its prompt-to-backend flow (/agentfield in Claude Code/Cursor) is a nice on-ramp, but evaluate the runtime, not the demo.","edges":[{"to":"paperclip","type":"alternative","why":"Both are open control planes for fleets of agents, attacking opposite ends: paperclip is the governance layer over agents you already have (BYO agent, goals, org charts, budgets, audited tickets); agentfield is the build-and-run backend where agents are SDK-written microservices with REST endpoints, queues and retries. Govern existing agents → paperclip; build the agent backend itself → agentfield.","confidence":0.6,"status":"approved"},{"to":"langgraph","type":"alternative","why":"Same job — production multi-agent systems — opposite shape. LangGraph is an in-process library: model your workflow as a stateful graph with durable execution. agentfield is an out-of-process control plane: plain functions become REST microservices and the platform handles fan-out, queues and retries, explicitly rejecting graph wiring. Library and embedded → LangGraph; platform and service-oriented → agentfield.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-08T16:13:44.080Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"scale-agentex","why":"Both are open control planes for building and deploying agents as long-running services with routing, queues and durable state. AgentField wraps plain Python/Go/TS functions as microservice agents; Agentex standardizes on the ACP protocol with Temporal underneath and a scaffolding CLI + dev UI.","dir":"in","confidence":0.7,"name":"scale-agentex"},{"slug":"paperclip","why":"Both are open control planes for fleets of agents, attacking opposite ends: paperclip is the governance layer over agents you already have (BYO agent, goals, org charts, budgets, audited tickets); agentfield is the build-and-run backend where agents are SDK-written microservices with REST endpoints, queues and retries. Govern existing agents → paperclip; build the agent backend itself → agentfield.","dir":"out","confidence":0.6,"name":"paperclip"},{"slug":"plano","why":"Same 'production infrastructure layer for agents' job from opposite ends: agentfield is a control plane that wraps your functions in routing/queues/retries, Plano is a data-plane proxy your unmodified HTTP agents sit behind.","dir":"in","confidence":0.6,"name":"plano"},{"slug":"langgraph","why":"Same job — production multi-agent systems — opposite shape. LangGraph is an in-process library: model your workflow as a stateful graph with durable execution. agentfield is an out-of-process control plane: plain functions become REST microservices and the platform handles fan-out, queues and retries, explicitly rejecting graph wiring. Library and embedded → LangGraph; platform and service-oriented → agentfield.","dir":"out","confidence":0.55,"name":"LangGraph"},{"slug":"sim","why":"Same 'run your AI workforce' ambition, opposite audiences: AgentField turns plain Python/Go/TS functions into agent microservices for engineers; Sim gives teams a visual workspace where agents are built and operated without touching the runtime.","dir":"in","confidence":0.55,"name":"sim"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Agent-Field/agentfield"},"agentic-context-engine":{"name":"agentic-context-engine","owner":"kayba-ai","slug":"agentic-context-engine","stars":2536,"image":"https://raw.githubusercontent.com/kayba-ai/agentic-context-engine/main/assets/kayba-banner.png","avatar":"https://avatars.githubusercontent.com/u/238284021?v=4&s=96","forks":301,"language":"Python","license":"Apache-2.0","updated":"16 days ago","topics":["memory","agents"],"summary":"Learning loop for any agent: reflect on failures, distill strategies into a Skillbook, inject them next run — 2x consistency on Tau2, 49% token cuts. LiteLLM-based, 100+ providers.","curator_note":"The in-process answer to 'my agent repeats the same mistakes': wrap your agent, feed it corrections, and ACE extracts reusable strategies it injects on later runs — no fine-tuning, no reward signals, and the numbers are concrete (2x pass^4 on Tau2, ~$1.50 to learn its way through a 14k-line translation). Pick it over a memory *service* when you want the learning inside your Python process rather than behind an HTTP API. NOT magic memory: strategies come from explicit feedback loops you wire up, quality follows the judge model, and the open-source engine is the on-ramp to the hosted Kayba product — check where the managed line lands before betting infra on it.","edges":[{"to":"litellm","type":"built_with","why":"ACELiteLLM is the primary entry point — the '100+ supported providers' run through LiteLLM's unified API.","confidence":0.75,"status":"approved"},{"to":"hindsight","type":"alternative","why":"Both sell 'agents that learn, not just remember' — opposite deployments: Hindsight is a self-hosted memory service (Docker + Postgres, retain/recall/reflect over HTTP); ACE is an embedded Python loop distilling strategies in-process. Service vs library.","confidence":0.65,"status":"approved"},{"to":"claude-reflect","type":"alternative","why":"Same learn-from-corrections loop at different scopes: claude-reflect is a Claude Code plugin syncing learnings into CLAUDE.md; ACE is a framework-level engine for any agent you build, with strategies as first-class objects.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-22T18:15:22.056Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"hindsight","why":"Both sell 'agents that learn, not just remember' — opposite deployments: Hindsight is a self-hosted memory service (Docker + Postgres, retain/recall/reflect over HTTP); ACE is an embedded Python loop distilling strategies in-process. Service vs library.","dir":"out","confidence":0.65,"name":"hindsight"},{"slug":"claude-reflect","why":"Same learn-from-corrections loop at different scopes: claude-reflect is a Claude Code plugin syncing learnings into CLAUDE.md; ACE is a framework-level engine for any agent you build, with strategies as first-class objects.","dir":"out","confidence":0.6,"name":"claude-reflect"},{"slug":"hivemind","why":"Same learn-from-experience loop, different scope: ACE is an embedded Python engine distilling strategies for the agent you're building; Hivemind is a service layer distilling skills across every off-the-shelf agent your team runs.","dir":"in","confidence":0.6,"name":"hivemind"}],"built_with":[{"slug":"litellm","why":"ACELiteLLM is the primary entry point — the '100+ supported providers' run through LiteLLM's unified API.","dir":"out","confidence":0.75,"name":"litellm"}]},"url":"https://stackmap.shipwithai.xyz/repos/kayba-ai/agentic-context-engine"},"agents-cli":{"name":"agents-cli","owner":"google","slug":"agents-cli","stars":5326,"image":"https://raw.githubusercontent.com/google/agents-cli/main/docs/src/assets/logo_sm.png","avatar":"https://avatars.githubusercontent.com/u/1342004?v=4&s=96","forks":563,"language":"Python","license":"Apache-2.0","updated":"14 days ago","topics":["coding","evals","skills"],"summary":"CLI + agent-skills layer that turns your coding assistant into a Google Cloud agent-lifecycle expert: scaffold ADK projects, run and evaluate them, then deploy and publish to Gemini Enterprise.","curator_note":"Reach for agents-cli when you're building and shipping agents ON Google Cloud / Gemini Enterprise and want one CLI for scaffold → eval → deploy, driven through a coding agent you already use. NOT the pick if you're not on Google Cloud — deploy targets and publish are GCP/Gemini-bound — or if you want a portable, vendor-neutral framework to embed in your own app; a framework like LangGraph or CrewAI fits that better.","edges":[{"to":"langgraph","type":"alternative","why":"Both are ways to build production agentic apps, but from opposite ends: LangGraph is an open, embeddable graph runtime that runs anywhere, while agents-cli is Google's CLI+skills around ADK that scaffolds, evaluates, and deploys agents specifically on Google Cloud / Gemini Enterprise. Choose by ecosystem and portability needs.","confidence":0.5,"status":"approved"},{"to":"ragas","type":"alternative","why":"agents-cli ships a full agent-eval suite — trace generation, metric grading, LLM-as-judge, failure-mode clustering, prompt optimization — which overlaps Ragas's job. Difference: Ragas is eval-only and framework-agnostic, while agents-cli's eval is one stage of a GCP-bound scaffold/deploy pipeline.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-05T14:49:26.000Z","linkCount":5,"related":{"complements":[{"slug":"claude-reflect","why":"Both extend a coding assistant via skills; reflect's correction-routing writes learnings back into whatever skill files you run — install it alongside any skill pack and the pack improves with use.","dir":"in","confidence":0.55,"name":"claude-reflect"},{"slug":"skillkit","why":"agents-cli ships as an agent-skills pack; skillkit is a skill package manager that installs packs like it and ports them across agents. Plausible pairing, unverified compatibility — agents-cli's skill format may not round-trip cleanly.","dir":"in","confidence":0.5,"name":"skillkit"}],"alternative":[{"slug":"squid","why":"Same shape — a skills layer that upgrades your coding assistant into a specialized workflow — different bet: agents-cli specializes it for the Google Cloud agent lifecycle; squid for a convention-enforcing spec→PR software factory.","dir":"in","confidence":0.55,"name":"squid"},{"slug":"langgraph","why":"Both are ways to build production agentic apps, but from opposite ends: LangGraph is an open, embeddable graph runtime that runs anywhere, while agents-cli is Google's CLI+skills around ADK that scaffolds, evaluates, and deploys agents specifically on Google Cloud / Gemini Enterprise. Choose by ecosystem and portability needs.","dir":"out","confidence":0.5,"name":"LangGraph"},{"slug":"ragas","why":"agents-cli ships a full agent-eval suite — trace generation, metric grading, LLM-as-judge, failure-mode clustering, prompt optimization — which overlaps Ragas's job. Difference: Ragas is eval-only and framework-agnostic, while agents-cli's eval is one stage of a GCP-bound scaffold/deploy pipeline.","dir":"out","confidence":0.45,"name":"Ragas"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/google/agents-cli"},"ai-engineering-hub":{"name":"ai-engineering-hub","owner":"patchy631","slug":"ai-engineering-hub","stars":36661,"image":"https://raw.githubusercontent.com/patchy631/ai-engineering-hub/main/assets/ai-eng-hub.gif","avatar":"https://avatars.githubusercontent.com/u/38653995?v=4&s=96","forks":6063,"language":"Jupyter Notebook","license":"MIT","updated":"9 days ago","topics":["coding"],"summary":"36k-star hub of runnable AI engineering tutorials — LLMs, RAG and agent apps as self-contained projects, including build-code-harness: a Claude-Code-style coding harness rebuilt on CrewAI + E2B.","curator_note":"A tutorial hub, not a tool — each folder is a complete, runnable project with real dependencies. The standout is build-code-harness: it rebuilds a coding-agent harness (planning, memory, checkpoints, sandbox, human gate) one layer at a time on CrewAI + E2B, the best way to understand what harnesses like Claude Code actually do. Content repo caveat: quality varies by folder, and there's a newsletter funnel attached.","edges":[{"to":"crewai","type":"built_with","why":"The flagship walkthroughs — including build-code-harness's hierarchical agent workflow — are built on CrewAI; the hub doubles as its largest applied-tutorial corpus.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:41.871Z","linkCount":1,"related":{"complements":[],"alternative":[],"built_with":[{"slug":"crewai","why":"The flagship walkthroughs — including build-code-harness's hierarchical agent workflow — are built on CrewAI; the hub doubles as its largest applied-tutorial corpus.","dir":"out","confidence":0.55,"name":"CrewAI"}]},"url":"https://stackmap.shipwithai.xyz/repos/patchy631/ai-engineering-hub"},"ai-job-search":{"name":"ai-job-search","owner":"MadsLorentzen","slug":"ai-job-search","stars":25490,"image":"https://raw.githubusercontent.com/MadsLorentzen/ai-job-search/master/assets/mascot/pip_flight_loop.gif","avatar":"https://avatars.githubusercontent.com/u/50207393?v=4&s=96","forks":8331,"language":"TypeScript","license":"MIT","updated":"yesterday","topics":["agents"],"summary":"A job-application framework built ON Claude Code: fork it, fill in your profile, and the agent evaluates postings, tailors CVs, writes cover letters and preps interviews — locally.","curator_note":"The best example yet of a 'fork-and-own' vertical agent: your data stays on your machine, the workflow is readable Markdown you adapt, and 22k stars say the shape resonates. Even if you never job-hunt, read it as a template for building personal agent workflows on Claude Code. NOT a service — you run and maintain it, and output quality tracks the effort you put into your profile files; garbage in, generic cover letters out.","edges":[],"status":"approved","added":"2026-07-14T16:04:54.715Z","linkCount":0,"related":{"complements":[],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/MadsLorentzen/ai-job-search"},"ai-marketing-claude":{"name":"ai-marketing-claude","owner":"zubair-trabzada","slug":"ai-marketing-claude","stars":2201,"image":"https://raw.githubusercontent.com/zubair-trabzada/ai-marketing-claude/main/banner.svg","avatar":"https://avatars.githubusercontent.com/u/13802400?v=4&s=96","forks":647,"language":"Python","license":"MIT","updated":"4 months ago","topics":["skills"],"summary":"15 marketing skills for Claude Code — /market audit runs 5 parallel agents scoring a site across 6 dimensions; copy, email sequences, ad creative, competitor intel and client-ready PDF reports.","curator_note":"The 'sell marketing services with Claude Code' starter kit: one command fans out five subagents that score a site on content, conversion, SEO, positioning, brand and growth, and the deliverables (weighted scores, before/after copy, client proposals, PDF reports) are shaped for agency work, not just self-serve tinkering. Templates for welcome/nurture/launch sequences included. Know what it is: prompts and scoring rubrics, not integrations — nothing talks to ad accounts or analytics APIs (claude-ads is the operations side of this coin). The scoring is LLM-judged, so treat numbers as directional. Repo hasn't moved since March and the README funnels to a paid Skool community — the MIT code stands alone fine.","edges":[{"to":"designer-skills","type":"alternative","why":"Same shape, different profession: both are domain-expertise skill packs that turn Claude Code into a practitioner — designer-skills for research→design→delivery, this one for audits→copy→campaigns→client reports.","confidence":0.5,"status":"approved"},{"to":"claude-ads","type":"complements","why":"Two halves of Claude-Code marketing work: ai-marketing-claude generates the strategy layer — audits, copy, sequences, client reports — while claude-ads operates the actual ad accounts across 12 platforms with gated, rollback-safe changes.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T15:39:30.646Z","linkCount":3,"related":{"complements":[{"slug":"claude-ads","why":"Two halves of Claude-Code marketing work: ai-marketing-claude generates the strategy layer — audits, copy, sequences, client reports — while claude-ads operates the actual ad accounts across 12 platforms with gated, rollback-safe changes.","dir":"out","confidence":0.55,"name":"claude-ads"}],"alternative":[{"slug":"aaron-marketing-skills","why":"Both turn Claude Code into a marketing practitioner: ai-marketing-claude is 15 skills with parallel audit agents and agency-shaped PDF deliverables; aaron-marketing-skills is 120 skills across 7 disciplines with benchmark-driven auditor gates and one shared contract.","dir":"in","confidence":0.65,"name":"aaron-marketing-skills"},{"slug":"designer-skills","why":"Same shape, different profession: both are domain-expertise skill packs that turn Claude Code into a practitioner — designer-skills for research→design→delivery, this one for audits→copy→campaigns→client reports.","dir":"out","confidence":0.5,"name":"designer-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/zubair-trabzada/ai-marketing-claude"},"ai-token-monitor":{"name":"ai-token-monitor","owner":"soulduse","slug":"ai-token-monitor","stars":288,"image":"https://raw.githubusercontent.com/soulduse/ai-token-monitor/main/docs/images/hero.png","avatar":"https://avatars.githubusercontent.com/u/11257459?v=4&s=96","forks":47,"language":"TypeScript","license":null,"updated":"6 days ago","topics":["coding"],"summary":"Menu-bar/tray app for macOS and Windows that reads Claude Code, Codex and OpenCode session logs and shows live token spend — per-model pricing with cache reads, plan-limit bars, webhook alerts.","curator_note":"Zero-setup observability for the coding-agent bill: no API keys, no proxy — it prices the session logs your CLIs already write (cache reads included) and puts today's spend next to the clock, with live 5-hour-session and weekly plan bars that ping Discord/Slack/Telegram before you hit the wall. Offline by default; the leaderboard and chat are strictly opt-in. NOT enforcement or routing — it measures, it doesn't reduce; pair it with a gateway or harness caching fixes when the number scares you. macOS/Windows only, and note the repo currently ships no license file.","edges":[{"to":"prompt-cache-skills","type":"complements","why":"Measure, then fix: the monitor prices cache reads separately, so it's the visible before/after for the prompt-caching patches — apply a skill, watch the same session logs show the spend drop.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.227Z","linkCount":5,"related":{"complements":[{"slug":"prompt-cache-skills","why":"Measure, then fix: the monitor prices cache reads separately, so it's the visible before/after for the prompt-caching patches — apply a skill, watch the same session logs show the spend drop.","dir":"out","confidence":0.5,"name":"prompt-cache-skills"},{"slug":"siliconscope","why":"Two menu-bar monitors covering the two halves of local AI work: ai-token-monitor watches what agents spend in tokens, SiliconScope watches what the hardware pays in compute, power and bandwidth.","dir":"in","confidence":0.5,"name":"SiliconScope"}],"alternative":[{"slug":"abtop","why":"Same watch-your-agents job, different surface: ai-token-monitor is a menu-bar app focused on token spend and plan limits; abtop is a terminal TUI adding context-window bars, child processes and port detection per live session.","dir":"in","confidence":0.75,"name":"abtop"},{"slug":"claude-code-karma","why":"Both read the session logs your coding CLI already writes, at opposite depths: ai-token-monitor is the glanceable menu-bar spend counter with plan-limit alerts; Karma is the full forensic dashboard — timelines, subagents, tools, tickets.","dir":"in","confidence":0.6,"name":"claude-code-karma"},{"slug":"cc-lens","why":"Both read local session logs to answer 'what is Claude Code costing me' — ai-token-monitor as a glanceable menu-bar counter with plan-limit alerts, cc-lens as a full browser dashboard with history, insights and budgets.","dir":"in","confidence":0.55,"name":"cc-lens"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/soulduse/ai-token-monitor"},"airllm":{"name":"airllm","owner":"lyogavin","slug":"airllm","stars":23923,"image":"https://raw.githubusercontent.com/lyogavin/airllm/main/assets/airllm_logo_sm.png?v=3&raw=true","avatar":"https://avatars.githubusercontent.com/u/1113905?v=4&s=96","forks":2700,"language":"Jupyter Notebook","license":"Apache-2.0","updated":"yesterday","topics":["local"],"summary":"Layer-by-layer inference that runs 70B models on a 4GB GPU — no quantization required; 405B on 8GB, DeepSeek-V3 671B on ~12GB. One AutoModel line for most open model families.","curator_note":"The trick is elegant and the tradeoff is brutal, and you should know both: only one layer lives on the GPU at a time, so VRAM scales with layer size instead of model size — that's how 671B fits on a hobbyist card — but every token streams the whole model from disk, so generation runs at seconds-per-token. Use it for batch/offline jobs where 'it fits' beats 'it's fast', for poking at frontier-scale open models on hardware you own, or with block-wise 4/8-bit compression for a ~3x claw-back. NOT a chat or serving solution: Ollama is the fits-in-VRAM daily driver, vLLM the throughput server. Apple-silicon Macs supported via MLX. README carries the author's sponsor/affiliate links — the library stands on its own.","edges":[{"to":"ollama","type":"alternative","why":"Both run open models on your own hardware, on opposite sides of one constraint: Ollama gives fast, polished local inference for models that fit your VRAM; AirLLM runs models that don't fit at all — 70B on 4GB — by streaming one layer at a time, at heavy latency cost.","confidence":0.6,"status":"approved"},{"to":"vllm","type":"alternative","why":"Opposite ends of the local-inference spectrum: vLLM maximizes throughput given abundant VRAM (production serving); AirLLM minimizes VRAM given abundant patience (frontier-size models on consumer cards).","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-16T16:07:34.645Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"colibri","why":"Same job — giant models on tiny hardware — opposite technique: AirLLM streams dense layers through a 4GB GPU from Python/HF, colibrì streams MoE experts from NVMe in pure C with a workload-learning pin cache.","dir":"in","confidence":0.85,"name":"colibri"},{"slug":"ollama","why":"Both run open models on your own hardware, on opposite sides of one constraint: Ollama gives fast, polished local inference for models that fit your VRAM; AirLLM runs models that don't fit at all — 70B on 4GB — by streaming one layer at a time, at heavy latency cost.","dir":"out","confidence":0.6,"name":"Ollama"},{"slug":"mesh-llm","why":"Opposite answers to 'the model doesn't fit': AirLLM streams layers from disk on one small GPU (slow, solo), mesh-llm splits stages across peers' GPUs (faster, needs friends).","dir":"in","confidence":0.55,"name":"mesh-llm"},{"slug":"vllm","why":"Opposite ends of the local-inference spectrum: vLLM maximizes throughput given abundant VRAM (production serving); AirLLM minimizes VRAM given abundant patience (frontier-size models on consumer cards).","dir":"out","confidence":0.5,"name":"vLLM"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/lyogavin/airllm"},"all-agentic-architectures":{"name":"all-agentic-architectures","owner":"FareedKhan-dev","slug":"all-agentic-architectures","stars":3915,"avatar":"https://avatars.githubusercontent.com/u/63067900?v=4&s=96","forks":675,"language":"Jupyter Notebook","license":"MIT","updated":"1 months ago","topics":["agents"],"summary":"35 agentic patterns (Reflexion, LATS, GraphRAG, MemGPT, Voyager…) as one Python library plus a runnable textbook — real LLM outputs, 9 providers, benchmark leaderboard, 283 tests.","curator_note":"The reference shelf for agent design, executable: every major pattern from the literature as an `Architecture` class with a uniform contract, each with a fully-run notebook (no mocked outputs) and a benchmark leaderboard ranking patterns per task — so 'which architecture for this job' gets an empirical answer, not a blog opinion. Read it before you invent your own loop. NOT a production framework: the uniform contract serves comparison and learning, and you'll re-implement the winner in your own stack — treat it as the textbook it says it is, not a dependency.","edges":[{"to":"langgraph","type":"built_with","why":"Several of the 35 architectures ship as runnable LangGraph implementations — the textbook teaches on the substrate.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-13T23:34:30.335Z","linkCount":1,"related":{"complements":[],"alternative":[],"built_with":[{"slug":"langgraph","why":"Several of the 35 architectures ship as runnable LangGraph implementations — the textbook teaches on the substrate.","dir":"out","confidence":0.7,"name":"LangGraph"}]},"url":"https://stackmap.shipwithai.xyz/repos/FareedKhan-dev/all-agentic-architectures"},"allama":{"name":"allama","owner":"digitranslab","slug":"allama","stars":181,"image":"https://raw.githubusercontent.com/digitranslab/allama/main/frontend/public/icon.svg","avatar":"https://avatars.githubusercontent.com/u/190535086?v=4&s=96","forks":14,"language":"Python","license":"AGPL-3.0","updated":"5 months ago","topics":["security","agents"],"summary":"Open-source AI security automation (SOAR): visual playbook builder, autonomous triage agents, and 80+ SIEM/EDR/identity/ticketing integrations. Self-hosted, multi-tenant.","curator_note":"The open answer to $100k SOAR contracts: drag-and-drop security playbooks, AI agents that enrich and prioritize the 500-alert firehose, and integrations across the stack you already run (Splunk, CrowdStrike, Okta, Jira…) — with local models via Ollama so nothing sensitive leaves your infra. The honesty check: it's young (~180 stars) and the commit graph has been quiet for months, so treat claims as a pilot to verify, not a deployment to trust — and AGPL-3.0 matters if you embed it in a service. Ingested over the curator's maturity objection: the niche (open AI-SOAR) is real and empty.","edges":[{"to":"ollama","type":"complements","why":"Security workloads are exactly where prompts can't leave the building — Allama's agents run against self-hosted models through Ollama for fully on-prem triage.","confidence":0.6,"status":"approved"},{"to":"opensre","type":"alternative","why":"The same pattern — AI agents automating incident response — pointed at different fires: OpenSRE at production outages, Allama at security alerts. Pick by which pager you carry.","confidence":0.55,"status":"approved"},{"to":"bytechef","type":"alternative","why":"Both are visual workflow-automation platforms with AI agents — ByteChef general-purpose across 200+ integrations, Allama specialized for SOC alert response.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:13.008Z","linkCount":4,"related":{"complements":[{"slug":"ollama","why":"Security workloads are exactly where prompts can't leave the building — Allama's agents run against self-hosted models through Ollama for fully on-prem triage.","dir":"out","confidence":0.6,"name":"Ollama"},{"slug":"flowsint","why":"Adjacent stages of self-hosted security operations: Flowsint is the investigation canvas where an analyst maps who/what during recon or fraud cases; allama automates the detection-and-response side with SOAR playbooks and triage agents.","dir":"in","confidence":0.5,"name":"flowsint"}],"alternative":[{"slug":"opensre","why":"The same pattern — AI agents automating incident response — pointed at different fires: OpenSRE at production outages, Allama at security alerts. Pick by which pager you carry.","dir":"out","confidence":0.55,"name":"opensre"},{"slug":"bytechef","why":"Both are visual workflow-automation platforms with AI agents — ByteChef general-purpose across 200+ integrations, Allama specialized for SOC alert response.","dir":"out","confidence":0.55,"name":"bytechef"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/digitranslab/allama"},"alook":{"name":"alook","owner":"alookai","slug":"alook","stars":971,"image":"https://raw.githubusercontent.com/alookai/alook/main/assets/readme-banner.png","avatar":"https://avatars.githubusercontent.com/u/273330576?v=4&s=96","forks":151,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["coding","orchestration"],"summary":"Self-hosted collaboration layer that turns local coding agents (Claude Code, Codex, OpenCode) into an always-on \"AI company\" — per-agent email, org chart, kanban, shared memory.","curator_note":"Pick alook when you already live in Claude Code or Codex and want those agents running as a persistent little team — email in and out, a kanban they work through, schedules, memory that compounds — without writing a line of orchestration code. It's a product, not a framework: BYO agent, you're the CEO. NOT for building agents programmatically (that's LangGraph/CrewAI territory), and if you need budgets, governance, and audit over a heterogeneous fleet, paperclip is the heavier, control-plane take on the same idea. Young project — expect sharp edges and a moving roadmap.","edges":[{"to":"paperclip","type":"alternative","why":"Same job — run your agents as a 'company' with an org chart and task system. paperclip bets on governance (budgets, audits, tickets) for heterogeneous fleets; alook bets on email-native simplicity for solo builders running coding agents.","confidence":0.85,"status":"approved"},{"to":"core","type":"alternative","why":"Both are self-hosted, always-on 'AI that works for you while you sleep' products with persistent memory — core shapes it as one personal AI OS watching your apps; alook shapes it as a team of role-assigned coding agents.","confidence":0.6,"status":"approved"},{"to":"squid","type":"complements","why":"alook gives Claude Code agents inboxes, roles, and an always-on runtime; squid gives each of them a disciplined spec→PR pipeline — the org chart and the SOP, together.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":10,"related":{"complements":[{"slug":"squid","why":"alook gives Claude Code agents inboxes, roles, and an always-on runtime; squid gives each of them a disciplined spec→PR pipeline — the org chart and the SOP, together.","dir":"out","confidence":0.65,"name":"squid"},{"slug":"loop-engineering","why":"alook is an always-on multi-agent 'AI company' runtime (per-agent email, org chart, kanban, shared memory) for local coding agents; loop-engineering supplies the loop-design methodology plus scheduling, state and budget tooling to drive agents autonomously. Design the loops with one, run them on the other.","dir":"in","confidence":0.55,"name":"loop-engineering"}],"alternative":[{"slug":"paperclip","why":"Same job — run your agents as a 'company' with an org chart and task system. paperclip bets on governance (budgets, audits, tickets) for heterogeneous fleets; alook bets on email-native simplicity for solo builders running coding agents.","dir":"out","confidence":0.85,"name":"paperclip"},{"slug":"squad","why":"Both coordinate multiple local coding agents into a team; alook builds the full always-on 'AI company' (email, org charts), squad strips it to slash commands + SQLite you can hold in your head.","dir":"in","confidence":0.65,"name":"squad"},{"slug":"core","why":"Both are self-hosted, always-on 'AI that works for you while you sleep' products with persistent memory — core shapes it as one personal AI OS watching your apps; alook shapes it as a team of role-assigned coding agents.","dir":"out","confidence":0.6,"name":"core"},{"slug":"auto-company","why":"Both stage an 'AI company' of local coding agents; alook is the collaboration layer you join (email, org charts), Auto-Company the hands-off experiment that runs the whole firm itself.","dir":"in","confidence":0.6,"name":"Auto-Company"},{"slug":"contrabass","why":"Both turn local coding agents into a coordinated team working a shared board; alook is an always-on collaboration layer (per-agent email, kanban, shared memory), Contrabass is a leaner issue-queue dispatcher with per-run verification and retries.","dir":"in","confidence":0.6,"name":"contrabass"},{"slug":"codenomad","why":"Both put a management surface over local coding agents — alook turns them into an always-on multi-agent 'AI company', CodeNomad gives one developer a rich single-cockpit for parallel OpenCode sessions.","dir":"in","confidence":0.55,"name":"CodeNomad"},{"slug":"fusion","why":"Both are collaboration layers that turn local coding agents into a coordinated team on a kanban board; alook leans on per-agent email and shared memory, Fusion on planned workflows, worktree isolation and merge gates.","dir":"in","confidence":0.55,"name":"Fusion"},{"slug":"ruflo","why":"Both layer an always-on multi-agent organization over the coding CLIs you already run. alook keeps it legible — email, org chart, kanban; Ruflo goes maximalist — swarms, adaptive memory, self-learning claims and a UI.","dir":"in","confidence":0.55,"name":"ruflo"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/alookai/alook"},"apm":{"name":"apm","owner":"microsoft","slug":"apm","stars":3326,"avatar":"https://avatars.githubusercontent.com/u/6154722?v=4&s=96","forks":299,"language":"Python","license":"MIT","updated":"4 days ago","topics":["coding","skills"],"summary":"Microsoft's dependency manager for agent context — declare skills, prompts, plugins and MCP servers in apm.yml; one install reproduces the setup across 8 clients with lockfile pinning and org policy.","curator_note":"The team/enterprise answer: apm.yml ships with the repo, so `git clone && apm install` gives every developer identical agent context — lockfile, transitive dependencies, drift detection, and an apm-policy.yml a security team can enforce with tighten-only inheritance. It's the only skill manager with a real governance story (SBOM export, MCP trust gates). NOT for quick personal use — manifest ceremony is overkill for 'just install this skill', where skillkit or asm is faster — and it targets 8 major clients, not the 40+ long tail. Young project; treat the roadmap as direction, not promise.","edges":[{"to":"skillkit","type":"alternative","why":"Overlapping job, different model: skillkit is an imperative installer + format translator chasing breadth (46 agents, 400K skills); apm is declarative — manifest, lockfile, transitive deps, org policy — chasing reproducibility. apm's README even ships a 'coming from npx skills add' migration.","confidence":0.8,"status":"approved"},{"to":"asm","type":"alternative","why":"Both manage agent skills across tools, at different layers: asm is an imperative, machine-driven CLI for installing and auditing skills on one workstation; apm is a declarative project manifest — apm.yml + lockfile ships with the repo so the whole team reproduces the same agent context.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-07T13:32:35.539Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"skillkit","why":"Overlapping job, different model: skillkit is an imperative installer + format translator chasing breadth (46 agents, 400K skills); apm is declarative — manifest, lockfile, transitive deps, org policy — chasing reproducibility. apm's README even ships a 'coming from npx skills add' migration.","dir":"out","confidence":0.8,"name":"skillkit"},{"slug":"asm","why":"Both manage agent skills across tools, at different layers: asm is an imperative, machine-driven CLI for installing and auditing skills on one workstation; apm is a declarative project manifest — apm.yml + lockfile ships with the repo so the whole team reproduces the same agent context.","dir":"out","confidence":0.75,"name":"asm"},{"slug":"ecc","why":"Both ship reproducible agent-context setups across many clients. apm is a lean dependency manager (declare skills/prompts/MCP in apm.yml, lockfile-pinned installs); ECC ships the content itself — a batteries-included catalog of skills, agents, hooks and rules.","dir":"in","confidence":0.55,"name":"ECC"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/microsoft/apm"},"asm":{"name":"asm","owner":"luongnv89","slug":"asm","stars":742,"image":"https://raw.githubusercontent.com/luongnv89/asm/main/assets/logo/logo-full.svg","avatar":"https://avatars.githubusercontent.com/u/3288457?v=4&s=96","forks":62,"language":"TypeScript","license":"MIT","updated":"2 days ago","topics":["coding","skills"],"summary":"Scriptable skill manager for AI coding agents — install, search, dedupe-audit and security-scan skills across 19 providers, with --json/--yes on every command so agents and CI can drive it.","curator_note":"Pick it when the CLI consumer is a machine: every command is non-interactive and JSON-emitting, so your agent or CI pipeline can inventory, install, and dedupe skills without a human — that plus `asm audit` (finds duplicate/stale skills scattered across 19 providers' hidden dirs) is the differentiator. NOT the breadth play: ~4.3K curated skills vs skillkit's 400K firehose, and no format translation between agent dialects. Solo-maintainer project — fine for personal/team tooling, but if you need org-level reproducibility and governance, apm's manifest+policy model is the grown-up answer.","edges":[{"to":"skillkit","type":"alternative","why":"Same job — install and manage skills across many coding agents — opposite bets: skillkit maximizes breadth (400K skills, 46 agents, format auto-translation); asm maximizes scriptability (agent-first --json/--yes CLI, dedupe audit, curated catalog, no accounts/telemetry).","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-07T13:32:05.187Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"skillkit","why":"Same job — install and manage skills across many coding agents — opposite bets: skillkit maximizes breadth (400K skills, 46 agents, format auto-translation); asm maximizes scriptability (agent-first --json/--yes CLI, dedupe audit, curated catalog, no accounts/telemetry).","dir":"out","confidence":0.8,"name":"skillkit"},{"slug":"apm","why":"Both manage agent skills across tools, at different layers: asm is an imperative, machine-driven CLI for installing and auditing skills on one workstation; apm is a declarative project manifest — apm.yml + lockfile ships with the repo so the whole team reproduces the same agent context.","dir":"in","confidence":0.75,"name":"apm"},{"slug":"skillnet","why":"Same job — search and install agent skills from a large catalog via CLI/SDK. asm is the scriptable ops tool (dedupe audits, security scans, --json everywhere for CI); SkillNet adds generation, scoring and graph analysis on top of discovery.","dir":"in","confidence":0.65,"name":"SkillNet"},{"slug":"autoskills","why":"Same category — skill installation tooling. asm is scriptable and provider-spanning with audit commands; autoskills trades all knobs for one command and a curated, hash-pinned registry.","dir":"in","confidence":0.6,"name":"autoskills"},{"slug":"n-skills","why":"Both install and manage agent skills across providers; asm is the scriptable power-tool with audit/dedupe, n-skills the curated storefront on the universal format.","dir":"in","confidence":0.55,"name":"n-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/luongnv89/asm"},"auto-company":{"name":"Auto-Company","owner":"MaxMiksa","slug":"auto-company","stars":2075,"image":"https://raw.githubusercontent.com/MaxMiksa/auto-company/main/presentation/dashboard-showcase.png","avatar":"https://avatars.githubusercontent.com/u/195391698?v=4&s=96","forks":317,"language":"Python","license":null,"updated":"2 months ago","topics":["agents","orchestration"],"summary":"A fully autonomous 'AI company' on your own PC: 14 expert-modeled agents ideate, decide, code, deploy and market 24/7 — driven by Claude Code or Codex CLI, with a local dashboard.","curator_note":"The maximalist experiment: give 14 role-agents a company charter and let them run around the clock on your hardware. As a living demo of agentic workflows it's genuinely instructive — watch where autonomy compounds and where it wanders. Treat the 'without human intervention' framing as aspiration, NOT audit: unattended agents produce unattended mistakes at 24/7 speed, token bills to match, and there's NO license file at review time — check before building on it.","edges":[{"to":"alook","type":"alternative","why":"Both stage an 'AI company' of local coding agents; alook is the collaboration layer you join (email, org charts), Auto-Company the hands-off experiment that runs the whole firm itself.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:54.818Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"alook","why":"Both stage an 'AI company' of local coding agents; alook is the collaboration layer you join (email, org charts), Auto-Company the hands-off experiment that runs the whole firm itself.","dir":"out","confidence":0.6,"name":"alook"},{"slug":"fusion","why":"Both run autonomous 'AI companies' on your hardware: Auto-Company ships 14 expert-modeled agents ideating and shipping 24/7; Fusion makes the company importable and inspectable — org charts, token share per agent, multi-agent chat rooms.","dir":"in","confidence":0.6,"name":"Fusion"},{"slug":"ruflo","why":"Same dream — an autonomous multi-agent operation driven through Claude Code/Codex — different scope: Auto-Company scripts 14 fixed expert roles; Ruflo is a general meta-harness you compose swarms from.","dir":"in","confidence":0.5,"name":"ruflo"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/MaxMiksa/auto-company"},"autogen":{"name":"AutoGen","owner":"microsoft","slug":"autogen","stars":59914,"image":"https://microsoft.github.io/autogen/0.2/img/ag.svg","avatar":"https://avatars.githubusercontent.com/u/6154722?v=4&s=96","forks":9019,"language":"Python","license":"CC-BY-4.0","updated":"3 months ago","topics":["agents","evals"],"summary":"Multi-agent conversation framework for building LLM applications with cooperating agents.","edges":[{"to":"ragas","type":"complements","why":"Score the outputs of your multi-agent conversations for faithfulness and relevance.","confidence":0.8,"status":"approved"},{"to":"ollama","type":"built_with","why":"Back your agents with a local model server.","confidence":0.7,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":9,"related":{"complements":[{"slug":"ragas","why":"Score the outputs of your multi-agent conversations for faithfulness and relevance.","dir":"out","confidence":0.8,"name":"Ragas"},{"slug":"cubesandbox","why":"AutoGen's code-executor step is the canonical sandbox case: run generated Python in an isolated microVM instead of on the host or a shared container.","dir":"in","confidence":0.75,"name":"CubeSandbox"}],"alternative":[{"slug":"crewai","why":"Both do multi-agent collaboration; CrewAI is more opinionated about roles, AutoGen about conversation.","dir":"in","confidence":0.8,"name":"CrewAI"},{"slug":"metagpt","why":"Same job — multi-agent LLM apps — opposite philosophy: AutoGen is a conversation substrate you shape freely; MetaGPT hard-codes the workflow as company SOPs.","dir":"in","confidence":0.8,"name":"MetaGPT"},{"slug":"praisonai","why":"Both orchestrate cooperating agents; AutoGen is the research-grade conversation framework, PraisonAI the low-code productized bundle.","dir":"in","confidence":0.7,"name":"PraisonAI"},{"slug":"deepagents","why":"Both are batteries-included multi-agent frameworks; AutoGen centers on agent-to-agent conversation, deepagents on a single planner with sub-agents, filesystem and context management.","dir":"in","confidence":0.6,"name":"deepagents"},{"slug":"council-of-high-intelligence","why":"Both stage multi-agent deliberation; AutoGen is the framework you build conversations with, Council is the finished product — personas, protocol and verdict included, one command away.","dir":"in","confidence":0.5,"name":"council-of-high-intelligence"},{"slug":"paperclip","why":"AutoGen is a multi-agent framework for cooperating LLM agents — same coordination job as Paperclip, but as a library you build with rather than a dashboard you run external agents under.","dir":"in","confidence":0.5,"name":"paperclip"}],"built_with":[{"slug":"ollama","why":"Back your agents with a local model server.","dir":"out","confidence":0.7,"name":"Ollama"}]},"url":"https://stackmap.shipwithai.xyz/repos/microsoft/autogen"},"autohedge":{"name":"AutoHedge","owner":"The-Swarm-Corporation","slug":"autohedge","stars":3900,"avatar":"https://avatars.githubusercontent.com/u/167784574?v=4&s=96","forks":659,"language":"Python","license":"MIT","updated":"2 months ago","topics":["finance","agents"],"summary":"Swarm-agent 'autonomous hedge fund': cooperating agents automate market analysis, risk management and trade execution. Python, from the Swarms ecosystem.","curator_note":"Multi-agent architecture applied to trading: director/analyst/risk agents deliberate before execution — as a reference architecture for agent-team decision pipelines it's worth reading. Now the cold water: 'enterprise-grade autonomous hedge fund' is marketing, not audit — no published live track record; the Swarms ecosystem runs hype-forward; the repo was quiet for two months at review. Paper-trade it, treat real capital as adversarial testing you pay for.","edges":[{"to":"vibe-trading","type":"alternative","why":"Same ambition — agents trading on your behalf — with opposite temperaments: Vibe-Trading (HKUDS) ships shadow accounts and an MCP surface for cautious integration; AutoHedge sells the autonomous hedge-fund dream. Both need your skepticism.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:54.855Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"vibe-trading","why":"Same ambition — agents trading on your behalf — with opposite temperaments: Vibe-Trading (HKUDS) ships shadow accounts and an MCP surface for cautious integration; AutoHedge sells the autonomous hedge-fund dream. Both need your skepticism.","dir":"out","confidence":0.6,"name":"Vibe-Trading"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/The-Swarm-Corporation/autohedge"},"autoresearch":{"name":"autoresearch","owner":"uditgoenka","slug":"autoresearch","stars":5374,"avatar":"https://avatars.githubusercontent.com/u/1478769?v=4&s=96","forks":399,"language":"Shell","license":"MIT","updated":"1 months ago","topics":["skills","coding"],"summary":"Karpathy's autoresearch loop as an installable skill for Claude Code, OpenCode and Codex: constraint + mechanical metric + autonomous modify→verify→keep/discard iteration.","curator_note":"The compounding-gains loop, packaged: pick a metric a machine can check, let the agent mutate-verify-keep against it for hours, and small wins stack — the Karpathy recipe without writing the harness yourself. Works across three agent CLIs. The discipline it demands is the catch: without a truly mechanical metric the loop optimizes noise, and unattended iteration burns real tokens — set budgets before you set it loose.","edges":[{"to":"loop-engineering","type":"alternative","why":"Both operationalize 'the loop is the unit of progress': loop-engineering is the design methodology and CLIs for building gated loops; autoresearch is one specific, proven loop shipped as an installable skill.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:54.890Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"evo","why":"Same lineage — Karpathy's autoresearch loop for coding agents. The skill is a single-branch keep/discard hill-climb; evo adds tree search, parallel worktree subagents, shared failure traces, gates and a dashboard. Start with the skill, graduate to evo.","dir":"in","confidence":0.85,"name":"evo"},{"slug":"loop-engineering","why":"Both operationalize 'the loop is the unit of progress': loop-engineering is the design methodology and CLIs for building gated loops; autoresearch is one specific, proven loop shipped as an installable skill.","dir":"out","confidence":0.6,"name":"loop-engineering"},{"slug":"sia","why":"Same lineage — the autoresearch loop: metric, modify, verify, keep or discard. The skill hill-climbs your code with a fixed agent; SIA makes the agent itself the thing that improves, up to and including its weights.","dir":"in","confidence":0.6,"name":"sia"},{"slug":"loopy","why":"Both distribute reusable agent loops: autoresearch is one loop (Karpathy's research iteration) as an installable skill; Loopy is a whole catalog of loops plus find/audit/run tooling.","dir":"in","confidence":0.5,"name":"loopy"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/uditgoenka/autoresearch"},"autoscraper":{"name":"autoscraper","owner":"alirezamika","slug":"autoscraper","stars":7677,"image":"https://user-images.githubusercontent.com/17881612/91968083-5ee92080-ed29-11ea-82ec-d99ec85367a5.png","avatar":"https://avatars.githubusercontent.com/u/17881612?v=4&s=96","forks":789,"language":"Python","license":"MIT","updated":"1 years ago","topics":["web"],"summary":"Learn-by-example Python scraper: give it a URL and sample values you want, it infers the extraction rules and reapplies them to similar pages. Tiny, fast, zero selectors.","curator_note":"The cleverest 500 lines in scraping: show it one example of what you want off a page and it figures out the rules — no selectors, no XPath, and the learned model reapplies across similar pages. Perfect for quick structured grabs and prototyping. But check the commit log before adopting: it's been quiet for over a year, so treat it as a finished small tool, NOT a maintained framework — no JS rendering, no anti-bot, no crawling infrastructure. When sites fight back or scale arrives, move to a real framework.","edges":[{"to":"scrapling","type":"alternative","why":"Both attack selector fragility by learning: autoscraper infers extraction rules from one example and stops there; Scrapling relearns elements after redesigns inside a full, maintained crawling framework.","confidence":0.6,"status":"approved"},{"to":"crawlee","type":"alternative","why":"Same scraping job, opposite shapes: autoscraper is 500 lines of learn-by-example rule inference in Python; Crawlee is the production TypeScript framework with queues, proxies and anti-blocking.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.235Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"scrapling","why":"Both attack selector fragility by learning: autoscraper infers extraction rules from one example and stops there; Scrapling relearns elements after redesigns inside a full, maintained crawling framework.","dir":"out","confidence":0.6,"name":"Scrapling"},{"slug":"crawlee","why":"Same scraping job, opposite shapes: autoscraper is 500 lines of learn-by-example rule inference in Python; Crawlee is the production TypeScript framework with queues, proxies and anti-blocking.","dir":"out","confidence":0.6,"name":"crawlee"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/alirezamika/autoscraper"},"autoskills":{"name":"autoskills","owner":"midudev","slug":"autoskills","stars":6564,"image":"https://autoskills.sh/og.jpg","avatar":"https://avatars.githubusercontent.com/u/1561955?v=4&s=96","forks":596,"language":"Ruby","license":"NOASSERTION","updated":"5 days ago","topics":["skills"],"summary":"One command installs your project's AI skill stack: scans package.json/Gradle/configs, detects the tech stack, and pulls matching skills from an audited, hash-verified registry.","curator_note":"The zero-config on-ramp for agent skills: `npx autoskills` detects your stack and installs only the matching skills — and the security model is the standout, a maintainer-synced registry scanned for prompt injection with SHA-256 manifests and a lockfile, instead of live-downloading from random repos. NOT for hand-picking: you get the registry's opinion of what a Next.js or Go project needs, and the registry's coverage is web/mobile-stack-shaped — niche stacks fall through. No standard license file at the time of review; check before corporate adoption.","edges":[{"to":"skillkit","type":"alternative","why":"Both end with skills installed in your agent; skillkit is the explicit package manager you drive, autoskills the zero-config detector that decides for you from a smaller audited registry.","confidence":0.7,"status":"approved"},{"to":"asm","type":"alternative","why":"Same category — skill installation tooling. asm is scriptable and provider-spanning with audit commands; autoskills trades all knobs for one command and a curated, hash-pinned registry.","confidence":0.6,"status":"approved"},{"to":"n-skills","type":"alternative","why":"Two curation-first answers to skill installation: autoskills detects your stack and decides for you from a hash-pinned registry; n-skills hands you a small human-curated marketplace on the universal SKILL.md format and lets you pick.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-12T23:46:16.065Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"skillkit","why":"Both end with skills installed in your agent; skillkit is the explicit package manager you drive, autoskills the zero-config detector that decides for you from a smaller audited registry.","dir":"out","confidence":0.7,"name":"skillkit"},{"slug":"asm","why":"Same category — skill installation tooling. asm is scriptable and provider-spanning with audit commands; autoskills trades all knobs for one command and a curated, hash-pinned registry.","dir":"out","confidence":0.6,"name":"asm"},{"slug":"n-skills","why":"Two curation-first answers to skill installation: autoskills detects your stack and decides for you from a hash-pinned registry; n-skills hands you a small human-curated marketplace on the universal SKILL.md format and lets you pick.","dir":"out","confidence":0.6,"name":"n-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/midudev/autoskills"},"bemyagent":{"name":"bemyagent","owner":"vitotafuni","slug":"bemyagent","stars":22,"avatar":"https://avatars.githubusercontent.com/u/286821?v=4&s=96","forks":6,"language":"HTML","license":"MIT","updated":"16 days ago","topics":["coding","memory"],"summary":"One markdown bootstrap file that scaffolds a structured workspace for any coding agent — persistent project docs plus a task tree run through a Think→Task→Execute→Verify cycle with human pacing gates.","curator_note":"Try it when your agent keeps drifting on multi-session projects and a flat CLAUDE.md isn't enough structure: one file bootstraps a disciplined docs+work tree, it's agent-agnostic (self-registers into .cursorrules/AGENTS.md), and INTERACTIVE mode gives you real plan/result approval gates. NOT for quick scripts or single-session work — the ceremony outweighs the benefit — and it's a young one-maintainer protocol: no enforcement layer exists, so compliance depends entirely on the model faithfully following a long ruleset; weaker models will ignore half of it. Expect to adapt conventions yourself rather than lean on tooling or community.","edges":[{"to":"claude-reflect","type":"complements","why":"Opposite halves of keeping agent context current: bemyagent scaffolds structured project memory up front and registers rules into AGENTS.md; claude-reflect mines your in-session corrections into those same rule files over time. Watch for both writing to AGENTS.md.","confidence":0.55,"status":"approved"},{"to":"squid","type":"alternative","why":"Same job — disciplined, human-gated task execution for coding agents — different bet: squid is a Claude Code plugin driving a 5-agent spec→PR pipeline; bemyagent is a tool-agnostic markdown protocol a single agent follows. Pick squid for enforced structure on Claude Code, bemyagent for portability.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-07T12:59:20.028Z","linkCount":4,"related":{"complements":[{"slug":"claude-reflect","why":"Opposite halves of keeping agent context current: bemyagent scaffolds structured project memory up front and registers rules into AGENTS.md; claude-reflect mines your in-session corrections into those same rule files over time. Watch for both writing to AGENTS.md.","dir":"out","confidence":0.55,"name":"claude-reflect"}],"alternative":[{"slug":"loop-engineering","why":"Both impose a disciplined, tool-agnostic operating structure on coding agents through scaffolding, an explicit work cycle and human gates. bemyagent structures a single human-paced session (Think→Task→Execute→Verify); loop-engineering structures scheduled, unattended loops that run over time. Pick bemyagent for interactive pacing, loop-engineering for autonomous automation.","dir":"in","confidence":0.6,"name":"loop-engineering"},{"slug":"three-man-team","why":"Both are markdown-first process scaffolds for coding agents. bemyagent bootstraps one structured workspace with a Think→Task→Execute→Verify cycle; three-man-team splits the discipline across three role personas with review as a hard gate.","dir":"in","confidence":0.6,"name":"three-man-team"},{"slug":"squid","why":"Same job — disciplined, human-gated task execution for coding agents — different bet: squid is a Claude Code plugin driving a 5-agent spec→PR pipeline; bemyagent is a tool-agnostic markdown protocol a single agent follows. Pick squid for enforced structure on Claude Code, bemyagent for portability.","dir":"out","confidence":0.55,"name":"squid"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/vitotafuni/bemyagent"},"browser-harness-js":{"name":"browser-harness-js","owner":"browser-use","slug":"browser-harness-js","stars":473,"image":"https://r2.browser-use.com/github/asbfgihsbfbaosfjla.png","avatar":"https://avatars.githubusercontent.com/u/192012301?v=4&s=96","forks":33,"language":"TypeScript","license":"MIT","updated":"3 months ago","topics":["web"],"summary":"browser-use's thinnest LLM-to-Chrome bridge: all 652 CDP methods as typed JS calls over one WebSocket — no click() helpers, no rails; the protocol is the API.","curator_note":"The bitter-lesson answer to browser frameworks, from the browser-use team themselves: no helpers — the agent writes raw CDP calls, with types as the docs. Brilliant for capable models on weird pages; frustrating for weak models that need click() rails — that's what browser-use itself is for.","edges":[{"to":"browser-use","type":"alternative","why":"Same team, opposite philosophy: browser-use is the batteries-included agent library with helpers; the harness strips it to 652 typed raw CDP calls and bets the model can write the protocol itself.","confidence":0.7,"status":"approved"},{"to":"stealth-browser-mcp","type":"alternative","why":"Both CDP-level browser control for agents: stealth-browser-mcp specializes in anti-bot evasion via nodriver; the harness is the thinnest generic typed bridge.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.312Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"browser-use","why":"Same team, opposite philosophy: browser-use is the batteries-included agent library with helpers; the harness strips it to 652 typed raw CDP calls and bets the model can write the protocol itself.","dir":"out","confidence":0.7,"name":"browser-use"},{"slug":"stealth-browser-mcp","why":"Both CDP-level browser control for agents: stealth-browser-mcp specializes in anti-bot evasion via nodriver; the harness is the thinnest generic typed bridge.","dir":"out","confidence":0.5,"name":"stealth-browser-mcp"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/browser-use/browser-harness-js"},"browser-use":{"name":"browser-use","owner":"browser-use","slug":"browser-use","stars":106216,"image":"https://github.com/user-attachments/assets/2ccdb752-22fb-41c7-8948-857fc1ad7e24","avatar":"https://avatars.githubusercontent.com/u/192012301?v=4&s=96","forks":11676,"language":"Python","license":"MIT","updated":"yesterday","topics":["agents","web"],"summary":"The standard library for letting AI agents drive a real browser — click, type, fill forms and complete tasks from a natural-language goal. 100k+ stars, Python.","curator_note":"When the task lives behind login walls, forms and JavaScript — 'book this', 'apply to that', 'put these in my cart' — this is the default tool: it feeds the agent a cleaned DOM, executes its clicks/typing, and recovers from the endless weirdness of real websites. Install-as-skill support means Claude Code/Cursor agents pick it up in one prompt. NOT for bulk data extraction — an LLM driving a browser is the slowest, most expensive way to scrape a thousand pages (use a crawler); and treat any agent-with-a-browser as having the keys to whatever it's logged into — sandbox accordingly.","edges":[{"to":"langgraph","type":"complements","why":"browser-use is the hands, LangGraph is the brain: wire it in as the browsing tool inside a stateful agent graph and keep planning, retries and human-in-the-loop where they belong.","confidence":0.65,"status":"approved"},{"to":"cubesandbox","type":"complements","why":"An agent driving a real browser is exactly the workload you want hardware-isolated: run browser-use sessions inside microVM sandboxes instead of on your own profile.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.307Z","linkCount":8,"related":{"complements":[{"slug":"langgraph","why":"browser-use is the hands, LangGraph is the brain: wire it in as the browsing tool inside a stateful agent graph and keep planning, retries and human-in-the-loop where they belong.","dir":"out","confidence":0.65,"name":"LangGraph"},{"slug":"cubesandbox","why":"An agent driving a real browser is exactly the workload you want hardware-isolated: run browser-use sessions inside microVM sandboxes instead of on your own profile.","dir":"out","confidence":0.55,"name":"CubeSandbox"},{"slug":"mails","why":"The missing half of automated signup flows: browser-use drives the registration form; mails receives the verification email in the agent's own mailbox and hands back the code — no human inbox in the loop.","dir":"in","confidence":0.55,"name":"mails"},{"slug":"gym-anything","why":"Build with one, grade with the other: browser-use is the standard library for constructing agents that drive real software; Gym-Anything supplies the standardized environments, tasks and automatic verifiers to benchmark exactly that kind of agent.","dir":"in","confidence":0.5,"name":"gym-anything"}],"alternative":[{"slug":"browser-harness-js","why":"Same team, opposite philosophy: browser-use is the batteries-included agent library with helpers; the harness strips it to 652 typed raw CDP calls and bets the model can write the protocol itself.","dir":"in","confidence":0.7,"name":"browser-harness-js"},{"slug":"browseros","why":"Same job — giving an AI agent a real browser — from opposite ends: browser-use is the Python library you embed in your agent; BrowserClaw is a shipped browser your MCP client drives, riding your existing logins instead of fresh sessions.","dir":"in","confidence":0.65,"name":"BrowserOS"},{"slug":"page-agent","why":"Same end state — AI operating a web interface — from opposite sides of the fence: browser-use is the agent's browser for any site; page-agent is embedded by the site itself, giving its own users an agent.","dir":"in","confidence":0.6,"name":"page-agent"},{"slug":"stealth-browser-mcp","why":"Both hand an AI agent a real browser; browser-use optimizes for task completion on the open web, stealth-browser-mcp for surviving anti-bot walls — pick by whether detection is your bottleneck.","dir":"in","confidence":0.6,"name":"stealth-browser-mcp"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/browser-use/browser-use"},"browseros":{"name":"BrowserOS","owner":"browseros-ai","slug":"browseros","stars":12541,"image":"https://github.com/user-attachments/assets/1e37941c-4dbc-4662-9c8c-3bbe9971301d","avatar":"https://avatars.githubusercontent.com/u/218857586?v=4&s=96","forks":1318,"language":"TypeScript","license":"AGPL-3.0","updated":"yesterday","topics":["agents","web"],"summary":"Open-source agentic browsing, twice: BrowserClaw — a browser your MCP agent drives using your real logged-in sessions — and BrowserOS, a Chromium fork with a built-in AI agent.","curator_note":"Two answers to 'AI needs a browser' in one repo: BrowserClaw lets Claude Code/Cursor/any MCP client drive a real browser with the accounts you're already signed into — watch live, replay every step — and BrowserOS is the privacy-first Chromium fork answering ChatGPT Atlas/Comet for humans. The logged-in-sessions model is the killer feature and the risk: your agent acts as YOU, so scope what it can reach. NOT a scraping library (that's browser-use's Python-SDK territory); AGPL-3.0, and a Chromium fork means trusting their patch cadence for security updates.","edges":[{"to":"browser-use","type":"alternative","why":"Same job — giving an AI agent a real browser — from opposite ends: browser-use is the Python library you embed in your agent; BrowserClaw is a shipped browser your MCP client drives, riding your existing logins instead of fresh sessions.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-13T23:34:30.417Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"browser-use","why":"Same job — giving an AI agent a real browser — from opposite ends: browser-use is the Python library you embed in your agent; BrowserClaw is a shipped browser your MCP client drives, riding your existing logins instead of fresh sessions.","dir":"out","confidence":0.65,"name":"browser-use"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/browseros-ai/browseros"},"bytechef":{"name":"bytechef","owner":"bytechefhq","slug":"bytechef","stars":916,"image":"https://raw.githubusercontent.com/bytechefhq/bytechef/master/static/bytechef_logo.png","avatar":"https://avatars.githubusercontent.com/u/90719245?v=4&s=96","forks":159,"language":"Java","license":"NOASSERTION","updated":"2 days ago","topics":["agents","orchestration"],"summary":"Open-source platform unifying AI agent orchestration with classic workflow automation — visual builder, 200+ integration components, self-hosted via Docker. Apache 2.0 + EE split.","curator_note":"The bet: agent autonomy and deterministic workflow automation belong in ONE platform, not two — let precise integration flows hand work to agents and vice versa. If your team already thinks in n8n/Zapier terms and wants agents in the same canvas, this is the natural home. NOT proven at scale yet (~900 stars, young community for a platform this ambitious), and mind the Apache-2.0 + Enterprise Edition split — check which features live behind the EE line before betting the roadmap.","edges":[{"to":"maxkb","type":"alternative","why":"Both are self-hosted no-code enterprise agent platforms with different centers of gravity: MaxKB starts from the knowledge base and RAG, ByteChef from integrations and workflow automation.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T15:35:46.061Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"sim","why":"Both are open, self-hostable platforms unifying AI-agent orchestration with classic workflow automation behind a visual builder and big integration catalogs — Sim leans agent-first with knowledge/tables/files built in; ByteChef leans integration-first with its 200+ components.","dir":"in","confidence":0.7,"name":"sim"},{"slug":"allama","why":"Both are visual workflow-automation platforms with AI agents — ByteChef general-purpose across 200+ integrations, Allama specialized for SOC alert response.","dir":"in","confidence":0.55,"name":"allama"},{"slug":"maxkb","why":"Both are self-hosted no-code enterprise agent platforms with different centers of gravity: MaxKB starts from the knowledge base and RAG, ByteChef from integrations and workflow automation.","dir":"out","confidence":0.5,"name":"MaxKB"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/bytechefhq/bytechef"},"cc-lens":{"name":"cc-lens","owner":"Arindam200","slug":"cc-lens","stars":572,"image":"https://raw.githubusercontent.com/Arindam200/cc-lens/main/public/cc-lens.png","avatar":"https://avatars.githubusercontent.com/u/109217591?v=4&s=96","forks":77,"language":"TypeScript","license":"MIT","updated":"27 days ago","topics":["coding"],"summary":"npx cc-lens: local analytics dashboard over ~/.claude — sessions with replay, cost and cache breakdowns, insights and budgets, team-adoption mode, a yearly Wrapped card. No cloud, no telemetry.","curator_note":"The zero-friction option among Claude Code dashboards: one npx command, loopback-bound by default, and v0.4's Insights tab actually recommends actions — cache, model choice, compaction, plan-fit and budget opportunities mined from your own usage — instead of just charting it. Team mode (shared-folder push hub + terminal digest) is a real differentiator for leads tracking adoption. Costs are estimates from a local pricing table (`~/.cc-lens/pricing.json` to override) — treat them as directional and keep the table current. Claude-only; for subagent forensics and ticket linking, claude-code-karma digs deeper; for a number next to your clock, ai-token-monitor.","edges":[{"to":"claude-code-karma","type":"alternative","why":"Same job — a local dashboard over ~/.claude JSONL. cc-lens is the one-command npx tool with insights, budgets and team adoption; Karma is the two-process forensic suite with subagent trees, hook/plugin inventories and Linear/Jira/GitHub ticket linking.","confidence":0.7,"status":"approved"},{"to":"ai-token-monitor","type":"alternative","why":"Both read local session logs to answer 'what is Claude Code costing me' — ai-token-monitor as a glanceable menu-bar counter with plan-limit alerts, cc-lens as a full browser dashboard with history, insights and budgets.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T16:07:34.704Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"claude-code-karma","why":"Same job — a local dashboard over ~/.claude JSONL. cc-lens is the one-command npx tool with insights, budgets and team adoption; Karma is the two-process forensic suite with subagent trees, hook/plugin inventories and Linear/Jira/GitHub ticket linking.","dir":"out","confidence":0.7,"name":"claude-code-karma"},{"slug":"abtop","why":"Live vs retrospective: abtop shows what your agents are doing right now; cc-lens analyzes what they did — replays, cost/cache breakdowns, budgets. Many desks end up running both.","dir":"in","confidence":0.6,"name":"abtop"},{"slug":"ai-token-monitor","why":"Both read local session logs to answer 'what is Claude Code costing me' — ai-token-monitor as a glanceable menu-bar counter with plan-limit alerts, cc-lens as a full browser dashboard with history, insights and budgets.","dir":"out","confidence":0.55,"name":"ai-token-monitor"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Arindam200/cc-lens"},"cc-thinking-skills":{"name":"cc-thinking-skills","owner":"tjboudreaux","slug":"cc-thinking-skills","stars":861,"image":"https://raw.githubusercontent.com/tjboudreaux/cc-thinking-skills/main/assets/readme-banner.png","avatar":"https://avatars.githubusercontent.com/u/11684?v=4&s=96","forks":126,"language":"JavaScript","license":"MIT","updated":"1 months ago","topics":["skills"],"summary":"18 mental models as Claude Code skills — First Principles, Bayesian reasoning, Systems Thinking, OODA, Pre-Mortem and more — invoked when a problem needs a thinking framework.","curator_note":"Decision frameworks as installable capability: instead of hoping the model reasons well, you hand it the explicit framework the situation calls for — pre-mortem before a launch, Bayesian updating on flaky evidence, OODA under time pressure. Cheap to adopt, zero infrastructure. The honest ceiling: a framework prompt shapes reasoning, it doesn't guarantee it — the model can still pattern-match its way past the discipline; treat the outputs as structured drafts for YOUR judgment.","edges":[{"to":"council-of-high-intelligence","type":"alternative","why":"Two upgrades for decision quality in Claude Code: Council convenes 18 disagreeing personas across providers; thinking-skills hands one model 18 explicit frameworks. Deliberation breadth vs reasoning discipline.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:54.921Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"council-of-high-intelligence","why":"Two upgrades for decision quality in Claude Code: Council convenes 18 disagreeing personas across providers; thinking-skills hands one model 18 explicit frameworks. Deliberation breadth vs reasoning discipline.","dir":"out","confidence":0.5,"name":"council-of-high-intelligence"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tjboudreaux/cc-thinking-skills"},"cc-wf-studio":{"name":"cc-wf-studio","owner":"breaking-brake","slug":"cc-wf-studio","stars":5328,"image":"https://raw.githubusercontent.com/breaking-brake/cc-wf-studio/main/packages/vscode/resources/icon-large.png","avatar":"https://avatars.githubusercontent.com/u/76818625?v=4&s=96","forks":569,"language":"TypeScript","license":"NOASSERTION","updated":"yesterday","topics":["skills","coding"],"summary":"Visual workflow canvas that exports to the Markdown your agent already understands — design on nodes, ship as skills/agents/commands for Claude Code, Copilot, Codex, Gemini and more.","curator_note":"'You think visually. AI thinks in .md.' — the whole product in one line. Draw the workflow on a canvas, export native formats for seven-plus agents (.claude/agents, .github/skills, .codex/skills…), stop hand-writing prompt files with guessed structure. For teams standardizing multi-step agent workflows across different assistants, the canvas is the shared language. NOT a runtime — it generates files, your agent does the work; and NO standard license file at review time — check before corporate adoption.","edges":[{"to":"skillkit","type":"complements","why":"Studio designs the skill, SkillKit distributes it — canvas-to-Markdown on one side, cross-agent packaging and security scanning on the other.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T15:35:46.158Z","linkCount":1,"related":{"complements":[{"slug":"skillkit","why":"Studio designs the skill, SkillKit distributes it — canvas-to-Markdown on one side, cross-agent packaging and security scanning on the other.","dir":"out","confidence":0.5,"name":"skillkit"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/breaking-brake/cc-wf-studio"},"chandra":{"name":"chandra","owner":"datalab-to","slug":"chandra","stars":11759,"image":"https://raw.githubusercontent.com/datalab-to/chandra/master/assets/datalab-logo.png","avatar":"https://avatars.githubusercontent.com/u/176631703?v=4&s=96","forks":1214,"language":"Python","license":"Apache-2.0","updated":"28 days ago","topics":["ocr"],"summary":"Datalab's SOTA open OCR model: images/PDFs to structured HTML/Markdown/JSON with layout, tables, forms, checkboxes, handwriting and math, in 90+ languages. Local HF or vLLM inference.","curator_note":"Currently the strongest open OCR weights on the olmocr benchmark (85.8, above olmOCR 2 and dots.ocr), with handwriting, filled forms and checkboxes as the real differentiators — plus a serious self-built 90-language benchmark where it averages 72.7% vs Gemini 2.5 Flash's 60.8%. `pip install chandra-ocr`, `chandra_vllm`, done; ~2 pages/s real-world on an H100. The catch is licensing: code is Apache-2.0 but the WEIGHTS are OpenRAIL-M — free for research, personal use and sub-$2M startups, commercial self-hosting needs a Datalab license, and their paid API deliberately stays ahead of the open weights. Pick olmOCR for a fully permissive stack, MinerU when you want a whole parsing pipeline rather than the model itself.","edges":[{"to":"olmocr","type":"alternative","why":"Head-to-head open OCR-VLM rivals on the same benchmark (Chandra 2 scores 85.8 vs olmOCR 2's 82.4). olmOCR is fully permissive and tuned for LLM-training-data linearization; Chandra leads on handwriting/forms/multilingual but carries an OpenRAIL-M weights license.","confidence":0.75,"status":"approved"},{"to":"unlimited-ocr","type":"alternative","why":"Both are open OCR vision-language models. Baidu's model bets on one-shot long-horizon parsing of entire multi-page documents; Chandra processes per-page with stronger layout/table/form structure and a broader language benchmark.","confidence":0.55,"status":"approved"},{"to":"mineru","type":"alternative","why":"Overlapping document-to-markdown job at different layers: MinerU is a full parsing pipeline (layout analysis + OCR + export) you deploy as tooling; Chandra is the single end-to-end OCR model you'd slot into such a pipeline.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.335Z","linkCount":6,"related":{"complements":[{"slug":"commonforms","why":"Opposite halves of the form pipeline: commonforms detects field locations and makes the PDF fillable, chandra reads forms — extracting labels, checkbox states and handwritten values from filled documents.","dir":"in","confidence":0.65,"name":"commonforms"},{"slug":"pdf-inspector","why":"pdf-inspector's whole design goal is smart routing: text-based pages extract locally in milliseconds, and the scanned remainder gets handed to an OCR model like Chandra. Triage first, VLM second.","dir":"in","confidence":0.6,"name":"pdf-inspector"}],"alternative":[{"slug":"olmocr","why":"Head-to-head open OCR-VLM rivals on the same benchmark (Chandra 2 scores 85.8 vs olmOCR 2's 82.4). olmOCR is fully permissive and tuned for LLM-training-data linearization; Chandra leads on handwriting/forms/multilingual but carries an OpenRAIL-M weights license.","dir":"out","confidence":0.75,"name":"olmocr"},{"slug":"unlimited-ocr","why":"Both are open OCR vision-language models. Baidu's model bets on one-shot long-horizon parsing of entire multi-page documents; Chandra processes per-page with stronger layout/table/form structure and a broader language benchmark.","dir":"out","confidence":0.55,"name":"Unlimited-OCR"},{"slug":"mineru","why":"Overlapping document-to-markdown job at different layers: MinerU is a full parsing pipeline (layout analysis + OCR + export) you deploy as tooling; Chandra is the single end-to-end OCR model you'd slot into such a pipeline.","dir":"out","confidence":0.55,"name":"MinerU"},{"slug":"chunkr","why":"Overlapping document-intelligence job at different layers: chunkr is the deployed pipeline API (layout + OCR + chunking); Chandra is the end-to-end OCR model you would slot into such a pipeline.","dir":"in","confidence":0.55,"name":"chunkr"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/datalab-to/chandra"},"chroma":{"name":"Chroma","owner":"chroma-core","slug":"chroma","stars":28858,"avatar":"https://avatars.githubusercontent.com/u/105881770?v=4&s=96","forks":2399,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["rag","memory"],"summary":"Open-source embedding database for building AI apps with retrieval.","edges":[],"added":"2026-06-22T23:27:20.000Z","linkCount":8,"related":{"complements":[{"slug":"xberg","why":"Natural RAG pairing: Xberg does the front half (parse → clean text → syntax-aware chunks → local or hosted embeddings) and Chroma stores and retrieves those vectors. Xberg feeds the index; Chroma serves the queries.","dir":"in","confidence":0.8,"name":"xberg"},{"slug":"chunkr","why":"Adjacent RAG pipeline stages: chunkr turns PDFs into RAG-ready semantic chunks; Chroma is the embedding store those chunks land in for retrieval.","dir":"in","confidence":0.55,"name":"chunkr"},{"slug":"olmocr","why":"olmocr's cleaned text is exactly what you embed and store for retrieval — feed its output into Chroma as the vector store behind a RAG app. Looser than the LlamaIndex pairing since any embedder/DB works, but Chroma is the natural open-source landing spot.","dir":"in","confidence":0.45,"name":"olmocr"}],"alternative":[{"slug":"zvec","why":"Same job — the embedding store under a RAG pipeline. Chroma is the Python-native developer default; Zvec is the in-process C++ engine betting on raw speed, DiskANN memory economics and built-in hybrid retrieval.","dir":"in","confidence":0.85,"name":"zvec"},{"slug":"omnigraph","why":"Same slot in the stack — the store your AI app's retrieval hits — opposite ends of the spectrum: Chroma is a lightweight embedding DB you outgrow; omnigraph fuses graph traversal, ANN and full-text with versioned branching, at the cost of running a declared-as-code server.","dir":"in","confidence":0.65,"name":"omnigraph"},{"slug":"memmachine","why":"Same slot — 'what my agent remembers' — different bets: Chroma is a general embedding store you shape into memory; MemMachine is purpose-built memory with episodic/profile/working tiers, at the cost of running Neo4j + SQL.","dir":"in","confidence":0.55,"name":"MemMachine"},{"slug":"supavec","why":"Different layers of the same job: Chroma gives you the embedding database and you assemble ingestion, chunking and chat around it; Supavec sells the whole assembled slice as one API.","dir":"in","confidence":0.45,"name":"supavec"}],"built_with":[{"slug":"llamaindex","why":"Use Chroma as the vector store behind your LlamaIndex retrievers.","dir":"in","confidence":0.85,"name":"LlamaIndex"}]},"url":"https://stackmap.shipwithai.xyz/repos/chroma-core/chroma"},"chunkr":{"name":"chunkr","owner":"lumina-ai-inc","slug":"chunkr","stars":4044,"image":"https://raw.githubusercontent.com/lumina-ai-inc/chunkr/main/images/logo.svg","avatar":"https://avatars.githubusercontent.com/u/133187075?v=4&s=96","forks":264,"language":"Rust","license":"AGPL-3.0","updated":"3 months ago","topics":["ocr","rag"],"summary":"Document intelligence API in Rust: layout analysis, OCR with bounding boxes, and semantic chunking that turn PDFs, PPTs and Word docs into RAG-ready structured chunks.","curator_note":"The RAG-ingestion specialist: where OCR tools stop at text, Chunkr does layout analysis and SEMANTIC chunking — the chunks arrive respecting document structure, with bounding boxes for citation-grounding. Self-hosted via Docker Compose. Read the split carefully though: the AGPL open-source version runs community models while the paid cloud runs proprietary ones — the README says plainly the accuracy differs; benchmark the OSS tier on YOUR documents before committing. Also ~3 months quiet at review time, and AGPL matters if you embed.","edges":[{"to":"mineru","type":"alternative","why":"Same job — complex documents into LLM-ready data. MinerU is the batteries-included extraction toolkit; Chunkr is an API-shaped service adding semantic chunking and bounding-box citations for RAG pipelines.","confidence":0.75,"status":"approved"},{"to":"olmocr","type":"alternative","why":"Both parse hard documents for AI consumption; olmocr bets on a VLM end-to-end, Chunkr on a layout-analysis pipeline with structured outputs and citations.","confidence":0.6,"status":"approved"},{"to":"llamaindex","type":"complements","why":"Chunkr slots into LlamaIndex ingestion as the parsing stage — hard documents become structured, citable chunks before indexing.","confidence":0.55,"status":"approved"},{"to":"chroma","type":"complements","why":"Adjacent RAG pipeline stages: chunkr turns PDFs into RAG-ready semantic chunks; Chroma is the embedding store those chunks land in for retrieval.","confidence":0.55,"status":"approved"},{"to":"chandra","type":"alternative","why":"Overlapping document-intelligence job at different layers: chunkr is the deployed pipeline API (layout + OCR + chunking); Chandra is the end-to-end OCR model you would slot into such a pipeline.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T15:35:46.259Z","linkCount":6,"related":{"complements":[{"slug":"llamaindex","why":"Chunkr slots into LlamaIndex ingestion as the parsing stage — hard documents become structured, citable chunks before indexing.","dir":"out","confidence":0.55,"name":"LlamaIndex"},{"slug":"chroma","why":"Adjacent RAG pipeline stages: chunkr turns PDFs into RAG-ready semantic chunks; Chroma is the embedding store those chunks land in for retrieval.","dir":"out","confidence":0.55,"name":"Chroma"},{"slug":"commonforms","why":"chunkr parses PDFs into RAG-ready structured chunks for reading pipelines; commonforms covers the write side of the same document stack — emitting an interactive fillable PDF instead of extracted text.","dir":"in","confidence":0.5,"name":"commonforms"}],"alternative":[{"slug":"mineru","why":"Same job — complex documents into LLM-ready data. MinerU is the batteries-included extraction toolkit; Chunkr is an API-shaped service adding semantic chunking and bounding-box citations for RAG pipelines.","dir":"out","confidence":0.75,"name":"MinerU"},{"slug":"olmocr","why":"Both parse hard documents for AI consumption; olmocr bets on a VLM end-to-end, Chunkr on a layout-analysis pipeline with structured outputs and citations.","dir":"out","confidence":0.6,"name":"olmocr"},{"slug":"chandra","why":"Overlapping document-intelligence job at different layers: chunkr is the deployed pipeline API (layout + OCR + chunking); Chandra is the end-to-end OCR model you would slot into such a pipeline.","dir":"out","confidence":0.55,"name":"chandra"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/lumina-ai-inc/chunkr"},"claude-ads":{"name":"claude-ads","owner":"AgriciDaniel","slug":"claude-ads","stars":7311,"image":"https://raw.githubusercontent.com/AgriciDaniel/claude-ads/main/assets/banner.svg","avatar":"https://avatars.githubusercontent.com/u/223140489?v=4&s=96","forks":1076,"language":"Python","license":"MIT","updated":"11 days ago","topics":["skills"],"summary":"Paid-media operations skill for Claude Code across 12 ad platforms — source-grounded audits with deterministic scoring, versioned JSON reports, and account changes gated behind approval and rollback.","curator_note":"The engineering discipline is the story here, rare in marketing tooling: read-only by default, and a live account change requires a tested capability, explicit IDs, a before/after diff with blast radius, owner approval within ceilings, an idempotency key, rollback and post-verification — miss one gate, no write. Scoring is honest too: unknown controls reduce evidence coverage instead of inflating health, and a failed platform makes the run explicitly partial. Twelve platforms from Google/Meta to Amazon/Reddit, each with its own skill, worker and capability manifest. For agencies and in-house performance teams who live in Claude Code. NOT plug-and-play: you bring real platform credentials, Python 3.11/3.12, and WeasyPrint for PDFs. Dual-homed with a paid community mirror — the public MIT release is the real, complete thing.","edges":[],"status":"approved","added":"2026-07-16T15:39:30.763Z","linkCount":2,"related":{"complements":[{"slug":"ai-marketing-claude","why":"Two halves of Claude-Code marketing work: ai-marketing-claude generates the strategy layer — audits, copy, sequences, client reports — while claude-ads operates the actual ad accounts across 12 platforms with gated, rollback-safe changes.","dir":"in","confidence":0.55,"name":"ai-marketing-claude"}],"alternative":[{"slug":"aaron-marketing-skills","why":"Overlapping on the paid-ads slice: claude-ads goes deep — 12 platforms, account reads and capability-gated writes; aaron's ads discipline stays at the strategy/creative level with a ROAS auditor gate and no account plumbing.","dir":"in","confidence":0.5,"name":"aaron-marketing-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/AgriciDaniel/claude-ads"},"claude-bug-bounty":{"name":"claude-bug-bounty","owner":"shuvonsec","slug":"claude-bug-bounty","stars":4021,"image":"https://raw.githubusercontent.com/shuvonsec/claude-bug-bounty/main/logo.png","avatar":"https://avatars.githubusercontent.com/u/83355567?v=4&s=96","forks":717,"language":"Python","license":"MIT","updated":"2 days ago","topics":["security"],"summary":"Autonomous bug-bounty agent for the terminal — recon, 20 vuln classes, a validation gate and submission-ready HackerOne/Bugcrowd reports. Runs as a Claude Code plugin or standalone on free providers.","curator_note":"For solo bounty hunters who want an agent to run recon→hunt→validate→report end to end: the strict validation gate before a finding becomes a report is the useful part (cuts false-positive noise reviewers hate), and standalone mode on Ollama means no subscription. NOT a replacement for skilled manual testing on serious targets, and point it ONLY at assets you're authorized to test — autonomous scanning of others' systems is illegal. Report quality still needs a human pass before submission.","edges":[{"to":"ollama","type":"complements","why":"Standalone mode runs on free local providers — `bughunter setup` offers Ollama as the offline, no-subscription backend so the whole recon/hunt loop works without a paid API.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-07T22:49:33.799Z","linkCount":1,"related":{"complements":[{"slug":"ollama","why":"Standalone mode runs on free local providers — `bughunter setup` offers Ollama as the offline, no-subscription backend so the whole recon/hunt loop works without a paid API.","dir":"out","confidence":0.45,"name":"Ollama"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/shuvonsec/claude-bug-bounty"},"claude-code-karma":{"name":"claude-code-karma","owner":"JayantDevkar","slug":"claude-code-karma","stars":245,"image":"https://raw.githubusercontent.com/JayantDevkar/claude-code-karma/main/docs/screenshots/banner.png","avatar":"https://avatars.githubusercontent.com/u/55962509?v=4&s=96","forks":29,"language":"Python","license":"Apache-2.0","updated":"3 days ago","topics":["coding"],"summary":"Local-first dashboard over ~/.claude — sessions, timelines, per-session costs, tool/agent/skill/plugin analytics, live activity via hooks, and ticket linking. No cloud, no telemetry.","curator_note":"Your ~/.claude directory is a goldmine Claude Code never shows you — Karma turns the JSONL into browsable sessions with full chronological timelines, cost and cache-hit breakdowns, subagent trees, and a surprisingly complete inventory of every plugin, skill, and hook you've installed with usage stats. Read-only ticket linking (Linear/Jira/GitHub) is the team-ready touch. Know the limits: Claude Code prunes session data after ~30 days, so history evaporates unless you keep Karma indexing; it's a two-process dev setup (FastAPI + SvelteKit), not a single binary; and it's Claude-only — no Codex/OpenCode. Pair with ai-token-monitor when all you want is the spend number next to your clock.","edges":[{"to":"ai-token-monitor","type":"alternative","why":"Both read the session logs your coding CLI already writes, at opposite depths: ai-token-monitor is the glanceable menu-bar spend counter with plan-limit alerts; Karma is the full forensic dashboard — timelines, subagents, tools, tickets.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.567Z","linkCount":3,"related":{"complements":[{"slug":"agent-flow","why":"Two halves of Claude Code observability over the same local data: Agent Flow is the live execution graph you watch while a session runs; Karma is the historical dashboard you consult afterwards — costs, timelines, tool and skill analytics.","dir":"in","confidence":0.6,"name":"agent-flow"}],"alternative":[{"slug":"cc-lens","why":"Same job — a local dashboard over ~/.claude JSONL. cc-lens is the one-command npx tool with insights, budgets and team adoption; Karma is the two-process forensic suite with subagent trees, hook/plugin inventories and Linear/Jira/GitHub ticket linking.","dir":"in","confidence":0.7,"name":"cc-lens"},{"slug":"ai-token-monitor","why":"Both read the session logs your coding CLI already writes, at opposite depths: ai-token-monitor is the glanceable menu-bar spend counter with plan-limit alerts; Karma is the full forensic dashboard — timelines, subagents, tools, tickets.","dir":"out","confidence":0.6,"name":"ai-token-monitor"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/JayantDevkar/claude-code-karma"},"claude-reflect":{"name":"claude-reflect","owner":"BayramAnnakov","slug":"claude-reflect","stars":1270,"image":"https://raw.githubusercontent.com/BayramAnnakov/claude-reflect/main/assets/reflect-demo.jpg","avatar":"https://avatars.githubusercontent.com/u/774131?v=4&s=96","forks":109,"language":"Python","license":"MIT","updated":"4 months ago","topics":["coding","memory"],"summary":"Claude Code plugin that learns from your corrections — hooks capture them in-session, /reflect syncs approved learnings to CLAUDE.md/AGENTS.md, /reflect-skills mines history into reusable commands.","curator_note":"Install claude-reflect the third time you catch yourself typing the same correction into Claude Code. It's the pragmatic take on agent memory: no vector DB, no service — hooks queue corrections, you review, markdown files get smarter, and the AGENTS.md sync means Codex/Cursor/Aider benefit too. The /reflect-skills pattern-mining is the sleeper feature: 15 similar requests become one command. When NOT: if you expect actual memory infrastructure (semantic recall, knowledge graphs) — this is disciplined note-taking with AI triage, personal-scale by design. Everything lands via human review, which is a feature, not friction.","edges":[{"to":"core","type":"alternative","why":"Both chase memory that compounds: core builds a full personal-AI OS with a persistent memory graph; claude-reflect does one narrow slice — your coding agent's lessons — with zero infrastructure.","confidence":0.5,"status":"approved"},{"to":"agents-cli","type":"complements","why":"Both extend a coding assistant via skills; reflect's correction-routing writes learnings back into whatever skill files you run — install it alongside any skill pack and the pack improves with use.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":11,"related":{"complements":[{"slug":"skillkit","why":"claude-reflect mines your sessions into reusable skills/commands for Claude Code; skillkit is the distribution layer that packages and translates them to 45 other agents. Minor overlap: skillkit also persists session learnings, but capture is claude-reflect's whole job.","dir":"in","confidence":0.6,"name":"skillkit"},{"slug":"squid","why":"Both ride Claude Code: squid runs the pipeline, claude-reflect captures your mid-run corrections and routes them back into skill and AGENTS.md files — the factory learns between features.","dir":"in","confidence":0.6,"name":"squid"},{"slug":"agents-cli","why":"Both extend a coding assistant via skills; reflect's correction-routing writes learnings back into whatever skill files you run — install it alongside any skill pack and the pack improves with use.","dir":"out","confidence":0.55,"name":"agents-cli"},{"slug":"bemyagent","why":"Opposite halves of keeping agent context current: bemyagent scaffolds structured project memory up front and registers rules into AGENTS.md; claude-reflect mines your in-session corrections into those same rule files over time. Watch for both writing to AGENTS.md.","dir":"in","confidence":0.55,"name":"bemyagent"}],"alternative":[{"slug":"pro-workflow","why":"Same core job — corrections that stick across Claude Code sessions. claude-reflect is the focused plugin (corrections → CLAUDE.md); pro-workflow builds a whole SQLite-backed memory-and-hooks platform around the idea.","dir":"in","confidence":0.75,"name":"pro-workflow"},{"slug":"agentic-context-engine","why":"Same learn-from-corrections loop at different scopes: claude-reflect is a Claude Code plugin syncing learnings into CLAUDE.md; ACE is a framework-level engine for any agent you build, with strategies as first-class objects.","dir":"in","confidence":0.6,"name":"agentic-context-engine"},{"slug":"ecc","why":"Both close the learning loop from your corrections: claude-reflect is a focused Claude Code plugin syncing approved learnings to CLAUDE.md; ECC's instinct system does the same continuous-learning job with confidence scoring inside a much larger harness framework.","dir":"in","confidence":0.6,"name":"ECC"},{"slug":"memsearch","why":"Both mine your sessions into reusable assets: claude-reflect captures corrections into CLAUDE.md and /reflect-skills commands for Claude Code; memsearch's skills-from-memory does the same distillation continuously, across four agent platforms.","dir":"in","confidence":0.6,"name":"memsearch"},{"slug":"skillx","why":"Both turn agent experience into reusable knowledge: claude-reflect captures your corrections into CLAUDE.md pragmatically, SkillX distills full trajectories into a transferable skill hierarchy — research-grade.","dir":"in","confidence":0.55,"name":"SkillX"},{"slug":"core","why":"Both chase memory that compounds: core builds a full personal-AI OS with a persistent memory graph; claude-reflect does one narrow slice — your coding agent's lessons — with zero infrastructure.","dir":"out","confidence":0.5,"name":"core"},{"slug":"opencode-mem","why":"Both close the cross-session learning loop for coding agents: claude-reflect distills your corrections into CLAUDE.md via /reflect; opencode-mem auto-captures work into a searchable vector store and injects relevant memories per session.","dir":"in","confidence":0.5,"name":"opencode-mem"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/BayramAnnakov/claude-reflect"},"claude-secrets":{"name":"claude-secrets","owner":"vaultry","slug":"claude-secrets","stars":6,"image":"https://raw.githubusercontent.com/vaultry/claude-secrets/main/docs/claude-input.gif?v=2","avatar":"https://avatars.githubusercontent.com/u/276987870?v=4&s=96","forks":2,"language":"JavaScript","license":"NOASSERTION","updated":"3 months ago","topics":["security"],"summary":"Encrypted secrets store for Claude Code: macOS-Keychain-backed vault with MCP server, CLI, commit-safe secret:// .env placeholders and a native input dialog.","curator_note":"Fixes a real leak: tokens pasted into chat live forever in transcripts and API logs — here they go dialog-to-Keychain-vault without the model ever seeing them, gated per-project by an allowlist. macOS-only, source-available (not OSI), 6 stars: adopt as a personal utility, not team infrastructure.","edges":[],"status":"approved","added":"2026-07-17T14:03:22.444Z","linkCount":0,"related":{"complements":[],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/vaultry/claude-secrets"},"cliproxyapi":{"name":"CLIProxyAPI","owner":"router-for-me","slug":"cliproxyapi","stars":44405,"image":"https://raw.githubusercontent.com/router-for-me/cliproxyapi/main/assets/logo/kimi.svg","avatar":"https://avatars.githubusercontent.com/u/233033915?v=4&s=96","forks":6957,"language":"Go","license":"MIT","updated":"yesterday","topics":["gateway"],"summary":"Turns your coding-CLI subscriptions (Claude Code, Codex, Antigravity, Kimi, Grok) into a local OpenAI/Gemini/Claude-compatible API — multi-account rotation in one Go proxy.","curator_note":"The inverse of a router: 9router points CLIs at providers, CLIProxyAPI exposes your subscription OAuth accounts AS a provider — any SDK can spend your Claude Code/Codex/Kimi quota. Multi-account balancing built in. ToS gray zone: unofficial client re-use; firewall it, don't resell it.","edges":[{"to":"9router","type":"alternative","why":"Two directions of the same arbitrage: 9router points coding CLIs at 40+ providers; CLIProxyAPI exposes your CLI subscriptions as an OpenAI/Claude/Gemini-compatible endpoint for any client.","confidence":0.6,"status":"approved"},{"to":"lynkr","type":"alternative","why":"Both self-hosted proxies between coding subscriptions and clients: lynkr optimizes (compression, caching, tier-routing); CLIProxyAPI multiplexes accounts and re-exposes them as standard APIs.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.500Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"9router","why":"Two directions of the same arbitrage: 9router points coding CLIs at 40+ providers; CLIProxyAPI exposes your CLI subscriptions as an OpenAI/Claude/Gemini-compatible endpoint for any client.","dir":"out","confidence":0.6,"name":"9router"},{"slug":"lynkr","why":"Both self-hosted proxies between coding subscriptions and clients: lynkr optimizes (compression, caching, tier-routing); CLIProxyAPI multiplexes accounts and re-exposes them as standard APIs.","dir":"out","confidence":0.55,"name":"Lynkr"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/router-for-me/cliproxyapi"},"cocoindex-code":{"name":"cocoindex-code","owner":"cocoindex-io","slug":"cocoindex-code","stars":2551,"image":"https://github.com/user-attachments/assets/d05961b4-0b7b-42ea-834a-59c3c01717ca","avatar":"https://avatars.githubusercontent.com/u/190812870?v=4&s=96","forks":203,"language":"Python","license":"Apache-2.0","updated":"6 days ago","topics":["coding","local"],"summary":"AST-based semantic code search for coding agents: pipx install, zero config, local embeddings out of the box — a CLI/skill/MCP that cuts agent context ~70% vs grepping. Built on CocoIndex.","curator_note":"The lowest-friction entry in the code-context-server category: one pipx install, zero config, and the [full] variant ships local embeddings so there's no API key before first search. AST-aware chunking plus the CocoIndex engine underneath means the index stays fresh incrementally instead of rebuilding. Integrates as a skill or MCP with Claude Code, Codex, Cursor. NOT the deepest graph: it's semantic search, not call-graph analysis — no callers/blast-radius queries (that's TokenSave or codegraph-mcp territory) — and the slim install quietly requires a cloud embedding key while [full] costs ~1GB of torch.","edges":[{"to":"cocoindex","type":"built_with","why":"Built on CocoIndex by the same team — the flagship application of the incremental indexing engine, which keeps the code index fresh by reprocessing only the delta.","confidence":0.95,"status":"approved"},{"to":"tokensave","type":"alternative","why":"Same job — replace agent grepping with a pre-built local index. TokenSave answers structural questions (symbols, callers, impact radius) from a semantic graph; cocoindex-code answers 'where is the code that does X' via AST-chunked embedding search.","confidence":0.8,"status":"approved"},{"to":"codegraph-mcp","type":"alternative","why":"Both serve codebase context to agents over MCP. codegraph-mcp builds a cross-language knowledge graph with audit logging for on-prem rigor; cocoindex-code trades graph depth for zero-config semantic search that installs in a minute.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-22T18:15:22.263Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"tokensave","why":"Same job — replace agent grepping with a pre-built local index. TokenSave answers structural questions (symbols, callers, impact radius) from a semantic graph; cocoindex-code answers 'where is the code that does X' via AST-chunked embedding search.","dir":"out","confidence":0.8,"name":"tokensave"},{"slug":"codegraph-mcp","why":"Both serve codebase context to agents over MCP. codegraph-mcp builds a cross-language knowledge graph with audit logging for on-prem rigor; cocoindex-code trades graph depth for zero-config semantic search that installs in a minute.","dir":"out","confidence":0.6,"name":"codegraph-mcp"},{"slug":"codebase-memory-mcp","why":"The two query styles of agent code context: cocoindex-code embeds AST chunks for semantic 'find the code that does X'; codebase-memory graphs symbols for structural 'who calls this and what breaks'. Different questions — many stacks want both.","dir":"in","confidence":0.6,"name":"codebase-memory-mcp"}],"built_with":[{"slug":"cocoindex","why":"Built on CocoIndex by the same team — the flagship application of the incremental indexing engine, which keeps the code index fresh by reprocessing only the delta.","dir":"out","confidence":0.95,"name":"cocoindex"}]},"url":"https://stackmap.shipwithai.xyz/repos/cocoindex-io/cocoindex-code"},"cocoindex":{"name":"cocoindex","owner":"cocoindex-io","slug":"cocoindex","stars":11008,"image":"https://cocoindex.io/blobs/github/homepage/enterprise-hero-light.svg","avatar":"https://avatars.githubusercontent.com/u/190812870?v=4&s=96","forks":848,"language":"Rust","license":"Apache-2.0","updated":"yesterday","topics":["rag"],"summary":"Rust-core incremental indexing engine: declare Target = F(Source) in Python and it keeps vector/graph/relational targets fresh forever, reprocessing only the delta — with per-row lineage.","curator_note":"The mental model sells it — 'React for data engineering': you declare what the index should contain, and the engine reconciles it against source changes forever, re-running only affected rows (cached by hash of input AND code, so editing your transform also invalidates precisely). That's the honest answer to stale agent context: sub-second freshness at a fraction of the re-embedding bill, with every vector traceable to its source byte. Sources span code, PDFs, Slack, audio; targets span pgvector, LanceDB, Neo4j, Kafka. The flagship application is cocoindex-code, an AST-aware incremental code-index MCP for coding agents. Use it when your corpus changes constantly and batch re-indexing is bleeding you; overkill for a static document pile — any one-shot RAG ingester handles that.","edges":[{"to":"llamaindex","type":"alternative","why":"Both connect LLMs to your data, at different layers: LlamaIndex is the retrieval framework (loaders, indexes, query engines) typically run as batch ingestion; CocoIndex is the incremental sync engine that keeps whatever store you target continuously fresh, delta-only, with lineage.","confidence":0.6,"status":"approved"},{"to":"tokensave","type":"alternative","why":"Head-to-head on the code-index-for-agents job: tokensave ships a pre-indexed libSQL+FTS5 semantic graph agents query over MCP; CocoIndex's flagship cocoindex-code does the same over MCP but incrementally re-indexed on every commit via the Δ engine.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.707Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"llamaindex","why":"Both connect LLMs to your data, at different layers: LlamaIndex is the retrieval framework (loaders, indexes, query engines) typically run as batch ingestion; CocoIndex is the incremental sync engine that keeps whatever store you target continuously fresh, delta-only, with lineage.","dir":"out","confidence":0.6,"name":"LlamaIndex"},{"slug":"tokensave","why":"Head-to-head on the code-index-for-agents job: tokensave ships a pre-indexed libSQL+FTS5 semantic graph agents query over MCP; CocoIndex's flagship cocoindex-code does the same over MCP but incrementally re-indexed on every commit via the Δ engine.","dir":"out","confidence":0.55,"name":"tokensave"}],"built_with":[{"slug":"cocoindex-code","why":"Built on CocoIndex by the same team — the flagship application of the incremental indexing engine, which keeps the code index fresh by reprocessing only the delta.","dir":"in","confidence":0.95,"name":"cocoindex-code"}]},"url":"https://stackmap.shipwithai.xyz/repos/cocoindex-io/cocoindex"},"code-review-graph":{"name":"code-review-graph","owner":"tirth8205","slug":"code-review-graph","stars":25590,"image":"https://raw.githubusercontent.com/tirth8205/code-review-graph/main/diagrams/diagram1_before_vs_after.png","avatar":"https://avatars.githubusercontent.com/u/68604113?v=4&s=96","forks":2406,"language":"Python","license":"MIT","updated":"3 days ago","topics":["coding"],"summary":"Local-first code intelligence graph for AI coding tools: Tree-sitter AST graph + blast-radius analysis served over MCP, so agents read ~82x fewer tokens per review question.","curator_note":"The rare benchmark-honest repo: it tells you the 528x number is the best case and its recall metric is circular. Install once, it configures 14 platforms (Codex, Claude Code, Cursor...). Reach for it on monorepos where review context is the token sink; skip on small repos where a grep costs less than the graph's own metadata.","edges":[{"to":"codegraph-mcp","type":"alternative","why":"Both pre-index a semantic code graph and serve it to agents over MCP; code-review-graph adds blast-radius/risk scoring and a CI review action, codegraph-mcp stays a leaner symbol/caller index.","confidence":0.7,"status":"approved"},{"to":"understand-anything","type":"alternative","why":"Code knowledge graphs for opposite consumers: understand-anything renders an explorable dashboard to teach humans the architecture; code-review-graph feeds minimal context to agents.","confidence":0.55,"status":"approved"},{"to":"repowise","type":"alternative","why":"Both build deterministic code intelligence over MCP; repowise scores health and plans refactors, code-review-graph computes the minimal read-set for a change.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:41.991Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"codegraph-mcp","why":"Both pre-index a semantic code graph and serve it to agents over MCP; code-review-graph adds blast-radius/risk scoring and a CI review action, codegraph-mcp stays a leaner symbol/caller index.","dir":"out","confidence":0.7,"name":"codegraph-mcp"},{"slug":"codeflow","why":"Both build a dependency graph with blast-radius analysis; codeflow renders it visually for humans with zero setup, code-review-graph serves a tree-sitter-accurate graph to AI agents over MCP.","dir":"in","confidence":0.7,"name":"codeflow"},{"slug":"codebase-memory-mcp","why":"Both build tree-sitter AST graphs served over MCP with blast-radius answers. code-review-graph is tuned for the review workflow (82x fewer tokens per review question); codebase-memory is the general-purpose engine with LSP-grade types.","dir":"in","confidence":0.65,"name":"codebase-memory-mcp"},{"slug":"understand-anything","why":"Code knowledge graphs for opposite consumers: understand-anything renders an explorable dashboard to teach humans the architecture; code-review-graph feeds minimal context to agents.","dir":"out","confidence":0.55,"name":"Understand-Anything"},{"slug":"repowise","why":"Both build deterministic code intelligence over MCP; repowise scores health and plans refactors, code-review-graph computes the minimal read-set for a change.","dir":"out","confidence":0.5,"name":"repowise"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tirth8205/code-review-graph"},"codebase-memory-mcp":{"name":"codebase-memory-mcp","owner":"DeusData","slug":"codebase-memory-mcp","stars":34159,"image":"https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/docs/graph-ui-screenshot.png","avatar":"https://avatars.githubusercontent.com/u/81762164?v=4&s=96","forks":2621,"language":"C","license":"MIT","updated":"2 days ago","topics":["coding","local"],"summary":"Code intelligence MCP in pure C: tree-sitter knowledge graph over 158 languages, average repo indexed in milliseconds, sub-ms queries, 10x fewer tokens. Single static binary, zero deps.","curator_note":"The performance ceiling of the code-context-server category: pure C, single static binary, the Linux kernel indexed in 3 minutes, structural queries under a millisecond — with a peer-reviewed preprint (83% answer quality, 10x fewer tokens across 31 repos) instead of vibes. Hybrid LSP adds real type resolution for the 12 languages that matter most, and 43 client surfaces means it plugs into whatever agent you run. NOT semantic search — it answers structural questions (call chains, routes, blast radius), not 'where's the code that does X'; pair it with an embedding tool for that. And note its own disclosure: it writes to your agent config files by design — audit posture is unusually good (SLSA 3, OpenSSF, per-release VirusTotal), use it.","edges":[{"to":"tokensave","type":"alternative","why":"The same exact job — pre-indexed structural code intelligence over MCP so agents stop grepping. TokenSave is a libSQL semantic graph across 50+ languages; codebase-memory is a zero-dependency C binary betting everything on speed: 158 languages, sub-ms.","confidence":0.8,"status":"approved"},{"to":"code-review-graph","type":"alternative","why":"Both build tree-sitter AST graphs served over MCP with blast-radius answers. code-review-graph is tuned for the review workflow (82x fewer tokens per review question); codebase-memory is the general-purpose engine with LSP-grade types.","confidence":0.65,"status":"approved"},{"to":"cocoindex-code","type":"alternative","why":"The two query styles of agent code context: cocoindex-code embeds AST chunks for semantic 'find the code that does X'; codebase-memory graphs symbols for structural 'who calls this and what breaks'. Different questions — many stacks want both.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T14:13:33.573Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"tokensave","why":"The same exact job — pre-indexed structural code intelligence over MCP so agents stop grepping. TokenSave is a libSQL semantic graph across 50+ languages; codebase-memory is a zero-dependency C binary betting everything on speed: 158 languages, sub-ms.","dir":"out","confidence":0.8,"name":"tokensave"},{"slug":"code-review-graph","why":"Both build tree-sitter AST graphs served over MCP with blast-radius answers. code-review-graph is tuned for the review workflow (82x fewer tokens per review question); codebase-memory is the general-purpose engine with LSP-grade types.","dir":"out","confidence":0.65,"name":"code-review-graph"},{"slug":"cocoindex-code","why":"The two query styles of agent code context: cocoindex-code embeds AST chunks for semantic 'find the code that does X'; codebase-memory graphs symbols for structural 'who calls this and what breaks'. Different questions — many stacks want both.","dir":"out","confidence":0.6,"name":"cocoindex-code"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/DeusData/codebase-memory-mcp"},"codeflow":{"name":"codeflow","owner":"braedonsaunders","slug":"codeflow","stars":4665,"image":"https://raw.githubusercontent.com/braedonsaunders/codeflow/main/screenshot.png","avatar":"https://avatars.githubusercontent.com/u/28276447?v=4&s=96","forks":719,"language":"HTML","license":null,"updated":"10 days ago","topics":["coding"],"summary":"Paste a GitHub URL or drop a local folder → interactive architecture map in the browser: dependency graph, blast radius, health grade, security scan. Single index.html, zero install.","curator_note":"The fastest 'what am I looking at?' tool for an unfamiliar repo: zero install, code never leaves the browser, and blast-radius answers 'what breaks if I touch this?' before you refactor. The JSON export and README-card GitHub Action make it more than a toy. NOT compiler-grade: dependency edges come from lightweight in-browser parsing across 30+ languages, so trust it for orientation, not for exhaustive call-graph truth — agents needing tree-sitter-accurate graphs should use code-review-graph or codegraph-mcp. Big repos also eat GitHub API rate limits fast without a token.","edges":[{"to":"code-review-graph","type":"alternative","why":"Both build a dependency graph with blast-radius analysis; codeflow renders it visually for humans with zero setup, code-review-graph serves a tree-sitter-accurate graph to AI agents over MCP.","confidence":0.7,"status":"approved"},{"to":"codegraph-mcp","type":"alternative","why":"Same 'understand the codebase structure' job: codegraph-mcp is an on-prem symbol/call-edge knowledge graph consumed by agents; codeflow is a zero-install visual map consumed by people.","confidence":0.6,"status":"approved"},{"to":"repowise","type":"alternative","why":"Both grade codebase health from graph structure. Repowise gives deterministic, defect-calibrated scores and agent-executable refactoring plans; codeflow gives an instant in-browser A–F and heatmaps.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-19T15:02:10.447Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"code-review-graph","why":"Both build a dependency graph with blast-radius analysis; codeflow renders it visually for humans with zero setup, code-review-graph serves a tree-sitter-accurate graph to AI agents over MCP.","dir":"out","confidence":0.7,"name":"code-review-graph"},{"slug":"codegraph-mcp","why":"Same 'understand the codebase structure' job: codegraph-mcp is an on-prem symbol/call-edge knowledge graph consumed by agents; codeflow is a zero-install visual map consumed by people.","dir":"out","confidence":0.6,"name":"codegraph-mcp"},{"slug":"repowise","why":"Both grade codebase health from graph structure. Repowise gives deterministic, defect-calibrated scores and agent-executable refactoring plans; codeflow gives an instant in-browser A–F and heatmaps.","dir":"out","confidence":0.55,"name":"repowise"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/braedonsaunders/codeflow"},"codegraph-mcp":{"name":"codegraph-mcp","owner":"cognis-digital","slug":"codegraph-mcp","stars":5,"image":"https://raw.githubusercontent.com/cognis-digital/codegraph-mcp/master/media/walkthrough-thumb.png","avatar":"https://avatars.githubusercontent.com/u/215970675?v=4&s=96","forks":0,"language":"Python","license":"NOASSERTION","updated":"20 days ago","topics":["coding","rag"],"summary":"No-train, on-prem code knowledge graph served to AI agents over MCP — symbols, call edges, cross-language links and blast-radius queries, with a hash-chained audit log of every read.","curator_note":"Niche but real: if agents must understand code that cannot leave your building AND compliance asks 'what exactly did the agent read', the tamper-evident audit chain is the only game in town; cross-language call edges (TS fetch → Go/Python handler) catch what one-file context misses. NOT for most teams yet: 5 stars, single vendor, and — deal-breaker until fixed — no clear open-source license (NOASSERTION on GitHub). Treat it as an evaluation candidate, not a dependency.","edges":[{"to":"omnigraph","type":"complements","why":"Two halves of a graph-context stack for agents: codegraph-mcp serves code topology (symbols, calls, blast radius) over MCP; omnigraph holds the org's general knowledge graph. Conceptual pairing — no packaged integration.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-07T15:50:11.276Z","linkCount":7,"related":{"complements":[{"slug":"omnigraph","why":"Two halves of a graph-context stack for agents: codegraph-mcp serves code topology (symbols, calls, blast radius) over MCP; omnigraph holds the org's general knowledge graph. Conceptual pairing — no packaged integration.","dir":"out","confidence":0.45,"name":"omnigraph"}],"alternative":[{"slug":"tokensave","why":"Same job — a code knowledge graph served to agents over MCP: codegraph-mcp bets on compliance (hash-chained audit of every read, cross-language HTTP edges); tokensave bets on breadth and token savings (50+ langs, 80+ tools, 12 agent integrations).","dir":"in","confidence":0.8,"name":"tokensave"},{"slug":"code-review-graph","why":"Both pre-index a semantic code graph and serve it to agents over MCP; code-review-graph adds blast-radius/risk scoring and a CI review action, codegraph-mcp stays a leaner symbol/caller index.","dir":"in","confidence":0.7,"name":"code-review-graph"},{"slug":"cocoindex-code","why":"Both serve codebase context to agents over MCP. codegraph-mcp builds a cross-language knowledge graph with audit logging for on-prem rigor; cocoindex-code trades graph depth for zero-config semantic search that installs in a minute.","dir":"in","confidence":0.6,"name":"cocoindex-code"},{"slug":"codeflow","why":"Same 'understand the codebase structure' job: codegraph-mcp is an on-prem symbol/call-edge knowledge graph consumed by agents; codeflow is a zero-install visual map consumed by people.","dir":"in","confidence":0.6,"name":"codeflow"},{"slug":"repowise","why":"Both serve a code knowledge graph to agents over MCP; codegraph-mcp goes deep on cross-language symbol/call/blast-radius queries, repowise trades some of that depth for defect-risk scores and refactoring plans.","dir":"in","confidence":0.6,"name":"repowise"},{"slug":"understand-anything","why":"Both build knowledge graphs of your codebase, for opposite consumers: codegraph-mcp serves symbols, call edges and blast-radius queries to AI agents over MCP with an audit chain; Understand Anything renders an interactive dashboard for human comprehension and onboarding.","dir":"in","confidence":0.6,"name":"Understand-Anything"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/cognis-digital/codegraph-mcp"},"codenomad":{"name":"CodeNomad","owner":"NeuralNomadsAI","slug":"codenomad","stars":2383,"image":"https://raw.githubusercontent.com/NeuralNomadsAI/codenomad/dev/docs/screenshots/newSession.png","avatar":"https://avatars.githubusercontent.com/u/244000212?v=4&s=96","forks":164,"language":"TypeScript","license":"MIT","updated":"3 days ago","topics":["coding"],"summary":"Desktop cockpit for OpenCode: multi-instance sessions, git worktrees, remote browser access, voice input and a command palette — a workspace for living in AI coding sessions.","curator_note":"For developers who run OpenCode all day and have outgrown the terminal: parallel sessions across projects in one window, git-worktree awareness for agent branches, a password-protected server mode for driving sessions from any browser, and quality-of-life extras (voice input, SideCars for embedding local web tools). NOT useful if OpenCode isn't your driver — it's a cockpit for that engine specifically, not a general agent UI; and it's young (2k stars), so expect rough edges alongside the fast release cadence.","edges":[{"to":"alook","type":"alternative","why":"Both put a management surface over local coding agents — alook turns them into an always-on multi-agent 'AI company', CodeNomad gives one developer a rich single-cockpit for parallel OpenCode sessions.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:51.991Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"codexmate","why":"Both are local cockpits for living in AI coding sessions. CodeNomad is a desktop app focused deeply on OpenCode (worktrees, voice, command palette); Codex Mate is a web dashboard spanning Codex, Claude Code, OpenCode and OpenClaw with provider management and session search.","dir":"in","confidence":0.7,"name":"codexmate"},{"slug":"omnigent","why":"Both are cockpits for living in AI coding sessions with remote access; CodeNomad is OpenCode-only, Omnigent spans many harnesses.","dir":"in","confidence":0.65,"name":"omnigent"},{"slug":"helmor","why":"Both local desktop workspaces for coding-agent sessions: codenomad is an OpenCode-deep cockpit; Helmor is agent-agnostic with per-repo workspaces and dispatchable actions.","dir":"in","confidence":0.6,"name":"helmor"},{"slug":"tailclaude","why":"Same itch — escaping the terminal for your coding agent. CodeNomad is a rich desktop/server cockpit for OpenCode; TailClaude is a zero-install browser UI for Claude Code over Tailscale.","dir":"in","confidence":0.6,"name":"tailclaude"},{"slug":"alook","why":"Both put a management surface over local coding agents — alook turns them into an always-on multi-agent 'AI company', CodeNomad gives one developer a rich single-cockpit for parallel OpenCode sessions.","dir":"out","confidence":0.55,"name":"alook"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/NeuralNomadsAI/codenomad"},"codexguide":{"name":"CodexGuide","owner":"freestylefly","slug":"codexguide","stars":2818,"image":"https://cdn.canghecode.com/codexguide/assets/banner.svg","avatar":"https://avatars.githubusercontent.com/u/43960064?v=4&s=96","forks":273,"language":"PowerShell","license":"MIT","updated":"2 days ago","topics":["coding"],"summary":"Community-maintained Chinese practice guide to OpenAI Codex — learning paths, CLI/App/Cloud/IDE setup, AGENTS.md templates, sandbox/approval safety and team playbooks, published at codexguide.ai.","curator_note":"Use it to onboard people — especially Chinese-speaking teams — onto Codex without the trial-and-error phase: real task flows (PPT, Obsidian, CI fixes, Feishu/Notion), AGENTS.md rule templates, sandbox and approval boundaries, and a team playbook for turning one successful run into reusable process. NOT a tool: nothing to install, it's a VuePress knowledge base. Content is Chinese-first (an English README mirror exists), the README carries a heavy sponsor block, and anything time-sensitive (pricing, availability) should be re-checked against OpenAI's own docs — the guide itself says so.","edges":[],"status":"approved","added":"2026-07-14T17:35:30.388Z","linkCount":0,"related":{"complements":[],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/freestylefly/codexguide"},"codexmate":{"name":"codexmate","owner":"SakuraByteCore","slug":"codexmate","stars":327,"image":"https://raw.githubusercontent.com/SakuraByteCore/codexmate/main/site/.vitepress/public/images/logo.png","avatar":"https://avatars.githubusercontent.com/u/269487193?v=4&s=96","forks":33,"language":"JavaScript","license":"Apache-2.0","updated":"4 days ago","topics":["coding"],"summary":"Local-first CLI + web dashboard for your coding agents — switch providers, browse sessions across Codex/Claude Code/Gemini CLI, share skills, queue tasks, and bridge Codex/Claude to any API.","curator_note":"Use it when you juggle several local agent CLIs and are tired of each one's config/session/skills silo: one web UI to switch providers (with health probes and bulk cleanup of dead configs), search and export sessions across four agents, and edit global/project CLAUDE.md and AGENTS.md with shared presets. The bridges are the sleeper feature — Codex's Responses API normalized to OpenAI-compatible (with official-looking fingerprint headers), and Claude Code pointed at any Chat Completions provider or Ollama. NOT an autonomous orchestrator despite the DAG task queue — for issue-driven unattended runs use contrabass. Early-stage, and it mutates ~/.codex and ~/.claude configs, so keep backups.","edges":[{"to":"codenomad","type":"alternative","why":"Both are local cockpits for living in AI coding sessions. CodeNomad is a desktop app focused deeply on OpenCode (worktrees, voice, command palette); Codex Mate is a web dashboard spanning Codex, Claude Code, OpenCode and OpenClaw with provider management and session search.","confidence":0.7,"status":"approved"},{"to":"9router","type":"alternative","why":"Overlapping job of pointing coding CLIs at arbitrary providers: 9router is a dedicated self-hosted gateway with auto-fallback across 40+ providers; Codex Mate does it via built-in Codex/Claude protocol bridges as one feature of its control plane.","confidence":0.55,"status":"approved"},{"to":"contrabass","type":"complements","why":"Same local-agent fleet, different modes: Codex Mate is the interactive dashboard for providers, sessions and skills; Contrabass takes over for unattended issue-driven runs.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T17:35:30.502Z","linkCount":3,"related":{"complements":[{"slug":"contrabass","why":"Same local-agent fleet, different modes: Codex Mate is the interactive dashboard for providers, sessions and skills; Contrabass takes over for unattended issue-driven runs.","dir":"out","confidence":0.5,"name":"contrabass"}],"alternative":[{"slug":"codenomad","why":"Both are local cockpits for living in AI coding sessions. CodeNomad is a desktop app focused deeply on OpenCode (worktrees, voice, command palette); Codex Mate is a web dashboard spanning Codex, Claude Code, OpenCode and OpenClaw with provider management and session search.","dir":"out","confidence":0.7,"name":"CodeNomad"},{"slug":"9router","why":"Overlapping job of pointing coding CLIs at arbitrary providers: 9router is a dedicated self-hosted gateway with auto-fallback across 40+ providers; Codex Mate does it via built-in Codex/Claude protocol bridges as one feature of its control plane.","dir":"out","confidence":0.55,"name":"9router"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/SakuraByteCore/codexmate"},"colibri":{"name":"colibri","owner":"JustVugg","slug":"colibri","stars":18073,"image":"https://raw.githubusercontent.com/JustVugg/colibri/main/assets/colibri.svg","avatar":"https://avatars.githubusercontent.com/u/13022503?v=4&s=96","forks":1745,"language":"C","license":"Apache-2.0","updated":"yesterday","topics":["local"],"summary":"Pure-C, zero-dep MoE runtime that runs GLM-5.2 (744B) on a 25GB-RAM consumer box by streaming experts from disk — VRAM/RAM/NVMe as one tiered hierarchy, never touching precision.","curator_note":"Use it when you want a frontier-scale open MoE (GLM-5.2, 744B) on hardware that can't hold it, and you care about fidelity: forward pass is token-exact against the transformers oracle, and placement only ever changes speed, never semantics. The learning cache pins the experts YOUR workload actually routes, so it genuinely gets faster with use. NOT a general runner — it's a one-model engine; for everyday multi-model local use pick ollama. And expect 0.05–6.8 tok/s depending on tier residency: this is patience-ware for correctness fanatics, not latency-sensitive serving.","edges":[{"to":"airllm","type":"alternative","why":"Same job — giant models on tiny hardware — opposite technique: AirLLM streams dense layers through a 4GB GPU from Python/HF, colibrì streams MoE experts from NVMe in pure C with a workload-learning pin cache.","confidence":0.85,"status":"approved"},{"to":"ollama","type":"alternative","why":"Both run open models locally. Ollama is the multi-model daily driver for models that fit; colibrì is a single-model specialist that makes a 744B MoE fit where nothing else will.","confidence":0.7,"status":"approved"},{"to":"mesh-llm","type":"alternative","why":"Both attack 'model bigger than your box': mesh-llm scales OUT by pooling GPUs across machines into one endpoint, colibrì scales DOWN by deepening one machine's memory hierarchy to disk.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-19T15:02:10.305Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"airllm","why":"Same job — giant models on tiny hardware — opposite technique: AirLLM streams dense layers through a 4GB GPU from Python/HF, colibrì streams MoE experts from NVMe in pure C with a workload-learning pin cache.","dir":"out","confidence":0.85,"name":"airllm"},{"slug":"ollama","why":"Both run open models locally. Ollama is the multi-model daily driver for models that fit; colibrì is a single-model specialist that makes a 744B MoE fit where nothing else will.","dir":"out","confidence":0.7,"name":"Ollama"},{"slug":"mesh-llm","why":"Both attack 'model bigger than your box': mesh-llm scales OUT by pooling GPUs across machines into one endpoint, colibrì scales DOWN by deepening one machine's memory hierarchy to disk.","dir":"out","confidence":0.65,"name":"mesh-llm"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/JustVugg/colibri"},"commonforms":{"name":"commonforms","owner":"jbarrow","slug":"commonforms","stars":1195,"image":"https://raw.githubusercontent.com/jbarrow/commonforms/main/assets/pipeline.png","avatar":"https://avatars.githubusercontent.com/u/942625?v=4&s=96","forks":151,"language":"Python","license":null,"updated":"1 months ago","topics":["ocr","vision"],"summary":"Turns any PDF into a fillable form: FFDNet models detect text, checkbox and signature fields; one CLI command writes the interactive PDF. Paper, dataset and weights all open.","curator_note":"The only open tool for this exact job: point `commonforms in.pdf out.pdf` at a flat or scanned form and get real AcroForm fields back — CPU works, --fast halves runtime. Use it for digitizing form backlogs or as training ground (the CommonForms dataset + FFDNet weights are on HF). NOT an extractor: it detects where fields go, it doesn't label them semantically or read filled-in values — pair with an OCR model for that. Mind the footprint (torch, transformers, ultralytics — install isolated via uv tool/pipx) and the licensing: no license file; author asks non-academic users to reach out.","edges":[{"to":"chandra","type":"complements","why":"Opposite halves of the form pipeline: commonforms detects field locations and makes the PDF fillable, chandra reads forms — extracting labels, checkbox states and handwritten values from filled documents.","confidence":0.65,"status":"approved"},{"to":"chunkr","type":"complements","why":"chunkr parses PDFs into RAG-ready structured chunks for reading pipelines; commonforms covers the write side of the same document stack — emitting an interactive fillable PDF instead of extracted text.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-19T13:40:10.465Z","linkCount":2,"related":{"complements":[{"slug":"chandra","why":"Opposite halves of the form pipeline: commonforms detects field locations and makes the PDF fillable, chandra reads forms — extracting labels, checkbox states and handwritten values from filled documents.","dir":"out","confidence":0.65,"name":"chandra"},{"slug":"chunkr","why":"chunkr parses PDFs into RAG-ready structured chunks for reading pipelines; commonforms covers the write side of the same document stack — emitting an interactive fillable PDF instead of extracted text.","dir":"out","confidence":0.5,"name":"chunkr"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/jbarrow/commonforms"},"contrabass":{"name":"contrabass","owner":"junhoyeo","slug":"contrabass","stars":202,"image":"https://raw.githubusercontent.com/junhoyeo/contrabass/main/.github/assets/contrabass.png","avatar":"https://avatars.githubusercontent.com/u/32605822?v=4&s=96","forks":24,"language":"Go","license":"Apache-2.0","updated":"7 days ago","topics":["coding","orchestration"],"summary":"Terminal-first orchestrator for issue-driven AI coding-agent runs — polls Linear/GitHub, runs Codex/OpenCode in git worktrees with retries and verification. Go/Charm rebuild of OpenAI's Symphony.","curator_note":"Pick it when your work already lives in Linear or GitHub issues and you want local coding agents (Codex, OpenCode, oh-my-*) burning down the backlog unattended — worktree per issue, branch-advance verification so 'success' means commits actually landed, stall detection, deterministic retries, and a Bubble Tea TUI plus embedded web dashboard for visibility. Skip it for single interactive sessions (just run the agent CLI) or if you want cloud-hosted execution — pullfrog is the GitHub-Actions version of this job. Young Symphony reimplementation: the workflow parser accepts more fields than the runtime consumes, and the default team mode needs tmux.","edges":[{"to":"squad","type":"alternative","why":"Both coordinate fleets of local coding-agent CLIs in the terminal; Squad does ad-hoc manager/worker collaboration over SQLite with no daemon, Contrabass dispatches from Linear/GitHub issues with worktrees, retries and branch-advance verification.","confidence":0.75,"status":"approved"},{"to":"pullfrog","type":"alternative","why":"Same job — turn tracker issues into coding-agent runs. Pullfrog executes per @-mention inside GitHub Actions; Contrabass polls continuously and runs agents on your own machine in tmux panes and git worktrees.","confidence":0.7,"status":"approved"},{"to":"alook","type":"alternative","why":"Both turn local coding agents into a coordinated team working a shared board; alook is an always-on collaboration layer (per-agent email, kanban, shared memory), Contrabass is a leaner issue-queue dispatcher with per-run verification and retries.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T17:21:39.832Z","linkCount":7,"related":{"complements":[{"slug":"codexmate","why":"Same local-agent fleet, different modes: Codex Mate is the interactive dashboard for providers, sessions and skills; Contrabass takes over for unattended issue-driven runs.","dir":"in","confidence":0.5,"name":"codexmate"}],"alternative":[{"slug":"squad","why":"Both coordinate fleets of local coding-agent CLIs in the terminal; Squad does ad-hoc manager/worker collaboration over SQLite with no daemon, Contrabass dispatches from Linear/GitHub issues with worktrees, retries and branch-advance verification.","dir":"out","confidence":0.75,"name":"squad"},{"slug":"pullfrog","why":"Same job — turn tracker issues into coding-agent runs. Pullfrog executes per @-mention inside GitHub Actions; Contrabass polls continuously and runs agents on your own machine in tmux panes and git worktrees.","dir":"out","confidence":0.7,"name":"pullfrog"},{"slug":"fabro","why":"Both run coding agents through gated, verified pipelines instead of a REPL. Contrabass is terminal-first and issue-driven (poll Linear/GitHub, worktrees, branch-advance checks); Fabro is a server with workflow graphs, ensemble models and cloud sandboxes.","dir":"in","confidence":0.65,"name":"fabro"},{"slug":"fusion","why":"Same job — dispatch coding agents into isolated git worktrees with review gates — opposite philosophies: Contrabass is a lean terminal orchestrator fed by Linear/GitHub issues; Fusion is a full board-driven factory with visual workflows, missions and agent chat.","dir":"in","confidence":0.65,"name":"Fusion"},{"slug":"alook","why":"Both turn local coding agents into a coordinated team working a shared board; alook is an always-on collaboration layer (per-agent email, kanban, shared memory), Contrabass is a leaner issue-queue dispatcher with per-run verification and retries.","dir":"out","confidence":0.6,"name":"alook"},{"slug":"helmor","why":"Contrabass drains an issue queue headlessly with verification; Helmor is the interactive GUI take on the same orchestrate-local-coding-agents job.","dir":"in","confidence":0.5,"name":"helmor"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/junhoyeo/contrabass"},"core":{"name":"core","owner":"RedPlanetHQ","slug":"core","stars":1924,"image":"https://raw.githubusercontent.com/RedPlanetHQ/core/main/docs/images/attentive.gif","avatar":"https://avatars.githubusercontent.com/u/157399902?v=4&s=96","forks":183,"language":"TypeScript","license":"NOASSERTION","updated":"2 days ago","topics":["agents","memory"],"summary":"Self-hosted, always-on \"personal AI OS\": watches your apps, keeps a persistent memory graph, and acts autonomously within guardrails — a product, not a library for building agents.","curator_note":"Reach for CORE when you want an event-driven, self-hosted personal assistant that notices things on its own, remembers across sessions via a memory graph, and acts across your apps with per-action approval gates. NOT the pick if you want a framework to embed agents inside your own software (it's a product/OS, not a toolkit), nor if you need to run a multi-agent team/company with org charts and budgets — that's paperclip's lane.","edges":[{"to":"paperclip","type":"alternative","why":"Both are the layer ABOVE agent frameworks rather than a framework themselves, and both spawn coding-agent sessions — but CORE is a personal, memory-driven, always-on assistant while paperclip manages fleets of agents as a team/company. Pick CORE for a personal AI OS, paperclip for multi-agent org governance.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-05T14:49:26.000Z","linkCount":8,"related":{"complements":[],"alternative":[{"slug":"exxperts","why":"Both are self-hosted 'personal AI with persistent memory' products, but they pick opposite defaults: core watches your apps and acts autonomously within guardrails, exxperts gates every single memory write behind your explicit approval.","dir":"in","confidence":0.7,"name":"exxperts"},{"slug":"rowboat","why":"The two strongest 'personal AI OS' plays in the catalog: both keep a persistent memory graph and act autonomously within guardrails. core watches your apps from the background; Rowboat ships its own work surfaces — email, browser, meeting notes — and stores the graph as editable local Markdown.","dir":"in","confidence":0.7,"name":"rowboat"},{"slug":"paperclip","why":"Both are the layer ABOVE agent frameworks rather than a framework themselves, and both spawn coding-agent sessions — but CORE is a personal, memory-driven, always-on assistant while paperclip manages fleets of agents as a team/company. Pick CORE for a personal AI OS, paperclip for multi-agent org governance.","dir":"out","confidence":0.6,"name":"paperclip"},{"slug":"alook","why":"Both are self-hosted, always-on 'AI that works for you while you sleep' products with persistent memory — core shapes it as one personal AI OS watching your apps; alook shapes it as a team of role-assigned coding agents.","dir":"in","confidence":0.6,"name":"alook"},{"slug":"personal_ai_infrastructure","why":"Both 'personal AI OS' plays: core is an always-on shipped product watching your apps with a memory graph; LifeOS is an open scaffold of intent, agents and commands you inhabit on top of Claude Code.","dir":"in","confidence":0.6,"name":"LifeOS"},{"slug":"my-brain-is-full-crew","why":"Both are self-hosted 'manage my life' AI systems with persistent memory: core is a product that watches your apps and acts within guardrails on its own memory graph; the Crew works entirely through your Obsidian vault, keeping the memory human-readable and human-editable.","dir":"in","confidence":0.55,"name":"My-Brain-Is-Full-Crew"},{"slug":"picobot","why":"Both are always-on self-hosted personal AI assistants with persistent memory, at opposite ends of the weight spectrum: core is a full 'personal AI OS' product watching your apps; Picobot is a 9MB binary you leave running on a $5 VPS and message from Telegram.","dir":"in","confidence":0.55,"name":"picobot"},{"slug":"claude-reflect","why":"Both chase memory that compounds: core builds a full personal-AI OS with a persistent memory graph; claude-reflect does one narrow slice — your coding agent's lessons — with zero infrastructure.","dir":"in","confidence":0.5,"name":"claude-reflect"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/RedPlanetHQ/core"},"coreai-models":{"name":"coreai-models","owner":"apple","slug":"coreai-models","stars":1435,"avatar":"https://avatars.githubusercontent.com/u/10639145?v=4&s=96","forks":128,"language":"Swift","license":"BSD-3-Clause","updated":"2 days ago","topics":["local"],"summary":"Apple's official Core AI toolkit: recipes exporting Hugging Face models to .aimodel, PyTorch primitives for authoring, Swift runtime for macOS/iOS apps — plus skills for coding agents.","curator_note":"The sanctioned path to shipping models inside Mac and iOS apps: export recipes take popular open models to Core AI format, the Swift package handles tokenizers and multi-model pipelines at runtime, and the bundled agent skills teach Claude Code/Codex the framework — Apple shipping skills for coding agents is itself a signal. NOT for experimentation velocity: this is the deploy-in-an-app stack, not the tinker stack (MLX is where Apple-silicon research lives), it requires macOS/iOS 27+ and Xcode 27+, and you're accepting Apple's format and release cadence as your foundation.","edges":[{"to":"mlx-lora-studio","type":"complements","why":"The two halves of Apple-silicon on-device AI: fine-tune your model in MLX LoRA Studio, then export through Core AI recipes to ship it inside a macOS/iOS app. Research stack in, deployment stack out.","confidence":0.55,"status":"approved"},{"to":"siliconscope","type":"complements","why":"Core AI models execute on the Neural Engine — SiliconScope is the monitor that shows ANE load and per-process ANE memory, so you can see what your .aimodel actually costs the silicon.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.809Z","linkCount":2,"related":{"complements":[{"slug":"mlx-lora-studio","why":"The two halves of Apple-silicon on-device AI: fine-tune your model in MLX LoRA Studio, then export through Core AI recipes to ship it inside a macOS/iOS app. Research stack in, deployment stack out.","dir":"out","confidence":0.55,"name":"MLX-LoRA-Studio"},{"slug":"siliconscope","why":"Core AI models execute on the Neural Engine — SiliconScope is the monitor that shows ANE load and per-process ANE memory, so you can see what your .aimodel actually costs the silicon.","dir":"out","confidence":0.5,"name":"SiliconScope"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/apple/coreai-models"},"council-of-high-intelligence":{"name":"council-of-high-intelligence","owner":"0xNyk","slug":"council-of-high-intelligence","stars":3722,"image":"https://raw.githubusercontent.com/0xNyk/council-of-high-intelligence/main/assets/header.jpeg","avatar":"https://avatars.githubusercontent.com/u/93952610?v=4&s=96","forks":15,"language":"Shell","license":"MIT","updated":"7 days ago","topics":["agents","skills"],"summary":"/council: 18 AI personas deliberate your hardest decisions across multiple LLM providers — structured multi-round disagreement, confidence-weighted verdicts, one slash command.","curator_note":"Structured disagreement as a product: 18 personas with genuinely different priors argue your decision across providers (real model diversity, not one model roleplaying), through quick/standard/deep deliberation modes, ending in a confidence-weighted verdict. Its own README has a 'When Not to Use It' section — our kind of project. Best for irreversible, ambiguous calls where you'd otherwise ask three friends. NOT for anything with a checkable answer (deliberation theater costs real tokens), and a council is still only as wise as its training data — it widens perspective, it doesn't add ground truth.","edges":[{"to":"autogen","type":"alternative","why":"Both stage multi-agent deliberation; AutoGen is the framework you build conversations with, Council is the finished product — personas, protocol and verdict included, one command away.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T15:35:46.199Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"autogen","why":"Both stage multi-agent deliberation; AutoGen is the framework you build conversations with, Council is the finished product — personas, protocol and verdict included, one command away.","dir":"out","confidence":0.5,"name":"AutoGen"},{"slug":"cc-thinking-skills","why":"Two upgrades for decision quality in Claude Code: Council convenes 18 disagreeing personas across providers; thinking-skills hands one model 18 explicit frameworks. Deliberation breadth vs reasoning discipline.","dir":"in","confidence":0.5,"name":"cc-thinking-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/0xNyk/council-of-high-intelligence"},"crawl4ai":{"name":"crawl4ai","owner":"unclecode","slug":"crawl4ai","stars":74558,"image":"https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/powered-by-disco.svg","avatar":"https://avatars.githubusercontent.com/u/12494079?v=4&s=96","forks":7675,"language":"Python","license":"Apache-2.0","updated":"2 days ago","topics":["web","rag"],"summary":"The 74k-star LLM-native crawler: turns any site into clean, RAG-ready Markdown — adaptive crawling, JS rendering, extraction strategies, Docker deploy. Python, Apache-2.0.","curator_note":"The default answer when the deliverable is Markdown for a model rather than structured data for a database: heuristic content filtering, LLM and CSS extraction strategies, deep-crawl dispatchers and a Dockerized API — battle-tested by the largest community in the category. NOT the stealth pick: for hostile anti-bot targets Scrapling's fetchers earn their keep, and for classic item-pipeline scraping at scale Crawlee's queue/proxy machinery is more mature. Watch the cloud beta — the open core is healthy, but the cost-effective-cloud pitch tells you where the roadmap's gravity is.","edges":[{"to":"crawlee","type":"alternative","why":"The two big crawl frameworks, split by output philosophy: Crawlee (Node/TS) is item-pipeline scraping with anti-blocking and storage; Crawl4AI (Python) optimizes everything toward clean Markdown for LLM consumption.","confidence":0.7,"status":"approved"},{"to":"scrapling","type":"alternative","why":"Same Python scraping job, different bets: Scrapling bets on resilience (self-healing selectors, Cloudflare-passing stealth); Crawl4AI bets on LLM-ready output and crawl orchestration. Hostile targets → Scrapling; RAG pipelines → Crawl4AI.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.593Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"crawlee","why":"The two big crawl frameworks, split by output philosophy: Crawlee (Node/TS) is item-pipeline scraping with anti-blocking and storage; Crawl4AI (Python) optimizes everything toward clean Markdown for LLM consumption.","dir":"out","confidence":0.7,"name":"crawlee"},{"slug":"scrapling","why":"Same Python scraping job, different bets: Scrapling bets on resilience (self-healing selectors, Cloudflare-passing stealth); Crawl4AI bets on LLM-ready output and crawl orchestration. Hostile targets → Scrapling; RAG pipelines → Crawl4AI.","dir":"out","confidence":0.7,"name":"Scrapling"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/unclecode/crawl4ai"},"crawlee":{"name":"crawlee","owner":"apify","slug":"crawlee","stars":24883,"image":"https://raw.githubusercontent.com/apify/crawlee/master/website/static/img/crawlee-light.svg?sanitize=true","avatar":"https://avatars.githubusercontent.com/u/24586296?v=4&s=96","forks":1571,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["web"],"summary":"Apify's web scraping and browser automation library for Node.js/TypeScript — HTTP and headless-browser crawlers with human-like anti-blocking defaults, queues, storage and proxies.","curator_note":"The TypeScript-native answer to production crawling: one API switches between cheap HTTP crawling and Playwright/Puppeteer when pages need a real browser, with anti-blocking fingerprints, request queues, proxy rotation and storage built in rather than bolted on. Reach for it when your stack is Node and the target fights back. NOT the pick for Python teams (its Python port is younger — Scrapy owns that ground), and mind the gravity: it's built by Apify and nudges toward their platform, though it runs fine standalone.","edges":[{"to":"scrapy","type":"alternative","why":"Same job — a production crawling framework. Scrapy is the fifteen-year Python veteran with the deepest ecosystem; Crawlee is TypeScript-native with headless-browser switching and anti-blocking as defaults rather than plugins.","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.331Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"scrapy","why":"Same job — a production crawling framework. Scrapy is the fifteen-year Python veteran with the deepest ecosystem; Crawlee is TypeScript-native with headless-browser switching and anti-blocking as defaults rather than plugins.","dir":"out","confidence":0.8,"name":"scrapy"},{"slug":"crawl4ai","why":"The two big crawl frameworks, split by output philosophy: Crawlee (Node/TS) is item-pipeline scraping with anti-blocking and storage; Crawl4AI (Python) optimizes everything toward clean Markdown for LLM consumption.","dir":"in","confidence":0.7,"name":"crawl4ai"},{"slug":"scrapling","why":"The two modern anti-bot-aware crawling frameworks, split by language: Scrapling for Python (adaptive parsing, MCP server), Crawlee for Node/TypeScript (browser switching, Apify ecosystem).","dir":"in","confidence":0.7,"name":"Scrapling"},{"slug":"autoscraper","why":"Same scraping job, opposite shapes: autoscraper is 500 lines of learn-by-example rule inference in Python; Crawlee is the production TypeScript framework with queues, proxies and anti-blocking.","dir":"in","confidence":0.6,"name":"autoscraper"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/apify/crawlee"},"crewai":{"name":"CrewAI","owner":"joaomdmoura","slug":"crewai","stars":55998,"image":"https://raw.githubusercontent.com/joaomdmoura/crewai/main/docs/images/crewai_logo.png","avatar":"https://avatars.githubusercontent.com/u/170677839?v=4&s=96","forks":7925,"language":"Python","license":"MIT","updated":"2 days ago","topics":["agents","orchestration"],"summary":"Orchestrate role-playing, autonomous AI agents that collaborate on tasks.","edges":[{"to":"autogen","type":"alternative","why":"Both do multi-agent collaboration; CrewAI is more opinionated about roles, AutoGen about conversation.","confidence":0.8,"status":"approved"},{"to":"ollama","type":"built_with","why":"Run your crew against local models for cost-free iteration.","confidence":0.75,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":11,"related":{"complements":[{"slug":"memmachine","why":"Ships a CrewAI integration: crews get shared persistent memory across sessions instead of CrewAI's per-run state.","dir":"in","confidence":0.75,"name":"MemMachine"},{"slug":"litellm","why":"Agent framework relies on a unified model layer; LiteLLM supplies the 100+-provider abstraction beneath the agents.","dir":"in","confidence":0.6,"name":"litellm"}],"alternative":[{"slug":"metagpt","why":"Both bet on role-playing agent teams. CrewAI gives you general-purpose crews you compose; MetaGPT ships one opinionated, SOP-encoded software company end to end.","dir":"in","confidence":0.85,"name":"MetaGPT"},{"slug":"autogen","why":"Both do multi-agent collaboration; CrewAI is more opinionated about roles, AutoGen about conversation.","dir":"out","confidence":0.8,"name":"AutoGen"},{"slug":"langgraph","why":"Higher-level, opinionated multi-agent API vs. LangGraph's low-level control. Trade control for speed.","dir":"in","confidence":0.8,"name":"LangGraph"},{"slug":"praisonai","why":"Same job — role-based multi-agent orchestration. PraisonAI trades CrewAI's focused API for a broader low-code bundle (memory, RAG, UI, MCP included) and even runs CrewAI-style configs.","dir":"in","confidence":0.8,"name":"PraisonAI"},{"slug":"deepagents","why":"Same job — a framework for multi-agent systems — different philosophy: CrewAI models role-playing crews; deepagents is one deep agent that plans and delegates to sub-agents with isolated contexts.","dir":"in","confidence":0.65,"name":"deepagents"},{"slug":"paperclip","why":"Both pitch orchestrating a team of agents toward a shared goal. Different altitude: CrewAI is a framework to build role-playing collaborating agents in-process; Paperclip is a BYO-agent management plane over external runtimes. Overlap in the multi-agent coordination job makes them substitutes for some users.","dir":"in","confidence":0.6,"name":"paperclip"},{"slug":"scale-agentex","why":"Overlapping 'build agents in Python' entry point, opposite emphasis: CrewAI is the in-process multi-agent collaboration framework; Agentex is the deployment platform — protocolized agents, dev sandbox, and durable async execution on Temporal.","dir":"in","confidence":0.55,"name":"scale-agentex"}],"built_with":[{"slug":"ollama","why":"Run your crew against local models for cost-free iteration.","dir":"out","confidence":0.75,"name":"Ollama"},{"slug":"ai-engineering-hub","why":"The flagship walkthroughs — including build-code-harness's hierarchical agent workflow — are built on CrewAI; the hub doubles as its largest applied-tutorial corpus.","dir":"in","confidence":0.55,"name":"ai-engineering-hub"}]},"url":"https://stackmap.shipwithai.xyz/repos/joaomdmoura/crewai"},"cubesandbox":{"name":"CubeSandbox","owner":"TencentCloud","slug":"cubesandbox","stars":10613,"image":"https://raw.githubusercontent.com/TencentCloud/cubesandbox/master/docs/assets/cube-sandbox-logo.png","avatar":"https://avatars.githubusercontent.com/u/20101770?v=4&s=96","forks":946,"language":"Rust","license":"NOASSERTION","updated":"yesterday","topics":["agents","local"],"summary":"Hardware-isolated microVM sandboxes for AI agents — sub-60ms boot, <5MB overhead, E2B-compatible API, self-hosted on your own KVM nodes.","curator_note":"Reach for CubeSandbox the moment your agents execute model-generated code and \"just run it in Docker\" stops feeling safe — it gives every tool call a disposable hardware-isolated microVM with E2B's SDK ergonomics, minus the SaaS bill, plus snapshot/rollback of any sandbox state. The catch: it's real infrastructure — you need KVM-capable Linux hosts and someone willing to operate them. Prototyping a single local agent? A container or E2B's hosted tier is less machinery. It's a runtime, not a framework — you still bring LangGraph/AutoGen/whatever on top.","edges":[{"to":"langgraph","type":"complements","why":"LangGraph nodes that run model-written code get a disposable, hardware-isolated microVM per execution via the E2B-compatible SDK — crash or escape attempts die with the VM.","confidence":0.8,"status":"approved"},{"to":"autogen","type":"complements","why":"AutoGen's code-executor step is the canonical sandbox case: run generated Python in an isolated microVM instead of on the host or a shared container.","confidence":0.75,"status":"approved"},{"to":"paperclip","type":"complements","why":"paperclip governs fleets of agents; CubeSandbox is the workload-runtime layer beneath — thousands of <5MB VMs per node, idle agents auto-pause to keep fleet cost sane.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":10,"related":{"complements":[{"slug":"langgraph","why":"LangGraph nodes that run model-written code get a disposable, hardware-isolated microVM per execution via the E2B-compatible SDK — crash or escape attempts die with the VM.","dir":"out","confidence":0.8,"name":"LangGraph"},{"slug":"autogen","why":"AutoGen's code-executor step is the canonical sandbox case: run generated Python in an isolated microVM instead of on the host or a shared container.","dir":"out","confidence":0.75,"name":"AutoGen"},{"slug":"paperclip","why":"paperclip governs fleets of agents; CubeSandbox is the workload-runtime layer beneath — thousands of <5MB VMs per node, idle agents auto-pause to keep fleet cost sane.","dir":"out","confidence":0.7,"name":"paperclip"},{"slug":"omnigent","why":"Omnigent launches sessions in E2B-compatible sandboxes; cubesandbox is a self-hosted E2B-compatible microVM backend to run them on your own KVM nodes.","dir":"in","confidence":0.65,"name":"omnigent"},{"slug":"browser-use","why":"An agent driving a real browser is exactly the workload you want hardware-isolated: run browser-use sessions inside microVM sandboxes instead of on your own profile.","dir":"in","confidence":0.55,"name":"browser-use"},{"slug":"evo","why":"evo dispatches experiments to E2B-compatible remote sandboxes; CubeSandbox self-hosts exactly that API on your own KVM nodes — a natural backend when experiment runs shouldn't leave your infrastructure.","dir":"in","confidence":0.5,"name":"evo"},{"slug":"mirage","why":"Two halves of the agent environment: CubeSandbox isolates the compute (hardware-isolated microVM per agent), mirage unifies the data plane (every external service as one mounted tree). Run the agent in the sandbox, hand it its world as files — neither replaces the other.","dir":"in","confidence":0.5,"name":"mirage"}],"alternative":[{"slug":"forkd","why":"Both are self-hosted Firecracker-class microVM sandboxes for agents. CubeSandbox sells a platform (E2B-compatible API, sub-60ms cold boots); forkd sells a primitive — copy-on-write fork from a warm parent, which wins precisely when 100 children share one expensive warm state.","dir":"in","confidence":0.75,"name":"forkd"},{"slug":"opensandbox","why":"Both self-hosted sandbox runtimes for agents, split by isolation bet: CubeSandbox is microVM-first (KVM boundary, E2B-compatible API); OpenSandbox is platform-first (Docker/K8s runtimes, SDKs, CLI, MCP) for teams standardizing on existing infra.","dir":"in","confidence":0.7,"name":"OpenSandbox"},{"slug":"superserve","why":"Both provide Firecracker/microVM sandbox infrastructure for AI agents. CubeSandbox is self-hosted on your own KVM nodes with an E2B-compatible API; Superserve is a hosted service with persistence as the headline and SDKs as the open-source surface.","dir":"in","confidence":0.7,"name":"superserve"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/TencentCloud/cubesandbox"},"cwc-long-running-agents":{"name":"cwc-long-running-agents","owner":"anthropics","slug":"cwc-long-running-agents","stars":589,"avatar":"https://avatars.githubusercontent.com/u/76263028?v=4&s=96","forks":61,"language":"Shell","license":"Apache-2.0","updated":"2 months ago","topics":["coding"],"summary":"Anthropic's harness primitives for long-running Claude agents: default-FAIL evidence gates, a fresh-context evaluator subagent and handoff hooks — each one standalone, readable file.","curator_note":"Read it, don't run it: an official worked example of WHY long runs fail (agents grading their own work, claiming success without evidence, losing state between sessions) with one hook per failure mode. The default-FAIL contract is the idea worth stealing. Explicitly an event demo — unmaintained, not accepting contributions — so copy the patterns into your own harness rather than depending on the repo.","edges":[{"to":"loop-engineering","type":"alternative","why":"Both are reference repos for keeping agents productive over long horizons: loop-engineering designs the control loops, Anthropic's primitives enforce evidence and evaluation inside Claude Code hooks.","confidence":0.6,"status":"approved"},{"to":"govctl","type":"alternative","why":"Same core move — structural gates instead of polite prompts: govctl makes governance a CLI with verification gates; these hooks make 'done' unclaimable without opened evidence.","confidence":0.5,"status":"approved"},{"to":"loopy","type":"alternative","why":"Loopy catalogs reusable agent loops as installable skills; this repo ships the raw hook/evaluator primitives you'd build such loops from.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:42.050Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"loop-engineering","why":"Both are reference repos for keeping agents productive over long horizons: loop-engineering designs the control loops, Anthropic's primitives enforce evidence and evaluation inside Claude Code hooks.","dir":"out","confidence":0.6,"name":"loop-engineering"},{"slug":"govctl","why":"Same core move — structural gates instead of polite prompts: govctl makes governance a CLI with verification gates; these hooks make 'done' unclaimable without opened evidence.","dir":"out","confidence":0.5,"name":"govctl"},{"slug":"loopy","why":"Loopy catalogs reusable agent loops as installable skills; this repo ships the raw hook/evaluator primitives you'd build such loops from.","dir":"out","confidence":0.45,"name":"loopy"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/anthropics/cwc-long-running-agents"},"dataflow":{"name":"DataFlow","owner":"OpenDCAI","slug":"dataflow","stars":6758,"image":"https://github.com/user-attachments/assets/a19865e5-221d-4c12-bb57-17421df87c8a","avatar":"https://avatars.githubusercontent.com/u/177615276?v=4&s=96","forks":832,"language":"Python","license":"Apache-2.0","updated":"9 days ago","topics":["training"],"summary":"Operator-based system for LLM data prep — 100+ operators composed into pipelines that generate, clean, evaluate and filter pretraining/SFT/RL data, with a WebUI and a pipeline-building agent.","curator_note":"Pick it when the model isn't the problem, the data is. Ready pipelines cover text/math/code synthesis, large-scale PDF→QA, Text2SQL and knowledge-base cleaning, all in a PyTorch-like Pipeline→Operator→Prompt hierarchy that makes data governance reproducible and shareable; the DataFlow-Agent assembles pipelines from a task description, and the WebUI gives non-coders drag-and-drop. Versus Data-Juicer/Nemo-Curator the differentiator is synthesis, with peer-reviewed pedigree (ICDE/KDD acceptances, arXiv report). NOT runtime data plumbing — this is offline training-data preparation, and budget for real LLM API burn since most interesting operators call models (vLLM/SGLang backends supported for local).","edges":[{"to":"llamafactory","type":"complements","why":"Adjacent stages of one workflow: DataFlow generates, cleans and filters the SFT/RL datasets; LLaMA-Factory is the fine-tuning framework that consumes them across 100+ open models.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.385Z","linkCount":2,"related":{"complements":[{"slug":"llamafactory","why":"Adjacent stages of one workflow: DataFlow generates, cleans and filters the SFT/RL datasets; LLaMA-Factory is the fine-tuning framework that consumes them across 100+ open models.","dir":"out","confidence":0.65,"name":"LlamaFactory"}],"alternative":[{"slug":"adala","why":"Both build LLM-powered training data at scale: DataFlow is operator pipelines you compose for generation/cleaning/filtering; Adala is agents that LEARN the labeling skill from ground truth and self-improve to a target accuracy.","dir":"in","confidence":0.55,"name":"Adala"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/OpenDCAI/dataflow"},"deepagents":{"name":"deepagents","owner":"langchain-ai","slug":"deepagents","stars":26702,"image":"https://raw.githubusercontent.com/langchain-ai/deepagents/main/.github/images/logo-dark.svg","avatar":"https://avatars.githubusercontent.com/u/126733545?v=4&s=96","forks":3738,"language":"Python","license":"MIT","updated":"yesterday","topics":["agents","orchestration"],"summary":"LangChain's batteries-included agent harness on LangGraph — planning, sub-agents with isolated context, filesystem, shell, skills, human-in-the-loop and persistent memory out of the box.","curator_note":"The fastest route to a serious long-horizon agent if you accept LangChain's stack: planning, sub-agents, context offloading and HITL gates work out of the box, and any LangGraph graph plugs in as a sub-agent, so custom orchestration composes instead of forking. NOT for simple tool-calling loops — LangChain's create_agent is lighter — and the opinions run deep: if you're fighting the harness, you wanted LangGraph directly. Model-agnostic in theory; tuned around frontier tool-callers in practice.","edges":[{"to":"langgraph","type":"built_with","why":"Built directly on LangGraph — streaming, persistence and checkpointing come from the runtime; deepagents is the opinionated harness layer above it, and CompiledStateGraphs drop in as sub-agents.","confidence":0.95,"status":"approved"},{"to":"langsmith","type":"complements","why":"The README's own production pairing: LangSmith supplies tracing, evaluation and monitoring for deepagents deployments.","confidence":0.7,"status":"approved"},{"to":"crewai","type":"alternative","why":"Same job — a framework for multi-agent systems — different philosophy: CrewAI models role-playing crews; deepagents is one deep agent that plans and delegates to sub-agents with isolated contexts.","confidence":0.65,"status":"approved"},{"to":"autogen","type":"alternative","why":"Both are batteries-included multi-agent frameworks; AutoGen centers on agent-to-agent conversation, deepagents on a single planner with sub-agents, filesystem and context management.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-07T15:50:15.952Z","linkCount":6,"related":{"complements":[{"slug":"langsmith","why":"The README's own production pairing: LangSmith supplies tracing, evaluation and monitoring for deepagents deployments.","dir":"out","confidence":0.7,"name":"LangSmith"},{"slug":"mirage","why":"deepagents deliberately gives its agents a filesystem + shell as core tools; mirage extends exactly that interface to real infrastructure — S3, Slack, Gmail, Postgres mounted as paths via its LangChain-family adapters. The agent keeps grep/cat/pipe semantics while reaching production data instead of a local scratch dir.","dir":"in","confidence":0.55,"name":"mirage"}],"alternative":[{"slug":"crewai","why":"Same job — a framework for multi-agent systems — different philosophy: CrewAI models role-playing crews; deepagents is one deep agent that plans and delegates to sub-agents with isolated contexts.","dir":"out","confidence":0.65,"name":"CrewAI"},{"slug":"openmanus","why":"Same job — a batteries-included general autonomous agent that plans, browses and executes multi-step tasks. deepagents is the actively maintained, LangGraph-backed harness; OpenManus is the leaner, quieter reference implementation.","dir":"in","confidence":0.65,"name":"OpenManus"},{"slug":"autogen","why":"Both are batteries-included multi-agent frameworks; AutoGen centers on agent-to-agent conversation, deepagents on a single planner with sub-agents, filesystem and context management.","dir":"out","confidence":0.6,"name":"AutoGen"}],"built_with":[{"slug":"langgraph","why":"Built directly on LangGraph — streaming, persistence and checkpointing come from the runtime; deepagents is the opinionated harness layer above it, and CompiledStateGraphs drop in as sub-agents.","dir":"out","confidence":0.95,"name":"LangGraph"}]},"url":"https://stackmap.shipwithai.xyz/repos/langchain-ai/deepagents"},"deepdive":{"name":"DeepDive","owner":"THUDM","slug":"deepdive","stars":333,"image":"https://raw.githubusercontent.com/THUDM/deepdive/main/assets/combine_head_figure.svg","avatar":"https://avatars.githubusercontent.com/u/48590610?v=4&s=96","forks":36,"language":"Python","license":null,"updated":"1 months ago","topics":["training"],"summary":"THUDM recipe for deep-search agents: synthesize hard multi-hop QA from knowledge-graph random walks, then multi-turn GRPO RL — DeepDive-32B hits 14.8% BrowseComp; data feeds GLM-4.5/4.6.","curator_note":"Read it for the data trick: random-walk paths through KILT/AMiner knowledge graphs, entity obfuscation into 'blurry entities', then difficulty-filtering that keeps only questions GPT-4o fails four times out of four. Strict binary rewards (format AND answer, else zero) resist reward hacking, and the test-time finding is counterintuitive gold — among 8 parallel trajectories, the answer reached with the FEWEST tool calls wins (24.8% vs 12.0% single-shot). Use it to train or study open-model search agents; NOT a runnable product — model checkpoints are still 'coming soon', and you bring Serper/Jina API keys plus slime training infra. The 4,108-entry dataset is open on HF and already went into GLM-4.5/4.6.","edges":[{"to":"slime","type":"built_with","why":"DeepDive's multi-turn RL runs on THUDM's slime framework — the released training code is a slime rollout setup, and the repo credits it directly.","confidence":0.85,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.428Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"openresearcher","why":"Rival open recipes for training deep-research agents: DeepDive synthesizes hard QA from knowledge-graph walks and trains with multi-turn RL on slime (checkpoints still pending); OpenResearcher distills 96K long trajectories from GPT-OSS-120B and ships data, 30B model and eval framework complete.","dir":"in","confidence":0.75,"name":"OpenResearcher"}],"built_with":[{"slug":"slime","why":"DeepDive's multi-turn RL runs on THUDM's slime framework — the released training code is a slime rollout setup, and the repo credits it directly.","dir":"out","confidence":0.85,"name":"slime"}]},"url":"https://stackmap.shipwithai.xyz/repos/THUDM/deepdive"},"deepeval":{"name":"deepeval","owner":"confident-ai","slug":"deepeval","stars":17065,"image":"https://raw.githubusercontent.com/confident-ai/deepeval/main/assets/hero/wordmark-light.svg","avatar":"https://avatars.githubusercontent.com/u/130858411?v=4&s=96","forks":1698,"language":"Python","license":"Apache-2.0","updated":"2 days ago","topics":["evals"],"summary":"Pytest for LLM apps: 40+ research-backed metrics — G-Eval, RAG suite, agent task completion, hallucination — as unit tests you run in CI, judged by any LLM including local ones.","curator_note":"The most complete general-purpose eval framework: pytest ergonomics ('assert_test' in CI), metrics with explanations grounded in the research they implement, component-level tracing for agents, plus benchmark harnesses (MMLU, HumanEval) when you need them. If you're evaluating LLM apps in Python and don't know where to start, start here. NOT vendor-neutral: the free framework feeds Confident AI's platform for reports and regression tracking — usable without it, but the gravity is real. And LLM-as-judge metrics inherit judge bias: treat scores as regression signals, not absolute truth. RAG-only shops may prefer Ragas' tighter focus.","edges":[{"to":"ragas","type":"alternative","why":"The two default pip installs for LLM evaluation: Ragas is RAG-centric (faithfulness, relevance); DeepEval covers the same RAG suite plus agents, chatbots, safety metrics and pytest-style CI integration. Broader tool vs sharper tool.","confidence":0.75,"status":"approved"},{"to":"deepteam","type":"complements","why":"Same team, one stack: DeepTeam is built ON DeepEval — red-team attacks and eval metrics share plumbing, so vulnerability findings land in the same workflow as your quality regressions instead of a separate report.","confidence":0.8,"status":"approved"},{"to":"giskard-oss","type":"alternative","why":"Both are modular Python eval frameworks with LLM-as-judge and safety scanning. Giskard leans scenario-based checks and OWASP-mapped vulnerability scanning; DeepEval leans metric breadth and pytest-in-CI ergonomics.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-20T10:48:48.808Z","linkCount":3,"related":{"complements":[{"slug":"deepteam","why":"Same team, one stack: DeepTeam is built ON DeepEval — red-team attacks and eval metrics share plumbing, so vulnerability findings land in the same workflow as your quality regressions instead of a separate report.","dir":"out","confidence":0.8,"name":"deepteam"}],"alternative":[{"slug":"ragas","why":"The two default pip installs for LLM evaluation: Ragas is RAG-centric (faithfulness, relevance); DeepEval covers the same RAG suite plus agents, chatbots, safety metrics and pytest-style CI integration. Broader tool vs sharper tool.","dir":"out","confidence":0.75,"name":"Ragas"},{"slug":"giskard-oss","why":"Both are modular Python eval frameworks with LLM-as-judge and safety scanning. Giskard leans scenario-based checks and OWASP-mapped vulnerability scanning; DeepEval leans metric breadth and pytest-in-CI ergonomics.","dir":"out","confidence":0.6,"name":"giskard-oss"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/confident-ai/deepeval"},"deepface":{"name":"deepface","owner":"serengil","slug":"deepface","stars":23138,"image":"https://raw.githubusercontent.com/serengil/deepface/master/icon/deepface-icon-labeled.png","avatar":"https://avatars.githubusercontent.com/u/18491038?v=4&s=96","forks":3143,"language":"Python","license":"MIT","updated":"25 days ago","topics":["vision"],"summary":"Lightweight Python face recognition and facial-attribute analysis (age, gender, emotion) wrapping VGG-Face, FaceNet, ArcFace and friends — pip install, self-hosted, battle-tested at 23k stars.","curator_note":"The de-facto open-source answer to cloud face APIs: pip install deepface and you get detection, alignment, verification, find-in-database and attribute analysis (age/gender/emotion) behind one call, wrapping VGG-Face, FaceNet, ArcFace, Dlib and more — self-hosted, no per-image API bill, with a REST server and Docker image for production. Reach for it when you need face verification or analysis on your own hardware in minutes. NOT an LLM/agent tool — it's classic computer vision, the catalog's first; and treat attribute analysis (race/emotion) with care: contested accuracy and serious GDPR/ethics exposure in production. GPU optional but strongly advised beyond hobby volume.","edges":[],"status":"approved","added":"2026-07-08T16:13:44.139Z","linkCount":1,"related":{"complements":[{"slug":"ultralytics","why":"The standard two-stage pipeline: YOLO finds the people in the frame, deepface tells you who they are — detector crops feed face verification directly.","dir":"in","confidence":0.75,"name":"ultralytics"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/serengil/deepface"},"deepteam":{"name":"deepteam","owner":"confident-ai","slug":"deepteam","stars":2289,"image":"https://raw.githubusercontent.com/confident-ai/deepteam/main/assets/hero/wordmark-light-v2.svg","avatar":"https://avatars.githubusercontent.com/u/130858411?v=4&s=96","forks":368,"language":"Python","license":"Apache-2.0","updated":"4 days ago","topics":["security","evals"],"summary":"Open-source red-teaming framework for LLM systems: 50+ vulnerabilities, jailbreak/injection/multi-turn attacks against agents, RAG pipelines and chatbots — plus guardrails. Runs locally.","curator_note":"Penetration testing for LLM apps with a framework's ergonomics: pick from 50+ documented vulnerabilities (bias, PII leakage, SQL injection via prompt) and simulated attacks including multi-turn exploitation, run locally, then wire the matching guardrails. Built on DeepEval, so red-team results plug into an eval workflow rather than dying in a report. NOT neutral infrastructure — it funnels toward the Confident AI platform for storing results (self-host your own reporting if that matters), and like every attack library it tests known patterns: a clean run is a floor, not a clearance.","edges":[{"to":"garak","type":"alternative","why":"Same job — adversarial testing of LLM systems. garak is the nmap-style scanner you point at a model endpoint; DeepTeam is the framework you embed in your eval suite, with attack simulation and guardrails in one loop.","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-13T10:36:42.150Z","linkCount":3,"related":{"complements":[{"slug":"deepeval","why":"Same team, one stack: DeepTeam is built ON DeepEval — red-team attacks and eval metrics share plumbing, so vulnerability findings land in the same workflow as your quality regressions instead of a separate report.","dir":"in","confidence":0.8,"name":"deepeval"}],"alternative":[{"slug":"garak","why":"Same job — adversarial testing of LLM systems. garak is the nmap-style scanner you point at a model endpoint; DeepTeam is the framework you embed in your eval suite, with attack simulation and guardrails in one loop.","dir":"out","confidence":0.8,"name":"garak"},{"slug":"giskard-oss","why":"Both red-team LLM systems with framework ergonomics: deepteam ships 50+ documented vulnerabilities and matching guardrails; giskard-scan auto-generates adversarial suites from a plain-language agent description and pairs with the same stack's eval checks.","dir":"in","confidence":0.7,"name":"giskard-oss"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/confident-ai/deepteam"},"designer-skills":{"name":"designer-skills","owner":"Owl-Listener","slug":"designer-skills","stars":1884,"avatar":"https://avatars.githubusercontent.com/u/200631076?v=4&s=96","forks":317,"language":"Markdown","license":"MIT","updated":"1 months ago","topics":["skills"],"summary":"239 design skills, 88 commands and 33 plugins for Claude Code and Gemini CLI — research, design systems, UI, interaction and delivery, written for agents to actually execute.","curator_note":"The largest coherent skill pack for design work: five collections covering research→systems→UI→interaction→delivery, installable straight from Claude Code's plugin marketplace with genuinely non-technical instructions. If you're a designer adopting agents (or an engineer faking design taste), this is the fastest on-ramp. NOT a design tool itself — it's prompts-as-skills, so output quality tracks the underlying model, and 239 skills means uneven depth: audit the ones you'll actually use rather than installing all five collections on day one.","edges":[{"to":"skillkit","type":"complements","why":"Content meets distribution: designer-skills is a curated skill pack, skillkit is the package manager that installs, translates and security-scans skills like these across 46 agent formats.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.351Z","linkCount":6,"related":{"complements":[{"slug":"skillkit","why":"Content meets distribution: designer-skills is a curated skill pack, skillkit is the package manager that installs, translates and security-scans skills like these across 46 agent formats.","dir":"out","confidence":0.55,"name":"skillkit"},{"slug":"npxskillui","why":"designer-skills teaches an agent design craft in general; npxskillui extracts YOUR specific design system into a skill — general literacy plus house style.","dir":"in","confidence":0.5,"name":"npxskillui"}],"alternative":[{"slug":"html-anything","why":"Both make coding agents produce designed deliverables via skill packs. designer-skills is 239 skills of design process for Claude Code/Gemini; HTML Anything is an editor product — 75 skills fused with preview, sandboxing and one-click publishing surfaces.","dir":"in","confidence":0.55,"name":"html-anything"},{"slug":"pm-skills","why":"Both are single-discipline skill packs that give Claude a professional's operating frameworks — design craft in one, product management in the other. Same install pattern, different seat at the table.","dir":"in","confidence":0.55,"name":"pm-skills"},{"slug":"ai-marketing-claude","why":"Same shape, different profession: both are domain-expertise skill packs that turn Claude Code into a practitioner — designer-skills for research→design→delivery, this one for audits→copy→campaigns→client reports.","dir":"in","confidence":0.5,"name":"ai-marketing-claude"},{"slug":"kami","why":"Two routes to design-literate agents: designer-skills is a broad pack of design-craft skills; Kami is one strict constraint system focused on making every shipped document consistent.","dir":"in","confidence":0.5,"name":"Kami"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Owl-Listener/designer-skills"},"dspy":{"name":"DSPy","owner":"stanfordnlp","slug":"dspy","stars":36321,"image":"https://raw.githubusercontent.com/stanfordnlp/dspy/main/docs/docs/static/img/dspy_logo.png","avatar":"https://avatars.githubusercontent.com/u/3046006?v=4&s=96","forks":3123,"language":"Python","license":"MIT","updated":"4 days ago","topics":["orchestration","rag"],"summary":"Program — don't prompt — your language models. Compile declarative pipelines into optimized prompts.","edges":[],"added":"2026-06-22T23:27:20.000Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"langgraph","why":"A different bet: optimize prompts as code rather than orchestrate them.","dir":"in","confidence":0.7,"name":"LangGraph"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/stanfordnlp/dspy"},"ecc":{"name":"ECC","owner":"affaan-m","slug":"ecc","stars":232350,"image":"https://raw.githubusercontent.com/affaan-m/ecc/main/assets/hero.png","avatar":"https://avatars.githubusercontent.com/u/124439313?v=4&s=96","forks":35418,"language":"JavaScript","license":"MIT","updated":"yesterday","topics":["coding","skills"],"summary":"Cross-harness 'operating system' for coding agents — 268 skills, 66 agents, hooks, rules, memory persistence, instinct-based continuous learning and AgentShield security scanning. MIT.","curator_note":"The kitchen-sink option, and the most popular one by a wide margin: one install gives Claude Code, Codex, Cursor, OpenCode and friends a full operating layer — session-memory hooks, learned instincts, quality gates, orch-* orchestrator commands, worktree lifecycle, security scanning — plus three genuinely good guides on token optimization, memory and agentic security. When NOT: it's maximalist. Hundreds of skills and global hooks add context weight, and the README's own top warning is about broken stacked installs — start with the minimal/no-hooks profile and never mix the plugin with the manual installer. If you only want one capability, take the surgical tool instead: claude-reflect for learning-from-corrections, pro-workflow for session memory, asm or skillkit for skill management. Watch the upsell surface (Pro app, sponsors) — the OSS core is MIT and complete.","edges":[{"to":"pro-workflow","type":"alternative","why":"Same job — a persistent operating layer under coding-agent sessions with hooks, learned rules and quality gates. pro-workflow is one SQLite store focused on Claude Code; ECC is the maximalist cross-harness system spanning 7+ harnesses with skills, agents and security tooling.","confidence":0.7,"status":"approved"},{"to":"claude-reflect","type":"alternative","why":"Both close the learning loop from your corrections: claude-reflect is a focused Claude Code plugin syncing approved learnings to CLAUDE.md; ECC's instinct system does the same continuous-learning job with confidence scoring inside a much larger harness framework.","confidence":0.6,"status":"approved"},{"to":"apm","type":"alternative","why":"Both ship reproducible agent-context setups across many clients. apm is a lean dependency manager (declare skills/prompts/MCP in apm.yml, lockfile-pinned installs); ECC ships the content itself — a batteries-included catalog of skills, agents, hooks and rules.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T17:53:59.877Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"pro-workflow","why":"Same job — a persistent operating layer under coding-agent sessions with hooks, learned rules and quality gates. pro-workflow is one SQLite store focused on Claude Code; ECC is the maximalist cross-harness system spanning 7+ harnesses with skills, agents and security tooling.","dir":"out","confidence":0.7,"name":"pro-workflow"},{"slug":"claude-reflect","why":"Both close the learning loop from your corrections: claude-reflect is a focused Claude Code plugin syncing approved learnings to CLAUDE.md; ECC's instinct system does the same continuous-learning job with confidence scoring inside a much larger harness framework.","dir":"out","confidence":0.6,"name":"claude-reflect"},{"slug":"apm","why":"Both ship reproducible agent-context setups across many clients. apm is a lean dependency manager (declare skills/prompts/MCP in apm.yml, lockfile-pinned installs); ECC ships the content itself — a batteries-included catalog of skills, agents, hooks and rules.","dir":"out","confidence":0.55,"name":"apm"},{"slug":"ponytail","why":"Opposite philosophies for shaping a coding agent: ECC installs an entire operating layer — hooks, gates, orchestrators; ponytail installs exactly one opinion. Maximalist vs minimalist, and the choice says more about your team than the tools.","dir":"in","confidence":0.5,"name":"ponytail"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/affaan-m/ecc"},"evo":{"name":"evo","owner":"evo-hq","slug":"evo","stars":1348,"image":"https://raw.githubusercontent.com/evo-hq/evo/main/assets/banner.png","avatar":"https://avatars.githubusercontent.com/u/274232282?v=4&s=96","forks":101,"language":"Python","license":"Apache-2.0","updated":"7 days ago","topics":["coding","skills"],"summary":"Karpathy-style autoresearch on any codebase: /evo:discover instruments the benchmark, /evo:optimize runs tree search with parallel subagents in worktrees. Plugin for Claude Code, Codex & co.","curator_note":"Where the autoresearch skill is a greedy hill-climb on one branch, evo is the industrialized version: tree search over committed nodes, parallel subagents in isolated git worktrees sharing failure traces, regression gates that discard bad experiments, frontier strategies (argmax to GEPA-inspired pareto), and a dashboard. Benchmark discovery is the sleeper feature — it figures out what to measure before optimizing. Runs on 8 agent hosts, local or remote sandboxes (Modal, E2B, Daytona, AWS). NOT for codebases without a mechanical metric — no benchmark, no loop — and unattended parallel agents editing your repo demand real gates; the hosted 'evo platform' beta hints where the business model lands.","edges":[{"to":"autoresearch","type":"alternative","why":"Same lineage — Karpathy's autoresearch loop for coding agents. The skill is a single-branch keep/discard hill-climb; evo adds tree search, parallel worktree subagents, shared failure traces, gates and a dashboard. Start with the skill, graduate to evo.","confidence":0.85,"status":"approved"},{"to":"cubesandbox","type":"complements","why":"evo dispatches experiments to E2B-compatible remote sandboxes; CubeSandbox self-hosts exactly that API on your own KVM nodes — a natural backend when experiment runs shouldn't leave your infrastructure.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-20T10:48:48.562Z","linkCount":3,"related":{"complements":[{"slug":"cubesandbox","why":"evo dispatches experiments to E2B-compatible remote sandboxes; CubeSandbox self-hosts exactly that API on your own KVM nodes — a natural backend when experiment runs shouldn't leave your infrastructure.","dir":"out","confidence":0.5,"name":"CubeSandbox"}],"alternative":[{"slug":"autoresearch","why":"Same lineage — Karpathy's autoresearch loop for coding agents. The skill is a single-branch keep/discard hill-climb; evo adds tree search, parallel worktree subagents, shared failure traces, gates and a dashboard. Start with the skill, graduate to evo.","dir":"out","confidence":0.85,"name":"autoresearch"},{"slug":"sia","why":"Both run autonomous improve-against-a-benchmark loops with paper-adjacent rigor. evo points coding agents at YOUR codebase (tree search over commits); SIA generates a task-specific agent and evolves the agent itself — harness and weights.","dir":"in","confidence":0.7,"name":"sia"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/evo-hq/evo"},"exxperts":{"name":"exxperts","owner":"EXXETA","slug":"exxperts","stars":225,"image":"https://raw.githubusercontent.com/EXXETA/exxperts/main/docs/assets/exxperts-logo.png","avatar":"https://avatars.githubusercontent.com/u/2116847?v=4&s=96","forks":21,"language":"TypeScript","license":"NOASSERTION","updated":"4 days ago","topics":["memory","agents","local"],"summary":"Local-first agentic runtime with persistent AI rooms and approval-gated memory: every memory write needs your OK; rooms, KB and artifacts are plain files on disk.","curator_note":"Use it when you want a personal AI that remembers you but on your terms — every memory write passes an approval gate, and all state is auditable files under ~/.exxperts. Two surfaces (web app + CLI/TUI) share the same rooms. NOT a chat-over-docs RAG frontend (that's Open WebUI territory) and not a coding agent — rooms never get an unrestricted shell. Single-user, local-only by design; and mind the PolyForm Noncommercial license: commercial use needs a separate deal with Exxeta.","edges":[{"to":"core","type":"alternative","why":"Both are self-hosted 'personal AI with persistent memory' products, but they pick opposite defaults: core watches your apps and acts autonomously within guardrails, exxperts gates every single memory write behind your explicit approval.","confidence":0.7,"status":"approved"},{"to":"litellm","type":"complements","why":"exxperts' AI setup has a first-class 'OpenAI-compatible gateway' path and names a company LiteLLM proxy as the canonical example — run rooms against your org's gateway instead of raw provider keys.","confidence":0.7,"status":"approved"},{"to":"vllm","type":"complements","why":"The same gateway setup explicitly supports a vLLM proxy as the model backend, so rooms can run entirely against self-served local models.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-19T13:40:10.262Z","linkCount":3,"related":{"complements":[{"slug":"litellm","why":"exxperts' AI setup has a first-class 'OpenAI-compatible gateway' path and names a company LiteLLM proxy as the canonical example — run rooms against your org's gateway instead of raw provider keys.","dir":"out","confidence":0.7,"name":"litellm"},{"slug":"vllm","why":"The same gateway setup explicitly supports a vLLM proxy as the model backend, so rooms can run entirely against self-served local models.","dir":"out","confidence":0.55,"name":"vLLM"}],"alternative":[{"slug":"core","why":"Both are self-hosted 'personal AI with persistent memory' products, but they pick opposite defaults: core watches your apps and acts autonomously within guardrails, exxperts gates every single memory write behind your explicit approval.","dir":"out","confidence":0.7,"name":"core"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/EXXETA/exxperts"},"fabro":{"name":"fabro","owner":"fabro-sh","slug":"fabro","stars":1442,"image":"https://raw.githubusercontent.com/fabro-sh/fabro/main/docs/public/logo/dark.svg","avatar":"https://avatars.githubusercontent.com/u/266531437?v=4&s=96","forks":155,"language":"Rust","license":"MIT","updated":"yesterday","topics":["coding","orchestration"],"summary":"Rust 'software factory' for coding agents: define the SDLC as a graph, agents execute it through verification gates, you intervene only at the stages that matter. Server, runs board, sandboxes.","curator_note":"For the team past the babysit-or-rubber-stamp dilemma: encode your process as a version-controlled workflow graph — implement with one model, cross-critique with another, gate on verification — and let the API server run it 24/7, locally or in isolated cloud sandboxes. CSS-like stylesheets for model routing is a genuinely clever cost lever. NOT a quick add-on: this is infrastructure with a server, a web wizard and a process philosophy — solo devs wanting lighter structure should look at terminal-first orchestrators first. Young project; the 'dark factory' ambition is ahead of the ecosystem's verification reality, so keep the gates strict.","edges":[{"to":"contrabass","type":"alternative","why":"Both run coding agents through gated, verified pipelines instead of a REPL. Contrabass is terminal-first and issue-driven (poll Linear/GitHub, worktrees, branch-advance checks); Fabro is a server with workflow graphs, ensemble models and cloud sandboxes.","confidence":0.65,"status":"approved"},{"to":"squad","type":"alternative","why":"Same goal — multi-agent software process with human gates — opposite weight classes: Squad coordinates AI CLIs through one-shot shell commands and SQLite with no daemon; Fabro runs a persistent server with a runs board and queued 24/7 execution.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.742Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"contrabass","why":"Both run coding agents through gated, verified pipelines instead of a REPL. Contrabass is terminal-first and issue-driven (poll Linear/GitHub, worktrees, branch-advance checks); Fabro is a server with workflow graphs, ensemble models and cloud sandboxes.","dir":"out","confidence":0.65,"name":"contrabass"},{"slug":"squad","why":"Same goal — multi-agent software process with human gates — opposite weight classes: Squad coordinates AI CLIs through one-shot shell commands and SQLite with no daemon; Fabro runs a persistent server with a runs board and queued 24/7 execution.","dir":"out","confidence":0.6,"name":"squad"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/fabro-sh/fabro"},"finance-skills":{"name":"finance-skills","owner":"himself65","slug":"finance-skills","stars":3054,"image":"https://raw.githubusercontent.com/fundamental-bottom/.github/main/profile/banner.png","avatar":"https://avatars.githubusercontent.com/u/14026360?v=4&s=96","forks":352,"language":"JavaScript","license":"MIT","updated":"3 days ago","topics":["finance","skills"],"summary":"Agent skills for financial analysis and trading on the agentskills.io open standard — installable into Claude Code and friends, with a documented demo site.","curator_note":"The skills-layer entry to finance agents: analysis and trading capabilities as portable, standard-format skills rather than a platform to adopt. The README leads with the right warning (educational, not financial advice) — and note the sponsorship: it's the open funnel for a commercial research product, so expect the deepest skills to live behind that door. Good on-ramp; audit each skill's data sources before trusting its numbers.","edges":[],"status":"approved","added":"2026-07-14T16:04:54.951Z","linkCount":1,"related":{"complements":[{"slug":"metatrader-mcp-server","why":"Skills and hands: finance-skills teaches Claude financial-analysis workflows on the agentskills standard; this MCP server is the execution layer that turns those conclusions into actual MT5 orders.","dir":"in","confidence":0.5,"name":"metatrader-mcp-server"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/himself65/finance-skills"},"flowsint":{"name":"flowsint","owner":"reconurge","slug":"flowsint","stars":7435,"image":"https://github.com/user-attachments/assets/01eb128e-bef4-486e-9276-c4da58f829ae","avatar":"https://avatars.githubusercontent.com/u/182449807?v=4&s=96","forks":934,"language":"TypeScript","license":"Apache-2.0","updated":"23 days ago","topics":["security"],"summary":"Self-hosted OSINT investigation platform — explore entities on a Neo4j-backed visual graph and expand them with 30+ enrichers: DNS/WHOIS/subdomains, breach checks, Maigret, crypto wallets.","curator_note":"The open-source Maltego-shaped option: drop an email, domain, wallet or username on the canvas and chain enrichers to grow the relationship graph — privacy-first (everything stays on your machine, API keys in an encrypted vault), team-deployable with one exposed port, and the UI stays responsive at thousands of nodes. Use it for structured recon and OSINT investigations: journalists, security analysts, fraud teams. NOT an AI tool — there is no LLM in the loop; the n8n connector is the automation escape hatch, and its place in this catalog is as the investigation surface your agents and workflows can feed. Early development, tests admittedly incomplete; read ETHICS.md — the responsible-use line is enforced culturally, not technically.","edges":[{"to":"allama","type":"complements","why":"Adjacent stages of self-hosted security operations: Flowsint is the investigation canvas where an analyst maps who/what during recon or fraud cases; allama automates the detection-and-response side with SOAR playbooks and triage agents.","confidence":0.5,"status":"approved"},{"to":"surfsense","type":"complements","why":"Both do open-source intelligence, at different tempos: surfsense's scheduled agents monitor live sources into cited briefs; Flowsint is where you take a lead from those briefs and investigate it hands-on as an entity graph.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-14T22:42:04.134Z","linkCount":2,"related":{"complements":[{"slug":"allama","why":"Adjacent stages of self-hosted security operations: Flowsint is the investigation canvas where an analyst maps who/what during recon or fraud cases; allama automates the detection-and-response side with SOAR playbooks and triage agents.","dir":"out","confidence":0.5,"name":"allama"},{"slug":"surfsense","why":"Both do open-source intelligence, at different tempos: surfsense's scheduled agents monitor live sources into cited briefs; Flowsint is where you take a lead from those briefs and investigate it hands-on as an entity graph.","dir":"out","confidence":0.45,"name":"SurfSense"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/reconurge/flowsint"},"forkd":{"name":"forkd","owner":"deeplethe","slug":"forkd","stars":2720,"image":"https://raw.githubusercontent.com/deeplethe/forkd/main/docs/logo.svg","avatar":"https://avatars.githubusercontent.com/u/282846749?v=4&s=96","forks":212,"language":"Rust","license":"Apache-2.0","updated":"2 days ago","topics":["agents","local"],"summary":"fork() for agent microVMs: children fork copy-on-write from a warm Firecracker parent — 100 KVM-isolated VMs in ~100ms, live-VM branching in ~56ms, portable snapshots from a hub.","curator_note":"The sandbox primitive agent fan-out actually wants: instead of cold-booting a kernel per child, a parent VM boots once with your runtime warm (deps imported, model loaded, JVM JITed) and children mmap its memory copy-on-write — so spawning 100 isolated environments costs ~100ms, and v0.4 can BRANCH a live running VM. Tree-search agents, parallel experiment runners and RL rollouts are the natural fits. NOT a managed platform: Linux/KVM only, you operate it, and there's no E2B-style API compatibility — this is a runtime primitive you build on, not a sandbox service you call.","edges":[{"to":"cubesandbox","type":"alternative","why":"Both are self-hosted Firecracker-class microVM sandboxes for agents. CubeSandbox sells a platform (E2B-compatible API, sub-60ms cold boots); forkd sells a primitive — copy-on-write fork from a warm parent, which wins precisely when 100 children share one expensive warm state.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.684Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"cubesandbox","why":"Both are self-hosted Firecracker-class microVM sandboxes for agents. CubeSandbox sells a platform (E2B-compatible API, sub-60ms cold boots); forkd sells a primitive — copy-on-write fork from a warm parent, which wins precisely when 100 children share one expensive warm state.","dir":"out","confidence":0.75,"name":"CubeSandbox"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/deeplethe/forkd"},"free-claude-code":{"name":"free-claude-code","owner":"Alishahryar1","slug":"free-claude-code","stars":41759,"image":"https://raw.githubusercontent.com/Alishahryar1/free-claude-code/main/assets/pic.png","avatar":"https://avatars.githubusercontent.com/u/20476625?v=4&s=96","forks":6797,"language":"Python","license":"MIT","updated":"2 days ago","topics":["gateway"],"summary":"Provider-backed proxy that runs Claude Code, Codex or Pi on 25 cloud and local providers — fcc-* launchers, local Admin UI with validation, per-tier model routing, IDE/Discord/Telegram hookups.","curator_note":"40k stars because it nails onboarding: one installer, `fcc-server` + `fcc-claude`, paste an NVIDIA NIM key and Claude Code runs for free — the Admin UI validates providers, and the agents' NATIVE /model pickers see the whole FCC catalog. Tier routing is the power feature: send Opus traffic to Kimi, Sonnet to an OpenRouter free route, Haiku to a local LM Studio model. Local models (Ollama, llama.cpp, LM Studio) are first-class. NOT a quota-stacker — you pick a provider rather than multiplex free tiers with failover (that's freellmapi's job), and expect capability drift vs real Anthropic models on complex agentic work. Discord/Telegram remote sessions with voice-note transcription are a genuinely odd, genuinely useful extra.","edges":[{"to":"freellmapi","type":"alternative","why":"Same slot — run coding CLIs on non-Anthropic providers for free — different philosophy: FCC is launcher + Admin UI where you pick and validate one provider (with per-tier overrides); FreeLLMAPI stacks 18 providers' free tiers behind a router with automatic failover and quota tracking.","confidence":0.75,"status":"approved"},{"to":"9router","type":"alternative","why":"Both are self-hosted gateways pointing Claude Code/Codex-class CLIs at many providers. 9router leans on routing policy (subscription→cheap→free fallback, token compression); FCC leans on UX — launchers, validation UI, native model-picker integration, IDE configs.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.470Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"freellmapi","why":"Same slot — run coding CLIs on non-Anthropic providers for free — different philosophy: FCC is launcher + Admin UI where you pick and validate one provider (with per-tier overrides); FreeLLMAPI stacks 18 providers' free tiers behind a router with automatic failover and quota tracking.","dir":"out","confidence":0.75,"name":"freellmapi"},{"slug":"9router","why":"Both are self-hosted gateways pointing Claude Code/Codex-class CLIs at many providers. 9router leans on routing policy (subscription→cheap→free fallback, token compression); FCC leans on UX — launchers, validation UI, native model-picker integration, IDE configs.","dir":"out","confidence":0.65,"name":"9router"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Alishahryar1/free-claude-code"},"freellmapi":{"name":"freellmapi","owner":"tashfeenahmed","slug":"freellmapi","stars":16733,"image":"https://raw.githubusercontent.com/tashfeenahmed/freellmapi/main/repo-assets/fallback-chain.png","avatar":"https://avatars.githubusercontent.com/u/9307356?v=4&s=96","forks":2424,"language":"TypeScript","license":"MIT","updated":"4 days ago","topics":["gateway"],"summary":"OpenAI-compatible proxy that stacks free tiers of 18 LLM providers (~1.7B tokens/mo) behind one /v1 — smart routing, failover, per-key quota tracking; Claude Code and Codex shims included.","curator_note":"Use it for personal experimentation when you want real working inference for $0: free tiers stacked across Google/Groq/Cerebras/Mistral/etc., a Thompson-sampling router with 20-attempt failover, and Anthropic Messages + Responses API shims so Claude Code and Codex CLI run straight against the free pool. Also does embeddings, images, TTS, and an MCP server agents can query for live model availability. NOT for production or teams — single-user by design, no billing, and several provider free tiers are ToS-gray (NVIDIA's is eval-only; the repo itself says personal experimentation only). Freemium catch: new models reach free installs 30 days late unless you pay $19/yr for the live catalog.","edges":[{"to":"9router","type":"alternative","why":"Both are self-hosted OpenAI-compatible gateways that multiplex providers with automatic fallback for coding CLIs. 9router optimizes paid usage (subscription→cheap→free routing, token compression); FreeLLMAPI exists purely to stack and stay under 18 providers' free-tier caps.","confidence":0.75,"status":"approved"},{"to":"lynkr","type":"alternative","why":"Same slot — a local proxy in front of Claude Code/Codex to cut LLM spend. lynkr's lever is compression, semantic caching and tier-routing of paid traffic; FreeLLMAPI's is routing everything to free-tier quota across providers.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T17:53:59.979Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"9router","why":"Both are self-hosted OpenAI-compatible gateways that multiplex providers with automatic fallback for coding CLIs. 9router optimizes paid usage (subscription→cheap→free routing, token compression); FreeLLMAPI exists purely to stack and stay under 18 providers' free-tier caps.","dir":"out","confidence":0.75,"name":"9router"},{"slug":"free-claude-code","why":"Same slot — run coding CLIs on non-Anthropic providers for free — different philosophy: FCC is launcher + Admin UI where you pick and validate one provider (with per-tier overrides); FreeLLMAPI stacks 18 providers' free tiers behind a router with automatic failover and quota tracking.","dir":"in","confidence":0.75,"name":"free-claude-code"},{"slug":"litellm","why":"Both expose many providers behind one OpenAI-compatible /v1 endpoint; freellmapi stacks free tiers, LiteLLM targets production apps and teams.","dir":"in","confidence":0.6,"name":"litellm"},{"slug":"lynkr","why":"Same slot — a local proxy in front of Claude Code/Codex to cut LLM spend. lynkr's lever is compression, semantic caching and tier-routing of paid traffic; FreeLLMAPI's is routing everything to free-tier quota across providers.","dir":"out","confidence":0.55,"name":"Lynkr"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tashfeenahmed/freellmapi"},"fusion":{"name":"Fusion","owner":"Runfusion","slug":"fusion","stars":999,"image":"https://raw.githubusercontent.com/Runfusion/fusion/main/demo/assets/fusion-logo-orange.svg","avatar":"https://avatars.githubusercontent.com/u/275577877?v=4&s=96","forks":128,"language":"TypeScript","license":"MIT","updated":"yesterday","topics":["coding","orchestration"],"summary":"A multi-agent software factory: describe a task, agents plan (PROMPT.md), build, review and merge in isolated worktrees — kanban + graph board, missions, agent chat rooms, any model. Early preview.","curator_note":"The most ambitious entry in the agent-factory wave: visual workflow authoring (plan→execute→review graphs you can edit), per-task oversight levels from observe to autonomous with human gates on merges, a multi-node mesh (fleet on a server, steered from your phone), importable 'agent companies' (440+ pre-built agents), and a Command Center with real fleet telemetry. Genuinely MIT and shipping weekly. When NOT: it wears its 'early preview' badge honestly — breadth currently outruns depth, so expect rough edges; if you want minimal-machinery unattended runs from an issue tracker, contrabass is the leaner tool, and a single interactive session needs none of this.","edges":[{"to":"contrabass","type":"alternative","why":"Same job — dispatch coding agents into isolated git worktrees with review gates — opposite philosophies: Contrabass is a lean terminal orchestrator fed by Linear/GitHub issues; Fusion is a full board-driven factory with visual workflows, missions and agent chat.","confidence":0.65,"status":"approved"},{"to":"auto-company","type":"alternative","why":"Both run autonomous 'AI companies' on your hardware: Auto-Company ships 14 expert-modeled agents ideating and shipping 24/7; Fusion makes the company importable and inspectable — org charts, token share per agent, multi-agent chat rooms.","confidence":0.6,"status":"approved"},{"to":"alook","type":"alternative","why":"Both are collaboration layers that turn local coding agents into a coordinated team on a kanban board; alook leans on per-agent email and shared memory, Fusion on planned workflows, worktree isolation and merge gates.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T16:07:34.753Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"contrabass","why":"Same job — dispatch coding agents into isolated git worktrees with review gates — opposite philosophies: Contrabass is a lean terminal orchestrator fed by Linear/GitHub issues; Fusion is a full board-driven factory with visual workflows, missions and agent chat.","dir":"out","confidence":0.65,"name":"contrabass"},{"slug":"auto-company","why":"Both run autonomous 'AI companies' on your hardware: Auto-Company ships 14 expert-modeled agents ideating and shipping 24/7; Fusion makes the company importable and inspectable — org charts, token share per agent, multi-agent chat rooms.","dir":"out","confidence":0.6,"name":"Auto-Company"},{"slug":"alook","why":"Both are collaboration layers that turn local coding agents into a coordinated team on a kanban board; alook leans on per-agent email and shared memory, Fusion on planned workflows, worktree isolation and merge gates.","dir":"out","confidence":0.55,"name":"alook"},{"slug":"helmor","why":"Factory vs workbench: Fusion runs autonomous plan-build-review-merge missions on a kanban board; Helmor keeps you as the lead dispatching tasks into workspaces.","dir":"in","confidence":0.55,"name":"helmor"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Runfusion/fusion"},"garak":{"name":"garak","owner":"NVIDIA","slug":"garak","stars":8539,"image":"https://i.imgur.com/8Dxf45N.png","avatar":"https://avatars.githubusercontent.com/u/1728152?v=4&s=96","forks":1109,"language":"Python","license":"Apache-2.0","updated":"3 days ago","topics":["security"],"summary":"NVIDIA's LLM vulnerability scanner: nmap-style probing for jailbreaks, prompt injection, data leakage, toxicity and hallucination across dozens of model endpoints.","curator_note":"Run it before anything LLM-shaped ships: static, dynamic and adaptive probes for jailbreaks, injection, leakage and toxicity, with reports that name the failing probe and prompt — nmap for language models. Speaks to HF, OpenAI-compatible, Ollama and REST endpoints, so it tests what you actually deploy. NOT a guardrail — it finds holes, it doesn't plug them (pair with a runtime defense); a full probe run takes hours, and a clean scan tests the MODEL, not your product logic around it. Authorized targets only.","edges":[{"to":"ollama","type":"complements","why":"garak ships an Ollama generator — point the scanner at your locally served model and red-team it before exposing it to users.","confidence":0.65,"status":"approved"},{"to":"nemo-guardrails","type":"complements","why":"Attack and defend: garak red-teams the model to find jailbreaks and injection holes, NeMo Guardrails is the runtime rail you deploy to police them — scan, patch rails, re-scan.","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.050Z","linkCount":4,"related":{"complements":[{"slug":"nemo-guardrails","why":"Attack and defend: garak red-teams the model to find jailbreaks and injection holes, NeMo Guardrails is the runtime rail you deploy to police them — scan, patch rails, re-scan.","dir":"out","confidence":0.8,"name":"Guardrails"},{"slug":"ollama","why":"garak ships an Ollama generator — point the scanner at your locally served model and red-team it before exposing it to users.","dir":"out","confidence":0.65,"name":"Ollama"}],"alternative":[{"slug":"deepteam","why":"Same job — adversarial testing of LLM systems. garak is the nmap-style scanner you point at a model endpoint; DeepTeam is the framework you embed in your eval suite, with attack simulation and guardrails in one loop.","dir":"in","confidence":0.8,"name":"deepteam"},{"slug":"giskard-oss","why":"Same scanning job, different shapes: garak is NVIDIA's nmap-style standalone prober across dozens of model endpoints; giskard-scan is a library-first scanner designed to live inside your Python test suite next to your evals.","dir":"in","confidence":0.6,"name":"giskard-oss"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/NVIDIA/garak"},"giskard-oss":{"name":"giskard-oss","owner":"Giskard-AI","slug":"giskard-oss","stars":5706,"image":"https://raw.githubusercontent.com/Giskard-AI/giskard-oss/main/readme/scan_updated.gif","avatar":"https://avatars.githubusercontent.com/u/71782571?v=4&s=96","forks":505,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["evals","security"],"summary":"Giskard v3: modular Python evals and red-teaming for agentic systems — scenario-based checks with LLM-as-judge, plus an automatic vulnerability scanner across OWASP LLM Top-10 categories.","curator_note":"One of the OG names in LLM testing, freshly rewritten: v3 drops the heavyweight monolith for focused packages — giskard-checks (scenario API for multi-turn agent evals: groundedness, conformity, LLM-judge) and giskard-scan (generates adversarial suites from a plain-language description of your agent, prompt-injection probes included). Async-first, wraps anything callable. Mind the transition, though: v3 is beta, RAG evaluation (RAGET) hasn't been ported yet, and v2 — where scan and RAG eval are mature — is no longer maintained. Python 3.12+ only, and libraries on giskard-core send aggregated telemetry (opt-out documented). Pick ragas for pure RAG metrics; Giskard earns its place on multi-turn agent scenarios plus red-teaming in one stack.","edges":[{"to":"ragas","type":"alternative","why":"Overlapping LLM-evaluation job: ragas is the specialist for RAG pipeline metrics; Giskard v3 covers multi-turn agent scenarios with built-in checks and LLM-as-judge, with RAG evaluation still pending its v3 port.","confidence":0.65,"status":"approved"},{"to":"deepteam","type":"alternative","why":"Both red-team LLM systems with framework ergonomics: deepteam ships 50+ documented vulnerabilities and matching guardrails; giskard-scan auto-generates adversarial suites from a plain-language agent description and pairs with the same stack's eval checks.","confidence":0.7,"status":"approved"},{"to":"garak","type":"alternative","why":"Same scanning job, different shapes: garak is NVIDIA's nmap-style standalone prober across dozens of model endpoints; giskard-scan is a library-first scanner designed to live inside your Python test suite next to your evals.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-16T15:39:30.817Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"deepteam","why":"Both red-team LLM systems with framework ergonomics: deepteam ships 50+ documented vulnerabilities and matching guardrails; giskard-scan auto-generates adversarial suites from a plain-language agent description and pairs with the same stack's eval checks.","dir":"out","confidence":0.7,"name":"deepteam"},{"slug":"ragas","why":"Overlapping LLM-evaluation job: ragas is the specialist for RAG pipeline metrics; Giskard v3 covers multi-turn agent scenarios with built-in checks and LLM-as-judge, with RAG evaluation still pending its v3 port.","dir":"out","confidence":0.65,"name":"Ragas"},{"slug":"garak","why":"Same scanning job, different shapes: garak is NVIDIA's nmap-style standalone prober across dozens of model endpoints; giskard-scan is a library-first scanner designed to live inside your Python test suite next to your evals.","dir":"out","confidence":0.6,"name":"garak"},{"slug":"deepeval","why":"Both are modular Python eval frameworks with LLM-as-judge and safety scanning. Giskard leans scenario-based checks and OWASP-mapped vulnerability scanning; DeepEval leans metric breadth and pytest-in-CI ergonomics.","dir":"in","confidence":0.6,"name":"deepeval"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Giskard-AI/giskard-oss"},"govctl":{"name":"govctl","owner":"govctl-org","slug":"govctl","stars":239,"image":"https://raw.githubusercontent.com/govctl-org/govctl/main/assets/logo.svg","avatar":"https://avatars.githubusercontent.com/u/255486183?v=4&s=96","forks":11,"language":"Rust","license":"MIT","updated":"3 days ago","topics":["coding"],"summary":"Governance-as-code CLI for AI-assisted development: prompts and patches become RFCs, ADRs and work items with executable verification gates — reviewable, traceable, phase-gated delivery.","curator_note":"The uncomfortable truth it addresses: with AI coding, 'done' drifts toward 'the agent stopped typing'. govctl makes governed artifacts part of the loop — RFCs state what must be true, ADRs record why, work items carry acceptance criteria, and verification guards gate completion. If your team ships AI-generated code into anything regulated or long-lived, this is the missing control plane. NOT lightweight: it's process-as-code, and process you don't enforce becomes decoration; also young (~240 stars) — expect to shape it as much as use it.","edges":[{"to":"loop-engineering","type":"complements","why":"Two halves of disciplined AI delivery: loop-engineering designs the gated control loops that drive coding agents; govctl supplies the governance artifacts and executable completion gates those loops should answer to.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T14:52:59.792Z","linkCount":2,"related":{"complements":[{"slug":"loop-engineering","why":"Two halves of disciplined AI delivery: loop-engineering designs the gated control loops that drive coding agents; govctl supplies the governance artifacts and executable completion gates those loops should answer to.","dir":"out","confidence":0.6,"name":"loop-engineering"}],"alternative":[{"slug":"cwc-long-running-agents","why":"Same core move — structural gates instead of polite prompts: govctl makes governance a CLI with verification gates; these hooks make 'done' unclaimable without opened evidence.","dir":"in","confidence":0.5,"name":"cwc-long-running-agents"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/govctl-org/govctl"},"gym-anything":{"name":"gym-anything","owner":"cmu-l3","slug":"gym-anything","stars":263,"avatar":"https://avatars.githubusercontent.com/u/151781139?v=4&s=96","forks":38,"language":"Shell","license":"MIT","updated":"yesterday","topics":["evals","agents"],"summary":"CMU framework that turns real software — browsers, IDEs, EMRs, CAD — into standardized agent environments: start the app, hand the agent a task, score it with automatic verifiers.","curator_note":"The missing middle layer for computer-use agents: a Core runtime (environment lifecycle, actions, observations, verifiers), a benchmark collection wrapping real applications like Moodle, and reference agents (Claude, Gemini, Qwen, Kimi) — three parts connected by contracts, each independently replaceable, one CLI to run it all. The `doctor` setup command and environment caching signal real operational care. Use it to evaluate your CUA agent beyond browser-only benchmarks, or to gym-ify internal software by adding a task folder with a setup script and a checker. NOT an agent framework — the agents are references, bring your own. Young (arXiv 2026, CMU L3): environment coverage is the current bottleneck.","edges":[{"to":"browser-use","type":"complements","why":"Build with one, grade with the other: browser-use is the standard library for constructing agents that drive real software; Gym-Anything supplies the standardized environments, tasks and automatic verifiers to benchmark exactly that kind of agent.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T23:32:32.800Z","linkCount":2,"related":{"complements":[{"slug":"browser-use","why":"Build with one, grade with the other: browser-use is the standard library for constructing agents that drive real software; Gym-Anything supplies the standardized environments, tasks and automatic verifiers to benchmark exactly that kind of agent.","dir":"out","confidence":0.5,"name":"browser-use"},{"slug":"opensandbox","why":"gym-anything defines the environments and verifiers; OpenSandbox provides the execution layer — RL training and agent evaluation are both first-class scenarios in its API.","dir":"in","confidence":0.5,"name":"OpenSandbox"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/cmu-l3/gym-anything"},"h2o-llmstudio":{"name":"h2o-llmstudio","owner":"h2oai","slug":"h2o-llmstudio","stars":5042,"image":"https://user-images.githubusercontent.com/1069138/233859311-32aa1f8c-4d68-47ac-8cd9-9313171ff9f9.png","avatar":"https://avatars.githubusercontent.com/u/1402695?v=4&s=96","forks":538,"language":"Python","license":"Apache-2.0","updated":"2 days ago","topics":["training"],"summary":"H2O's no-code GUI and framework for fine-tuning LLMs — LoRA, 8-bit, DPO and experiment tracking behind a web UI, with CLI and Docker paths for the same configs.","curator_note":"Fine-tuning for teams where not everyone writes training loops: pick a base model, upload data, tune LoRA/quantization/DPO hyperparameters in a web UI, compare runs visually, export to the Hub. The CLI runs the same configs headless, so GUI experiments graduate to scripted jobs. NOT for RL post-training at scale (its RL is experimental — that's verl/slime territory) and not for frontier-size models; and it's opinionated toward the H2O ecosystem — if your workflow is already Hub-native code, a library fits better than a studio.","edges":[{"to":"slime","type":"alternative","why":"Both post-train LLMs, from opposite ends: slime is Megatron-scale RL for frontier runs, LLM Studio is no-code LoRA/DPO fine-tuning on models a single node can hold.","confidence":0.5,"status":"approved"},{"to":"trl","type":"alternative","why":"The same LoRA/DPO fine-tuning jobs behind different interfaces: TRL is the code-first library for Hub-native workflows, LLM Studio the no-code GUI for teams that don't write training loops.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.070Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"trl","why":"The same LoRA/DPO fine-tuning jobs behind different interfaces: TRL is the code-first library for Hub-native workflows, LLM Studio the no-code GUI for teams that don't write training loops.","dir":"out","confidence":0.75,"name":"trl"},{"slug":"llamafactory","why":"Both put a GUI on fine-tuning; LLM Studio optimizes for the no-code experience, LlamaFactory for maximum model/method coverage with the GUI as one of several front doors.","dir":"in","confidence":0.6,"name":"LlamaFactory"},{"slug":"mlx-lora-studio","why":"Same promise — fine-tune without writing code — different homes: H2O LLM Studio is a web UI/Docker framework for GPU boxes; MLX LoRA Studio is a native macOS app for the machine on your desk.","dir":"in","confidence":0.6,"name":"MLX-LoRA-Studio"},{"slug":"slime","why":"Both post-train LLMs, from opposite ends: slime is Megatron-scale RL for frontier runs, LLM Studio is no-code LoRA/DPO fine-tuning on models a single node can hold.","dir":"out","confidence":0.5,"name":"slime"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/h2oai/h2o-llmstudio"},"handy":{"name":"Handy","owner":"cjpais","slug":"handy","stars":27234,"image":"https://www.raycast.com/mattiacolombomc/handy/install_button@2x.png?v=1.1","avatar":"https://avatars.githubusercontent.com/u/1559480?v=4&s=96","forks":2352,"language":"Rust","license":"MIT","updated":"2 days ago","topics":["voice","local"],"summary":"Push-to-talk offline dictation: hotkey, speak, text lands in whatever field has focus. Whisper or Parakeet fully on-device; cross-platform Rust/Tauri, built to be forked.","curator_note":"The one-job dictation tool done right: a global hotkey, Silero VAD, local Whisper (GPU) or Parakeet (CPU), and the transcript pastes into any app — nothing leaves the machine. The 'most forkable, not best' positioning is honest and is the reason to pick it: small Rust/Tauri codebase, CLI flags for scripting a running instance, Homebrew/winget installs, a Raycast extension. NOT for meetings (no diarization, no summaries — that's Meetily) and not a voice suite (no TTS/cloning — that's Voicebox). Known rough edges: Whisper crashes on some Windows/Linux configs, and Wayland needs wtype/dotool for text injection.","edges":[{"to":"voicebox","type":"alternative","why":"Both give you system-wide local Whisper dictation via a hotkey in a Tauri app. Voicebox bundles it into a full voice studio (cloning, 7 TTS engines, MCP agent voices); Handy is dictation only — lighter, simpler, and deliberately forkable.","confidence":0.8,"status":"approved"},{"to":"meetily","type":"alternative","why":"Same fully-local Whisper transcription core, different job: Handy dictates what YOU say into the focused text field; Meetily captures system audio of meetings with diarization and Ollama summaries. Pick by whether the audio is your voice or a call.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-20T09:28:52.168Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"voicebox","why":"Both give you system-wide local Whisper dictation via a hotkey in a Tauri app. Voicebox bundles it into a full voice studio (cloning, 7 TTS engines, MCP agent voices); Handy is dictation only — lighter, simpler, and deliberately forkable.","dir":"out","confidence":0.8,"name":"voicebox"},{"slug":"meetily","why":"Same fully-local Whisper transcription core, different job: Handy dictates what YOU say into the focused text field; Meetily captures system audio of meetings with diarization and Ollama summaries. Pick by whether the audio is your voice or a call.","dir":"out","confidence":0.55,"name":"meetily"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/cjpais/handy"},"headroom":{"name":"headroom","owner":"headroomlabs-ai","slug":"headroom","stars":61328,"image":"https://raw.githubusercontent.com/headroomlabs-ai/headroom/main/HeadroomDemo-Fast.gif","avatar":"https://avatars.githubusercontent.com/u/294291659?v=4&s=96","forks":4606,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["coding","gateway"],"summary":"Context compression layer for agents: squeezes tool outputs, logs, files and RAG chunks 20-95% before the LLM — reversible, local-first; library, proxy, one-command agent wrap, or MCP.","curator_note":"The dedicated answer to context bloat: content-aware compressors route JSON, logs and code differently (60-95% on JSON, 15-20% on real coding sessions), originals stay cached so the model can retrieve what compression dropped — the reversibility is what makes aggressive ratios safe. Adoption cost is near zero: `headroom wrap claude` and you're running, or use it as proxy/library/MCP. Prompt-cache-aware alignment avoids torching your cache hit rate. NOT free lunch: a lossy-in-context layer between agent and model is another thing to debug when the model 'misses' something — budget for retrieval round-trips — and the 20% coding figure is the honest number, not the 95% headline.","edges":[{"to":"9router","type":"alternative","why":"Same wire, same goal — cut coding-agent token spend at a local proxy. 9router is a multi-provider routing gateway with tool_result compression as a feature; Headroom is compression-first with content-aware routers, reversibility and no routing opinions.","confidence":0.6,"status":"approved"},{"to":"tokensave","type":"complements","why":"Two layers of the same diet: TokenSave stops tokens at the source (indexed code queries instead of grep dumps), Headroom compresses whatever still flows through. Stackable — one shrinks what the agent asks for, the other what it receives.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-20T10:48:48.753Z","linkCount":3,"related":{"complements":[{"slug":"tokensave","why":"Two layers of the same diet: TokenSave stops tokens at the source (indexed code queries instead of grep dumps), Headroom compresses whatever still flows through. Stackable — one shrinks what the agent asks for, the other what it receives.","dir":"out","confidence":0.55,"name":"tokensave"},{"slug":"ponytail","why":"Two ends of the same token diet: Headroom compresses what the agent reads, ponytail shrinks what it writes — less generated code is fewer output tokens and less context in every later turn. Stack them.","dir":"in","confidence":0.55,"name":"ponytail"}],"alternative":[{"slug":"9router","why":"Same wire, same goal — cut coding-agent token spend at a local proxy. 9router is a multi-provider routing gateway with tool_result compression as a feature; Headroom is compression-first with content-aware routers, reversibility and no routing opinions.","dir":"out","confidence":0.6,"name":"9router"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/headroomlabs-ai/headroom"},"helmor":{"name":"helmor","owner":"dohooo","slug":"helmor","stars":1276,"image":"https://raw.githubusercontent.com/dohooo/helmor/main/src/assets/helmor-logo-light.png","avatar":"https://avatars.githubusercontent.com/u/32405058?v=4&s=96","forks":115,"language":"TypeScript","license":"Apache-2.0","updated":"2 days ago","topics":["coding","orchestration"],"summary":"Local-first desktop workbench for orchestrating coding agents: per-repo workspaces, task dispatch, live status and runnable actions — GUI plus CLI, Apache-2.0.","curator_note":"Workbench, not factory: you stay the lead and dispatch agents into per-repo workspaces, watching status like a build queue. Cleaner mental model than kanban swarms for a solo dev. Skip for headless issue-queue automation (contrabass) or if you live inside one OpenCode session (codenomad).","edges":[{"to":"codenomad","type":"alternative","why":"Both local desktop workspaces for coding-agent sessions: codenomad is an OpenCode-deep cockpit; Helmor is agent-agnostic with per-repo workspaces and dispatchable actions.","confidence":0.6,"status":"approved"},{"to":"fusion","type":"alternative","why":"Factory vs workbench: Fusion runs autonomous plan-build-review-merge missions on a kanban board; Helmor keeps you as the lead dispatching tasks into workspaces.","confidence":0.55,"status":"approved"},{"to":"contrabass","type":"alternative","why":"Contrabass drains an issue queue headlessly with verification; Helmor is the interactive GUI take on the same orchestrate-local-coding-agents job.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.548Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"omnigent","why":"Both orchestrate coding agents from a desktop control plane; Helmor is local-first per-repo workspaces, Omnigent adds multi-device sessions, policies and cloud sandboxes.","dir":"in","confidence":0.8,"name":"omnigent"},{"slug":"codenomad","why":"Both local desktop workspaces for coding-agent sessions: codenomad is an OpenCode-deep cockpit; Helmor is agent-agnostic with per-repo workspaces and dispatchable actions.","dir":"out","confidence":0.6,"name":"CodeNomad"},{"slug":"fusion","why":"Factory vs workbench: Fusion runs autonomous plan-build-review-merge missions on a kanban board; Helmor keeps you as the lead dispatching tasks into workspaces.","dir":"out","confidence":0.55,"name":"Fusion"},{"slug":"contrabass","why":"Contrabass drains an issue queue headlessly with verification; Helmor is the interactive GUI take on the same orchestrate-local-coding-agents job.","dir":"out","confidence":0.5,"name":"contrabass"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/dohooo/helmor"},"hindsight":{"name":"hindsight","owner":"vectorize-io","slug":"hindsight","stars":18716,"image":"https://raw.githubusercontent.com/vectorize-io/hindsight/main/hindsight-docs/static/img/hindsight-github-banner.png","avatar":"https://avatars.githubusercontent.com/u/81837244?v=4&s=96","forks":1161,"language":"Python","license":"MIT","updated":"yesterday","topics":["memory","agents"],"summary":"Agent memory that learns, not just recalls: retain/recall/reflect API over Postgres, SOTA on LongMemEval. Self-host via Docker with UI; Python/TS clients, any LLM provider.","curator_note":"Pick it when you want a deployable memory *service* whose pitch is learning — agents that get better over time, not a transcript search. The LongMemEval lead was independently reproduced (Virginia Tech, Washington Post), which is more than most memory vendors offer, and the LLM side is pluggable down to Ollama/LM Studio for fully-local stacks. NOT an embedded library: you run a Docker service with Postgres and talk to it over HTTP — overkill for a single coding agent wanting session notes. The ™ and Hindsight Cloud signal a commercial trajectory; watch where the open/paid line lands.","edges":[{"to":"memmachine","type":"alternative","why":"Both self-hostable long-term memory services for agents behind Python/TS SDKs and REST. MemMachine splits episodic/profile/working memory with framework integrations; Hindsight bets on retain/recall/reflect and benchmark-topping learned memory.","confidence":0.7,"status":"approved"},{"to":"lightmem","type":"alternative","why":"Both chase LongMemEval leadership from opposite ends: LightMem is a research framework optimizing token cost via pre-compression; Hindsight is a packaged service optimizing accuracy. Benchmark rivals, different deployment realities.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-20T09:51:35.359Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"memmachine","why":"Both self-hostable long-term memory services for agents behind Python/TS SDKs and REST. MemMachine splits episodic/profile/working memory with framework integrations; Hindsight bets on retain/recall/reflect and benchmark-topping learned memory.","dir":"out","confidence":0.7,"name":"MemMachine"},{"slug":"agentic-context-engine","why":"Both sell 'agents that learn, not just remember' — opposite deployments: Hindsight is a self-hosted memory service (Docker + Postgres, retain/recall/reflect over HTTP); ACE is an embedded Python loop distilling strategies in-process. Service vs library.","dir":"in","confidence":0.65,"name":"agentic-context-engine"},{"slug":"lightmem","why":"Both chase LongMemEval leadership from opposite ends: LightMem is a research framework optimizing token cost via pre-compression; Hindsight is a packaged service optimizing accuracy. Benchmark rivals, different deployment realities.","dir":"out","confidence":0.6,"name":"LightMem"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/vectorize-io/hindsight"},"hivemind":{"name":"hivemind","owner":"activeloopai","slug":"hivemind","stars":1505,"image":"https://raw.githubusercontent.com/activeloopai/hivemind/main/docs/public/hivemind-logo.svg","avatar":"https://avatars.githubusercontent.com/u/34816118?v=4&s=96","forks":94,"language":"TypeScript","license":"Apache-2.0","updated":"3 days ago","topics":["memory","coding"],"summary":"Activeloop's shared brain for agent TEAMS: traces from Claude Code, Codex, Cursor & co become reusable skills every teammate's agent can execute — cloud-backed, 25% cheaper on LoCoMo.","curator_note":"The pitch is the org-level version of agent memory: one engineer's agent figures out the tricky migration on Monday, every agent on the team executes the pattern Tuesday. Auto-learning from traces across seven agent hosts, with real LoCoMo receipts (25% cheaper, 1.7x fewer tokens vs no shared memory). Reach for it when the unit of learning is the team, not the seat. NOT local-first: cloud-backed on Deeplake is the architecture AND the business model (Activeloop, YC) — traces of your engineers' sessions leave the machine, so clear it with whoever owns your IP policy before the whole team wires in.","edges":[{"to":"memsearch","type":"alternative","why":"Both distill agent sessions into shared, reusable skills across multiple agent hosts. memsearch is local-first, one machine, Markdown+Milvus; Hivemind is cloud-backed and team-scoped — your colleague's agent learns from yours.","confidence":0.75,"status":"approved"},{"to":"agentic-context-engine","type":"alternative","why":"Same learn-from-experience loop, different scope: ACE is an embedded Python engine distilling strategies for the agent you're building; Hivemind is a service layer distilling skills across every off-the-shelf agent your team runs.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.526Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"memsearch","why":"Both distill agent sessions into shared, reusable skills across multiple agent hosts. memsearch is local-first, one machine, Markdown+Milvus; Hivemind is cloud-backed and team-scoped — your colleague's agent learns from yours.","dir":"out","confidence":0.75,"name":"memsearch"},{"slug":"agentic-context-engine","why":"Same learn-from-experience loop, different scope: ACE is an embedded Python engine distilling strategies for the agent you're building; Hivemind is a service layer distilling skills across every off-the-shelf agent your team runs.","dir":"out","confidence":0.6,"name":"agentic-context-engine"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/activeloopai/hivemind"},"horizon":{"name":"Horizon","owner":"Thysrael","slug":"horizon","stars":8404,"image":"https://abroad.hellogithub.com/v1/widgets/recommend.svg?rid=7a4b606e28e4477998d35851cf4fdddf&claim_uid=rtjnLkYT7ziQJUG","avatar":"https://avatars.githubusercontent.com/u/72613958?v=4&s=96","forks":1246,"language":"Python","license":"MIT","updated":"2 days ago","topics":["web","agents"],"summary":"Your own AI news radar: monitors the sources you choose and generates daily briefings in English and Chinese — self-hosted, personal, scheduled.","curator_note":"The personal-scale answer to information overload: point it at your sources, get one daily briefing instead of forty tabs — bilingual EN/中文 out of the box, run on your own machine on a schedule. NOT a research platform: it monitors and summarizes, it doesn't build a queryable knowledge base or serve agents — when the briefing needs to become infrastructure, you've outgrown it.","edges":[{"to":"surfsense","type":"alternative","why":"Both watch live sources and produce briefs; SurfSense is the platform play (connectors, API/MCP, cited knowledge base for agents), Horizon the personal radar — one human, one daily briefing.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:54.981Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"surfsense","why":"Both watch live sources and produce briefs; SurfSense is the platform play (connectors, API/MCP, cited knowledge base for agents), Horizon the personal radar — one human, one daily briefing.","dir":"out","confidence":0.6,"name":"SurfSense"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Thysrael/horizon"},"html-anything":{"name":"html-anything","owner":"nexu-io","slug":"html-anything","stars":7921,"image":"https://raw.githubusercontent.com/nexu-io/html-anything/main/docs/assets/banner.png","avatar":"https://avatars.githubusercontent.com/u/263625318?v=4&s=96","forks":774,"language":"HTML","license":"Apache-2.0","updated":"10 days ago","topics":["skills","web"],"summary":"The agentic HTML editor: your local coding-agent CLI (9 auto-detected, zero API keys) writes magazine pages, decks, posters and tweet cards via 75 skills — sandboxed preview, 1-click export.","curator_note":"A sharp thesis executed well: in the agent era you don't hand-edit documents, so stop settling for Markdown — let the agent you already pay for write real HTML across 9 deliverable surfaces, previewed sandboxed, exported to WeChat/X/Zhihu/PNG in a click. Reusing your logged-in CLI session (no API key) is the friction-killer. NOT a design system: output quality rides entirely on the skill templates and your agent's taste, the export targets skew China-platform-first (WeChat, XHS, Zhihu), and the team openly funnels you toward their larger Open Design project — this is the focused spin-off, not the flagship.","edges":[{"to":"designer-skills","type":"alternative","why":"Both make coding agents produce designed deliverables via skill packs. designer-skills is 239 skills of design process for Claude Code/Gemini; HTML Anything is an editor product — 75 skills fused with preview, sandboxing and one-click publishing surfaces.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:19.009Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"designer-skills","why":"Both make coding agents produce designed deliverables via skill packs. designer-skills is 239 skills of design process for Claude Code/Gemini; HTML Anything is an editor product — 75 skills fused with preview, sandboxing and one-click publishing surfaces.","dir":"out","confidence":0.55,"name":"designer-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/nexu-io/html-anything"},"hyper-extract":{"name":"Hyper-Extract","owner":"yifanfeng97","slug":"hyper-extract","stars":3144,"image":"https://raw.githubusercontent.com/yifanfeng97/hyper-extract/main/docs/assets/logo/logo-horizontal.svg","avatar":"https://avatars.githubusercontent.com/u/15007108?v=4&s=96","forks":374,"language":"Python","license":"NOASSERTION","updated":"4 days ago","topics":["rag"],"summary":"Knowledge-extraction CLI: LLMs turn documents into structured graphs, hypergraphs and spatio-temporal knowledge — with an MCP server for agents and Obsidian vault export.","curator_note":"The pitch beyond ordinary KG extraction is the hypergraph: relations that connect MORE than two entities survive instead of being flattened into pairwise triples. One command per document, query the abstracts over MCP from Claude Desktop or your IDE, export to Obsidian wikilinks. NOT a graph database (it extracts, storage stays simple) and no standard license resolution at review time — verify before building on it; extraction quality tracks the LLM you plug in.","edges":[{"to":"omnigraph","type":"complements","why":"Extract with one, store and retrieve with the other: Hyper-Extract turns documents into structured graph knowledge; OmniGraph is the lakehouse runtime that serves graph+vector context to agent fleets.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.011Z","linkCount":1,"related":{"complements":[{"slug":"omnigraph","why":"Extract with one, store and retrieve with the other: Hyper-Extract turns documents into structured graph knowledge; OmniGraph is the lakehouse runtime that serves graph+vector context to agent fleets.","dir":"out","confidence":0.5,"name":"omnigraph"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/yifanfeng97/hyper-extract"},"insforge":{"name":"InsForge","owner":"InsForge","slug":"insforge","stars":12422,"image":"https://raw.githubusercontent.com/InsForge/insforge/main/assets/logo-dark.svg","avatar":"https://avatars.githubusercontent.com/u/198419463?v=4&s=96","forks":1085,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["coding"],"summary":"All-in-one open-source backend for agentic coding: Postgres, auth, storage, edge functions, model gateway and site hosting — your coding agent operates it over MCP.","curator_note":"Supabase reshaped for agents: the MCP server plus fetch-docs loop means your coding agent provisions auth/db/storage without you reading dashboards. Choose it when the agent is the operator; if humans run your backend, plain Supabase has the deeper ecosystem.","edges":[],"status":"approved","added":"2026-07-17T14:03:22.593Z","linkCount":0,"related":{"complements":[],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/InsForge/insforge"},"kami":{"name":"Kami","owner":"tw93","slug":"kami","stars":10054,"image":"https://raw.githubusercontent.com/tw93/kami/main/assets/images/logo.svg","avatar":"https://avatars.githubusercontent.com/u/8736212?v=4&s=96","forks":468,"language":"HTML","license":"MIT","updated":"5 days ago","topics":["skills"],"summary":"A document design system for AI agents: one constraint language and eight templates (plus a landing-page system) so agent-produced documents ship consistent instead of generic gray.","curator_note":"Built on a sharp diagnosis: AI can produce documents, but without a design system every session drifts into inconsistent, generic output. Kami is the constraint language — strict enough that agents produce shippable pages, simple enough that they follow it reliably. Part of tw93's agent trilogy (Kaku writes code, Waza drills habits, Kami delivers documents). NOT a generator itself — it's the discipline your agent works within; you still bring the agent and the content.","edges":[{"to":"designer-skills","type":"alternative","why":"Two routes to design-literate agents: designer-skills is a broad pack of design-craft skills; Kami is one strict constraint system focused on making every shipped document consistent.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.042Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"designer-skills","why":"Two routes to design-literate agents: designer-skills is a broad pack of design-craft skills; Kami is one strict constraint system focused on making every shipped document consistent.","dir":"out","confidence":0.5,"name":"designer-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tw93/kami"},"karakeep":{"name":"karakeep","owner":"karakeep-app","slug":"karakeep","stars":27663,"image":"https://raw.githubusercontent.com/karakeep-app/karakeep/main/screenshots/logo.png","avatar":"https://avatars.githubusercontent.com/u/202258986?v=4&s=96","forks":1361,"language":"TypeScript","license":"AGPL-3.0","updated":"6 days ago","topics":["storage"],"summary":"Self-hostable bookmark-everything app — links, notes, images, PDFs — with AI auto-tagging, OCR, full-text search, full-page archival, RSS auto-hoarding, and a CLI plus skills for LLM agents.","curator_note":"The mature choice for self-hosted read-it-later and data hoarding: browser extensions, iOS/Android apps, SSO, importers from Pocket/Omnivore/Linkwarden, monolith full-page archival against link rot, yt-dlp video capture. The AI is seasoning, not the dish — LLM auto-tagging and summarization (works with local Ollama) plus an official CLI and agent skills that turn it into a searchable personal knowledge store OpenClaw/Hermes-class agents can read and write. NOT a RAG framework or team wiki — it's a personal app with collaboration on lists at most. AGPL-3.0, and the stack is real: Meilisearch for search, Puppeteer workers for crawling.","edges":[{"to":"telegram-drive","type":"alternative","why":"Same shelf — self-hosted personal hoarding with an API agents can drive — different content: Telegram-Drive stores raw files on Telegram's servers; Karakeep hoards bookmarks, notes and pages with AI tagging, search and archival.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-14T17:54:00.028Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"open-notebook","why":"Two self-hosted personal knowledge hoards with AI on top: Karakeep is capture-first (bookmark everything, auto-tag, archive), Open Notebook is synthesis-first (research a topic, chat over sources, produce podcasts). Collector vs study desk.","dir":"in","confidence":0.55,"name":"open-notebook"},{"slug":"telegram-drive","why":"Same shelf — self-hosted personal hoarding with an API agents can drive — different content: Telegram-Drive stores raw files on Telegram's servers; Karakeep hoards bookmarks, notes and pages with AI tagging, search and archival.","dir":"out","confidence":0.45,"name":"Telegram-Drive"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/karakeep-app/karakeep"},"ktx":{"name":"ktx","owner":"Kaelio","slug":"ktx","stars":1508,"image":"https://raw.githubusercontent.com/Kaelio/ktx/main/assets/ktx-lockup.svg","avatar":"https://avatars.githubusercontent.com/u/189464497?v=4&s=96","forks":95,"language":"TypeScript","license":"Apache-2.0","updated":"5 days ago","topics":["rag","agents"],"summary":"Self-improving context layer for data agents — ingests dbt/Looker/wikis, maps your warehouse, builds a semantic layer with approved metrics, and serves Claude Code/Codex via CLI and MCP.","curator_note":"Reach for it when agents re-explore your warehouse on every question and invent their own metric logic: ktx samples tables, detects joinable columns (resolving chasm/fan traps), absorbs dbt/MetricFlow/LookML/Notion knowledge into one searchable surface, and flags contradictions for human review. Read-only by design; runs locally on your own LLM keys or your Claude Code / Codex login. Skip it if you have no SQL warehouse to sit on, or for one ad-hoc query. It ingests your existing semantic layers rather than replacing them. YC-backed (Kaelio); telemetry is on by default with opt-out.","edges":[{"to":"mirage","type":"alternative","why":"Opposite philosophies for the same job — letting agents work with your company's data. Mirage mounts ~50 backends as a raw virtual filesystem to grep and pipe; ktx curates a semantic layer with approved metric definitions and compiled read-only SQL.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T17:35:30.547Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"wrenai","why":"Both are governed semantic layers that make agents trustworthy over company data: ktx curates warehouse metrics and serves read-only SQL context; Wren goes further into governed execution and agent-deployed dashboards.","dir":"in","confidence":0.65,"name":"WrenAI"},{"slug":"mirage","why":"Opposite philosophies for the same job — letting agents work with your company's data. Mirage mounts ~50 backends as a raw virtual filesystem to grep and pipe; ktx curates a semantic layer with approved metric definitions and compiled read-only SQL.","dir":"out","confidence":0.55,"name":"mirage"},{"slug":"scout","why":"Two ways to hand agents company context: ktx curates a semantic layer over your warehouse with approved metrics; Scout navigates Slack/Drive/wiki/CRM live and writes its own wiki+CRM as it learns.","dir":"in","confidence":0.55,"name":"scout"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Kaelio/ktx"},"langgraph":{"name":"LangGraph","owner":"langchain-ai","slug":"langgraph","stars":37913,"image":"https://raw.githubusercontent.com/langchain-ai/langgraph/main/.github/images/logo-dark.svg","avatar":"https://avatars.githubusercontent.com/u/126733545?v=4&s=96","forks":6368,"language":"Python","license":"MIT","updated":"2 days ago","topics":["agents","orchestration"],"summary":"Build stateful, multi-actor LLM apps as graphs — durable execution, human-in-the-loop, streaming.","editors_pick":true,"curator_note":"You reach for LangGraph the moment a simple agent loop stops being enough — when you need state that survives a crash, a human approving a step mid-run, or a flow that can loop back on itself. Most teams arrive here from plain LangChain and don't leave. If all you want is a quick tool-calling agent, this is more machinery than you need — start lighter and come back when you hit the wall.\n","edges":[{"to":"langsmith","type":"complements","why":"Once your graph runs in production you'll want to see why a run failed — LangSmith traces every node.","confidence":0.9,"status":"approved"},{"to":"crewai","type":"alternative","why":"Higher-level, opinionated multi-agent API vs. LangGraph's low-level control. Trade control for speed.","confidence":0.8,"status":"approved"},{"to":"dspy","type":"alternative","why":"A different bet: optimize prompts as code rather than orchestrate them.","confidence":0.7,"status":"approved"},{"to":"ollama","type":"built_with","why":"Point your graph nodes at a local model — no API keys while iterating.","confidence":0.8,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":18,"related":{"complements":[{"slug":"langsmith","why":"Once your graph runs in production you'll want to see why a run failed — LangSmith traces every node.","dir":"out","confidence":0.9,"name":"LangSmith"},{"slug":"cubesandbox","why":"LangGraph nodes that run model-written code get a disposable, hardware-isolated microVM per execution via the E2B-compatible SDK — crash or escape attempts die with the VM.","dir":"in","confidence":0.8,"name":"CubeSandbox"},{"slug":"memmachine","why":"First-class documented integration: MemMachine plugs in as the persistent, cross-session memory behind LangGraph workflows — LangGraph checkpoints the graph state, MemMachine remembers the user.","dir":"in","confidence":0.8,"name":"MemMachine"},{"slug":"nemo-guardrails","why":"RunnableRails wraps LangChain/LangGraph runnables — the agent graph does the work, the rails police what goes in and out of every LLM call.","dir":"in","confidence":0.7,"name":"Guardrails"},{"slug":"browser-use","why":"browser-use is the hands, LangGraph is the brain: wire it in as the browsing tool inside a stateful agent graph and keep planning, retries and human-in-the-loop where they belong.","dir":"in","confidence":0.65,"name":"browser-use"},{"slug":"plano","why":"Keep intra-agent graph logic in LangGraph and let Plano handle what's outside the framework: cross-agent routing, model failover, guardrails and OTEL traces — Plano is explicitly framework-agnostic about what runs behind each agent URL.","dir":"in","confidence":0.6,"name":"plano"},{"slug":"shepherd","why":"LangGraph checkpoints state inside the graph; Shepherd wraps the whole run in a reversible trace — fork and replay an agent from any step, review outputs before they apply, instead of re-running the graph and hoping.","dir":"in","confidence":0.55,"name":"shepherd"},{"slug":"litellm","why":"LangGraph orchestrates the control flow; LiteLLM can serve as the swappable multi-provider model backend behind it.","dir":"in","confidence":0.5,"name":"litellm"},{"slug":"omnigraph","why":"LangGraph orchestrates the multi-actor workflow; omnigraph is built as the durable state those actors share — branch per agent or task, merged on review. Natural pairing for fleet-scale memory, though no packaged integration exists yet.","dir":"in","confidence":0.5,"name":"omnigraph"},{"slug":"paperclip","why":"LangGraph builds stateful, durable multi-actor agent apps; such an app can be 'hired' into Paperclip via HTTP/heartbeat and managed (budget, goals, audit) at the org level. Build the agent with LangGraph, govern the fleet with Paperclip.","dir":"in","confidence":0.5,"name":"paperclip"}],"alternative":[{"slug":"crewai","why":"Higher-level, opinionated multi-agent API vs. LangGraph's low-level control. Trade control for speed.","dir":"out","confidence":0.8,"name":"CrewAI"},{"slug":"dspy","why":"A different bet: optimize prompts as code rather than orchestrate them.","dir":"out","confidence":0.7,"name":"DSPy"},{"slug":"metagpt","why":"Trade-off runs the other way: LangGraph is low-level graph control you assemble; MetaGPT is a pre-built company you configure. Outgrowing MetaGPT's opinions usually lands you here.","dir":"in","confidence":0.6,"name":"MetaGPT"},{"slug":"agentfield","why":"Same job — production multi-agent systems — opposite shape. LangGraph is an in-process library: model your workflow as a stateful graph with durable execution. agentfield is an out-of-process control plane: plain functions become REST microservices and the platform handles fan-out, queues and retries, explicitly rejecting graph wiring. Library and embedded → LangGraph; platform and service-oriented → agentfield.","dir":"in","confidence":0.55,"name":"agentfield"},{"slug":"agents-cli","why":"Both are ways to build production agentic apps, but from opposite ends: LangGraph is an open, embeddable graph runtime that runs anywhere, while agents-cli is Google's CLI+skills around ADK that scaffolds, evaluates, and deploys agents specifically on Google Cloud / Gemini Enterprise. Choose by ecosystem and portability needs.","dir":"in","confidence":0.5,"name":"agents-cli"}],"built_with":[{"slug":"deepagents","why":"Built directly on LangGraph — streaming, persistence and checkpointing come from the runtime; deepagents is the opinionated harness layer above it, and CompiledStateGraphs drop in as sub-agents.","dir":"in","confidence":0.95,"name":"deepagents"},{"slug":"ollama","why":"Point your graph nodes at a local model — no API keys while iterating.","dir":"out","confidence":0.8,"name":"Ollama"},{"slug":"all-agentic-architectures","why":"Several of the 35 architectures ship as runnable LangGraph implementations — the textbook teaches on the substrate.","dir":"in","confidence":0.7,"name":"all-agentic-architectures"}]},"url":"https://stackmap.shipwithai.xyz/repos/langchain-ai/langgraph"},"langsmith":{"name":"LangSmith","owner":"langchain-ai","slug":"langsmith","stars":0,"closed_source":true,"language":"TypeScript","license":"Proprietary","updated":"1 week ago","topics":["evals"],"summary":"Trace, test and monitor LLM apps in production.","edges":[],"added":"2026-06-22T23:27:20.000Z","linkCount":2,"related":{"complements":[{"slug":"langgraph","why":"Once your graph runs in production you'll want to see why a run failed — LangSmith traces every node.","dir":"in","confidence":0.9,"name":"LangGraph"},{"slug":"deepagents","why":"The README's own production pairing: LangSmith supplies tracing, evaluation and monitoring for deepagents deployments.","dir":"in","confidence":0.7,"name":"deepagents"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/langchain-ai/langsmith"},"last30days-skill":{"name":"last30days-skill","owner":"mvanhorn","slug":"last30days-skill","stars":53233,"image":"https://raw.githubusercontent.com/mvanhorn/last30days-skill/main/media/pr-assets/last30days-ad.gif","avatar":"https://avatars.githubusercontent.com/u/455140?v=4&s=96","forks":4606,"language":"Python","license":"MIT","updated":"3 days ago","topics":["skills","web"],"summary":"The /last30days skill: researches any topic across Reddit, X, YouTube, HN, TikTok and Polymarket in parallel, scores by real engagement, and synthesizes one grounded brief. 50+ agent hosts.","curator_note":"A search engine scored by upvotes, likes and prediction-market money instead of editors — and the insight that made it #1 trending: no single AI can search all these walled gardens, but YOUR agent with YOUR keys can. Reddit, HN, Polymarket and GitHub work with zero config; a 30-second wizard unlocks X, YouTube, TikTok, arXiv. Installs into Claude Code via marketplace or 50+ hosts via npx skills. NOT neutral research: engagement-weighted synthesis inherits every platform's biases and hype cycles — it tells you what people are excited about, which is not the same as what's true. BYO keys means BYO rate limits and ToS exposure on X/TikTok.","edges":[{"to":"surfsense","type":"alternative","why":"Same signal sources (Reddit, YouTube, social, search), different shapes: SurfSense is a self-hosted platform where scheduled agents build a cited knowledge base; /last30days is a skill inside your own agent — on-demand briefs, BYO keys, no infrastructure.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:11.049Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"surfsense","why":"Same signal sources (Reddit, YouTube, social, search), different shapes: SurfSense is a self-hosted platform where scheduled agents build a cited knowledge base; /last30days is a skill inside your own agent — on-demand briefs, BYO keys, no infrastructure.","dir":"out","confidence":0.65,"name":"SurfSense"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/mvanhorn/last30days-skill"},"lightmem":{"name":"LightMem","owner":"zjunlp","slug":"lightmem","stars":1029,"image":"https://raw.githubusercontent.com/zjunlp/lightmem/main/figs/lightmem_logo_light.png","avatar":"https://avatars.githubusercontent.com/u/41887875?v=4&s=96","forks":94,"language":"Python","license":"MIT","updated":"8 days ago","topics":["memory"],"summary":"ICLR 2026 memory framework for LLMs/agents: LLMLingua pre-compression, topic segmentation and offline memory updates — leading LoCoMo/LongMemEval results at lower token cost.","curator_note":"Research-grade memory with receipts: reproduction scripts for LoCoMo/LongMemEval plus a baseline harness that benchmarks Mem0, A-MEM and LangMem side by side — useful even if you adopt none of them. The compression-first pipeline (LLMLingua before storage) is the differentiating idea. Expect paper-adjacent ergonomics: manual model downloads and config dicts, not a polished product.","edges":[{"to":"memmachine","type":"alternative","why":"Both are agent memory layers with MCP servers; MemMachine ships product-shaped SDKs and framework integrations, LightMem ships a compression-first pipeline with published benchmark wins.","confidence":0.6,"status":"approved"},{"to":"memory-os","type":"alternative","why":"Layered memory operating systems from opposite cultures: memory-os is a 7-layer production system for one agent stack, LightMem is a modular research framework you assemble per experiment.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:42.099Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"memmachine","why":"Both are agent memory layers with MCP servers; MemMachine ships product-shaped SDKs and framework integrations, LightMem ships a compression-first pipeline with published benchmark wins.","dir":"out","confidence":0.6,"name":"MemMachine"},{"slug":"hindsight","why":"Both chase LongMemEval leadership from opposite ends: LightMem is a research framework optimizing token cost via pre-compression; Hindsight is a packaged service optimizing accuracy. Benchmark rivals, different deployment realities.","dir":"in","confidence":0.6,"name":"hindsight"},{"slug":"memory-os","why":"Layered memory operating systems from opposite cultures: memory-os is a 7-layer production system for one agent stack, LightMem is a modular research framework you assemble per experiment.","dir":"out","confidence":0.5,"name":"memory-os"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/zjunlp/lightmem"},"litellm":{"name":"litellm","owner":"BerriAI","slug":"litellm","stars":54447,"image":"https://render.com/images/deploy-to-render-button.svg","avatar":"https://avatars.githubusercontent.com/u/121462774?v=4&s=96","forks":10003,"language":"Python","license":"NOASSERTION","updated":"yesterday","topics":["gateway"],"summary":"Open-source AI gateway: call 100+ LLM providers in OpenAI format via a Python SDK or self-hosted proxy — with cost tracking, virtual keys, guardrails, load balancing and logging.","curator_note":"The default answer to \"one API for every LLM.\" Reach for it when app code or an agent fleet must hit many providers without per-SDK glue, or when a team needs a central proxy with spend caps, virtual keys and logging. The SDK is a thin drop-in; the proxy is the real value (dashboard, budgets, rate limits). NOT an inference engine — it routes to backends like vLLM/Ollama, doesn't run models. Overkill if you only ever call one provider. Guardrails/evals exist but are lighter than dedicated tools.","edges":[{"to":"vllm","type":"complements","why":"LiteLLM fronts a self-hosted vLLM server as one more OpenAI-compatible backend — vLLM does the serving, LiteLLM does routing, keys and spend tracking.","confidence":0.72,"status":"approved"},{"to":"ollama","type":"complements","why":"Native Ollama provider lets local models be called through the same unified interface as cloud APIs.","confidence":0.65,"status":"approved"},{"to":"freellmapi","type":"alternative","why":"Both expose many providers behind one OpenAI-compatible /v1 endpoint; freellmapi stacks free tiers, LiteLLM targets production apps and teams.","confidence":0.6,"status":"approved"},{"to":"crewai","type":"complements","why":"Agent framework relies on a unified model layer; LiteLLM supplies the 100+-provider abstraction beneath the agents.","confidence":0.6,"status":"approved"},{"to":"9router","type":"alternative","why":"Both are self-hosted gateways routing to many providers with fallback; 9router aims at coding CLIs, LiteLLM at app/SDK developers.","confidence":0.55,"status":"approved"},{"to":"praisonai","type":"complements","why":"PraisonAI runs agents across 100+ LLMs — LiteLLM is the standard provider-abstraction layer for exactly that reach.","confidence":0.52,"status":"approved"},{"to":"langgraph","type":"complements","why":"LangGraph orchestrates the control flow; LiteLLM can serve as the swappable multi-provider model backend behind it.","confidence":0.5,"status":"approved"},{"to":"nemo-guardrails","type":"complements","why":"Pair a dedicated programmable-guardrails layer with LiteLLM's gateway when input/output rails need to be richer than the built-in checks.","confidence":0.48,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:42.149Z","linkCount":12,"related":{"complements":[{"slug":"vllm","why":"LiteLLM fronts a self-hosted vLLM server as one more OpenAI-compatible backend — vLLM does the serving, LiteLLM does routing, keys and spend tracking.","dir":"out","confidence":0.72,"name":"vLLM"},{"slug":"exxperts","why":"exxperts' AI setup has a first-class 'OpenAI-compatible gateway' path and names a company LiteLLM proxy as the canonical example — run rooms against your org's gateway instead of raw provider keys.","dir":"in","confidence":0.7,"name":"exxperts"},{"slug":"ollama","why":"Native Ollama provider lets local models be called through the same unified interface as cloud APIs.","dir":"out","confidence":0.65,"name":"Ollama"},{"slug":"crewai","why":"Agent framework relies on a unified model layer; LiteLLM supplies the 100+-provider abstraction beneath the agents.","dir":"out","confidence":0.6,"name":"CrewAI"},{"slug":"praisonai","why":"PraisonAI runs agents across 100+ LLMs — LiteLLM is the standard provider-abstraction layer for exactly that reach.","dir":"out","confidence":0.52,"name":"PraisonAI"},{"slug":"langgraph","why":"LangGraph orchestrates the control flow; LiteLLM can serve as the swappable multi-provider model backend behind it.","dir":"out","confidence":0.5,"name":"LangGraph"},{"slug":"nemo-guardrails","why":"Pair a dedicated programmable-guardrails layer with LiteLLM's gateway when input/output rails need to be richer than the built-in checks.","dir":"out","confidence":0.48,"name":"Guardrails"}],"alternative":[{"slug":"plano","why":"Both sit between your app and 100+ LLM providers as a self-hosted gateway; LiteLLM is the lighter Python proxy for provider unification and cost control, Plano adds agent orchestration, guardrail filter chains and signals on an Envoy data plane.","dir":"in","confidence":0.85,"name":"plano"},{"slug":"freellmapi","why":"Both expose many providers behind one OpenAI-compatible /v1 endpoint; freellmapi stacks free tiers, LiteLLM targets production apps and teams.","dir":"out","confidence":0.6,"name":"freellmapi"},{"slug":"9router","why":"Both are self-hosted gateways routing to many providers with fallback; 9router aims at coding CLIs, LiteLLM at app/SDK developers.","dir":"out","confidence":0.55,"name":"9router"}],"built_with":[{"slug":"agentic-context-engine","why":"ACELiteLLM is the primary entry point — the '100+ supported providers' run through LiteLLM's unified API.","dir":"in","confidence":0.75,"name":"agentic-context-engine"},{"slug":"mini-swe-agent","why":"The model layer is a LiteLLM wrapper — one small file gives the 100-line agent every provider LiteLLM speaks, which is exactly the minimalism the project preaches.","dir":"in","confidence":0.7,"name":"mini-swe-agent"}]},"url":"https://stackmap.shipwithai.xyz/repos/BerriAI/litellm"},"llamafactory":{"name":"LlamaFactory","owner":"hiyouga","slug":"llamafactory","stars":73465,"image":"https://raw.githubusercontent.com/hiyouga/llamafactory/main/assets/logo.png","avatar":"https://avatars.githubusercontent.com/u/16256802?v=4&s=96","forks":8974,"language":"Python","license":"Apache-2.0","updated":"7 days ago","topics":["training"],"summary":"The unified fine-tuning framework: 100+ LLMs and VLMs via LoRA/QLoRA/full-parameter, config-driven or through the LlamaBoard GUI. ACL 2024, 1000+ citations, 73k stars.","curator_note":"The default answer to 'how do I fine-tune model X': whatever the architecture (Llama, Qwen, Mistral, VLMs…), whatever the method (LoRA, QLoRA, DPO, PPO, full), one YAML config or the LlamaBoard GUI runs it — with the broadest model-coverage matrix in open source and academic citation weight behind it. If TRL is the library you code against, LlamaFactory is the trainer you configure. NOT for frontier-scale RL dataflows (verl/slime territory), and the kitchen-sink coverage means version bumps occasionally break niche model+method combos — pin versions for anything long-running.","edges":[{"to":"trl","type":"alternative","why":"The two dominant open fine-tuning stacks: TRL is the code-first Hugging Face library; LlamaFactory the config-driven unified trainer with a GUI and a wider model-coverage matrix.","confidence":0.8,"status":"approved"},{"to":"h2o-llmstudio","type":"alternative","why":"Both put a GUI on fine-tuning; LLM Studio optimizes for the no-code experience, LlamaFactory for maximum model/method coverage with the GUI as one of several front doors.","confidence":0.6,"status":"approved"},{"to":"vllm","type":"complements","why":"Fine-tune here, serve fast there — LlamaFactory ships a vLLM backend for high-throughput inference on the models it trains.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T15:35:46.228Z","linkCount":5,"related":{"complements":[{"slug":"dataflow","why":"Adjacent stages of one workflow: DataFlow generates, cleans and filters the SFT/RL datasets; LLaMA-Factory is the fine-tuning framework that consumes them across 100+ open models.","dir":"in","confidence":0.65,"name":"DataFlow"},{"slug":"vllm","why":"Fine-tune here, serve fast there — LlamaFactory ships a vLLM backend for high-throughput inference on the models it trains.","dir":"out","confidence":0.6,"name":"vLLM"}],"alternative":[{"slug":"trl","why":"The two dominant open fine-tuning stacks: TRL is the code-first Hugging Face library; LlamaFactory the config-driven unified trainer with a GUI and a wider model-coverage matrix.","dir":"out","confidence":0.8,"name":"trl"},{"slug":"mlx-lora-studio","why":"Both GUI-driven fine-tuning across many methods. LLaMA-Factory is the CUDA-world standard (100+ models, LlamaBoard, cluster-ready); MLX LoRA Studio trades that breadth for a native Mac app that trains entirely on Apple Silicon.","dir":"in","confidence":0.7,"name":"MLX-LoRA-Studio"},{"slug":"h2o-llmstudio","why":"Both put a GUI on fine-tuning; LLM Studio optimizes for the no-code experience, LlamaFactory for maximum model/method coverage with the GUI as one of several front doors.","dir":"out","confidence":0.6,"name":"h2o-llmstudio"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/hiyouga/llamafactory"},"llamaindex":{"name":"LlamaIndex","owner":"run-llama","slug":"llamaindex","github":"run-llama/llama_index","stars":51028,"avatar":"https://avatars.githubusercontent.com/u/130722866?v=4&s=96","forks":7798,"language":"Python","license":"MIT","updated":"yesterday","topics":["rag","agents"],"summary":"Data framework for connecting custom data sources to LLMs — ingestion, indexing, retrieval.","edges":[{"to":"chroma","type":"built_with","why":"Use Chroma as the vector store behind your LlamaIndex retrievers.","confidence":0.85,"status":"approved"},{"to":"ragas","type":"complements","why":"Evaluate your RAG pipeline's retrieval and answer quality.","confidence":0.8,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":8,"related":{"complements":[{"slug":"ragas","why":"Evaluate your RAG pipeline's retrieval and answer quality.","dir":"out","confidence":0.8,"name":"Ragas"},{"slug":"memmachine","why":"Documented LlamaIndex integration — LlamaIndex retrieves from your data, MemMachine recalls what the user said and prefers across sessions; the two retrieval planes compose.","dir":"in","confidence":0.7,"name":"MemMachine"},{"slug":"xberg","why":"Xberg is a stronger document reader/extraction layer than LlamaIndex's built-in loaders; plug it in as the ingestion front-end (96 formats, OCR, transcription, chunking) and let LlamaIndex handle indexing, retrieval and query orchestration.","dir":"in","confidence":0.65,"name":"xberg"},{"slug":"olmocr","why":"olmocr is the document-parsing front-end; LlamaIndex is the indexing/retrieval layer. Run messy PDFs through olmocr to get clean text, then hand that text to LlamaIndex to chunk, embed, and retrieve — a natural two-stage RAG-ingestion pipeline.","dir":"in","confidence":0.6,"name":"olmocr"},{"slug":"chunkr","why":"Chunkr slots into LlamaIndex ingestion as the parsing stage — hard documents become structured, citable chunks before indexing.","dir":"in","confidence":0.55,"name":"chunkr"}],"alternative":[{"slug":"cocoindex","why":"Both connect LLMs to your data, at different layers: LlamaIndex is the retrieval framework (loaders, indexes, query engines) typically run as batch ingestion; CocoIndex is the incremental sync engine that keeps whatever store you target continuously fresh, delta-only, with lineage.","dir":"in","confidence":0.6,"name":"cocoindex"},{"slug":"searchbox","why":"Both answer questions over a private corpus, but by opposite paradigms. LlamaIndex is a production RAG framework: build a structured index, then query it. searchbox is an airgapped research harness where an agent explores the raw corpus with grep/embed/rerank tools under a budget — no prebuilt index. Use LlamaIndex to ship RAG, searchbox to study agentic retrieval.","dir":"in","confidence":0.5,"name":"searchbox"}],"built_with":[{"slug":"chroma","why":"Use Chroma as the vector store behind your LlamaIndex retrievers.","dir":"out","confidence":0.85,"name":"Chroma"}]},"url":"https://stackmap.shipwithai.xyz/repos/run-llama/llamaindex"},"llm-d":{"name":"llm-d","owner":"llm-d","slug":"llm-d","stars":3861,"image":"https://raw.githubusercontent.com/llm-d/llm-d/main/docs/assets/images/llm-d-logo.png","avatar":"https://avatars.githubusercontent.com/u/211385051?v=4&s=96","forks":626,"language":"Shell","license":"Apache-2.0","updated":"yesterday","topics":["local"],"summary":"Distributed inference stack for Kubernetes from Red Hat, Google and IBM (CNCF) — prefix-cache-aware routing, tiered KV-cache, prefill/decode disaggregation and SLO autoscaling above vLLM/SGLang.","curator_note":"The Kubernetes answer when one vLLM box stops scaling: prefix-cache-aware routing, disaggregated prefill/decode and tiered KV offloading deliver real, benchmarked wins (3x throughput, 2x TTFT in partner numbers) at fleet scale, with Red Hat/Google/IBM/NVIDIA behind it. NOT for a single GPU or a laptop — that's Ollama or plain vLLM territory — and not turnkey: you are operating Kubernetes, Helm charts and gateways. If you don't already run K8s, don't start here.","edges":[{"to":"vllm","type":"built_with","why":"llm-d is explicitly the orchestration layer above model servers: vLLM does the on-accelerator inference, llm-d adds cluster-level routing, KV-cache management, disaggregation and autoscaling.","confidence":0.9,"status":"approved"},{"to":"ollama","type":"alternative","why":"Same job — serve open models on your own hardware — at opposite scales: Ollama is one command on one machine; llm-d is a CNCF stack for multi-node GPU fleets. Outgrow one, reach for the other.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-07T15:50:14.462Z","linkCount":4,"related":{"complements":[{"slug":"lmcache","why":"llm-d's tiered KV-cache subsystem in Kubernetes is built around LMCache for offload; deploy together for cluster-scale prefix caching.","dir":"in","confidence":0.85,"name":"LMCache"}],"alternative":[{"slug":"ollama","why":"Same job — serve open models on your own hardware — at opposite scales: Ollama is one command on one machine; llm-d is a CNCF stack for multi-node GPU fleets. Outgrow one, reach for the other.","dir":"out","confidence":0.5,"name":"Ollama"},{"slug":"mesh-llm","why":"Distributed inference at opposite trust levels: llm-d orchestrates vLLM across a Kubernetes cluster you own; mesh-llm federates volunteer boxes over the open internet.","dir":"in","confidence":0.5,"name":"mesh-llm"}],"built_with":[{"slug":"vllm","why":"llm-d is explicitly the orchestration layer above model servers: vLLM does the on-accelerator inference, llm-d adds cluster-level routing, KV-cache management, disaggregation and autoscaling.","dir":"out","confidence":0.9,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/llm-d/llm-d"},"lmcache":{"name":"LMCache","owner":"LMCache","slug":"lmcache","stars":10824,"image":"https://raw.githubusercontent.com/LMCache/lmcache/dev/asset/logo.png","avatar":"https://avatars.githubusercontent.com/u/171091289?v=4&s=96","forks":1603,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["local","storage"],"summary":"KV-cache layer for scalable LLM serving: offload and reuse KV across GPU/CPU/disk/remote tiers to cut TTFT and prefill cost. vLLM-first; used by NVIDIA Dynamo and llm-d.","curator_note":"Use when serving LLMs at scale on vLLM with long or repeated contexts — multi-turn agents, RAG, shared system prompts — where prefill dominates and prefix cache hits are money. It moves KV out of GPU into CPU/disk/Redis-style backends and shares it across instances. NOT for single-box hobby setups (Ollama-class) or workloads with no prompt overlap: you'd add a storage tier and ops surface for zero hit rate. PyTorch Foundation project, Apache-2.0.","edges":[{"to":"vllm","type":"built_with","why":"LMCache plugs into vLLM as its KV-connector; vLLM is the primary serving engine it accelerates.","confidence":0.95,"status":"approved"},{"to":"llm-d","type":"complements","why":"llm-d's tiered KV-cache subsystem in Kubernetes is built around LMCache for offload; deploy together for cluster-scale prefix caching.","confidence":0.85,"status":"approved"}],"status":"approved","added":"2026-07-20T07:55:01.458Z","linkCount":2,"related":{"complements":[{"slug":"llm-d","why":"llm-d's tiered KV-cache subsystem in Kubernetes is built around LMCache for offload; deploy together for cluster-scale prefix caching.","dir":"out","confidence":0.85,"name":"llm-d"}],"alternative":[],"built_with":[{"slug":"vllm","why":"LMCache plugs into vLLM as its KV-connector; vLLM is the primary serving engine it accelerates.","dir":"out","confidence":0.95,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/LMCache/lmcache"},"loop-engineering":{"name":"loop-engineering","owner":"cobusgreyling","slug":"loop-engineering","stars":9204,"image":"https://raw.githubusercontent.com/cobusgreyling/loop-engineering/main/assets/visuals/loop-engineering-logo.svg","avatar":"https://avatars.githubusercontent.com/u/7868717?v=4&s=96","forks":1254,"language":"JavaScript","license":"MIT","updated":"yesterday","topics":["coding","orchestration"],"summary":"Reference repo plus npm CLIs (loop-init/audit/cost) for loop engineering: designing scheduled, gated control loops that prompt and orchestrate AI coding agents — Grok, Claude Code, Codex — over time.","curator_note":"The methodology layer for autonomous coding agents — read it once you've stopped hand-prompting and want to design the loop that prompts the agent instead: scheduled triage, worktree isolation, maker/checker sub-agents, MCP connectors, human gates, and a phased L1-report → L3-unattended rollout. Ships 7 clone-and-run patterns, tool-agnostic starters (Grok/Claude Code/Codex/OpenCode), and CLIs that score loop-readiness (loop-audit), estimate token spend (loop-cost) and scaffold state/budget (loop-init). It's a patterns + tooling reference, NOT a runtime — it won't execute or host your agents, so bring your own harness (an alook-style AI-company layer or a squid-style pipeline). Overkill if you just want one-shot interactive help; heed its own caveat that unattended loops make unattended mistakes, so verification stays on you.","edges":[{"to":"bemyagent","type":"alternative","why":"Both impose a disciplined, tool-agnostic operating structure on coding agents through scaffolding, an explicit work cycle and human gates. bemyagent structures a single human-paced session (Think→Task→Execute→Verify); loop-engineering structures scheduled, unattended loops that run over time. Pick bemyagent for interactive pacing, loop-engineering for autonomous automation.","confidence":0.6,"status":"approved"},{"to":"squid","type":"complements","why":"squid is a concrete, productized maker/checker agent pipeline (PA→SWE→Tester→PR-Reviewer→On-Call) with human gates; loop-engineering is the cross-tool methodology and CLI tooling (PR-babysitter/CI-sweeper patterns, loop-audit, loop-worktree, design checklist) behind such pipelines. Use loop-engineering to design the loop, squid as one ready-made implementation.","confidence":0.6,"status":"approved"},{"to":"alook","type":"complements","why":"alook is an always-on multi-agent 'AI company' runtime (per-agent email, org chart, kanban, shared memory) for local coding agents; loop-engineering supplies the loop-design methodology plus scheduling, state and budget tooling to drive agents autonomously. Design the loops with one, run them on the other.","confidence":0.55,"status":"approved"},{"to":"lynkr","type":"complements","why":"loop-engineering's loop-cost flags the token blowup that long, sub-agent-heavy loops create; lynkr is a drop-in gateway around Claude Code/Cursor/Codex that compresses tool results and tier-routes cheap work to local models — cutting exactly that runtime spend without touching your loop design.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-08T00:23:51.035Z","linkCount":8,"related":{"complements":[{"slug":"squid","why":"squid is a concrete, productized maker/checker agent pipeline (PA→SWE→Tester→PR-Reviewer→On-Call) with human gates; loop-engineering is the cross-tool methodology and CLI tooling (PR-babysitter/CI-sweeper patterns, loop-audit, loop-worktree, design checklist) behind such pipelines. Use loop-engineering to design the loop, squid as one ready-made implementation.","dir":"out","confidence":0.6,"name":"squid"},{"slug":"govctl","why":"Two halves of disciplined AI delivery: loop-engineering designs the gated control loops that drive coding agents; govctl supplies the governance artifacts and executable completion gates those loops should answer to.","dir":"in","confidence":0.6,"name":"govctl"},{"slug":"alook","why":"alook is an always-on multi-agent 'AI company' runtime (per-agent email, org chart, kanban, shared memory) for local coding agents; loop-engineering supplies the loop-design methodology plus scheduling, state and budget tooling to drive agents autonomously. Design the loops with one, run them on the other.","dir":"out","confidence":0.55,"name":"alook"},{"slug":"lynkr","why":"loop-engineering's loop-cost flags the token blowup that long, sub-agent-heavy loops create; lynkr is a drop-in gateway around Claude Code/Cursor/Codex that compresses tool results and tier-routes cheap work to local models — cutting exactly that runtime spend without touching your loop design.","dir":"out","confidence":0.55,"name":"Lynkr"}],"alternative":[{"slug":"bemyagent","why":"Both impose a disciplined, tool-agnostic operating structure on coding agents through scaffolding, an explicit work cycle and human gates. bemyagent structures a single human-paced session (Think→Task→Execute→Verify); loop-engineering structures scheduled, unattended loops that run over time. Pick bemyagent for interactive pacing, loop-engineering for autonomous automation.","dir":"out","confidence":0.6,"name":"bemyagent"},{"slug":"autoresearch","why":"Both operationalize 'the loop is the unit of progress': loop-engineering is the design methodology and CLIs for building gated loops; autoresearch is one specific, proven loop shipped as an installable skill.","dir":"in","confidence":0.6,"name":"autoresearch"},{"slug":"cwc-long-running-agents","why":"Both are reference repos for keeping agents productive over long horizons: loop-engineering designs the control loops, Anthropic's primitives enforce evidence and evaluation inside Claude Code hooks.","dir":"in","confidence":0.6,"name":"cwc-long-running-agents"},{"slug":"loopy","why":"Same conviction — the loop is the unit of agent work — expressed differently: loop-engineering teaches you to design gated control loops; Loopy catalogs finished loops so you can borrow instead of design.","dir":"in","confidence":0.6,"name":"loopy"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/cobusgreyling/loop-engineering"},"loopy":{"name":"loopy","owner":"Forward-Future","slug":"loopy","stars":2837,"avatar":"https://avatars.githubusercontent.com/u/259970374?v=4&s=96","forks":253,"language":"JavaScript","license":"MIT","updated":"17 days ago","topics":["skills"],"summary":"A public library of reusable AI-agent loops plus Loopy, an installable skill that helps agents find, audit, adapt, run and publish loops from the live catalog.","curator_note":"The insight: most agent work is a repeatable loop someone already designed — so catalog them. The website is browsable by humans AND agents (llms.txt, JSON catalog, agent guide), and the Loopy skill turns your agent into a loop librarian: discover, audit, repair, debrief, publish. Genuinely useful for not reinventing the same research/review/refactor loop weekly. NOT a runtime — loops are prompts and procedure, not executable infrastructure; quality varies by contributor, so audit before you adopt (the skill's audit step exists for a reason).","edges":[{"to":"loop-engineering","type":"alternative","why":"Same conviction — the loop is the unit of agent work — expressed differently: loop-engineering teaches you to design gated control loops; Loopy catalogs finished loops so you can borrow instead of design.","confidence":0.6,"status":"approved"},{"to":"autoresearch","type":"alternative","why":"Both distribute reusable agent loops: autoresearch is one loop (Karpathy's research iteration) as an installable skill; Loopy is a whole catalog of loops plus find/audit/run tooling.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:12.788Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"loop-engineering","why":"Same conviction — the loop is the unit of agent work — expressed differently: loop-engineering teaches you to design gated control loops; Loopy catalogs finished loops so you can borrow instead of design.","dir":"out","confidence":0.6,"name":"loop-engineering"},{"slug":"autoresearch","why":"Both distribute reusable agent loops: autoresearch is one loop (Karpathy's research iteration) as an installable skill; Loopy is a whole catalog of loops plus find/audit/run tooling.","dir":"out","confidence":0.5,"name":"autoresearch"},{"slug":"cwc-long-running-agents","why":"Loopy catalogs reusable agent loops as installable skills; this repo ships the raw hook/evaluator primitives you'd build such loops from.","dir":"in","confidence":0.45,"name":"cwc-long-running-agents"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Forward-Future/loopy"},"luxtts":{"name":"LuxTTS","owner":"ysharma3501","slug":"luxtts","stars":4844,"avatar":"https://avatars.githubusercontent.com/u/244096243?v=4&s=96","forks":628,"language":"Python","license":"Apache-2.0","updated":"1 months ago","topics":["voice","local"],"summary":"Lightweight voice-cloning TTS — 48kHz speech at 150x realtime, fits in 1GB VRAM and runs on CPU or MPS. SOTA cloning from a ~3s reference sample, rivaling models 10x larger.","curator_note":"Reach for it when you need fast, local voice cloning that fits anywhere: 48kHz output (most open TTS caps at 24kHz), sub-1GB VRAM, and faster-than-realtime even on CPU make it practical for on-device apps and batch narration. NOT for you if you need many built-in voices or languages out of the box — it clones a reference, it doesn't ship a voice library — and cloning any real person's voice without consent is an ethics/legal minefield. Quality rides on a clean 3s+ reference.","edges":[{"to":"pocket-tts","type":"alternative","why":"Both are tiny local TTS engines: pocket-tts is Kyutai's 100M CPU-only streaming model with a fixed voice set; LuxTTS targets high-fidelity 48kHz voice cloning from a reference sample. Pick by whether you need cloning or streaming.","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-07T22:49:34.309Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"pocket-tts","why":"Both are tiny local TTS engines: pocket-tts is Kyutai's 100M CPU-only streaming model with a fixed voice set; LuxTTS targets high-fidelity 48kHz voice cloning from a reference sample. Pick by whether you need cloning or streaming.","dir":"out","confidence":0.8,"name":"pocket-tts"}],"built_with":[{"slug":"voicebox","why":"LuxTTS ships inside Voicebox as one of its seven selectable TTS engines — the lightweight English option (~1GB VRAM, 150x realtime on CPU).","dir":"in","confidence":0.8,"name":"voicebox"}]},"url":"https://stackmap.shipwithai.xyz/repos/ysharma3501/luxtts"},"lynkr":{"name":"Lynkr","owner":"Fast-Editor","slug":"lynkr","stars":535,"avatar":"https://avatars.githubusercontent.com/u/249706325?v=4&s=96","forks":57,"language":"JavaScript","license":"Apache-2.0","updated":"yesterday","topics":["coding","local","gateway"],"summary":"Self-hosted LLM gateway wrapping Claude Code, Cursor or Codex with zero code changes — strips unused tools, compresses JSON tool results ~88%, semantic-caches, tier-routes easy work to local models.","curator_note":"Use it when your Claude Code/Cursor subscription limits or API bill hurt: `lynkr wrap claude` is a one-liner, and compression + routing SIMPLE-tier traffic to a free Ollama model genuinely stretches quotas. Also the escape hatch when corporate policy forces traffic through Databricks/Azure/Bedrock. NOT a model server — it only routes, so pair it with Ollama or another backend. Skip it for light usage: a proxy is one more moving part with a big config surface (tiers, budgets, cache), and semantic caching can serve stale hits on near-duplicate prompts. If you only want provider switching without the token tricks, plain LiteLLM is the boring default.","edges":[{"to":"ollama","type":"complements","why":"Lynkr's flagship trick is tier routing: SIMPLE/MEDIUM requests go to a local Ollama endpoint (TIER_SIMPLE=ollama:qwen2.5-coder) so only hard tasks hit your paid subscription — the README's recommended free setup is Ollama-first.","confidence":0.85,"status":"approved"},{"to":"vllm","type":"complements","why":"For team deployments, Lynkr fronts any OpenAI-compatible backend — vLLM is the standard self-hosted server to put behind it when a laptop-class Ollama isn't enough. Indirect (generic OpenAI-compatible path, vLLM not named in the provider table).","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-07T12:53:33.057Z","linkCount":8,"related":{"complements":[{"slug":"ollama","why":"Lynkr's flagship trick is tier routing: SIMPLE/MEDIUM requests go to a local Ollama endpoint (TIER_SIMPLE=ollama:qwen2.5-coder) so only hard tasks hit your paid subscription — the README's recommended free setup is Ollama-first.","dir":"out","confidence":0.85,"name":"Ollama"},{"slug":"prompt-cache-skills","why":"Both attack the coding-agent API bill and stack cleanly: lynkr compresses and semantic-caches at a gateway in front of the model, while these skills fix the harness's own cache_control / prompt_cache_key bugs so provider-side prompt caching actually engages.","dir":"in","confidence":0.6,"name":"prompt-cache-skills"},{"slug":"vllm","why":"For team deployments, Lynkr fronts any OpenAI-compatible backend — vLLM is the standard self-hosted server to put behind it when a laptop-class Ollama isn't enough. Indirect (generic OpenAI-compatible path, vLLM not named in the provider table).","dir":"out","confidence":0.55,"name":"vLLM"},{"slug":"loop-engineering","why":"loop-engineering's loop-cost flags the token blowup that long, sub-agent-heavy loops create; lynkr is a drop-in gateway around Claude Code/Cursor/Codex that compresses tool results and tier-routes cheap work to local models — cutting exactly that runtime spend without touching your loop design.","dir":"in","confidence":0.55,"name":"loop-engineering"},{"slug":"tokensave","why":"Two token-savers at different layers that stack: lynkr compresses and routes at the gateway; tokensave cuts the exploration calls at the MCP layer so the agent asks the graph instead of scanning files.","dir":"in","confidence":0.55,"name":"tokensave"}],"alternative":[{"slug":"9router","why":"Both are self-hosted LLM gateways that wrap coding agents (Claude Code/Codex/Cursor/Cline) with zero code changes and cut tokens by compressing tool_result payloads, then route across backends. 9router optimizes for cost/uptime — 40+ providers, multi-account round-robin, subscription→cheap→free auto-fallback. lynkr optimizes for efficiency — strips unused tools, semantic-caches, and tier-routes easy work to local models. Pick by whether you need free-provider breadth or local-model routing.","dir":"in","confidence":0.85,"name":"9router"},{"slug":"cliproxyapi","why":"Both self-hosted proxies between coding subscriptions and clients: lynkr optimizes (compression, caching, tier-routing); CLIProxyAPI multiplexes accounts and re-exposes them as standard APIs.","dir":"in","confidence":0.55,"name":"CLIProxyAPI"},{"slug":"freellmapi","why":"Same slot — a local proxy in front of Claude Code/Codex to cut LLM spend. lynkr's lever is compression, semantic caching and tier-routing of paid traffic; FreeLLMAPI's is routing everything to free-tier quota across providers.","dir":"in","confidence":0.55,"name":"freellmapi"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Fast-Editor/lynkr"},"mails":{"name":"mails","owner":"chekusu","slug":"mails","stars":362,"avatar":"https://avatars.githubusercontent.com/u/253923356?v=4&s=96","forks":27,"language":"TypeScript","license":null,"updated":"18 days ago","topics":["agents"],"summary":"Email infrastructure for AI agents — send, receive, search and auto-extract verification codes via CLI/SDK, backed by a Cloudflare Worker; free hosted @mails.dev mailboxes or self-host.","curator_note":"Solves the unglamorous half of agent autonomy: an agent that signs up for services needs its own inbox, and `mails code --to agent@mails.dev` long-polls until the verification email lands and prints the extracted code straight to stdout for piping. The hosted tier is genuinely usable (free mailboxes, 100 sends/month); self-hosting is one wrangler deploy with mailbox-scoped bearer tokens, and everything syncs to local SQLite for offline access. Zero runtime dependencies. Caveats: young project, the hosted service rides one maintainer's Cloudflare account — self-host for anything production — and the advanced query API (attachment/sender/time filters) exists only on the hosted side, not in the CLI/SDK yet.","edges":[{"to":"browser-use","type":"complements","why":"The missing half of automated signup flows: browser-use drives the registration form; mails receives the verification email in the agent's own mailbox and hands back the code — no human inbox in the loop.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T16:07:34.801Z","linkCount":1,"related":{"complements":[{"slug":"browser-use","why":"The missing half of automated signup flows: browser-use drives the registration form; mails receives the verification email in the agent's own mailbox and hands back the code — no human inbox in the loop.","dir":"out","confidence":0.55,"name":"browser-use"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/chekusu/mails"},"maxkb":{"name":"MaxKB","owner":"1Panel-dev","slug":"maxkb","stars":22186,"avatar":"https://avatars.githubusercontent.com/u/109613420?v=4&s=96","forks":3016,"language":"Python","license":"GPL-3.0","updated":"yesterday","topics":["agents","rag"],"summary":"Open-source enterprise agent platform: RAG pipelines (upload or crawl docs), a visual workflow engine with MCP tool-use, and zero-code embedding into existing business systems.","curator_note":"The self-hosted answer to 'we need an internal AI assistant this quarter': upload or crawl your docs, get a RAG-grounded Q&A agent with a real workflow engine and MCP tool-use, then embed it into existing systems without code. Model-agnostic including fully private deployments. Battle-tested at 22k stars, mostly in enterprise support/knowledge-base roles. NOT a developer framework — you orchestrate in its UI, not your codebase (LangGraph territory); GPL-3.0 matters if you redistribute; and the project's center of gravity is the Chinese enterprise ecosystem — English docs and community trail the code.","edges":[{"to":"ollama","type":"complements","why":"MaxKB's private-model story runs through Ollama — point the platform at a local DeepSeek/Llama/Qwen and the whole RAG + agent stack stays on-prem.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-14T14:52:59.690Z","linkCount":3,"related":{"complements":[{"slug":"ollama","why":"MaxKB's private-model story runs through Ollama — point the platform at a local DeepSeek/Llama/Qwen and the whole RAG + agent stack stays on-prem.","dir":"out","confidence":0.65,"name":"Ollama"}],"alternative":[{"slug":"supavec","why":"Both are open platforms for shipping RAG: MaxKB is the batteries-included enterprise product — visual workflows, UI, zero-code embedding into business systems; Supavec is the developer primitive — a clean REST API for ingestion and chat you build your own product on.","dir":"in","confidence":0.55,"name":"supavec"},{"slug":"bytechef","why":"Both are self-hosted no-code enterprise agent platforms with different centers of gravity: MaxKB starts from the knowledge base and RAG, ByteChef from integrations and workflow automation.","dir":"in","confidence":0.5,"name":"bytechef"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/1Panel-dev/maxkb"},"meetily":{"name":"meetily","owner":"Zackriya-Solutions","slug":"meetily","stars":26175,"image":"https://raw.githubusercontent.com/Zackriya-Solutions/meetily/main/docs/Meetily-6.png","avatar":"https://avatars.githubusercontent.com/u/82556810?v=4&s=96","forks":2638,"language":"Rust","license":"MIT","updated":"1 months ago","topics":["local","voice"],"summary":"Privacy-first meeting note-taker that runs 100% on-device: live Whisper/Parakeet transcription, speaker diarization, local Ollama summaries. Desktop app for macOS & Windows — no cloud, no call bots.","curator_note":"The default pick when meeting content must not leave the machine — legal, health, finance, anything under NDA: transcription (Whisper/Parakeet + Sortformer diarization) and summarization (Ollama) both run locally, and it captures system audio directly instead of injecting a bot into the call, so it works with any meeting app. NOT a team product: notes live on one desktop — no shared workspace, no CRM hooks — and summary quality is capped by whatever model your hardware can run. If you want a searchable team archive and don't mind the cloud, a hosted note-taker will serve better. Budget real RAM/CPU; live diarization on a weak laptop will struggle.","edges":[{"to":"ollama","type":"built_with","why":"Meeting summarization runs on Ollama — Meetily pipes finished transcripts to a local Ollama model to produce structured notes; it's the headline local path (cloud APIs are the optional fallback, not the default).","confidence":0.85,"status":"approved"}],"status":"approved","added":"2026-07-07T12:59:19.620Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"handy","why":"Same fully-local Whisper transcription core, different job: Handy dictates what YOU say into the focused text field; Meetily captures system audio of meetings with diarization and Ollama summaries. Pick by whether the audio is your voice or a call.","dir":"in","confidence":0.55,"name":"Handy"},{"slug":"rowboat","why":"Overlapping on the meeting slice: Meetily is the dedicated local meeting note-taker; Rowboat bundles an equivalent mic+speaker transcriber whose summaries feed its knowledge graph rather than standing alone.","dir":"in","confidence":0.45,"name":"rowboat"}],"built_with":[{"slug":"ollama","why":"Meeting summarization runs on Ollama — Meetily pipes finished transcripts to a local Ollama model to produce structured notes; it's the headline local path (cloud APIs are the optional fallback, not the default).","dir":"out","confidence":0.85,"name":"Ollama"}]},"url":"https://stackmap.shipwithai.xyz/repos/Zackriya-Solutions/meetily"},"memanto":{"name":"memanto","owner":"moorcheh-ai","slug":"memanto","stars":1677,"image":"https://raw.githubusercontent.com/moorcheh-ai/memanto/main/assets/memanto-logo.svg","avatar":"https://avatars.githubusercontent.com/u/220075209?v=4&s=96","forks":539,"language":"Python","license":"MIT","updated":"yesterday","topics":["memory","local"],"summary":"Fully-local persistent memory for 14+ coding agents, built on an information-theoretic search engine — no vector DB, no API keys, no backend. pip install and your agents remember.","curator_note":"The zero-infrastructure entry in the agent-memory category: no embedding provider, no vector database, no server — the information-theoretic search bet means everything runs on-device from a pip install, and 14+ agent integrations cover the usual suspects. If 'no API keys, nothing leaves the machine' is your constraint, this is the shortest path to cross-agent memory. NOT battle-hardened at team scale: single-machine by design, and the novel search engine is the differentiator AND the risk — benchmark recall on YOUR corpus before trusting it over boring embeddings; memanto.ai signals a company forming behind it.","edges":[{"to":"memsearch","type":"alternative","why":"Same job — one persistent memory shared across your coding agents — opposite infrastructure bets: memsearch runs Markdown + Milvus with embedding models; Memanto runs an information-theoretic engine with no vector DB and no keys at all.","confidence":0.7,"status":"approved"},{"to":"memmolt","type":"alternative","why":"Both are local, no-cloud agent memory stores: MemMolt enforces a strict bucket-thread-memo hierarchy over MCP; Memanto bets on companion-agent ergonomics and keyless information-theoretic search.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.755Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"memsearch","why":"Same job — one persistent memory shared across your coding agents — opposite infrastructure bets: memsearch runs Markdown + Milvus with embedding models; Memanto runs an information-theoretic engine with no vector DB and no keys at all.","dir":"out","confidence":0.7,"name":"memsearch"},{"slug":"memmolt","why":"Both are local, no-cloud agent memory stores: MemMolt enforces a strict bucket-thread-memo hierarchy over MCP; Memanto bets on companion-agent ergonomics and keyless information-theoretic search.","dir":"out","confidence":0.55,"name":"MemMolt"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/moorcheh-ai/memanto"},"memmachine":{"name":"MemMachine","owner":"MemMachine","slug":"memmachine","stars":3341,"image":"https://raw.githubusercontent.com/MemMachine/MemMachine/main/assets/img/MemMachine_Hero_Banner.png","avatar":"https://avatars.githubusercontent.com/u/226739620?v=4&s=96","forks":197,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["memory"],"summary":"Long-term memory layer for AI agents — episodic (graph), profile (SQL) and working memory behind Python/TS SDKs, REST and MCP; ships LangChain, LangGraph, CrewAI and LlamaIndex integrations.","curator_note":"Pick it when memory is a product requirement, not a cache: separating episodic (graph) from profile (SQL) from working memory maps to how assistants actually personalize, and the documented LangGraph/CrewAI/LlamaIndex integrations mean you don't write the glue. NOT worth the footprint for a single-user tool — it wants a server plus Neo4j and SQL; a vector store or a JSON file gets a prototype further. Watch the open-core boundary: the managed platform is the business model.","edges":[{"to":"langgraph","type":"complements","why":"First-class documented integration: MemMachine plugs in as the persistent, cross-session memory behind LangGraph workflows — LangGraph checkpoints the graph state, MemMachine remembers the user.","confidence":0.8,"status":"approved"},{"to":"crewai","type":"complements","why":"Ships a CrewAI integration: crews get shared persistent memory across sessions instead of CrewAI's per-run state.","confidence":0.75,"status":"approved"},{"to":"llamaindex","type":"complements","why":"Documented LlamaIndex integration — LlamaIndex retrieves from your data, MemMachine recalls what the user said and prefers across sessions; the two retrieval planes compose.","confidence":0.7,"status":"approved"},{"to":"chroma","type":"alternative","why":"Same slot — 'what my agent remembers' — different bets: Chroma is a general embedding store you shape into memory; MemMachine is purpose-built memory with episodic/profile/working tiers, at the cost of running Neo4j + SQL.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-07T15:50:13.069Z","linkCount":8,"related":{"complements":[{"slug":"langgraph","why":"First-class documented integration: MemMachine plugs in as the persistent, cross-session memory behind LangGraph workflows — LangGraph checkpoints the graph state, MemMachine remembers the user.","dir":"out","confidence":0.8,"name":"LangGraph"},{"slug":"crewai","why":"Ships a CrewAI integration: crews get shared persistent memory across sessions instead of CrewAI's per-run state.","dir":"out","confidence":0.75,"name":"CrewAI"},{"slug":"llamaindex","why":"Documented LlamaIndex integration — LlamaIndex retrieves from your data, MemMachine recalls what the user said and prefers across sessions; the two retrieval planes compose.","dir":"out","confidence":0.7,"name":"LlamaIndex"}],"alternative":[{"slug":"hindsight","why":"Both self-hostable long-term memory services for agents behind Python/TS SDKs and REST. MemMachine splits episodic/profile/working memory with framework integrations; Hindsight bets on retain/recall/reflect and benchmark-topping learned memory.","dir":"in","confidence":0.7,"name":"hindsight"},{"slug":"lightmem","why":"Both are agent memory layers with MCP servers; MemMachine ships product-shaped SDKs and framework integrations, LightMem ships a compression-first pipeline with published benchmark wins.","dir":"in","confidence":0.6,"name":"LightMem"},{"slug":"chroma","why":"Same slot — 'what my agent remembers' — different bets: Chroma is a general embedding store you shape into memory; MemMachine is purpose-built memory with episodic/profile/working tiers, at the cost of running Neo4j + SQL.","dir":"out","confidence":0.55,"name":"Chroma"},{"slug":"memory-os","why":"Both give agents durable, structured memory; MemMachine is the framework-agnostic layer with SDKs and MCP, memory-os the deeper 7-layer architecture wedded to Hermes Agent.","dir":"in","confidence":0.55,"name":"memory-os"},{"slug":"memmolt","why":"Both agent memory layers over MCP: MemMachine ships episodic/profile/working memory with framework SDKs; MemMolt bets on one enforced bucket-thread-memo hierarchy in a single SQLite file.","dir":"in","confidence":0.5,"name":"MemMolt"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/MemMachine/memmachine"},"memmolt":{"name":"MemMolt","owner":"rituraj-io","slug":"memmolt","stars":4,"image":"https://raw.githubusercontent.com/rituraj-io/memmolt/main/assets/memory-architecture.png","avatar":"https://avatars.githubusercontent.com/u/49026287?v=4&s=96","forks":0,"language":"JavaScript","license":"MIT","updated":"3 months ago","topics":["memory"],"summary":"Structured long-term memory over MCP: an enforced bucket-thread-memo hierarchy in one SQLite file, hybrid FTS5 + vector search fused with RRF, local embeddings.","curator_note":"The anti-sprawl memory play: a forced 3-level hierarchy the agent can't turn into a jungle, with ~10ms hybrid search from a single SQLite file and zero cloud calls. Young and tiny (4 stars) — the schema idea is worth studying even if you don't adopt it. Skip if you want auto-capture; this is deliberate, curated memory.","edges":[{"to":"memmachine","type":"alternative","why":"Both agent memory layers over MCP: MemMachine ships episodic/profile/working memory with framework SDKs; MemMolt bets on one enforced bucket-thread-memo hierarchy in a single SQLite file.","confidence":0.5,"status":"approved"},{"to":"opencode-mem","type":"alternative","why":"Local SQLite+vector agent memory either way: opencode-mem auto-captures per-session for OpenCode; MemMolt is deliberate, structured recall for any MCP client.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.638Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"memanto","why":"Both are local, no-cloud agent memory stores: MemMolt enforces a strict bucket-thread-memo hierarchy over MCP; Memanto bets on companion-agent ergonomics and keyless information-theoretic search.","dir":"in","confidence":0.55,"name":"memanto"},{"slug":"memsearch","why":"Both are structured long-term memory with hybrid search for agents: MemMolt enforces a bucket-thread-memo hierarchy in one SQLite file over MCP; memsearch bets on free-form Markdown + Milvus and per-agent plugins.","dir":"in","confidence":0.55,"name":"memsearch"},{"slug":"memmachine","why":"Both agent memory layers over MCP: MemMachine ships episodic/profile/working memory with framework SDKs; MemMolt bets on one enforced bucket-thread-memo hierarchy in a single SQLite file.","dir":"out","confidence":0.5,"name":"MemMachine"},{"slug":"opencode-mem","why":"Local SQLite+vector agent memory either way: opencode-mem auto-captures per-session for OpenCode; MemMolt is deliberate, structured recall for any MCP client.","dir":"out","confidence":0.5,"name":"opencode-mem"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/rituraj-io/memmolt"},"memory-os":{"name":"memory-os","owner":"ClaudioDrews","slug":"memory-os","stars":1294,"image":"https://raw.githubusercontent.com/ClaudioDrews/memory-os/main/assets/banner.jpg","avatar":"https://avatars.githubusercontent.com/u/242654945?v=4&s=96","forks":120,"language":"Python","license":"MIT","updated":"1 months ago","topics":["memory"],"summary":"A 7-layer memory operating system for Hermes Agent: Qdrant vectors, structured facts, fabric recall, an auto-curated wiki, and surgical context injection.","curator_note":"The most architecturally ambitious take on agent memory we've mapped: seven distinct layers from raw vectors to an auto-curated wiki, each with its own recall path, so the agent gets the RIGHT kind of memory injected rather than a similarity dump. The catch is coupling: it's built FOR Hermes Agent — adopting the architecture elsewhere means porting, not installing; and 7 layers is real operational surface for a ~1.3k-star project.","edges":[{"to":"memmachine","type":"alternative","why":"Both give agents durable, structured memory; MemMachine is the framework-agnostic layer with SDKs and MCP, memory-os the deeper 7-layer architecture wedded to Hermes Agent.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.073Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"memmachine","why":"Both give agents durable, structured memory; MemMachine is the framework-agnostic layer with SDKs and MCP, memory-os the deeper 7-layer architecture wedded to Hermes Agent.","dir":"out","confidence":0.55,"name":"MemMachine"},{"slug":"lightmem","why":"Layered memory operating systems from opposite cultures: memory-os is a 7-layer production system for one agent stack, LightMem is a modular research framework you assemble per experiment.","dir":"in","confidence":0.5,"name":"LightMem"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/ClaudioDrews/memory-os"},"memsearch":{"name":"memsearch","owner":"zilliztech","slug":"memsearch","stars":2345,"image":"https://raw.githubusercontent.com/zilliztech/memsearch/main/assets/logo-icon.jpg","avatar":"https://avatars.githubusercontent.com/u/18416694?v=4&s=96","forks":206,"language":"Python","license":"MIT","updated":"2 days ago","topics":["memory","coding"],"summary":"Zilliz's unified memory for coding agents: one Markdown + Milvus store shared across Claude Code, Codex, OpenCode and OpenClaw — hybrid search, plus repeated workflows distilled into skills.","curator_note":"The cross-platform play is the point: a conversation in Claude Code becomes searchable context in Codex, OpenCode and OpenClaw — one memory, four plugins, zero per-agent setup. Memories live in readable Markdown (greppable, versionable) with Milvus doing hybrid search, and the standout feature is procedural: it watches for workflows you repeat and distills them into installable skills, maintained in the background. NOT for single-agent loyalists — if you only run Claude Code, opencode-mem-style plugins are lighter — and it's a Zilliz project: the Milvus dependency is also the funnel; check what 'backed by Milvus' costs you operationally before teams adopt.","edges":[{"to":"opencode-mem","type":"alternative","why":"Same job — persistent cross-session memory for coding agents — different reach: opencode-mem goes deep on one host (OpenCode, local SQLite, nothing leaves the machine); memsearch spans four hosts with one shared Markdown+Milvus store.","confidence":0.75,"status":"approved"},{"to":"claude-reflect","type":"alternative","why":"Both mine your sessions into reusable assets: claude-reflect captures corrections into CLAUDE.md and /reflect-skills commands for Claude Code; memsearch's skills-from-memory does the same distillation continuously, across four agent platforms.","confidence":0.6,"status":"approved"},{"to":"memmolt","type":"alternative","why":"Both are structured long-term memory with hybrid search for agents: MemMolt enforces a bucket-thread-memo hierarchy in one SQLite file over MCP; memsearch bets on free-form Markdown + Milvus and per-agent plugins.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T14:13:33.388Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"opencode-mem","why":"Same job — persistent cross-session memory for coding agents — different reach: opencode-mem goes deep on one host (OpenCode, local SQLite, nothing leaves the machine); memsearch spans four hosts with one shared Markdown+Milvus store.","dir":"out","confidence":0.75,"name":"opencode-mem"},{"slug":"hivemind","why":"Both distill agent sessions into shared, reusable skills across multiple agent hosts. memsearch is local-first, one machine, Markdown+Milvus; Hivemind is cloud-backed and team-scoped — your colleague's agent learns from yours.","dir":"in","confidence":0.75,"name":"hivemind"},{"slug":"memanto","why":"Same job — one persistent memory shared across your coding agents — opposite infrastructure bets: memsearch runs Markdown + Milvus with embedding models; Memanto runs an information-theoretic engine with no vector DB and no keys at all.","dir":"in","confidence":0.7,"name":"memanto"},{"slug":"claude-reflect","why":"Both mine your sessions into reusable assets: claude-reflect captures corrections into CLAUDE.md and /reflect-skills commands for Claude Code; memsearch's skills-from-memory does the same distillation continuously, across four agent platforms.","dir":"out","confidence":0.6,"name":"claude-reflect"},{"slug":"memmolt","why":"Both are structured long-term memory with hybrid search for agents: MemMolt enforces a bucket-thread-memo hierarchy in one SQLite file over MCP; memsearch bets on free-form Markdown + Milvus and per-agent plugins.","dir":"out","confidence":0.55,"name":"MemMolt"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/zilliztech/memsearch"},"mesh-llm":{"name":"mesh-llm","owner":"Mesh-LLM","slug":"mesh-llm","stars":2835,"image":"https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/docs/mesh-llm-wordmark.png","avatar":"https://avatars.githubusercontent.com/u/273933736?v=4&s=96","forks":338,"language":"Rust","license":"Apache-2.0","updated":"yesterday","topics":["local"],"summary":"Distributed LLM inference in Rust: pool GPUs across machines into one OpenAI-compatible endpoint — local fit first, mesh routing, and stage splits for models too large for any single box.","curator_note":"The 'LLM for the people' play: friends or a homelab pool mid-range GPUs and serve models none of them could run alone, with public meshes discoverable via Nostr. The Skippy stage-split design is genuinely clever. Experimental distributed systems — expect rough edges, and never treat a public mesh as private infrastructure. One box that fits your model? Just run Ollama.","edges":[{"to":"ollama","type":"alternative","why":"Both give you a local OpenAI-compatible endpoint for GGUF-family models; Ollama serves from one machine, mesh-llm pools many and routes to whoever can serve.","confidence":0.6,"status":"approved"},{"to":"airllm","type":"alternative","why":"Opposite answers to 'the model doesn't fit': AirLLM streams layers from disk on one small GPU (slow, solo), mesh-llm splits stages across peers' GPUs (faster, needs friends).","confidence":0.55,"status":"approved"},{"to":"llm-d","type":"alternative","why":"Distributed inference at opposite trust levels: llm-d orchestrates vLLM across a Kubernetes cluster you own; mesh-llm federates volunteer boxes over the open internet.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:42.202Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"colibri","why":"Both attack 'model bigger than your box': mesh-llm scales OUT by pooling GPUs across machines into one endpoint, colibrì scales DOWN by deepening one machine's memory hierarchy to disk.","dir":"in","confidence":0.65,"name":"colibri"},{"slug":"ollama","why":"Both give you a local OpenAI-compatible endpoint for GGUF-family models; Ollama serves from one machine, mesh-llm pools many and routes to whoever can serve.","dir":"out","confidence":0.6,"name":"Ollama"},{"slug":"airllm","why":"Opposite answers to 'the model doesn't fit': AirLLM streams layers from disk on one small GPU (slow, solo), mesh-llm splits stages across peers' GPUs (faster, needs friends).","dir":"out","confidence":0.55,"name":"airllm"},{"slug":"llm-d","why":"Distributed inference at opposite trust levels: llm-d orchestrates vLLM across a Kubernetes cluster you own; mesh-llm federates volunteer boxes over the open internet.","dir":"out","confidence":0.5,"name":"llm-d"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Mesh-LLM/mesh-llm"},"metagpt":{"name":"MetaGPT","owner":"FoundationAgents","slug":"metagpt","stars":69484,"image":"https://raw.githubusercontent.com/FoundationAgents/metagpt/main/docs/resources/MetaGPT-new-log.png","avatar":"https://avatars.githubusercontent.com/u/198047230?v=4&s=96","forks":8862,"language":"Python","license":"MIT","updated":"6 months ago","topics":["agents","orchestration"],"summary":"The original \"AI software company\" multi-agent framework — role-assigned agents (PM, architect, engineer) turn a one-line requirement into PRD, design, and code.","curator_note":"MetaGPT is the canonical demonstration that multi-agent choreography works: encode a software company's SOPs as roles and watch one prompt become a PRD, a design doc, and running code. Study it for that — the role/SOP pattern shows up everywhere now. But be honest about 2026: agentic coding tools (Claude Code, Codex) have eaten the \"write my app from a prompt\" job, and the repo's cadence has slowed noticeably (last push Jan 2026) while the team focuses on its commercial MGX product. For production pipelines you control, reach for LangGraph; for lightweight role crews, CrewAI. Reach for MetaGPT to learn from the most complete SOP-driven design in the wild.","edges":[{"to":"crewai","type":"alternative","why":"Both bet on role-playing agent teams. CrewAI gives you general-purpose crews you compose; MetaGPT ships one opinionated, SOP-encoded software company end to end.","confidence":0.85,"status":"approved"},{"to":"autogen","type":"alternative","why":"Same job — multi-agent LLM apps — opposite philosophy: AutoGen is a conversation substrate you shape freely; MetaGPT hard-codes the workflow as company SOPs.","confidence":0.8,"status":"approved"},{"to":"langgraph","type":"alternative","why":"Trade-off runs the other way: LangGraph is low-level graph control you assemble; MetaGPT is a pre-built company you configure. Outgrowing MetaGPT's opinions usually lands you here.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"crewai","why":"Both bet on role-playing agent teams. CrewAI gives you general-purpose crews you compose; MetaGPT ships one opinionated, SOP-encoded software company end to end.","dir":"out","confidence":0.85,"name":"CrewAI"},{"slug":"autogen","why":"Same job — multi-agent LLM apps — opposite philosophy: AutoGen is a conversation substrate you shape freely; MetaGPT hard-codes the workflow as company SOPs.","dir":"out","confidence":0.8,"name":"AutoGen"},{"slug":"langgraph","why":"Trade-off runs the other way: LangGraph is low-level graph control you assemble; MetaGPT is a pre-built company you configure. Outgrowing MetaGPT's opinions usually lands you here.","dir":"out","confidence":0.6,"name":"LangGraph"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/FoundationAgents/metagpt"},"metatrader-mcp-server":{"name":"metatrader-mcp-server","owner":"ariadng","slug":"metatrader-mcp-server","stars":658,"image":"https://raw.githubusercontent.com/ariadng/metatrader-mcp-server/main/docs/media/hero.webp","avatar":"https://avatars.githubusercontent.com/u/57926750?v=4&s=96","forks":221,"language":"Python","license":"MIT","updated":"3 months ago","topics":["finance"],"summary":"MCP server that puts MetaTrader 5 in an LLM's hands — natural-language trading, account and market data, order management over stdio/SSE, plus REST and WebSocket quotes. 32 tools, MIT.","curator_note":"The shortest path from 'Claude, close all profitable positions' to it actually happening: pip install, point Claude Desktop/Code at it — or run the SSE server on the Windows VPS where MT5 lives — and the bundled /trading skill teaches the agent all 32 tools plus MT5 domain conventions. Treat it like a loaded weapon: the MCP layer has NO authentication (firewall it, tunnel it, or keep it on localhost), MT5 is Windows-only, and an LLM with live order-execution powers is a risk decision, not a convenience — start on a demo account and read the repo's own disclaimer. NOT a strategy or an analysis engine; it's the broker bridge — the hands, not the brain.","edges":[{"to":"vibe-trading","type":"alternative","why":"Same job — give your agent real trading capability over MCP — different scope: Vibe-Trading is a full personal trading-agent stack with market analysis and a shadow-account safety mode; metatrader-mcp-server is the raw MT5 broker bridge, bring your own judgment.","confidence":0.55,"status":"approved"},{"to":"finance-skills","type":"complements","why":"Skills and hands: finance-skills teaches Claude financial-analysis workflows on the agentskills standard; this MCP server is the execution layer that turns those conclusions into actual MT5 orders.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.620Z","linkCount":2,"related":{"complements":[{"slug":"finance-skills","why":"Skills and hands: finance-skills teaches Claude financial-analysis workflows on the agentskills standard; this MCP server is the execution layer that turns those conclusions into actual MT5 orders.","dir":"out","confidence":0.5,"name":"finance-skills"}],"alternative":[{"slug":"vibe-trading","why":"Same job — give your agent real trading capability over MCP — different scope: Vibe-Trading is a full personal trading-agent stack with market analysis and a shadow-account safety mode; metatrader-mcp-server is the raw MT5 broker bridge, bring your own judgment.","dir":"out","confidence":0.55,"name":"Vibe-Trading"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/ariadng/metatrader-mcp-server"},"mindshub":{"name":"mindshub","owner":"mindsdb","slug":"mindshub","stars":39473,"image":"https://github.com/user-attachments/assets/048761b8-aa77-4506-9c4d-32e2fdecbb60","avatar":"https://avatars.githubusercontent.com/u/31035808?v=4&s=96","forks":6223,"language":"Makefile","license":"MIT","updated":"14 days ago","topics":["agents"],"summary":"MindsDB's pivot: a unified 'Cowork' workspace where you delegate whole projects — apps, research, analysis, scheduled operations — to open-source models you can swap anytime.","curator_note":"The pitch is model-agnostic delegation: hand it a project, keep your artifacts and workflows when you swap the underlying model — a real answer to model-vendor lock-in anxiety. Read the star count honestly though: ~39k stars are inherited from this repo's previous life as the MindsDB database product; the Cowork workspace is a young pivot wearing an old repo's reputation. Expect open-core gravity (console, pricing links throughout) and a product still finding its shape. NOT a building block — it's a destination workspace; if you're composing your own stack, the frameworks here compose, this doesn't.","edges":[{"to":"openmanus","type":"alternative","why":"Both answer 'delegate the whole task to an agent' — OpenManus as a minimal open reference you run yourself, MindsHub as a productized workspace with scheduling and a console.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:12.884Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"openmanus","why":"Both answer 'delegate the whole task to an agent' — OpenManus as a minimal open reference you run yourself, MindsHub as a productized workspace with scheduling and a console.","dir":"out","confidence":0.55,"name":"OpenManus"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/mindsdb/mindshub"},"mineru":{"name":"MinerU","owner":"opendatalab","slug":"mineru","stars":75523,"image":"https://gcore.jsdelivr.net/gh/opendatalab/MinerU@master/docs/images/MinerU-logo.png","avatar":"https://avatars.githubusercontent.com/u/97503431?v=4&s=96","forks":6342,"language":"Python","license":"NOASSERTION","updated":"yesterday","topics":["ocr","rag"],"summary":"Heavyweight document-to-markdown/JSON parser — PDFs plus Office (docx/pptx/xlsx) through layout analysis and OCR into LLM-ready output for RAG and agentic pipelines. 73k stars, self-hostable.","curator_note":"The incumbent when document variety is the problem: beyond PDFs it handles Office formats, with mature layout analysis (reading order, tables, formulas) and a huge user base shaking out edge cases. NOT the lightest option — it's a full pipeline with model downloads and real hardware appetite; for a handful of clean PDFs a smaller tool is faster to stand up. Check the license (NOASSERTION on GitHub — AGPL-family, matters for commercial use).","edges":[{"to":"olmocr","type":"alternative","why":"Same job — self-hosted document→markdown for LLM ingestion: olmocr is a focused VLM PDF-linearizer; MinerU is the broader pipeline (PDF + Office, layout analysis) with far more traction.","confidence":0.8,"status":"approved"},{"to":"unlimited-ocr","type":"alternative","why":"Both turn documents into LLM-ready text: Unlimited-OCR is a one-shot long-horizon VLM; MinerU is a staged layout-analysis + OCR pipeline that also ingests Office formats.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-07T22:49:34.749Z","linkCount":9,"related":{"complements":[],"alternative":[{"slug":"olmocr","why":"Same job — self-hosted document→markdown for LLM ingestion: olmocr is a focused VLM PDF-linearizer; MinerU is the broader pipeline (PDF + Office, layout analysis) with far more traction.","dir":"out","confidence":0.8,"name":"olmocr"},{"slug":"xberg","why":"Both turn documents (PDFs, Office) into LLM-ready text/JSON for RAG. MinerU is a heavyweight Python/VLM parser tuned for max-fidelity layout + OCR; Xberg is a lightweight polyglot engine spanning 96 formats and 15 language bindings with pluggable OCR. Pick MinerU for the hardest scanned/complex PDFs, Xberg for breadth and multi-language embedding.","dir":"in","confidence":0.8,"name":"xberg"},{"slug":"unlimited-ocr","why":"Both turn documents into LLM-ready text: Unlimited-OCR is a one-shot long-horizon VLM; MinerU is a staged layout-analysis + OCR pipeline that also ingests Office formats.","dir":"out","confidence":0.75,"name":"Unlimited-OCR"},{"slug":"chunkr","why":"Same job — complex documents into LLM-ready data. MinerU is the batteries-included extraction toolkit; Chunkr is an API-shaped service adding semantic chunking and bounding-box citations for RAG pipelines.","dir":"in","confidence":0.75,"name":"chunkr"},{"slug":"opendataloader-pdf","why":"Same heavyweight PDF-to-structured-data slot: MinerU throws ML layout analysis and OCR at everything; OpenDataLoader is deterministic-first with an optional hybrid AI mode — and beats it on extraction accuracy in its published bench at a fraction of the compute.","dir":"in","confidence":0.7,"name":"opendataloader-pdf"},{"slug":"pixelrag","why":"Two answers to the same RAG-ingestion problem: MinerU parses PDFs/Office through layout analysis and OCR into LLM-ready markdown; PixelRAG skips parsing entirely and retrieves over rendered screenshot tiles. Parse-to-text vs stay-in-pixels.","dir":"in","confidence":0.7,"name":"PixelRAG"},{"slug":"chandra","why":"Overlapping document-to-markdown job at different layers: MinerU is a full parsing pipeline (layout analysis + OCR + export) you deploy as tooling; Chandra is the single end-to-end OCR model you'd slot into such a pipeline.","dir":"in","confidence":0.55,"name":"chandra"},{"slug":"pdf-inspector","why":"Same PDF-to-Markdown slot, opposite weight: MinerU runs full layout analysis and OCR on everything; pdf-inspector is the 200ms no-ML path for documents that don't need it. Many pipelines should front MinerU with exactly this triage.","dir":"in","confidence":0.55,"name":"pdf-inspector"}],"built_with":[{"slug":"sie","why":"MinerU ships in SIE's model catalog as a backend for the document-to-markdown task — SIE serves it (alongside GLM-OCR, PaddleOCR-VL and docling) behind its unified API.","dir":"in","confidence":0.6,"name":"sie"}]},"url":"https://stackmap.shipwithai.xyz/repos/opendatalab/mineru"},"mini-swe-agent":{"name":"mini-swe-agent","owner":"SWE-agent","slug":"mini-swe-agent","stars":5986,"image":"https://raw.githubusercontent.com/SWE-agent/mini-swe-agent/main/docs/assets/mini-swe-agent-banner.svg","avatar":"https://avatars.githubusercontent.com/u/166046056?v=4&s=96","forks":831,"language":"Python","license":"MIT","updated":"yesterday","topics":["coding","agents"],"summary":"The 100-line agent from the SWE-bench team: >74% on SWE-bench Verified with no tools but bash, no config sprawl — the reference minimal harness, adopted by Meta, NVIDIA and Ramp.","curator_note":"The existence proof that most harness complexity is optional: the team that built SWE-bench and SWE-agent asked what a 100x simpler agent loses — the answer is almost nothing (>74% Verified), which is why it became the standard baseline harness for benchmarking models (Ramp's SWE-bench, DeepSWE — where it beats Claude Code and Codex as a harness). Read it to understand agents; use it to evaluate models fairly. NOT a daily driver: no MCP, no skills, no IDE plumbing — by design. If you're choosing a tool to ship features with, this is the control group, not the product.","edges":[{"to":"litellm","type":"built_with","why":"The model layer is a LiteLLM wrapper — one small file gives the 100-line agent every provider LiteLLM speaks, which is exactly the minimalism the project preaches.","confidence":0.7,"status":"approved"},{"to":"openmanus","type":"alternative","why":"Both are open, minimal-core autonomous agents you can actually read. OpenManus goes general (browsing, tool use, multi-step tasks); mini-swe-agent goes narrow and measurable — software engineering with bash only, scored on SWE-bench.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.817Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"openmanus","why":"Both are open, minimal-core autonomous agents you can actually read. OpenManus goes general (browsing, tool use, multi-step tasks); mini-swe-agent goes narrow and measurable — software engineering with bash only, scored on SWE-bench.","dir":"out","confidence":0.55,"name":"OpenManus"}],"built_with":[{"slug":"litellm","why":"The model layer is a LiteLLM wrapper — one small file gives the 100-line agent every provider LiteLLM speaks, which is exactly the minimalism the project preaches.","dir":"out","confidence":0.7,"name":"litellm"}]},"url":"https://stackmap.shipwithai.xyz/repos/SWE-agent/mini-swe-agent"},"mirage":{"name":"mirage","owner":"strukto-ai","slug":"mirage","stars":3346,"image":"https://raw.githubusercontent.com/strukto-ai/mirage/main/assets/mirage-og-light@2x.png","avatar":"https://avatars.githubusercontent.com/u/248286111?v=4&s=96","forks":242,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["agents","coding"],"summary":"Unified virtual filesystem for AI agents — mounts S3, Slack, Gmail, Postgres and ~50 backends as one tree so any bash-speaking LLM can grep and pipe across services. Snapshotable, embeddable.","curator_note":"One interface instead of N SDKs and M MCP servers: mount S3, Slack, Gmail, GitHub, Postgres, Redis (~50 backends) side-by-side as one filesystem, and any LLM that already knows bash can cat, grep and pipe across all of them with zero new tool vocabulary — plus per-resource command overrides (cat a Parquet as JSON rows), snapshotable/portable workspaces, and in-process Python/TS SDKs with OpenAI Agents/LangChain/Pydantic AI adapters and a CLI+daemon for Claude Code/Codex. Reach for it when your agent touches many external systems and tool-schema sprawl is eating your context window. NOT for a single data source (just use its SDK), heavy binary/streaming workloads, or Windows — FUSE mounts need macOS/Linux.","edges":[{"to":"deepagents","type":"complements","why":"deepagents deliberately gives its agents a filesystem + shell as core tools; mirage extends exactly that interface to real infrastructure — S3, Slack, Gmail, Postgres mounted as paths via its LangChain-family adapters. The agent keeps grep/cat/pipe semantics while reaching production data instead of a local scratch dir.","confidence":0.55,"status":"approved"},{"to":"cubesandbox","type":"complements","why":"Two halves of the agent environment: CubeSandbox isolates the compute (hardware-isolated microVM per agent), mirage unifies the data plane (every external service as one mounted tree). Run the agent in the sandbox, hand it its world as files — neither replaces the other.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-08T16:13:44.129Z","linkCount":3,"related":{"complements":[{"slug":"deepagents","why":"deepagents deliberately gives its agents a filesystem + shell as core tools; mirage extends exactly that interface to real infrastructure — S3, Slack, Gmail, Postgres mounted as paths via its LangChain-family adapters. The agent keeps grep/cat/pipe semantics while reaching production data instead of a local scratch dir.","dir":"out","confidence":0.55,"name":"deepagents"},{"slug":"cubesandbox","why":"Two halves of the agent environment: CubeSandbox isolates the compute (hardware-isolated microVM per agent), mirage unifies the data plane (every external service as one mounted tree). Run the agent in the sandbox, hand it its world as files — neither replaces the other.","dir":"out","confidence":0.5,"name":"CubeSandbox"}],"alternative":[{"slug":"ktx","why":"Opposite philosophies for the same job — letting agents work with your company's data. Mirage mounts ~50 backends as a raw virtual filesystem to grep and pipe; ktx curates a semantic layer with approved metric definitions and compiled read-only SQL.","dir":"in","confidence":0.55,"name":"ktx"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/strukto-ai/mirage"},"mlx-lora-studio":{"name":"MLX-LoRA-Studio","owner":"Goekdeniz-Guelmez","slug":"mlx-lora-studio","stars":250,"image":"https://raw.githubusercontent.com/Goekdeniz-Guelmez/mlx-lora-studio/main/Sources/Media/logo_ultra-wide.png","avatar":"https://avatars.githubusercontent.com/u/60228478?v=4&s=96","forks":25,"language":"Swift","license":"MIT","updated":"8 days ago","topics":["training","local"],"summary":"Native Mac app for on-device LLM fine-tuning via mlx-lm-lora: pick a model, choose SFT/LoRA/DPO-family algorithms, watch loss fall live, push to Hugging Face. No cloud, no code.","curator_note":"The shortest path from base model to tuned adapter on a Mac: no Python env, no GPU rental — a native SwiftUI app driving mlx-lm-lora, with an algorithm menu unusually deep for a GUI (SFT through DPO-family preference methods, FTPO, Dynamic Fine-Tuning, memory-bounded Chunked NLL), live loss curves and direct HF upload. NOT for serious scale: your ceiling is unified memory on one machine — multi-GPU and large post-training runs belong in LLaMA-Factory/TRL territory. Young single-maintainer project (~250 stars) with a quarantine-xattr dance on first launch; Apple Silicon only, by design.","edges":[{"to":"llamafactory","type":"alternative","why":"Both GUI-driven fine-tuning across many methods. LLaMA-Factory is the CUDA-world standard (100+ models, LlamaBoard, cluster-ready); MLX LoRA Studio trades that breadth for a native Mac app that trains entirely on Apple Silicon.","confidence":0.7,"status":"approved"},{"to":"h2o-llmstudio","type":"alternative","why":"Same promise — fine-tune without writing code — different homes: H2O LLM Studio is a web UI/Docker framework for GPU boxes; MLX LoRA Studio is a native macOS app for the machine on your desk.","confidence":0.6,"status":"approved"},{"to":"siliconscope","type":"complements","why":"Fine-tuning on Apple Silicon is a unified-memory balancing act — SiliconScope shows the GPU/ANE load, memory pressure and bandwidth of a Studio training run live, so you can size batch and quantization to the machine instead of guessing.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-20T09:51:35.584Z","linkCount":4,"related":{"complements":[{"slug":"siliconscope","why":"Fine-tuning on Apple Silicon is a unified-memory balancing act — SiliconScope shows the GPU/ANE load, memory pressure and bandwidth of a Studio training run live, so you can size batch and quantization to the machine instead of guessing.","dir":"out","confidence":0.65,"name":"SiliconScope"},{"slug":"coreai-models","why":"The two halves of Apple-silicon on-device AI: fine-tune your model in MLX LoRA Studio, then export through Core AI recipes to ship it inside a macOS/iOS app. Research stack in, deployment stack out.","dir":"in","confidence":0.55,"name":"coreai-models"}],"alternative":[{"slug":"llamafactory","why":"Both GUI-driven fine-tuning across many methods. LLaMA-Factory is the CUDA-world standard (100+ models, LlamaBoard, cluster-ready); MLX LoRA Studio trades that breadth for a native Mac app that trains entirely on Apple Silicon.","dir":"out","confidence":0.7,"name":"LlamaFactory"},{"slug":"h2o-llmstudio","why":"Same promise — fine-tune without writing code — different homes: H2O LLM Studio is a web UI/Docker framework for GPU boxes; MLX LoRA Studio is a native macOS app for the machine on your desk.","dir":"out","confidence":0.6,"name":"h2o-llmstudio"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Goekdeniz-Guelmez/mlx-lora-studio"},"my-brain-is-full-crew":{"name":"My-Brain-Is-Full-Crew","owner":"gnekt","slug":"my-brain-is-full-crew","stars":3241,"avatar":"https://avatars.githubusercontent.com/u/39567665?v=4&s=96","forks":334,"language":"Shell","license":"NOASSERTION","updated":"27 days ago","topics":["agents","skills"],"summary":"8 AI agents + 14 skills that run your Obsidian vault through chat — capture, triage, search, linking, vault health, transcription, email and calendar. One codebase, four agent platforms, any language.","curator_note":"Built from lived need — a PhD researcher whose memory was slipping — and it shows in the design: this is for people who are drowning, not optimizing. The chat IS the interface (no folder dragging), a dispatcher chains agents automatically (transcribe a meeting → Architect creates the project structure), it answers in whatever language you speak, and 'create a new agent' is a guided conversation, not a config file. Installs onto Claude Code, Gemini CLI, OpenCode or Codex from one source. NOT a note-taking plugin: it assumes you hand vault management over entirely, and email/calendar agents mean third-party data — read the disclaimers, they're unusually serious (GitHub reports no standard license; usage is Terms-of-Use gated). Pair judgment: rowboat is the app-shaped version of this idea; this is the vault-shaped one.","edges":[{"to":"core","type":"alternative","why":"Both are self-hosted 'manage my life' AI systems with persistent memory: core is a product that watches your apps and acts within guardrails on its own memory graph; the Crew works entirely through your Obsidian vault, keeping the memory human-readable and human-editable.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.763Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"core","why":"Both are self-hosted 'manage my life' AI systems with persistent memory: core is a product that watches your apps and acts within guardrails on its own memory graph; the Crew works entirely through your Obsidian vault, keeping the memory human-readable and human-editable.","dir":"out","confidence":0.55,"name":"core"},{"slug":"rowboat","why":"Two shapes of the same idea — an AI that manages your work life on top of a persistent, human-readable memory: Rowboat is a desktop app with its own surfaces (email, browser, meetings) over a Markdown graph; the Crew is a team of agents living inside your existing Obsidian vault, driven purely by chat.","dir":"in","confidence":0.55,"name":"rowboat"},{"slug":"personal_ai_infrastructure","why":"Manage-your-life AI at different scopes: an 8-agent crew over an Obsidian vault vs a whole-life operating system organized around your goals (TELOS).","dir":"in","confidence":0.5,"name":"LifeOS"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/gnekt/my-brain-is-full-crew"},"n-skills":{"name":"n-skills","owner":"numman-ali","slug":"n-skills","stars":1031,"image":"https://raw.githubusercontent.com/numman-ali/n-skills/main/assets/logo.svg","avatar":"https://avatars.githubusercontent.com/u/35879093?v=4&s=96","forks":106,"language":"TypeScript","license":"Apache-2.0","updated":"8 days ago","topics":["skills"],"summary":"Curated skill marketplace for AI coding agents on the universal SKILL.md/AGENTS.md format — write a skill once, install it into Claude Code, Codex, Copilot, Cursor and friends.","curator_note":"The 'write once, run everywhere' bet applied to agent skills: one curated marketplace on the SKILL.md + AGENTS.md conventions, installed via openskills into whichever agent you drive. Smaller and more opinionated than the mega-registries — curation IS the pitch, every skill is hand-picked. NOT a discovery engine: if you want breadth across hundreds of sources, a package manager with security scanning covers more ground; and it's one person's taste with ~1k stars — check the skills you'd rely on exist before adopting the workflow.","edges":[{"to":"skillkit","type":"alternative","why":"Same job — getting skills into any agent's format. skillkit is the breadth play (400K+ skills, 31 sources, security scans); n-skills is the depth play: one small marketplace where a human curated every entry.","confidence":0.65,"status":"approved"},{"to":"asm","type":"alternative","why":"Both install and manage agent skills across providers; asm is the scriptable power-tool with audit/dedupe, n-skills the curated storefront on the universal format.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-12T23:46:15.906Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"skillkit","why":"Same job — getting skills into any agent's format. skillkit is the breadth play (400K+ skills, 31 sources, security scans); n-skills is the depth play: one small marketplace where a human curated every entry.","dir":"out","confidence":0.65,"name":"skillkit"},{"slug":"autoskills","why":"Two curation-first answers to skill installation: autoskills detects your stack and decides for you from a hash-pinned registry; n-skills hands you a small human-curated marketplace on the universal SKILL.md format and lets you pick.","dir":"in","confidence":0.6,"name":"autoskills"},{"slug":"asm","why":"Both install and manage agent skills across providers; asm is the scriptable power-tool with audit/dedupe, n-skills the curated storefront on the universal format.","dir":"out","confidence":0.55,"name":"asm"},{"slug":"skillnet","why":"Both distribute reusable skills on the SKILL.md format; n-skills is a small curated marketplace, SkillNet is a 500K+ crawled-and-deduplicated index with quality ranking.","dir":"in","confidence":0.55,"name":"SkillNet"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/numman-ali/n-skills"},"nemo-guardrails":{"name":"Guardrails","owner":"NVIDIA-NeMo","slug":"nemo-guardrails","stars":6774,"image":"https://raw.githubusercontent.com/NVIDIA-NeMo/Guardrails/develop/docs/_static/images/programmable_guardrails.png","avatar":"https://avatars.githubusercontent.com/u/213689629?v=4&s=96","forks":780,"language":"Python","license":"NOASSERTION","updated":"yesterday","topics":["security"],"summary":"NVIDIA's programmable guardrails for LLM apps: input, output, dialog and retrieval rails defined in Colang, wrapping any model or LangChain runnable.","curator_note":"For when policy needs to be programmable, not a blocklist: topic bans, jailbreak checks, tool-use constraints and dialog flows written in Colang, enforced as input/output/dialog/retrieval rails around any LLM — RunnableRails drops it straight into a LangChain pipeline. NOT free at runtime: every rail is extra LLM calls and latency, Colang is its own language to learn, and rails mitigate rather than guarantee — you still red-team the result (that's garak's job).","edges":[{"to":"langgraph","type":"complements","why":"RunnableRails wraps LangChain/LangGraph runnables — the agent graph does the work, the rails police what goes in and out of every LLM call.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.088Z","linkCount":4,"related":{"complements":[{"slug":"garak","why":"Attack and defend: garak red-teams the model to find jailbreaks and injection holes, NeMo Guardrails is the runtime rail you deploy to police them — scan, patch rails, re-scan.","dir":"in","confidence":0.8,"name":"garak"},{"slug":"langgraph","why":"RunnableRails wraps LangChain/LangGraph runnables — the agent graph does the work, the rails police what goes in and out of every LLM call.","dir":"out","confidence":0.7,"name":"LangGraph"},{"slug":"litellm","why":"Pair a dedicated programmable-guardrails layer with LiteLLM's gateway when input/output rails need to be richer than the built-in checks.","dir":"in","confidence":0.48,"name":"litellm"}],"alternative":[{"slug":"plano","why":"Two places to enforce guardrails: NeMo Guardrails runs Colang rails in-process around a model or chain, Plano enforces moderation/jailbreak filters at the proxy so every agent inherits them without code changes.","dir":"in","confidence":0.6,"name":"plano"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/NVIDIA-NeMo/nemo-guardrails"},"npxskillui":{"name":"npxskillui","owner":"amaancoderx","slug":"npxskillui","stars":1080,"image":"https://raw.githubusercontent.com/amaancoderx/npxskillui/main/skillui.png","avatar":"https://avatars.githubusercontent.com/u/89069951?v=4&s=96","forks":113,"language":"TypeScript","license":null,"updated":"2 months ago","topics":["skills"],"summary":"Reverse-engineers any design system into a Claude-ready skill via pure static analysis — no AI, no API keys: point it at your codebase, get a skill that teaches your agent your UI.","curator_note":"Solves the 'my agent writes someone else's design system' problem mechanically: static analysis over your components and tokens produces a skill file that teaches the agent YOUR conventions — deterministic, free, no model in the loop. The no-AI approach is the feature AND the ceiling: it captures what's statically visible (tokens, component APIs), not the taste and intent behind them; regenerate when the design system moves or the skill quietly rots.","edges":[{"to":"designer-skills","type":"complements","why":"designer-skills teaches an agent design craft in general; npxskillui extracts YOUR specific design system into a skill — general literacy plus house style.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.106Z","linkCount":1,"related":{"complements":[{"slug":"designer-skills","why":"designer-skills teaches an agent design craft in general; npxskillui extracts YOUR specific design system into a skill — general literacy plus house style.","dir":"out","confidence":0.5,"name":"designer-skills"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/amaancoderx/npxskillui"},"ollama":{"name":"Ollama","owner":"ollama","slug":"ollama","stars":176685,"avatar":"https://avatars.githubusercontent.com/u/151674099?v=4&s=96","forks":17072,"language":"Go","license":"MIT","updated":"yesterday","topics":["local"],"summary":"Run Llama, Mistral and other open models locally with a single command and a clean API.","edges":[{"to":"vllm","type":"alternative","why":"Both serve open models locally; vLLM optimizes for throughput, Ollama for one-command simplicity.","confidence":0.8,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":21,"related":{"complements":[{"slug":"lynkr","why":"Lynkr's flagship trick is tier routing: SIMPLE/MEDIUM requests go to a local Ollama endpoint (TIER_SIMPLE=ollama:qwen2.5-coder) so only hard tasks hit your paid subscription — the README's recommended free setup is Ollama-first.","dir":"in","confidence":0.85,"name":"Lynkr"},{"slug":"pocket-tts","why":"The fully-offline voice assistant stack: Ollama runs the brain on your machine, pocket-tts gives it a voice on two CPU cores — no GPU, no cloud, community integrations already wire the two together.","dir":"in","confidence":0.75,"name":"pocket-tts"},{"slug":"garak","why":"garak ships an Ollama generator — point the scanner at your locally served model and red-team it before exposing it to users.","dir":"in","confidence":0.65,"name":"garak"},{"slug":"litellm","why":"Native Ollama provider lets local models be called through the same unified interface as cloud APIs.","dir":"in","confidence":0.65,"name":"litellm"},{"slug":"maxkb","why":"MaxKB's private-model story runs through Ollama — point the platform at a local DeepSeek/Llama/Qwen and the whole RAG + agent stack stays on-prem.","dir":"in","confidence":0.65,"name":"MaxKB"},{"slug":"allama","why":"Security workloads are exactly where prompts can't leave the building — Allama's agents run against self-hosted models through Ollama for fully on-prem triage.","dir":"in","confidence":0.6,"name":"allama"},{"slug":"open-llm-vtuber","why":"The fully-local promise runs through a local model server — Ollama is the standard backend that keeps the whole voice-avatar loop on your machine.","dir":"in","confidence":0.6,"name":"Open-LLM-VTuber"},{"slug":"open-notebook","why":"The privacy pitch is only real with a local backend — Ollama is the canonical provider that keeps notebook content, chat and search on your machine end to end.","dir":"in","confidence":0.6,"name":"open-notebook"},{"slug":"searchbox","why":"searchbox drives its agent LLM over any OpenAI-compatible endpoint via LLAMA_URL/MODEL_ID (the reference deployment is a GGUF Qwen on llama.cpp). ollama is the easy way to serve that local chat model behind it — point LLAMA_URL at ollama's OpenAI-compatible API and swap models without touching searchbox.","dir":"in","confidence":0.6,"name":"searchbox"},{"slug":"siliconscope","why":"The question SiliconScope was built to answer — what is my on-device model actually doing to the silicon — comes up the moment Ollama is serving: watch GPU, ANE and bandwidth load per model and quantization instead of guessing.","dir":"in","confidence":0.6,"name":"SiliconScope"},{"slug":"claude-bug-bounty","why":"Standalone mode runs on free local providers — `bughunter setup` offers Ollama as the offline, no-subscription backend so the whole recon/hunt loop works without a paid API.","dir":"in","confidence":0.45,"name":"claude-bug-bounty"}],"alternative":[{"slug":"vllm","why":"Both serve open models locally; vLLM optimizes for throughput, Ollama for one-command simplicity.","dir":"out","confidence":0.8,"name":"vLLM"},{"slug":"colibri","why":"Both run open models locally. Ollama is the multi-model daily driver for models that fit; colibrì is a single-model specialist that makes a 744B MoE fit where nothing else will.","dir":"in","confidence":0.7,"name":"colibri"},{"slug":"airllm","why":"Both run open models on your own hardware, on opposite sides of one constraint: Ollama gives fast, polished local inference for models that fit your VRAM; AirLLM runs models that don't fit at all — 70B on 4GB — by streaming one layer at a time, at heavy latency cost.","dir":"in","confidence":0.6,"name":"airllm"},{"slug":"mesh-llm","why":"Both give you a local OpenAI-compatible endpoint for GGUF-family models; Ollama serves from one machine, mesh-llm pools many and routes to whoever can serve.","dir":"in","confidence":0.6,"name":"mesh-llm"},{"slug":"sie","why":"Same 'serve open models behind one local API' job at different scales: Ollama is the single-machine developer runner; SIE is the multi-model production cluster with autoscaling, gateway and Terraform.","dir":"in","confidence":0.55,"name":"sie"},{"slug":"llm-d","why":"Same job — serve open models on your own hardware — at opposite scales: Ollama is one command on one machine; llm-d is a CNCF stack for multi-node GPU fleets. Outgrow one, reach for the other.","dir":"in","confidence":0.5,"name":"llm-d"}],"built_with":[{"slug":"meetily","why":"Meeting summarization runs on Ollama — Meetily pipes finished transcripts to a local Ollama model to produce structured notes; it's the headline local path (cloud APIs are the optional fallback, not the default).","dir":"in","confidence":0.85,"name":"meetily"},{"slug":"langgraph","why":"Point your graph nodes at a local model — no API keys while iterating.","dir":"in","confidence":0.8,"name":"LangGraph"},{"slug":"crewai","why":"Run your crew against local models for cost-free iteration.","dir":"in","confidence":0.75,"name":"CrewAI"},{"slug":"autogen","why":"Back your agents with a local model server.","dir":"in","confidence":0.7,"name":"AutoGen"}]},"url":"https://stackmap.shipwithai.xyz/repos/ollama/ollama"},"olmocr":{"name":"olmocr","owner":"allenai","slug":"olmocr","stars":19161,"image":"https://github.com/user-attachments/assets/24f1b596-4059-46f1-8130-5d72dcc0b02e","avatar":"https://avatars.githubusercontent.com/u/5667695?v=4&s=96","forks":1578,"language":"Python","license":"Apache-2.0","updated":"4 months ago","topics":["rag","ocr"],"summary":"Open toolkit that linearizes messy PDFs — scans, tables, equations, handwriting — into clean ordered Markdown with a self-hosted vision-language model. Built for LLM training data and RAG ingestion.","curator_note":"Reach for olmocr when you have lots of messy PDFs to convert into high-quality text for a training corpus or RAG index and you have GPU to run the VLM. It is the parsing FRONT-END of a pipeline, not a RAG system itself — pair it with an indexer/retriever (LlamaIndex) and a vector store (Chroma). NOT worth it for clean, digital-native PDFs where a cheap text extractor does the job — running a vision-language model for those is overkill.","edges":[{"to":"llamaindex","type":"complements","why":"olmocr is the document-parsing front-end; LlamaIndex is the indexing/retrieval layer. Run messy PDFs through olmocr to get clean text, then hand that text to LlamaIndex to chunk, embed, and retrieve — a natural two-stage RAG-ingestion pipeline.","confidence":0.6,"status":"approved"},{"to":"vllm","type":"built_with","why":"olmocr runs its OCR vision-language model through a high-throughput inference backend — vLLM (or SGLang) — to batch-process PDFs at scale, so vLLM is the serving engine under olmocr's pipeline.","confidence":0.5,"status":"approved"},{"to":"chroma","type":"complements","why":"olmocr's cleaned text is exactly what you embed and store for retrieval — feed its output into Chroma as the vector store behind a RAG app. Looser than the LlamaIndex pairing since any embedder/DB works, but Chroma is the natural open-source landing spot.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-05T14:49:26.000Z","linkCount":9,"related":{"complements":[{"slug":"llamaindex","why":"olmocr is the document-parsing front-end; LlamaIndex is the indexing/retrieval layer. Run messy PDFs through olmocr to get clean text, then hand that text to LlamaIndex to chunk, embed, and retrieve — a natural two-stage RAG-ingestion pipeline.","dir":"out","confidence":0.6,"name":"LlamaIndex"},{"slug":"chroma","why":"olmocr's cleaned text is exactly what you embed and store for retrieval — feed its output into Chroma as the vector store behind a RAG app. Looser than the LlamaIndex pairing since any embedder/DB works, but Chroma is the natural open-source landing spot.","dir":"out","confidence":0.45,"name":"Chroma"}],"alternative":[{"slug":"mineru","why":"Same job — self-hosted document→markdown for LLM ingestion: olmocr is a focused VLM PDF-linearizer; MinerU is the broader pipeline (PDF + Office, layout analysis) with far more traction.","dir":"in","confidence":0.8,"name":"MinerU"},{"slug":"unlimited-ocr","why":"Same job — self-hosted VLM that linearizes messy PDFs into LLM-ready text: olmocr is AllenAI's page-pipeline toolkit built for training-data ingestion; Unlimited-OCR bets on one-shot long-horizon parsing that preserves cross-page structure.","dir":"in","confidence":0.8,"name":"Unlimited-OCR"},{"slug":"chandra","why":"Head-to-head open OCR-VLM rivals on the same benchmark (Chandra 2 scores 85.8 vs olmOCR 2's 82.4). olmOCR is fully permissive and tuned for LLM-training-data linearization; Chandra leads on handwriting/forms/multilingual but carries an OpenRAIL-M weights license.","dir":"in","confidence":0.75,"name":"chandra"},{"slug":"xberg","why":"Same end goal — clean, ordered Markdown from PDFs for RAG/training ingestion. olmocr is a dedicated self-hosted vision-language model that excels on scans, tables and handwriting; Xberg is a general extraction framework where OCR is one pluggable backend. Use olmocr when OCR quality is the bottleneck, Xberg when you need many formats and language bindings.","dir":"in","confidence":0.7,"name":"xberg"},{"slug":"chunkr","why":"Both parse hard documents for AI consumption; olmocr bets on a VLM end-to-end, Chunkr on a layout-analysis pipeline with structured outputs and citations.","dir":"in","confidence":0.6,"name":"chunkr"},{"slug":"pixelrag","why":"olmocr linearizes messy PDFs into clean ordered Markdown so a text pipeline can index them; PixelRAG argues the linearization step is the loss — embed the page image and let a VLM read the tile directly.","dir":"in","confidence":0.55,"name":"PixelRAG"}],"built_with":[{"slug":"vllm","why":"olmocr runs its OCR vision-language model through a high-throughput inference backend — vLLM (or SGLang) — to batch-process PDFs at scale, so vLLM is the serving engine under olmocr's pipeline.","dir":"out","confidence":0.5,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/allenai/olmocr"},"omnigent":{"name":"omnigent","owner":"omnigent-ai","slug":"omnigent","stars":7654,"image":"https://raw.githubusercontent.com/omnigent-ai/omnigent/main/docs/images/omnigent-logo.svg","avatar":"https://avatars.githubusercontent.com/u/292215228?v=4&s=96","forks":1095,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["agents","orchestration","coding"],"summary":"Open-source meta-harness over Claude Code, Codex, Cursor, OpenCode, Hermes, Pi and custom agents: swap harnesses without rewriting, enforce policies and sandboxes, follow sessions from any device.","curator_note":"Use when you run several coding harnesses and want one control plane: mix agents in a session, gate risky actions with policies, run in Modal/Daytona/E2B/K8s sandboxes, and pick a session up from phone or browser. Team sharing (co-drive, fork) is a real differentiator. NOT if you live in one CLI on one laptop — it's a server + desktop app stack, and it's alpha; expect churn. Apache-2.0.","edges":[{"to":"helmor","type":"alternative","why":"Both orchestrate coding agents from a desktop control plane; Helmor is local-first per-repo workspaces, Omnigent adds multi-device sessions, policies and cloud sandboxes.","confidence":0.8,"status":"approved"},{"to":"squad","type":"alternative","why":"Same job — coordinating multiple AI coding CLIs — opposite philosophy: squad is daemonless one-shot shell + SQLite, Omnigent is a full server/meta-harness.","confidence":0.7,"status":"approved"},{"to":"codenomad","type":"alternative","why":"Both are cockpits for living in AI coding sessions with remote access; CodeNomad is OpenCode-only, Omnigent spans many harnesses.","confidence":0.65,"status":"approved"},{"to":"tailclaude","type":"alternative","why":"Covers tailclaude's use case — driving your coding agent from any browser/phone — but for many harnesses, at the cost of a much heavier stack.","confidence":0.6,"status":"approved"},{"to":"cubesandbox","type":"complements","why":"Omnigent launches sessions in E2B-compatible sandboxes; cubesandbox is a self-hosted E2B-compatible microVM backend to run them on your own KVM nodes.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-20T07:55:01.629Z","linkCount":5,"related":{"complements":[{"slug":"cubesandbox","why":"Omnigent launches sessions in E2B-compatible sandboxes; cubesandbox is a self-hosted E2B-compatible microVM backend to run them on your own KVM nodes.","dir":"out","confidence":0.65,"name":"CubeSandbox"}],"alternative":[{"slug":"helmor","why":"Both orchestrate coding agents from a desktop control plane; Helmor is local-first per-repo workspaces, Omnigent adds multi-device sessions, policies and cloud sandboxes.","dir":"out","confidence":0.8,"name":"helmor"},{"slug":"squad","why":"Same job — coordinating multiple AI coding CLIs — opposite philosophy: squad is daemonless one-shot shell + SQLite, Omnigent is a full server/meta-harness.","dir":"out","confidence":0.7,"name":"squad"},{"slug":"codenomad","why":"Both are cockpits for living in AI coding sessions with remote access; CodeNomad is OpenCode-only, Omnigent spans many harnesses.","dir":"out","confidence":0.65,"name":"CodeNomad"},{"slug":"tailclaude","why":"Covers tailclaude's use case — driving your coding agent from any browser/phone — but for many harnesses, at the cost of a much heavier stack.","dir":"out","confidence":0.6,"name":"tailclaude"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/omnigent-ai/omnigent"},"omnigraph":{"name":"omnigraph","owner":"ModernRelay","slug":"omnigraph","stars":954,"image":"https://raw.githubusercontent.com/ModernRelay/omnigraph/main/assets/omnigraph-wordmark.svg","avatar":"https://avatars.githubusercontent.com/u/206314331?v=4&s=96","forks":138,"language":"Rust","license":"MIT","updated":"2 days ago","topics":["memory","rag"],"summary":"Lakehouse graph database for agent context — graph, vector and full-text retrieval fused in one runtime on branchable Lance/S3 storage; agent fleets write on isolated branches and merge Git-style.","curator_note":"Reach for it when many agents must share one evolving knowledge store and you need blame, rollback and review on their writes — branch-per-agent with merge gates is the feature nothing else in this space has; also strong when retrieval genuinely needs graph + vector + full-text fused, not a vector store with metadata filters. NOT a drop-in vector DB: you take on a server, cluster.yaml, schemas and Cedar policy — for plain RAG recall, Chroma is answering queries before you've finished omnigraph's docs. Young project on a credible Rust/Lance foundation; expect sharp edges and a moving API.","edges":[{"to":"chroma","type":"alternative","why":"Same slot in the stack — the store your AI app's retrieval hits — opposite ends of the spectrum: Chroma is a lightweight embedding DB you outgrow; omnigraph fuses graph traversal, ANN and full-text with versioned branching, at the cost of running a declared-as-code server.","confidence":0.65,"status":"approved"},{"to":"langgraph","type":"complements","why":"LangGraph orchestrates the multi-actor workflow; omnigraph is built as the durable state those actors share — branch per agent or task, merged on review. Natural pairing for fleet-scale memory, though no packaged integration exists yet.","confidence":0.5,"status":"approved"},{"to":"skillkit","type":"complements","why":"Omnigraph ships its operational playbook as an installable agent skill (`npx skills add ModernRelay/omnigraph@omnigraph`) — exactly the artifact skill managers like skillkit install and translate into whatever agent operates your cluster.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-07T14:14:24.102Z","linkCount":5,"related":{"complements":[{"slug":"langgraph","why":"LangGraph orchestrates the multi-actor workflow; omnigraph is built as the durable state those actors share — branch per agent or task, merged on review. Natural pairing for fleet-scale memory, though no packaged integration exists yet.","dir":"out","confidence":0.5,"name":"LangGraph"},{"slug":"skillkit","why":"Omnigraph ships its operational playbook as an installable agent skill (`npx skills add ModernRelay/omnigraph@omnigraph`) — exactly the artifact skill managers like skillkit install and translate into whatever agent operates your cluster.","dir":"out","confidence":0.5,"name":"skillkit"},{"slug":"hyper-extract","why":"Extract with one, store and retrieve with the other: Hyper-Extract turns documents into structured graph knowledge; OmniGraph is the lakehouse runtime that serves graph+vector context to agent fleets.","dir":"in","confidence":0.5,"name":"Hyper-Extract"},{"slug":"codegraph-mcp","why":"Two halves of a graph-context stack for agents: codegraph-mcp serves code topology (symbols, calls, blast radius) over MCP; omnigraph holds the org's general knowledge graph. Conceptual pairing — no packaged integration.","dir":"in","confidence":0.45,"name":"codegraph-mcp"}],"alternative":[{"slug":"chroma","why":"Same slot in the stack — the store your AI app's retrieval hits — opposite ends of the spectrum: Chroma is a lightweight embedding DB you outgrow; omnigraph fuses graph traversal, ANN and full-text with versioned branching, at the cost of running a declared-as-code server.","dir":"out","confidence":0.65,"name":"Chroma"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/ModernRelay/omnigraph"},"open-code-review":{"name":"open-code-review","owner":"alibaba","slug":"open-code-review","stars":11018,"image":"https://raw.githubusercontent.com/alibaba/open-code-review/main/imgs/logo-core.svg","avatar":"https://avatars.githubusercontent.com/u/1961952?v=4&s=96","forks":765,"language":"Go","license":"Apache-2.0","updated":"yesterday","topics":["coding"],"summary":"Alibaba's battle-tested AI code-review CLI: deterministic pipelines + LLM agent, line-level comments, tuned rulesets (NPE, XSS, SQLi) — higher precision than general agents at ~1/9 the tokens.","curator_note":"Two years as Alibaba's internal review assistant before open-sourcing, and it shows: the hybrid architecture (deterministic checks + tool-using LLM agent) is tuned for precision over recall — fewer, better line-level comments instead of noise — and their 200-PR benchmark honestly reports the trade-off, including losing to general agents on recall. `ocr scan` reviews whole files, useful for auditing unfamiliar code with no diff. Bring your own endpoint (OpenAI/Anthropic-compatible). NOT a replacement for your coding agent's judgment on architecture — it's a defect-finder (NPE, thread-safety, XSS, SQLi rulesets), strongest on the bug classes it was fine-tuned for; the deliberate low-recall bias means a clean review is not a clean PR.","edges":[{"to":"pullfrog","type":"alternative","why":"Both put an AI agent on your GitHub PRs. pullfrog is the generalist — @mention it and your own agent runs any task in Actions; Open Code Review is the specialist — reviews only, hybrid deterministic+LLM, precision-tuned line comments.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-22T18:15:22.202Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"pullfrog","why":"Both put an AI agent on your GitHub PRs. pullfrog is the generalist — @mention it and your own agent runs any task in Actions; Open Code Review is the specialist — reviews only, hybrid deterministic+LLM, precision-tuned line comments.","dir":"out","confidence":0.55,"name":"pullfrog"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/alibaba/open-code-review"},"open-llm-vtuber":{"name":"Open-LLM-VTuber","owner":"Open-LLM-VTuber","slug":"open-llm-vtuber","stars":12710,"image":"https://raw.githubusercontent.com/Open-LLM-VTuber/open-llm-vtuber/main/assets/banner.jpg","avatar":"https://avatars.githubusercontent.com/u/193920693?v=4&s=96","forks":1477,"language":"Python","license":"NOASSERTION","updated":"2 months ago","topics":["voice"],"summary":"Hands-free voice conversation with any LLM behind a Live2D animated face — voice interruption included, running fully local and cross-platform.","curator_note":"The most complete open 'talking companion' stack: speech in, LLM of your choice, voice out, and a Live2D avatar that reacts — with real-time interruption, which is the feature that makes voice feel alive and that most stacks skip. Fully local is the point: pair with a local model and nothing leaves your machine. NOT a components library — it's an integrated app; if you only need STT or TTS pieces, take those directly. License resolution was unclear at review time and the repo had a quiet spell — check both before shipping on it.","edges":[{"to":"ollama","type":"complements","why":"The fully-local promise runs through a local model server — Ollama is the standard backend that keeps the whole voice-avatar loop on your machine.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.138Z","linkCount":1,"related":{"complements":[{"slug":"ollama","why":"The fully-local promise runs through a local model server — Ollama is the standard backend that keeps the whole voice-avatar loop on your machine.","dir":"out","confidence":0.6,"name":"Ollama"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Open-LLM-VTuber/open-llm-vtuber"},"open-notebook":{"name":"open-notebook","owner":"lfnovo","slug":"open-notebook","stars":35917,"image":"https://raw.githubusercontent.com/lfnovo/open-notebook/main/docs/assets/hero.svg","avatar":"https://avatars.githubusercontent.com/u/579178?v=4&s=96","forks":4090,"language":"TypeScript","license":"MIT","updated":"2 days ago","topics":["rag","storage"],"summary":"Self-hosted NotebookLM alternative: multi-modal sources, vector plus full-text search, context-aware chat and multi-speaker podcast generation — 18+ model providers incl. Ollama, full REST API.","curator_note":"The default answer to 'NotebookLM but private': drop in PDFs, videos, audio and web pages, search and chat across them, and generate podcasts with 1-4 custom speakers where Google caps you at two. Provider freedom is the real lever — 18+ backends down to Ollama/LM Studio keeps sensitive research fully local, and the REST API makes it automatable where NotebookLM is a closed app. NOT ahead on citations — it concedes NotebookLM's source-grounding is stronger (theirs is 'basic, will improve'), which matters if research integrity is the whole point. A product you deploy (Docker), not a library you embed.","edges":[{"to":"ollama","type":"complements","why":"The privacy pitch is only real with a local backend — Ollama is the canonical provider that keeps notebook content, chat and search on your machine end to end.","confidence":0.6,"status":"approved"},{"to":"karakeep","type":"alternative","why":"Two self-hosted personal knowledge hoards with AI on top: Karakeep is capture-first (bookmark everything, auto-tag, archive), Open Notebook is synthesis-first (research a topic, chat over sources, produce podcasts). Collector vs study desk.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-20T10:48:48.692Z","linkCount":2,"related":{"complements":[{"slug":"ollama","why":"The privacy pitch is only real with a local backend — Ollama is the canonical provider that keeps notebook content, chat and search on your machine end to end.","dir":"out","confidence":0.6,"name":"Ollama"}],"alternative":[{"slug":"karakeep","why":"Two self-hosted personal knowledge hoards with AI on top: Karakeep is capture-first (bookmark everything, auto-tag, archive), Open Notebook is synthesis-first (research a topic, chat over sources, produce podcasts). Collector vs study desk.","dir":"out","confidence":0.55,"name":"karakeep"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/lfnovo/open-notebook"},"opencode-mem":{"name":"opencode-mem","owner":"tickernelz","slug":"opencode-mem","stars":1205,"image":"https://raw.githubusercontent.com/tickernelz/opencode-mem/main/.github/banner.png","avatar":"https://avatars.githubusercontent.com/u/10543415?v=4&s=96","forks":121,"language":"TypeScript","license":null,"updated":"yesterday","topics":["memory","coding"],"summary":"OpenCode plugin giving coding agents persistent cross-session memory — local SQLite + vector search, automatic memory capture, user-profile learning, and a web UI. Nothing leaves your machine.","curator_note":"For OpenCode users tired of re-explaining their architecture every session: auto-capture summarizes each prompt's work via a background structured-output call that reuses your existing opencode provider auth, memories inject into the first chat message, and the `.opencode-mem-project` marker file solves multi-repo workspaces properly (directory-driven identity, not env vars). A real web UI at :4747 for browsing what it learned. NOT for Claude Code — this is OpenCode-specific; pro-workflow and claude-reflect are the equivalents on that side. Auto-capture needs a provider that speaks structured output, and the default local embedding model downloads on first use. Inspired by opencode-supermemory, but local-first.","edges":[{"to":"pro-workflow","type":"alternative","why":"Same job on rival harnesses: a persistent local store under every coding-agent session. pro-workflow puts SQLite + FTS5 rules under Claude Code; opencode-mem puts SQLite + vector recall with auto-capture under OpenCode.","confidence":0.6,"status":"approved"},{"to":"claude-reflect","type":"alternative","why":"Both close the cross-session learning loop for coding agents: claude-reflect distills your corrections into CLAUDE.md via /reflect; opencode-mem auto-captures work into a searchable vector store and injects relevant memories per session.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.666Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"memsearch","why":"Same job — persistent cross-session memory for coding agents — different reach: opencode-mem goes deep on one host (OpenCode, local SQLite, nothing leaves the machine); memsearch spans four hosts with one shared Markdown+Milvus store.","dir":"in","confidence":0.75,"name":"memsearch"},{"slug":"pro-workflow","why":"Same job on rival harnesses: a persistent local store under every coding-agent session. pro-workflow puts SQLite + FTS5 rules under Claude Code; opencode-mem puts SQLite + vector recall with auto-capture under OpenCode.","dir":"out","confidence":0.6,"name":"pro-workflow"},{"slug":"claude-reflect","why":"Both close the cross-session learning loop for coding agents: claude-reflect distills your corrections into CLAUDE.md via /reflect; opencode-mem auto-captures work into a searchable vector store and injects relevant memories per session.","dir":"out","confidence":0.5,"name":"claude-reflect"},{"slug":"memmolt","why":"Local SQLite+vector agent memory either way: opencode-mem auto-captures per-session for OpenCode; MemMolt is deliberate, structured recall for any MCP client.","dir":"in","confidence":0.5,"name":"MemMolt"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tickernelz/opencode-mem"},"opendataloader-pdf":{"name":"opendataloader-pdf","owner":"opendataloader-project","slug":"opendataloader-pdf","stars":27759,"image":"https://raw.githubusercontent.com/opendataloader-project/opendataloader-pdf/main/samples/image/example_annotated_pdf.png","avatar":"https://avatars.githubusercontent.com/u/211280852?v=4&s=96","forks":2663,"language":"Java","license":"Apache-2.0","updated":"yesterday","topics":["ocr","rag"],"summary":"Deterministic PDF parser for AI pipelines: #1 extraction accuracy (0.907) on its public bench, bounding boxes on every element, 0.015s/page — plus the first open PDF auto-tagging for accessibility.","curator_note":"Two products in one repo, both rare: a benchmark-topping deterministic parser (Markdown/JSON/HTML with bounding boxes, XY-Cut++ reading order, 0.015s/page, hybrid AI mode when you want it) and the first open-source auto-tagging to Tagged PDF — the accessibility path, built with the PDF Association and validated by veraPDF. Java core with Python/Node SDKs. NOT fully open at the edges: PDF/UA export and the accessibility studio are the enterprise add-on, and Java 11+ is a heavier runtime than the Rust competition. Benchmark caveat: the #1 score is on their own (public, reproducible) bench — verify on your documents.","edges":[{"to":"mineru","type":"alternative","why":"Same heavyweight PDF-to-structured-data slot: MinerU throws ML layout analysis and OCR at everything; OpenDataLoader is deterministic-first with an optional hybrid AI mode — and beats it on extraction accuracy in its published bench at a fraction of the compute.","confidence":0.7,"status":"approved"},{"to":"pdf-inspector","type":"alternative","why":"Kindred no-ML philosophy, different depth — pdf-inspector triages and extracts in 200ms; OpenDataLoader does full layout analysis with bounding boxes. Notably, pdf-inspector benchmarks itself on opendataloader-bench: same yardstick, acknowledged rivalry.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.944Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"mineru","why":"Same heavyweight PDF-to-structured-data slot: MinerU throws ML layout analysis and OCR at everything; OpenDataLoader is deterministic-first with an optional hybrid AI mode — and beats it on extraction accuracy in its published bench at a fraction of the compute.","dir":"out","confidence":0.7,"name":"MinerU"},{"slug":"pdf-inspector","why":"Kindred no-ML philosophy, different depth — pdf-inspector triages and extracts in 200ms; OpenDataLoader does full layout analysis with bounding boxes. Notably, pdf-inspector benchmarks itself on opendataloader-bench: same yardstick, acknowledged rivalry.","dir":"out","confidence":0.6,"name":"pdf-inspector"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/opendataloader-project/opendataloader-pdf"},"openmanus":{"name":"OpenManus","owner":"FoundationAgents","slug":"openmanus","stars":57560,"image":"https://raw.githubusercontent.com/FoundationAgents/openmanus/main/assets/logo.jpg","avatar":"https://avatars.githubusercontent.com/u/198047230?v=4&s=96","forks":10015,"language":"Python","license":"MIT","updated":"5 months ago","topics":["agents"],"summary":"Open-source general autonomous agent from the MetaGPT team — the 'Manus without an invite code': browsing, tool use and multi-step task execution from a simple Python core.","curator_note":"The fastest way to feel what a general autonomous agent does: clone, add a model key, hand it a task — the 3-hour-prototype energy that earned 57k stars keeps the codebase small enough to actually read, which makes it a great learning skeleton. But look at the commit graph before betting on it: activity has been quiet for months while the team's focus moved on (OpenManus-RL and beyond), so treat it as a reference implementation, NOT a maintained production framework — for durable agent infrastructure reach for an actively developed harness instead.","edges":[{"to":"deepagents","type":"alternative","why":"Same job — a batteries-included general autonomous agent that plans, browses and executes multi-step tasks. deepagents is the actively maintained, LangGraph-backed harness; OpenManus is the leaner, quieter reference implementation.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.104Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"deepagents","why":"Same job — a batteries-included general autonomous agent that plans, browses and executes multi-step tasks. deepagents is the actively maintained, LangGraph-backed harness; OpenManus is the leaner, quieter reference implementation.","dir":"out","confidence":0.65,"name":"deepagents"},{"slug":"mindshub","why":"Both answer 'delegate the whole task to an agent' — OpenManus as a minimal open reference you run yourself, MindsHub as a productized workspace with scheduling and a console.","dir":"in","confidence":0.55,"name":"mindshub"},{"slug":"mini-swe-agent","why":"Both are open, minimal-core autonomous agents you can actually read. OpenManus goes general (browsing, tool use, multi-step tasks); mini-swe-agent goes narrow and measurable — software engineering with bash only, scored on SWE-bench.","dir":"in","confidence":0.55,"name":"mini-swe-agent"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/FoundationAgents/openmanus"},"openresearcher":{"name":"OpenResearcher","owner":"TIGER-AI-Lab","slug":"openresearcher","stars":1094,"image":"https://raw.githubusercontent.com/TIGER-AI-Lab/openresearcher/main/assets/imgs/or-logo1.png","avatar":"https://avatars.githubusercontent.com/u/144196744?v=4&s=96","forks":101,"language":"Python","license":null,"updated":"1 months ago","topics":["training"],"summary":"TIGER-AI-Lab's fully open deep-research recipe: 96K long-horizon trajectories (adopted by NVIDIA Nemotron), a 30B-A3B model hitting 54.8% BrowseComp-Plus, training code and eval harness.","curator_note":"The open counterpoint to closed Deep Research products — and the data is the crown jewel: 100+-turn research trajectories distilled from GPT-OSS-120B over a self-built 11B-token retriever corpus (no search-API bills at generation scale), good enough that NVIDIA folded it into Nemotron 3 Ultra. The 30B-A3B model beats GPT-4.1, Claude-Opus-4 and Gemini-2.5-Pro on BrowseComp-Plus. Reproducing anything is a real commitment: the setup assumes 8×A100, training lives in a Megatron-LM fork, and the local retriever needs Java + tevatron. ⚠ No LICENSE file in the repo — clarify terms before commercial use. Pick DeepDive for the KG-synthesis + multi-turn-RL recipe; OpenResearcher for large-scale SFT distillation with everything — data, model, eval — actually released.","edges":[{"to":"deepdive","type":"alternative","why":"Rival open recipes for training deep-research agents: DeepDive synthesizes hard QA from knowledge-graph walks and trains with multi-turn RL on slime (checkpoints still pending); OpenResearcher distills 96K long trajectories from GPT-OSS-120B and ships data, 30B model and eval framework complete.","confidence":0.75,"status":"approved"},{"to":"vllm","type":"built_with","why":"The whole serving path runs on vLLM — the 30B-A3B deploys via bundled vLLM server scripts and the trajectory generation leans on vLLM's native browser-tool support for GPT-OSS.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.709Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"deepdive","why":"Rival open recipes for training deep-research agents: DeepDive synthesizes hard QA from knowledge-graph walks and trains with multi-turn RL on slime (checkpoints still pending); OpenResearcher distills 96K long trajectories from GPT-OSS-120B and ships data, 30B model and eval framework complete.","dir":"out","confidence":0.75,"name":"DeepDive"}],"built_with":[{"slug":"vllm","why":"The whole serving path runs on vLLM — the 30B-A3B deploys via bundled vLLM server scripts and the trajectory generation leans on vLLM's native browser-tool support for GPT-OSS.","dir":"out","confidence":0.6,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/TIGER-AI-Lab/openresearcher"},"opensandbox":{"name":"OpenSandbox","owner":"opensandbox-group","slug":"opensandbox","stars":12141,"image":"https://raw.githubusercontent.com/opensandbox-group/opensandbox/main/docs/public/images/logo.svg","avatar":"https://avatars.githubusercontent.com/u/290567599?v=4&s=96","forks":1017,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["agents","local"],"summary":"CNCF-landscape sandbox platform for AI agents: multi-language SDKs, unified API, CLI and MCP over Docker/Kubernetes runtimes — coding agents, GUI agents, evals and RL training.","curator_note":"The platform play in agent sandboxing: one API over Docker and Kubernetes runtimes, SDKs in multiple languages, an MCP server, and OpenSSF/CNCF hygiene — built for the org that needs sandboxes as shared infrastructure across coding agents, GUI agents, eval harnesses and RL training, not a per-project tool. NOT the isolation ceiling: container runtimes trade the hard KVM boundary microVM sandboxes give you for operational familiarity — if untrusted code is the threat model, weigh a Firecracker-class runtime instead; if platform ergonomics on your existing K8s is the goal, this is the mature option.","edges":[{"to":"cubesandbox","type":"alternative","why":"Both self-hosted sandbox runtimes for agents, split by isolation bet: CubeSandbox is microVM-first (KVM boundary, E2B-compatible API); OpenSandbox is platform-first (Docker/K8s runtimes, SDKs, CLI, MCP) for teams standardizing on existing infra.","confidence":0.7,"status":"approved"},{"to":"gym-anything","type":"complements","why":"gym-anything defines the environments and verifiers; OpenSandbox provides the execution layer — RL training and agent evaluation are both first-class scenarios in its API.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:18.881Z","linkCount":2,"related":{"complements":[{"slug":"gym-anything","why":"gym-anything defines the environments and verifiers; OpenSandbox provides the execution layer — RL training and agent evaluation are both first-class scenarios in its API.","dir":"out","confidence":0.5,"name":"gym-anything"}],"alternative":[{"slug":"cubesandbox","why":"Both self-hosted sandbox runtimes for agents, split by isolation bet: CubeSandbox is microVM-first (KVM boundary, E2B-compatible API); OpenSandbox is platform-first (Docker/K8s runtimes, SDKs, CLI, MCP) for teams standardizing on existing infra.","dir":"out","confidence":0.7,"name":"CubeSandbox"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/opensandbox-group/opensandbox"},"opensre":{"name":"opensre","owner":"Tracer-Cloud","slug":"opensre","stars":9001,"image":"https://raw.githubusercontent.com/Tracer-Cloud/opensre/main/docs/logo/opensre-logo-white.svg","avatar":"https://avatars.githubusercontent.com/u/196393879?v=4&s=96","forks":1246,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["agents","evals"],"summary":"Open-source framework for AI SRE agents plus the RL training and evaluation environment they need — connect 60+ tools you already run and investigate incidents on your own infra.","curator_note":"The most serious open attempt at 'SWE-bench for production incidents': not just an agent that greps logs, but a reinforcement-learning environment where SRE agents can actually be trained and scored on distributed failures. Connect the observability stack you already run (60+ integrations) and investigate on your own infrastructure. NOT production-stable — it's a public alpha and says so, APIs will move; and mind the telemetry section plus the vendor (Tracer Cloud) gravity. Adopt the benchmark and ideas today, bet the pager on it later.","edges":[{"to":"verl","type":"complements","why":"OpenSRE supplies the missing piece verl-style RL training needs for infrastructure agents: a realistic environment with scalable feedback for incident-response rollouts.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-12T23:46:16.016Z","linkCount":2,"related":{"complements":[{"slug":"verl","why":"OpenSRE supplies the missing piece verl-style RL training needs for infrastructure agents: a realistic environment with scalable feedback for incident-response rollouts.","dir":"out","confidence":0.55,"name":"verl"}],"alternative":[{"slug":"allama","why":"The same pattern — AI agents automating incident response — pointed at different fires: OpenSRE at production outages, Allama at security alerts. Pick by which pager you carry.","dir":"in","confidence":0.55,"name":"allama"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Tracer-Cloud/opensre"},"oransim":{"name":"oransim","owner":"OranAi-Ltd","slug":"oransim","stars":1168,"image":"https://raw.githubusercontent.com/OranAi-Ltd/oransim/main/assets/wordmark.svg","avatar":"https://avatars.githubusercontent.com/u/237740540?v=4&s=96","forks":152,"language":"Python","license":"Apache-2.0","updated":"7 days ago","topics":["agents"],"summary":"Open causal engine for marketing simulation: a virtual consumer society with LLM personas answers do()-style counterfactuals — rank campaign combos, swap KOLs mid-flight, replay spend. Apache-2.0.","curator_note":"Rare thing: a production causal-ML engine open-sourced end to end. Pearl-style SCM (64 nodes/117 edges), per-arm counterfactual heads in the TARNet/Dragonnet lineage, causal Neural Hawkes rollouts, and LLM 'soul personas' that read your actual creatives — every layer carries inline paper citations and any prediction is traceable through the graph. Know what you're holding: the audit artifact of a Shenzhen martech company's open-core play. The OSS ships a 21k-note demo corpus plus synthetic data; real predictive power needs their licensed 小红书 panel or your own DataProvider, and the research-grade models are code-complete but weights-pending. Use it to study causal agent-based simulation or as a BYO-data scaffold — don't expect turnkey marketing truth.","edges":[],"status":"approved","added":"2026-07-14T23:32:32.935Z","linkCount":0,"related":{"complements":[],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/OranAi-Ltd/oransim"},"page-agent":{"name":"page-agent","owner":"alibaba","slug":"page-agent","stars":27542,"image":"https://page-agent.github.io/assets/readme/banner-light.png","avatar":"https://avatars.githubusercontent.com/u/1961952?v=4&s=96","forks":2413,"language":"TypeScript","license":"MIT","updated":"2 days ago","topics":["agents","web"],"summary":"Alibaba's in-page GUI agent: one script tag gives any webpage its own AI agent — users drive the interface in natural language. TypeScript, tiny bundle, Chrome extension available.","curator_note":"The inverted take on browser agents: instead of YOUR agent driving SOMEONE's site, the site ships its own — one script tag and visitors can say 'file an expense report for Tuesday's taxi' at your UI. For products with deep, form-heavy interfaces this is the cheapest 'AI feature' with real utility, and 26k stars say the demand is real. NOT for automating third-party sites (that's browser-use's job — this requires the site owner to opt in), and handing an LLM the ability to click your own UI means auditing what it can reach: same-origin power cuts both ways.","edges":[{"to":"browser-use","type":"alternative","why":"Same end state — AI operating a web interface — from opposite sides of the fence: browser-use is the agent's browser for any site; page-agent is embedded by the site itself, giving its own users an agent.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T14:52:59.829Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"browser-use","why":"Same end state — AI operating a web interface — from opposite sides of the fence: browser-use is the agent's browser for any site; page-agent is embedded by the site itself, giving its own users an agent.","dir":"out","confidence":0.6,"name":"browser-use"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/alibaba/page-agent"},"paperclip":{"name":"paperclip","owner":"paperclipai","slug":"paperclip","stars":74539,"skip_image":true,"avatar":"https://avatars.githubusercontent.com/u/264498616?v=4&s=96","forks":13873,"language":"TypeScript","license":"MIT","updated":"yesterday","topics":["agents","orchestration"],"summary":"Open-source control plane for running fleets of heterogeneous AI agents as a \"company\" — bring your own agent, assign goals, org charts, budgets, governance, and an audited ticket system.","curator_note":"Reach for Paperclip when you run many agents across providers 24/7 and need a boss layer — budgets that hard-stop, org charts, goal alignment, audit trail, mobile monitoring. Sweet spot: \"20 Claude Code tabs, lost track of who does what.\" NOT the tool to build an agent (it orchestrates ones you already have), and overkill for a single agent or one linear pipeline. It's a control plane, not a framework.","edges":[{"to":"crewai","type":"alternative","why":"Both pitch orchestrating a team of agents toward a shared goal. Different altitude: CrewAI is a framework to build role-playing collaborating agents in-process; Paperclip is a BYO-agent management plane over external runtimes. Overlap in the multi-agent coordination job makes them substitutes for some users.","confidence":0.6,"status":"approved"},{"to":"autogen","type":"alternative","why":"AutoGen is a multi-agent framework for cooperating LLM agents — same coordination job as Paperclip, but as a library you build with rather than a dashboard you run external agents under.","confidence":0.5,"status":"approved"},{"to":"langgraph","type":"complements","why":"LangGraph builds stateful, durable multi-actor agent apps; such an app can be 'hired' into Paperclip via HTTP/heartbeat and managed (budget, goals, audit) at the org level. Build the agent with LangGraph, govern the fleet with Paperclip.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-06-29T16:24:39.000Z","linkCount":7,"related":{"complements":[{"slug":"cubesandbox","why":"paperclip governs fleets of agents; CubeSandbox is the workload-runtime layer beneath — thousands of <5MB VMs per node, idle agents auto-pause to keep fleet cost sane.","dir":"in","confidence":0.7,"name":"CubeSandbox"},{"slug":"langgraph","why":"LangGraph builds stateful, durable multi-actor agent apps; such an app can be 'hired' into Paperclip via HTTP/heartbeat and managed (budget, goals, audit) at the org level. Build the agent with LangGraph, govern the fleet with Paperclip.","dir":"out","confidence":0.5,"name":"LangGraph"}],"alternative":[{"slug":"alook","why":"Same job — run your agents as a 'company' with an org chart and task system. paperclip bets on governance (budgets, audits, tickets) for heterogeneous fleets; alook bets on email-native simplicity for solo builders running coding agents.","dir":"in","confidence":0.85,"name":"alook"},{"slug":"crewai","why":"Both pitch orchestrating a team of agents toward a shared goal. Different altitude: CrewAI is a framework to build role-playing collaborating agents in-process; Paperclip is a BYO-agent management plane over external runtimes. Overlap in the multi-agent coordination job makes them substitutes for some users.","dir":"out","confidence":0.6,"name":"CrewAI"},{"slug":"agentfield","why":"Both are open control planes for fleets of agents, attacking opposite ends: paperclip is the governance layer over agents you already have (BYO agent, goals, org charts, budgets, audited tickets); agentfield is the build-and-run backend where agents are SDK-written microservices with REST endpoints, queues and retries. Govern existing agents → paperclip; build the agent backend itself → agentfield.","dir":"in","confidence":0.6,"name":"agentfield"},{"slug":"core","why":"Both are the layer ABOVE agent frameworks rather than a framework themselves, and both spawn coding-agent sessions — but CORE is a personal, memory-driven, always-on assistant while paperclip manages fleets of agents as a team/company. Pick CORE for a personal AI OS, paperclip for multi-agent org governance.","dir":"in","confidence":0.6,"name":"core"},{"slug":"autogen","why":"AutoGen is a multi-agent framework for cooperating LLM agents — same coordination job as Paperclip, but as a library you build with rather than a dashboard you run external agents under.","dir":"out","confidence":0.5,"name":"AutoGen"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/paperclipai/paperclip"},"pdf-inspector":{"name":"pdf-inspector","owner":"firecrawl","slug":"pdf-inspector","stars":1641,"avatar":"https://avatars.githubusercontent.com/u/135057108?v=4&s=96","forks":149,"language":"Rust","license":"MIT","updated":"7 days ago","topics":["ocr","rag"],"summary":"Firecrawl's Rust PDF triage: classifies text-based vs scanned in ~10-50ms, extracts positioned text and clean Markdown without OCR — routing the ~54% of PDFs that never needed a model.","curator_note":"The router your document pipeline is missing: most stacks OCR everything, but ~54% of PDFs are text-based — this classifies in tens of milliseconds (with confidence and per-page routing), extracts locally in under 200ms, and only the genuinely scanned pages go to an expensive OCR model. Pure Rust, one dependency, bindings for Python, Node and browser WASM. NOT an OCR engine and NOT a layout-analysis heavyweight: scanned documents still need a model downstream, and complex-layout fidelity trails ML parsers — its job is knowing when you don't need them.","edges":[{"to":"chandra","type":"complements","why":"pdf-inspector's whole design goal is smart routing: text-based pages extract locally in milliseconds, and the scanned remainder gets handed to an OCR model like Chandra. Triage first, VLM second.","confidence":0.6,"status":"approved"},{"to":"xberg","type":"alternative","why":"Both are deterministic Rust document engines with multi-language bindings. xberg goes wide — 96 formats into RAG-ready chunks; pdf-inspector goes deep on one format with classification, confidence scores and OCR routing.","confidence":0.6,"status":"approved"},{"to":"mineru","type":"alternative","why":"Same PDF-to-Markdown slot, opposite weight: MinerU runs full layout analysis and OCR on everything; pdf-inspector is the 200ms no-ML path for documents that don't need it. Many pipelines should front MinerU with exactly this triage.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.989Z","linkCount":4,"related":{"complements":[{"slug":"chandra","why":"pdf-inspector's whole design goal is smart routing: text-based pages extract locally in milliseconds, and the scanned remainder gets handed to an OCR model like Chandra. Triage first, VLM second.","dir":"out","confidence":0.6,"name":"chandra"}],"alternative":[{"slug":"xberg","why":"Both are deterministic Rust document engines with multi-language bindings. xberg goes wide — 96 formats into RAG-ready chunks; pdf-inspector goes deep on one format with classification, confidence scores and OCR routing.","dir":"out","confidence":0.6,"name":"xberg"},{"slug":"opendataloader-pdf","why":"Kindred no-ML philosophy, different depth — pdf-inspector triages and extracts in 200ms; OpenDataLoader does full layout analysis with bounding boxes. Notably, pdf-inspector benchmarks itself on opendataloader-bench: same yardstick, acknowledged rivalry.","dir":"in","confidence":0.6,"name":"opendataloader-pdf"},{"slug":"mineru","why":"Same PDF-to-Markdown slot, opposite weight: MinerU runs full layout analysis and OCR on everything; pdf-inspector is the 200ms no-ML path for documents that don't need it. Many pipelines should front MinerU with exactly this triage.","dir":"out","confidence":0.55,"name":"MinerU"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/firecrawl/pdf-inspector"},"personal_ai_infrastructure":{"name":"LifeOS","owner":"danielmiessler","slug":"personal_ai_infrastructure","stars":16880,"image":"https://raw.githubusercontent.com/danielmiessler/personal_ai_infrastructure/main/images/lifeos-logo-full.png","avatar":"https://avatars.githubusercontent.com/u/50654?v=4&s=96","forks":2291,"language":"TypeScript","license":"MIT","updated":"2 days ago","topics":["agents"],"summary":"Daniel Miessler's LifeOS: an AI 'life operating system' that carries your goals and context into every task — an intent engineering platform with dashboard, agents and installer.","curator_note":"The most-starred personal-AI-OS scaffold: markdown, commands and agents on top of Claude Code, organized around TELOS (your goals) so every task starts from what you actually want. A philosophy with an installer — adopt the structure, not just the files. NOT a library; expect to live inside its conventions.","edges":[{"to":"core","type":"alternative","why":"Both 'personal AI OS' plays: core is an always-on shipped product watching your apps with a memory graph; LifeOS is an open scaffold of intent, agents and commands you inhabit on top of Claude Code.","confidence":0.6,"status":"approved"},{"to":"my-brain-is-full-crew","type":"alternative","why":"Manage-your-life AI at different scopes: an 8-agent crew over an Obsidian vault vs a whole-life operating system organized around your goals (TELOS).","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.682Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"core","why":"Both 'personal AI OS' plays: core is an always-on shipped product watching your apps with a memory graph; LifeOS is an open scaffold of intent, agents and commands you inhabit on top of Claude Code.","dir":"out","confidence":0.6,"name":"core"},{"slug":"my-brain-is-full-crew","why":"Manage-your-life AI at different scopes: an 8-agent crew over an Obsidian vault vs a whole-life operating system organized around your goals (TELOS).","dir":"out","confidence":0.5,"name":"My-Brain-Is-Full-Crew"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/danielmiessler/personal_ai_infrastructure"},"picobot":{"name":"picobot","owner":"louisho5","slug":"picobot","stars":1298,"image":"https://raw.githubusercontent.com/louisho5/picobot/main/docs/logo.png","avatar":"https://avatars.githubusercontent.com/u/24430165?v=4&s=96","forks":161,"language":"Go","license":"MIT","updated":"3 months ago","topics":["agents"],"summary":"A self-hosted personal AI agent in a single ~9MB Go binary — persistent memory, 16 tools + MCP, skills, cron/heartbeat, and Telegram/Discord/Slack/WhatsApp channels. Runs on a $5 VPS.","curator_note":"The anti-bloat statement piece: zero dependencies, ~10MB RAM idle, instant cold start — an always-on agent with ranked memory recall, background subagents, a natural-language HEARTBEAT.md cron, and self-authored skills ('create a skill for checking weather' → it writes the markdown), reachable from your phone via four chat channels. Runs on a Raspberry Pi or Termux on an old Android. Any OpenAI-compatible endpoint, including Ollama. Its own README names the target: OpenClaw's power without the 500MB container. NOT a coding harness and single-user by design — this is your personal daemon, not a team platform; complex multi-agent workflows will outgrow it fast, which is rather the point.","edges":[{"to":"core","type":"alternative","why":"Both are always-on self-hosted personal AI assistants with persistent memory, at opposite ends of the weight spectrum: core is a full 'personal AI OS' product watching your apps; Picobot is a 9MB binary you leave running on a $5 VPS and message from Telegram.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.810Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"core","why":"Both are always-on self-hosted personal AI assistants with persistent memory, at opposite ends of the weight spectrum: core is a full 'personal AI OS' product watching your apps; Picobot is a 9MB binary you leave running on a $5 VPS and message from Telegram.","dir":"out","confidence":0.55,"name":"core"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/louisho5/picobot"},"pixelrag":{"name":"PixelRAG","owner":"StarTrail-org","slug":"pixelrag","stars":7080,"image":"https://raw.githubusercontent.com/StarTrail-org/pixelrag/main/docs/assets/banner.png","avatar":"https://avatars.githubusercontent.com/u/288858980?v=4&s=96","forks":589,"language":"Python","license":"Apache-2.0","updated":"8 days ago","topics":["rag","vision"],"summary":"Berkeley's visual RAG: render pages and PDFs to screenshot tiles and retrieve with a VLM embedder — tables, charts and layout survive. pixelshot CLI plus a hosted 8.28M-page Wikipedia index.","curator_note":"The paper's claim — screenshots beat parsed text for RAG — matters when your answers live in visual structure: tables, charts, infographics, layout-heavy PDFs that text chunkers flatten into noise. Zero-setup on-ramp is real: a hosted 8.28M-page Wikipedia index with no API key, and the pixelbrowse Claude Code plugin gives the agent screenshot-based page reading (pixelshot via Playwright/CDP, no MCP server). NOT a drop-in for text RAG stacks: your reader must be a VLM, tile indexes cost more storage/compute than text embeddings, and retrieval quality rides on their LoRA-tuned Qwen3-VL-Embedding model. Research codebase — expect pipeline assembly, not a product.","edges":[{"to":"mineru","type":"alternative","why":"Two answers to the same RAG-ingestion problem: MinerU parses PDFs/Office through layout analysis and OCR into LLM-ready markdown; PixelRAG skips parsing entirely and retrieves over rendered screenshot tiles. Parse-to-text vs stay-in-pixels.","confidence":0.7,"status":"approved"},{"to":"olmocr","type":"alternative","why":"olmocr linearizes messy PDFs into clean ordered Markdown so a text pipeline can index them; PixelRAG argues the linearization step is the loss — embed the page image and let a VLM read the tile directly.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-20T09:28:52.324Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"mineru","why":"Two answers to the same RAG-ingestion problem: MinerU parses PDFs/Office through layout analysis and OCR into LLM-ready markdown; PixelRAG skips parsing entirely and retrieves over rendered screenshot tiles. Parse-to-text vs stay-in-pixels.","dir":"out","confidence":0.7,"name":"MinerU"},{"slug":"olmocr","why":"olmocr linearizes messy PDFs into clean ordered Markdown so a text pipeline can index them; PixelRAG argues the linearization step is the loss — embed the page image and let a VLM read the tile directly.","dir":"out","confidence":0.55,"name":"olmocr"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/StarTrail-org/pixelrag"},"plano":{"name":"plano","owner":"katanemo","slug":"plano","stars":6887,"image":"https://raw.githubusercontent.com/katanemo/plano/main/docs/source/_static/img/PlanoTagline.svg","avatar":"https://avatars.githubusercontent.com/u/112724757?v=4&s=96","forks":469,"language":"Rust","license":"Apache-2.0","updated":"2 days ago","topics":["gateway","orchestration"],"summary":"AI-native Envoy-based proxy for agentic apps: agent orchestration via a 4B routing model, smart LLM routing, guardrail filter chains and zero-code OTEL traces. Rust, framework-agnostic.","curator_note":"Reach for it when multi-agent code is drowning in hidden middleware — intent routing, provider quirks, guardrail hooks, tracing glue. Plano moves all of that out-of-process: agents are plain OpenAI-compatible HTTP servers in any language, orchestration is YAML plus a purpose-built 4B router model. NOT for a quick single-agent demo (it adds an infra hop and Envoy operational surface), and note the catch: the hosted Plano-Orchestrator LLM is free-tier only — production means running the routing models yourself or getting API keys. If you only need provider unification, LiteLLM is lighter.","edges":[{"to":"litellm","type":"alternative","why":"Both sit between your app and 100+ LLM providers as a self-hosted gateway; LiteLLM is the lighter Python proxy for provider unification and cost control, Plano adds agent orchestration, guardrail filter chains and signals on an Envoy data plane.","confidence":0.85,"status":"approved"},{"to":"agentfield","type":"alternative","why":"Same 'production infrastructure layer for agents' job from opposite ends: agentfield is a control plane that wraps your functions in routing/queues/retries, Plano is a data-plane proxy your unmodified HTTP agents sit behind.","confidence":0.6,"status":"approved"},{"to":"nemo-guardrails","type":"alternative","why":"Two places to enforce guardrails: NeMo Guardrails runs Colang rails in-process around a model or chain, Plano enforces moderation/jailbreak filters at the proxy so every agent inherits them without code changes.","confidence":0.6,"status":"approved"},{"to":"langgraph","type":"complements","why":"Keep intra-agent graph logic in LangGraph and let Plano handle what's outside the framework: cross-agent routing, model failover, guardrails and OTEL traces — Plano is explicitly framework-agnostic about what runs behind each agent URL.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-19T13:40:10.405Z","linkCount":4,"related":{"complements":[{"slug":"langgraph","why":"Keep intra-agent graph logic in LangGraph and let Plano handle what's outside the framework: cross-agent routing, model failover, guardrails and OTEL traces — Plano is explicitly framework-agnostic about what runs behind each agent URL.","dir":"out","confidence":0.6,"name":"LangGraph"}],"alternative":[{"slug":"litellm","why":"Both sit between your app and 100+ LLM providers as a self-hosted gateway; LiteLLM is the lighter Python proxy for provider unification and cost control, Plano adds agent orchestration, guardrail filter chains and signals on an Envoy data plane.","dir":"out","confidence":0.85,"name":"litellm"},{"slug":"agentfield","why":"Same 'production infrastructure layer for agents' job from opposite ends: agentfield is a control plane that wraps your functions in routing/queues/retries, Plano is a data-plane proxy your unmodified HTTP agents sit behind.","dir":"out","confidence":0.6,"name":"agentfield"},{"slug":"nemo-guardrails","why":"Two places to enforce guardrails: NeMo Guardrails runs Colang rails in-process around a model or chain, Plano enforces moderation/jailbreak filters at the proxy so every agent inherits them without code changes.","dir":"out","confidence":0.6,"name":"Guardrails"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/katanemo/plano"},"pm-skills":{"name":"pm-skills","owner":"phuryn","slug":"pm-skills","stars":24252,"image":"https://raw.githubusercontent.com/phuryn/pm-skills/main/.docs/images/plugins.png","avatar":"https://avatars.githubusercontent.com/u/7837354?v=4&s=96","forks":2493,"language":null,"license":"MIT","updated":"21 days ago","topics":["skills"],"summary":"PM Skills Marketplace: 68 skills and 42 chained workflows in 9 plugins — discovery, strategy, PRDs, launch, growth — encoding Torres/Cagan-style frameworks for Claude Code and Cowork.","curator_note":"The product-management counterpart to the marketing skill packs: each skill encodes a named PM framework (continuous discovery, assumption mapping, north-star metrics) and walks the model through it, so /write-prd produces structured thinking, not just faster documents. Chained commands (/discover → /strategy → /plan-launch) are the differentiator over loose skill piles, and it's tested in CI. NOT a substitute for talking to users — frameworks encode the questions, not the answers — and it's Claude Code/Cowork-first: other assistants get the skills but lose the command plumbing. From an indie maintainer; gauge bus factor before making it team infrastructure.","edges":[{"to":"aaron-marketing-skills","type":"alternative","why":"Same shape, adjacent discipline: large curated skill marketplaces for Claude Code that encode a profession's frameworks — 120 marketing skills there, 68 PM skills with chained workflows here. Run both if you own both functions.","confidence":0.65,"status":"approved"},{"to":"designer-skills","type":"alternative","why":"Both are single-discipline skill packs that give Claude a professional's operating frameworks — design craft in one, product management in the other. Same install pattern, different seat at the table.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-23T14:13:33.514Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"aaron-marketing-skills","why":"Same shape, adjacent discipline: large curated skill marketplaces for Claude Code that encode a profession's frameworks — 120 marketing skills there, 68 PM skills with chained workflows here. Run both if you own both functions.","dir":"out","confidence":0.65,"name":"aaron-marketing-skills"},{"slug":"designer-skills","why":"Both are single-discipline skill packs that give Claude a professional's operating frameworks — design craft in one, product management in the other. Same install pattern, different seat at the table.","dir":"out","confidence":0.55,"name":"designer-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/phuryn/pm-skills"},"pocket-tts":{"name":"pocket-tts","owner":"kyutai-labs","slug":"pocket-tts","stars":7850,"image":"https://github.com/user-attachments/assets/637b5ed6-831f-4023-9b4c-741be21ab238","avatar":"https://avatars.githubusercontent.com/u/151010778?v=4&s=96","forks":789,"language":"Python","license":"MIT","updated":"8 days ago","topics":["local","voice"],"summary":"Kyutai's 100M-parameter CPU-only TTS — streaming audio in ~200ms, ~6× real-time on two laptop cores, voice cloning, six languages. pip install and it talks.","curator_note":"This is the TTS you add when the rest of your stack is already local: no GPU, no API key, ~200ms to first audio on two CPU cores, and a port ecosystem (WASM, MLX, ONNX, C++, Home Assistant) that means it runs basically anywhere — the natural voice for an Ollama-powered assistant. When NOT: you need studio-grade expressiveness or high-concurrency serving — it's a batch-of-1, small-model design, not a production voice API; and the polished voice catalog is English-heaviest. Voice cloning needs 20s of audio — use it with consent.","edges":[{"to":"ollama","type":"complements","why":"The fully-offline voice assistant stack: Ollama runs the brain on your machine, pocket-tts gives it a voice on two CPU cores — no GPU, no cloud, community integrations already wire the two together.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":3,"related":{"complements":[{"slug":"ollama","why":"The fully-offline voice assistant stack: Ollama runs the brain on your machine, pocket-tts gives it a voice on two CPU cores — no GPU, no cloud, community integrations already wire the two together.","dir":"out","confidence":0.75,"name":"Ollama"}],"alternative":[{"slug":"luxtts","why":"Both are tiny local TTS engines: pocket-tts is Kyutai's 100M CPU-only streaming model with a fixed voice set; LuxTTS targets high-fidelity 48kHz voice cloning from a reference sample. Pick by whether you need cloning or streaming.","dir":"in","confidence":0.8,"name":"LuxTTS"},{"slug":"voicebox","why":"Same say-it-locally job, different shape: pocket-tts is a pip-installable 100M CPU library for embedding speech in your own code; Voicebox is a full desktop studio with cloning, dictation, effects and MCP agent integration.","dir":"in","confidence":0.5,"name":"voicebox"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/kyutai-labs/pocket-tts"},"ponytail":{"name":"ponytail","owner":"DietrichGebert","slug":"ponytail","stars":87872,"image":"https://raw.githubusercontent.com/DietrichGebert/ponytail/main/assets/logo.png","avatar":"https://avatars.githubusercontent.com/u/137048761?v=4&s=96","forks":4809,"language":"JavaScript","license":"MIT","updated":"9 days ago","topics":["skills","coding"],"summary":"A skill that makes your agent code like the laziest senior dev: YAGNI enforced — ~54% less code, ~20% cheaper, ~27% faster on measured Claude Code sessions. Works with 20 agents.","curator_note":"One opinion, installed: the best code is the code you never wrote. Where every other add-on gives your agent MORE — more context, more tools, more process — ponytail gives it restraint, and the benchmarking is unusually honest: ~54% mean code reduction across 12 real tasks (94% is the over-building ceiling, not the average), with the safety-guard regression of a naive 'write one-liners' prompt explicitly tested and avoided. Works across 20 agents. NOT for codebases where verbosity is the convention (enterprise Java won't thank you), and watch the trajectory: 88k stars, a waitlist banner and ponytail.dev — the skill is free today, the product is coming.","edges":[{"to":"headroom","type":"complements","why":"Two ends of the same token diet: Headroom compresses what the agent reads, ponytail shrinks what it writes — less generated code is fewer output tokens and less context in every later turn. Stack them.","confidence":0.55,"status":"approved"},{"to":"ecc","type":"alternative","why":"Opposite philosophies for shaping a coding agent: ECC installs an entire operating layer — hooks, gates, orchestrators; ponytail installs exactly one opinion. Maximalist vs minimalist, and the choice says more about your team than the tools.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-23T14:13:33.454Z","linkCount":2,"related":{"complements":[{"slug":"headroom","why":"Two ends of the same token diet: Headroom compresses what the agent reads, ponytail shrinks what it writes — less generated code is fewer output tokens and less context in every later turn. Stack them.","dir":"out","confidence":0.55,"name":"headroom"}],"alternative":[{"slug":"ecc","why":"Opposite philosophies for shaping a coding agent: ECC installs an entire operating layer — hooks, gates, orchestrators; ponytail installs exactly one opinion. Maximalist vs minimalist, and the choice says more about your team than the tools.","dir":"out","confidence":0.5,"name":"ECC"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/DietrichGebert/ponytail"},"praisonai":{"name":"PraisonAI","owner":"MervinPraison","slug":"praisonai","stars":8509,"image":"https://raw.githubusercontent.com/MervinPraison/praisonai/main/.github/images/logo_light.png","avatar":"https://avatars.githubusercontent.com/u/454862?v=4&s=96","forks":1327,"language":"Python","license":"MIT","updated":"2 days ago","topics":["agents"],"summary":"Low-code multi-agent framework: autonomous agents with built-in memory, RAG and MCP support across 100+ LLMs — from one agent to an 'AI workforce' in a few lines or YAML.","curator_note":"The kitchen-sink take on multi-agent: agents, memory, RAG, UI, MCP registry entry and 100+ LLM backends in one package, configurable from YAML or ~5 lines of Python — genuinely fast for getting a working crew today, and it interoperates with CrewAI/AG2 patterns rather than fighting them. The breadth is also the caution: surface area this wide runs shallower per feature than dedicated tools, and the marketing-forward README means you should verify each capability against your use case. NOT for teams who need one deeply engineered abstraction — that's LangGraph territory.","edges":[{"to":"crewai","type":"alternative","why":"Same job — role-based multi-agent orchestration. PraisonAI trades CrewAI's focused API for a broader low-code bundle (memory, RAG, UI, MCP included) and even runs CrewAI-style configs.","confidence":0.8,"status":"approved"},{"to":"autogen","type":"alternative","why":"Both orchestrate cooperating agents; AutoGen is the research-grade conversation framework, PraisonAI the low-code productized bundle.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.121Z","linkCount":3,"related":{"complements":[{"slug":"litellm","why":"PraisonAI runs agents across 100+ LLMs — LiteLLM is the standard provider-abstraction layer for exactly that reach.","dir":"in","confidence":0.52,"name":"litellm"}],"alternative":[{"slug":"crewai","why":"Same job — role-based multi-agent orchestration. PraisonAI trades CrewAI's focused API for a broader low-code bundle (memory, RAG, UI, MCP included) and even runs CrewAI-style configs.","dir":"out","confidence":0.8,"name":"CrewAI"},{"slug":"autogen","why":"Both orchestrate cooperating agents; AutoGen is the research-grade conversation framework, PraisonAI the low-code productized bundle.","dir":"out","confidence":0.7,"name":"AutoGen"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/MervinPraison/praisonai"},"pro-workflow":{"name":"pro-workflow","owner":"rohitg00","slug":"pro-workflow","stars":2637,"image":"https://raw.githubusercontent.com/rohitg00/pro-workflow/main/assets/banner.svg","avatar":"https://avatars.githubusercontent.com/u/48523873?v=4&s=96","forks":255,"language":"JavaScript","license":null,"updated":"4 days ago","topics":["coding","memory"],"summary":"One SQLite store under every Claude Code session: corrections become FTS5-searchable rules that auto-load, research grows persistent wikis, and 37 hook scripts add quality gates.","curator_note":"The most complete attack on Claude Code amnesia we've mapped: correct it once and the correction becomes a durable, searchable rule; research lands in wikis that persist and even grow via an auto-research loop; hooks add git/secret guards and cost tracking. After 50 sessions the compounding is real. Two cautions: NO license file at review time (all rights reserved by default — same gap as this author's other tools), and 34 skills + 37 hooks is a lot of surface — adopt the memory core first, audit the hooks before letting them gate your commits.","edges":[{"to":"claude-reflect","type":"alternative","why":"Same core job — corrections that stick across Claude Code sessions. claude-reflect is the focused plugin (corrections → CLAUDE.md); pro-workflow builds a whole SQLite-backed memory-and-hooks platform around the idea.","confidence":0.75,"status":"approved"},{"to":"skillkit","type":"complements","why":"pro-workflow ships its 34 skills through SkillKit — the package manager is its distribution channel to 32+ agent formats.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:12.979Z","linkCount":4,"related":{"complements":[{"slug":"skillkit","why":"pro-workflow ships its 34 skills through SkillKit — the package manager is its distribution channel to 32+ agent formats.","dir":"out","confidence":0.6,"name":"skillkit"}],"alternative":[{"slug":"claude-reflect","why":"Same core job — corrections that stick across Claude Code sessions. claude-reflect is the focused plugin (corrections → CLAUDE.md); pro-workflow builds a whole SQLite-backed memory-and-hooks platform around the idea.","dir":"out","confidence":0.75,"name":"claude-reflect"},{"slug":"ecc","why":"Same job — a persistent operating layer under coding-agent sessions with hooks, learned rules and quality gates. pro-workflow is one SQLite store focused on Claude Code; ECC is the maximalist cross-harness system spanning 7+ harnesses with skills, agents and security tooling.","dir":"in","confidence":0.7,"name":"ECC"},{"slug":"opencode-mem","why":"Same job on rival harnesses: a persistent local store under every coding-agent session. pro-workflow puts SQLite + FTS5 rules under Claude Code; opencode-mem puts SQLite + vector recall with auto-capture under OpenCode.","dir":"in","confidence":0.6,"name":"opencode-mem"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/rohitg00/pro-workflow"},"prompt-cache-skills":{"name":"prompt-cache-skills","owner":"OnlyTerp","slug":"prompt-cache-skills","stars":108,"image":"https://raw.githubusercontent.com/OnlyTerp/prompt-cache-skills/main/assets/banner.svg","avatar":"https://avatars.githubusercontent.com/u/121772140?v=4&s=96","forks":7,"language":"Python","license":"NOASSERTION","updated":"14 days ago","topics":["skills","coding"],"summary":"13 drop-in skills that fix prompt-caching bugs in OSS agent harnesses (Cline, Roo, Continue, OpenCode, Aider) — point your coding agent at the repo, it patches and verifies on the wire.","curator_note":"If you run Cline, Roo Code, Continue, Aider or OpenCode against your own API keys, this likely pays for itself in an afternoon: audited findings with file:line citations, each fix a 5-15 line diff, and a wire-level verify step (check_cache.py fires the request cold+warm and diffs the cache_* token fields — trust but verify is built in). Skip it on Claude Code or Codex CLI, both audited as already-correct reference implementations, and on managed backends (Devin, Windsurf, Antigravity) where nothing is patchable. Audits are dated 2026-05-27; harnesses move fast, so re-run the verify step after upgrading yours.","edges":[{"to":"lynkr","type":"complements","why":"Both attack the coding-agent API bill and stack cleanly: lynkr compresses and semantic-caches at a gateway in front of the model, while these skills fix the harness's own cache_control / prompt_cache_key bugs so provider-side prompt caching actually engages.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T17:35:30.592Z","linkCount":2,"related":{"complements":[{"slug":"lynkr","why":"Both attack the coding-agent API bill and stack cleanly: lynkr compresses and semantic-caches at a gateway in front of the model, while these skills fix the harness's own cache_control / prompt_cache_key bugs so provider-side prompt caching actually engages.","dir":"out","confidence":0.6,"name":"Lynkr"},{"slug":"ai-token-monitor","why":"Measure, then fix: the monitor prices cache reads separately, so it's the visible before/after for the prompt-caching patches — apply a skill, watch the same session logs show the spend drop.","dir":"in","confidence":0.5,"name":"ai-token-monitor"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/OnlyTerp/prompt-cache-skills"},"pullfrog":{"name":"pullfrog","owner":"pullfrog","slug":"pullfrog","stars":828,"image":"https://pullfrog.com/frog-green-200px.png","avatar":"https://avatars.githubusercontent.com/u/223181266?v=4&s=96","forks":47,"language":"TypeScript","license":"MIT","updated":"yesterday","topics":["coding"],"summary":"Open-source, model-agnostic GitHub bot: tag @pullfrog on any issue or PR and your own coding agent (BYOK) runs the task inside GitHub Actions, context via an internal MCP server.","curator_note":"The missing GitHub-native surface for coding agents: instead of copying issues into a terminal session, tag the bot and the work happens where the work lives — Actions runners, your keys, your choice of model, automated triggers for recurring flows. NOT for latency-sensitive interactive coding (Actions cold-starts apply), and 'agent with repo write access via bot' deserves the same permission audit you'd give any CI credential. Young (~800 stars) but moving fast.","edges":[{"to":"squid","type":"alternative","why":"Both turn repo intent into reviewed changes: squid pipelines spec→PR inside Claude Code; pullfrog embeds the agent in GitHub itself — tag a comment and Actions does the rest, any model.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T16:04:55.171Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"contrabass","why":"Same job — turn tracker issues into coding-agent runs. Pullfrog executes per @-mention inside GitHub Actions; Contrabass polls continuously and runs agents on your own machine in tmux panes and git worktrees.","dir":"in","confidence":0.7,"name":"contrabass"},{"slug":"squid","why":"Both turn repo intent into reviewed changes: squid pipelines spec→PR inside Claude Code; pullfrog embeds the agent in GitHub itself — tag a comment and Actions does the rest, any model.","dir":"out","confidence":0.55,"name":"squid"},{"slug":"open-code-review","why":"Both put an AI agent on your GitHub PRs. pullfrog is the generalist — @mention it and your own agent runs any task in Actions; Open Code Review is the specialist — reviews only, hybrid deterministic+LLM, precision-tuned line comments.","dir":"in","confidence":0.55,"name":"open-code-review"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/pullfrog/pullfrog"},"ragas":{"name":"Ragas","owner":"explodinggradients","slug":"ragas","stars":14967,"image":"https://raw.githubusercontent.com/vibrantlabsai/ragas/main/docs/_static/imgs/logo.png","avatar":"https://avatars.githubusercontent.com/u/122604797?v=4&s=96","forks":1579,"language":"Python","license":"Apache-2.0","updated":"5 months ago","topics":["evals","rag"],"summary":"Evaluation toolkit for your RAG and agent pipelines — faithfulness, relevance, and more.","edges":[{"to":"llamaindex","type":"complements","why":"Turns 'looks right' into a number — score LlamaIndex pipelines before you ship.","confidence":0.85,"status":"approved"}],"added":"2026-06-22T23:27:20.000Z","linkCount":6,"related":{"complements":[{"slug":"llamaindex","why":"Turns 'looks right' into a number — score LlamaIndex pipelines before you ship.","dir":"out","confidence":0.85,"name":"LlamaIndex"},{"slug":"autogen","why":"Score the outputs of your multi-agent conversations for faithfulness and relevance.","dir":"in","confidence":0.8,"name":"AutoGen"},{"slug":"searchbox","why":"searchbox produces answers plus ablation rows (which tools, local vs API retrieval, budget) but has no answer-quality rubric of its own. Ragas scores faithfulness/relevance of RAG and agent answers, so it grades the ANSWER.md a searchbox run emits — turning a tool/retrieval ablation sweep into a measured quality comparison.","dir":"in","confidence":0.55,"name":"searchbox"}],"alternative":[{"slug":"deepeval","why":"The two default pip installs for LLM evaluation: Ragas is RAG-centric (faithfulness, relevance); DeepEval covers the same RAG suite plus agents, chatbots, safety metrics and pytest-style CI integration. Broader tool vs sharper tool.","dir":"in","confidence":0.75,"name":"deepeval"},{"slug":"giskard-oss","why":"Overlapping LLM-evaluation job: ragas is the specialist for RAG pipeline metrics; Giskard v3 covers multi-turn agent scenarios with built-in checks and LLM-as-judge, with RAG evaluation still pending its v3 port.","dir":"in","confidence":0.65,"name":"giskard-oss"},{"slug":"agents-cli","why":"agents-cli ships a full agent-eval suite — trace generation, metric grading, LLM-as-judge, failure-mode clustering, prompt optimization — which overlaps Ragas's job. Difference: Ragas is eval-only and framework-agnostic, while agents-cli's eval is one stage of a GCP-bound scaffold/deploy pipeline.","dir":"in","confidence":0.45,"name":"agents-cli"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/explodinggradients/ragas"},"repowise":{"name":"repowise","owner":"repowise-dev","slug":"repowise","stars":4150,"image":"https://raw.githubusercontent.com/repowise-dev/repowise/main/.github/assets/banner.png","avatar":"https://avatars.githubusercontent.com/u/271941743?v=4&s=96","forks":467,"language":"Python","license":"AGPL-3.0","updated":"yesterday","topics":["coding"],"summary":"Codebase intelligence for AI and humans: deterministic code-health scores calibrated on real defects, graph-aware refactoring plans agents can execute, auto-docs and git analytics over 9 MCP tools.","curator_note":"The interesting bet: defect-risk scoring with NO LLM — 25 deterministic markers calibrated against a real defect corpus (published ROC AUC 0.74), indexed in seconds, then the same dependency graph generates concrete refactoring plans (split the god class, break the cycle) your coding agent executes. Health→locate→fix as one loop is what linters and dashboards don't do. NOT a pure-open play: AGPL-3.0 with a hosted-teams funnel, and benchmark claims are self-published — reproduce them on your repo before quoting them; overlaps only partially with symbol-level code-graph MCPs.","edges":[{"to":"tokensave","type":"alternative","why":"Both are code-intelligence MCP servers that cut agent context burn; tokensave answers structural questions (symbols, callers, impact) from a semantic graph, repowise layers health scoring and executable refactoring plans on top of its own graph.","confidence":0.6,"status":"approved"},{"to":"codegraph-mcp","type":"alternative","why":"Both serve a code knowledge graph to agents over MCP; codegraph-mcp goes deep on cross-language symbol/call/blast-radius queries, repowise trades some of that depth for defect-risk scores and refactoring plans.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-13T23:34:30.447Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"tokensave","why":"Both are code-intelligence MCP servers that cut agent context burn; tokensave answers structural questions (symbols, callers, impact) from a semantic graph, repowise layers health scoring and executable refactoring plans on top of its own graph.","dir":"out","confidence":0.6,"name":"tokensave"},{"slug":"codegraph-mcp","why":"Both serve a code knowledge graph to agents over MCP; codegraph-mcp goes deep on cross-language symbol/call/blast-radius queries, repowise trades some of that depth for defect-risk scores and refactoring plans.","dir":"out","confidence":0.6,"name":"codegraph-mcp"},{"slug":"codeflow","why":"Both grade codebase health from graph structure. Repowise gives deterministic, defect-calibrated scores and agent-executable refactoring plans; codeflow gives an instant in-browser A–F and heatmaps.","dir":"in","confidence":0.55,"name":"codeflow"},{"slug":"code-review-graph","why":"Both build deterministic code intelligence over MCP; repowise scores health and plans refactors, code-review-graph computes the minimal read-set for a change.","dir":"in","confidence":0.5,"name":"code-review-graph"},{"slug":"understand-anything","why":"Overlapping codebase-intelligence job: repowise leans quantitative — defect-calibrated health scores, refactoring plans, git analytics over MCP; Understand Anything leans pedagogical — teach the architecture through an explorable graph.","dir":"in","confidence":0.5,"name":"Understand-Anything"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/repowise-dev/repowise"},"rowboat":{"name":"rowboat","owner":"rowboatlabs","slug":"rowboat","stars":16723,"image":"https://raw.githubusercontent.com/rowboatlabs/rowboat/main/assets/readme-dark/hero-video.png","avatar":"https://avatars.githubusercontent.com/u/172591271?v=4&s=96","forks":1662,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["agents","memory"],"summary":"Desktop AI coworker (YC S24) that indexes email, meetings and Slack into a living backlinked knowledge graph, then acts on it — email client, browser, meeting notes, background agents, code mode.","curator_note":"The thesis is memory that compounds vs retrieval that starts cold: everything you touch gets indexed into an Obsidian-style backlinked graph stored as plain local Markdown — inspectable, editable, yours — and the work surfaces (email client that drafts replies with full context, isolated browser, local meeting transcriber, event/schedule-triggered background agents) all read from it. Code mode drives Claude Code or Codex with that context. BYO model including Ollama. Use it if you want one desktop app to BE the AI layer over your work life. When NOT: it's maximalist by nature — email + browser + meetings + code in one young app means breadth outruns polish in places, and Google-service setup is a manual OAuth dance. Data stays local; that part they got exactly right.","edges":[{"to":"core","type":"alternative","why":"The two strongest 'personal AI OS' plays in the catalog: both keep a persistent memory graph and act autonomously within guardrails. core watches your apps from the background; Rowboat ships its own work surfaces — email, browser, meeting notes — and stores the graph as editable local Markdown.","confidence":0.7,"status":"approved"},{"to":"meetily","type":"alternative","why":"Overlapping on the meeting slice: Meetily is the dedicated local meeting note-taker; Rowboat bundles an equivalent mic+speaker transcriber whose summaries feed its knowledge graph rather than standing alone.","confidence":0.45,"status":"approved"},{"to":"my-brain-is-full-crew","type":"alternative","why":"Two shapes of the same idea — an AI that manages your work life on top of a persistent, human-readable memory: Rowboat is a desktop app with its own surfaces (email, browser, meetings) over a Markdown graph; the Crew is a team of agents living inside your existing Obsidian vault, driven purely by chat.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.854Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"core","why":"The two strongest 'personal AI OS' plays in the catalog: both keep a persistent memory graph and act autonomously within guardrails. core watches your apps from the background; Rowboat ships its own work surfaces — email, browser, meeting notes — and stores the graph as editable local Markdown.","dir":"out","confidence":0.7,"name":"core"},{"slug":"my-brain-is-full-crew","why":"Two shapes of the same idea — an AI that manages your work life on top of a persistent, human-readable memory: Rowboat is a desktop app with its own surfaces (email, browser, meetings) over a Markdown graph; the Crew is a team of agents living inside your existing Obsidian vault, driven purely by chat.","dir":"out","confidence":0.55,"name":"My-Brain-Is-Full-Crew"},{"slug":"meetily","why":"Overlapping on the meeting slice: Meetily is the dedicated local meeting note-taker; Rowboat bundles an equivalent mic+speaker transcriber whose summaries feed its knowledge graph rather than standing alone.","dir":"out","confidence":0.45,"name":"meetily"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/rowboatlabs/rowboat"},"ruflo":{"name":"ruflo","owner":"ruvnet","slug":"ruflo","stars":65714,"image":"https://raw.githubusercontent.com/ruvnet/ruflo/main/ruflo/assets/ruflo-small.jpeg","avatar":"https://avatars.githubusercontent.com/u/2934394?v=4&s=96","forks":7806,"language":"TypeScript","license":"MIT","updated":"yesterday","topics":["agents","orchestration"],"summary":"ruvnet's 65k-star 'agent meta-harness': multi-agent swarms, adaptive memory and RAG layered over Claude Code, Codex and Hermes — npx ruflo, a UI beta, and a sprawling plugin ecosystem.","curator_note":"The maximalist swarm layer (formerly claude-flow): hive-mind coordination, neural-sounding memory, 8M+ ecosystem downloads and a feature list that outruns most vendors — when it clicks, you get parallel specialist agents on your existing CLI subscriptions with one npx. Approach it like a curator: the marketing is louder than the documentation, benchmark claims are self-reported, and the ecosystem's size (plugins, sub-projects, rebrands) makes auditing what actually runs non-trivial. NOT for teams that need boring, inspectable orchestration — pick a small gated pipeline instead and add swarm ambition later. Run it in a sandbox first; form your own opinion of the hype-to-substance ratio.","edges":[{"to":"alook","type":"alternative","why":"Both layer an always-on multi-agent organization over the coding CLIs you already run. alook keeps it legible — email, org chart, kanban; Ruflo goes maximalist — swarms, adaptive memory, self-learning claims and a UI.","confidence":0.55,"status":"approved"},{"to":"auto-company","type":"alternative","why":"Same dream — an autonomous multi-agent operation driven through Claude Code/Codex — different scope: Auto-Company scripts 14 fixed expert roles; Ruflo is a general meta-harness you compose swarms from.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-23T23:40:19.071Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"alook","why":"Both layer an always-on multi-agent organization over the coding CLIs you already run. alook keeps it legible — email, org chart, kanban; Ruflo goes maximalist — swarms, adaptive memory, self-learning claims and a UI.","dir":"out","confidence":0.55,"name":"alook"},{"slug":"auto-company","why":"Same dream — an autonomous multi-agent operation driven through Claude Code/Codex — different scope: Auto-Company scripts 14 fixed expert roles; Ruflo is a general meta-harness you compose swarms from.","dir":"out","confidence":0.5,"name":"Auto-Company"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/ruvnet/ruflo"},"sam2":{"name":"sam2","owner":"facebookresearch","slug":"sam2","stars":19576,"image":"https://raw.githubusercontent.com/facebookresearch/sam2/main/assets/model_diagram.png?raw=true","avatar":"https://avatars.githubusercontent.com/u/16943930?v=4&s=96","forks":2510,"language":"Jupyter Notebook","license":"Apache-2.0","updated":"1 months ago","topics":["vision"],"summary":"Meta's Segment Anything 2: promptable zero-shot segmentation for images and video with streaming memory — click and box prompts become tracked masks in real time.","curator_note":"Zero-shot masks with no training data: click a thing or pass a detector's box and SAM 2 segments it — and in video, tracks it across frames via streaming memory. Unbeatable for annotation pipelines, data engines and any 'segment whatever the user points at' feature. NOT a classifier — it draws masks but never names them (pair it with a detector for labels), and it's a research-grade repo: checkpoints and notebooks, not a serving stack; plan your own inference infrastructure and a real GPU for video work.","edges":[],"status":"approved","added":"2026-07-11T13:28:52.139Z","linkCount":2,"related":{"complements":[{"slug":"ultralytics","why":"The standard detect-then-segment pipeline: YOLO boxes become SAM 2 prompts — the detector names and localizes, SAM 2 delivers pixel-perfect masks and video tracking.","dir":"in","confidence":0.8,"name":"ultralytics"},{"slug":"supervision","why":"supervision speaks SAM 2 natively — masks become the same Detections object as boxes, so segmentation results plug into the identical annotate/filter/count pipeline.","dir":"in","confidence":0.6,"name":"supervision"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/facebookresearch/sam2"},"scale-agentex":{"name":"scale-agentex","owner":"scaleapi","slug":"scale-agentex","stars":457,"image":"https://github.com/user-attachments/assets/68beed69-9737-4ae5-96fc-47f3477dd1f3","avatar":"https://avatars.githubusercontent.com/u/21693938?v=4&s=96","forks":51,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["agents","orchestration"],"summary":"Scale AI's open agent platform: scaffold agents with a CLI, run them behind the ACP protocol with a dev UI, and graduate from sync chat to durable Temporal-backed long-running workflows.","curator_note":"Pick it when agents outgrow request/response: the async tier runs on Temporal, so long-running autonomous work gets durability, retries and resumability without changing your agent code's architecture — that L1→L5 'same framework at every level' pitch is the real differentiator. The local story is genuinely turnkey: ./dev.sh boots Postgres, Redis, Mongo and Temporal plus a dev UI, and `agentex init` scaffolds a working agent. NOT for a simple chatbot — that stack is heavy, and Python 3.12+/Docker are hard requirements. If you want multi-agent conversation patterns rather than deploy-and-scale infrastructure, a framework like CrewAI or AutoGen is the lighter tool. The enterprise 'zero-ops' path funnels into Scale's hosted SGP platform — fine, but know it.","edges":[{"to":"agentfield","type":"alternative","why":"Both are open control planes for building and deploying agents as long-running services with routing, queues and durable state. AgentField wraps plain Python/Go/TS functions as microservice agents; Agentex standardizes on the ACP protocol with Temporal underneath and a scaffolding CLI + dev UI.","confidence":0.7,"status":"approved"},{"to":"crewai","type":"alternative","why":"Overlapping 'build agents in Python' entry point, opposite emphasis: CrewAI is the in-process multi-agent collaboration framework; Agentex is the deployment platform — protocolized agents, dev sandbox, and durable async execution on Temporal.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T17:54:00.068Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"agentfield","why":"Both are open control planes for building and deploying agents as long-running services with routing, queues and durable state. AgentField wraps plain Python/Go/TS functions as microservice agents; Agentex standardizes on the ACP protocol with Temporal underneath and a scaffolding CLI + dev UI.","dir":"out","confidence":0.7,"name":"agentfield"},{"slug":"crewai","why":"Overlapping 'build agents in Python' entry point, opposite emphasis: CrewAI is the in-process multi-agent collaboration framework; Agentex is the deployment platform — protocolized agents, dev sandbox, and durable async execution on Temporal.","dir":"out","confidence":0.55,"name":"CrewAI"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/scaleapi/scale-agentex"},"scout":{"name":"scout","owner":"agno-agi","slug":"scout","stars":642,"avatar":"https://avatars.githubusercontent.com/u/104874993?v=4&s=96","forks":60,"language":"Python","license":"Apache-2.0","updated":"14 days ago","topics":["agents","memory"],"summary":"Company intelligence agent that navigates Slack, Drive, wiki and CRM live — no ingest/embed pipeline — and builds its own wiki + CRM as it learns your company.","curator_note":"The navigation-over-search thesis applied to company knowledge: ls/grep/follow-the-link across live sources instead of chunk-embed-pray. Built on Agno's AgentOS. Pick it to prototype a 'company brain' on live sources; skip if you need strict access-control auditing or a classic ingested RAG corpus.","edges":[{"to":"ktx","type":"alternative","why":"Two ways to hand agents company context: ktx curates a semantic layer over your warehouse with approved metrics; Scout navigates Slack/Drive/wiki/CRM live and writes its own wiki+CRM as it learns.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-17T14:03:22.728Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"ktx","why":"Two ways to hand agents company context: ktx curates a semantic layer over your warehouse with approved metrics; Scout navigates Slack/Drive/wiki/CRM live and writes its own wiki+CRM as it learns.","dir":"out","confidence":0.55,"name":"ktx"},{"slug":"wrenai","why":"Two shapes of 'company context for agents': Scout navigates unstructured sources (Slack/Drive/wiki) building its own wiki; Wren governs the structured-data side with semantic models and SQL.","dir":"in","confidence":0.5,"name":"WrenAI"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/agno-agi/scout"},"scrapling":{"name":"Scrapling","owner":"D4Vinci","slug":"scrapling","stars":70855,"image":"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/docs/assets/cover_light.svg?sanitize=true","avatar":"https://avatars.githubusercontent.com/u/20604835?v=4&s=96","forks":7025,"language":"Python","license":"BSD-3-Clause","updated":"4 days ago","topics":["web"],"summary":"Adaptive Python scraping framework: selectors that relearn when sites redesign, stealth fetchers that pass Cloudflare, spiders with proxy rotation and an MCP server — request to full crawl.","curator_note":"The modern-web answer to scraping's two chronic pains: selectors break (its parser relearns elements after redesigns via auto-save/auto-match) and bots get blocked (StealthyFetcher passes Cloudflare Turnstile out of the box). The MCP server is a quiet killer feature — your coding agent can scrape through it directly. NOT the veteran choice: younger ecosystem than Scrapy with fewer third-party answers when you're deep in the weeds, and the adaptive magic needs its cache warmed — first-run breakage still lands on you.","edges":[{"to":"scrapy","type":"alternative","why":"Both are Python scraping frameworks; Scrapling rebuilds the stack for the hostile modern web — adaptive selectors that survive redesigns, Cloudflare-passing stealth fetchers — where Scrapy relies on its ecosystem for both.","confidence":0.75,"status":"approved"},{"to":"crawlee","type":"alternative","why":"The two modern anti-bot-aware crawling frameworks, split by language: Scrapling for Python (adaptive parsing, MCP server), Crawlee for Node/TypeScript (browser switching, Apify ecosystem).","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.371Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"scrapy","why":"Both are Python scraping frameworks; Scrapling rebuilds the stack for the hostile modern web — adaptive selectors that survive redesigns, Cloudflare-passing stealth fetchers — where Scrapy relies on its ecosystem for both.","dir":"out","confidence":0.75,"name":"scrapy"},{"slug":"crawlee","why":"The two modern anti-bot-aware crawling frameworks, split by language: Scrapling for Python (adaptive parsing, MCP server), Crawlee for Node/TypeScript (browser switching, Apify ecosystem).","dir":"out","confidence":0.7,"name":"crawlee"},{"slug":"crawl4ai","why":"Same Python scraping job, different bets: Scrapling bets on resilience (self-healing selectors, Cloudflare-passing stealth); Crawl4AI bets on LLM-ready output and crawl orchestration. Hostile targets → Scrapling; RAG pipelines → Crawl4AI.","dir":"in","confidence":0.7,"name":"crawl4ai"},{"slug":"autoscraper","why":"Both attack selector fragility by learning: autoscraper infers extraction rules from one example and stops there; Scrapling relearns elements after redesigns inside a full, maintained crawling framework.","dir":"in","confidence":0.6,"name":"autoscraper"},{"slug":"stealth-browser-mcp","why":"Two stealth approaches to protected pages: Scrapling is a scraping framework with Cloudflare-passing fetchers for extraction at scale; stealth-browser-mcp is interactive browser control over MCP for agent-driven sessions.","dir":"in","confidence":0.55,"name":"stealth-browser-mcp"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/D4Vinci/scrapling"},"scrapy":{"name":"scrapy","owner":"scrapy","slug":"scrapy","stars":63314,"avatar":"https://avatars.githubusercontent.com/u/733635?v=4&s=96","forks":11805,"language":"Python","license":"BSD-3-Clause","updated":"yesterday","topics":["web"],"summary":"The veteran Python web crawling framework: spiders, middlewares, pipelines and battle-tested scheduling — 60k+ stars and still the reference architecture for structured scraping.","curator_note":"Fifteen-plus years of production hardening in one framework: spiders declare what to extract, middlewares/pipelines handle retries, throttling, dedup and export, and the ecosystem has an answer for everything. For large structured crawls in Python it's still the default. NOT a browser — JS-heavy or anti-bot-protected sites need Playwright bolted on or a different tool (Scrapling's stealth fetchers, Crawlee's browser mode), and the framework's inversion of control feels heavy when you just need one page: for that, requests + a parser beats a Scrapy project.","edges":[],"status":"approved","added":"2026-07-11T15:07:43.390Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"crawlee","why":"Same job — a production crawling framework. Scrapy is the fifteen-year Python veteran with the deepest ecosystem; Crawlee is TypeScript-native with headless-browser switching and anti-blocking as defaults rather than plugins.","dir":"in","confidence":0.8,"name":"crawlee"},{"slug":"scrapling","why":"Both are Python scraping frameworks; Scrapling rebuilds the stack for the hostile modern web — adaptive selectors that survive redesigns, Cloudflare-passing stealth fetchers — where Scrapy relies on its ecosystem for both.","dir":"in","confidence":0.75,"name":"Scrapling"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/scrapy/scrapy"},"searchbox":{"name":"searchbox","owner":"hanxiao","slug":"searchbox","stars":49,"image":"https://raw.githubusercontent.com/hanxiao/searchbox/main/docs/img/banner.png","avatar":"https://avatars.githubusercontent.com/u/2041322?v=4&s=96","forks":6,"language":"Python","license":"MIT","updated":"28 days ago","topics":["local","rag","agents"],"summary":"Airgapped closed-corpus QA testbed: a local Qwen agent in a Pi harness explores a .zip dataroom with grep/embeddings/rerankers under a token budget — a bed to study search as test-time compute.","curator_note":"A research testbed, not a product: it locks a local Qwen3.6 agent in an airgapped box with a .zip 'dataroom' and makes it answer purely from what's inside, using grep, embeddings (jina-v5) and a cross-encoder reranker (jina-v3) as tools under a turn budget. It exists to ablate one question — is search test-time compute? — so every axis (which tools, local-weights vs Jina-API retrieval, budget, base LLM) is an env knob and each run emits per-turn token/tool telemetry. Reach for it to study or benchmark agentic retrieval over your own corpus, self-hosted and offline. NOT a drop-in RAG library or production QA service: no persistent index (it embeds lazily per job), one job at a time on a single model slot, and you bring your own OpenAI-compatible LLM endpoint (llama.cpp/ollama) plus ideally a GPU.","edges":[{"to":"ollama","type":"complements","why":"searchbox drives its agent LLM over any OpenAI-compatible endpoint via LLAMA_URL/MODEL_ID (the reference deployment is a GGUF Qwen on llama.cpp). ollama is the easy way to serve that local chat model behind it — point LLAMA_URL at ollama's OpenAI-compatible API and swap models without touching searchbox.","confidence":0.6,"status":"approved"},{"to":"ragas","type":"complements","why":"searchbox produces answers plus ablation rows (which tools, local vs API retrieval, budget) but has no answer-quality rubric of its own. Ragas scores faithfulness/relevance of RAG and agent answers, so it grades the ANSWER.md a searchbox run emits — turning a tool/retrieval ablation sweep into a measured quality comparison.","confidence":0.55,"status":"approved"},{"to":"llamaindex","type":"alternative","why":"Both answer questions over a private corpus, but by opposite paradigms. LlamaIndex is a production RAG framework: build a structured index, then query it. searchbox is an airgapped research harness where an agent explores the raw corpus with grep/embed/rerank tools under a budget — no prebuilt index. Use LlamaIndex to ship RAG, searchbox to study agentic retrieval.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-08T00:38:16.972Z","linkCount":3,"related":{"complements":[{"slug":"ollama","why":"searchbox drives its agent LLM over any OpenAI-compatible endpoint via LLAMA_URL/MODEL_ID (the reference deployment is a GGUF Qwen on llama.cpp). ollama is the easy way to serve that local chat model behind it — point LLAMA_URL at ollama's OpenAI-compatible API and swap models without touching searchbox.","dir":"out","confidence":0.6,"name":"Ollama"},{"slug":"ragas","why":"searchbox produces answers plus ablation rows (which tools, local vs API retrieval, budget) but has no answer-quality rubric of its own. Ragas scores faithfulness/relevance of RAG and agent answers, so it grades the ANSWER.md a searchbox run emits — turning a tool/retrieval ablation sweep into a measured quality comparison.","dir":"out","confidence":0.55,"name":"Ragas"}],"alternative":[{"slug":"llamaindex","why":"Both answer questions over a private corpus, but by opposite paradigms. LlamaIndex is a production RAG framework: build a structured index, then query it. searchbox is an airgapped research harness where an agent explores the raw corpus with grep/embed/rerank tools under a budget — no prebuilt index. Use LlamaIndex to ship RAG, searchbox to study agentic retrieval.","dir":"out","confidence":0.5,"name":"LlamaIndex"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/hanxiao/searchbox"},"shepherd":{"name":"shepherd","owner":"shepherd-agents","slug":"shepherd","stars":1525,"image":"https://shepherd-agents.ai/assets/logo-shepherd.png","avatar":"https://avatars.githubusercontent.com/u/296595227?v=4&s=96","forks":120,"language":"Python","license":"MIT","updated":"3 days ago","topics":["agents","orchestration"],"summary":"Runtime substrate that records agent runs as reversible, Git-like execution traces — meta-agents can observe, fork, replay and revert any run before outputs are applied or released.","curator_note":"The missing 'version control for agent execution': every run becomes a durable trace you can inspect, fork from any step, replay deterministically, or revert — and workspace outputs are held for review before they're applied. If you're building supervision or meta-agents (agents that manage agents), this is the substrate that makes it tractable. NOT mature: early alpha with a moving API and a fresh paper behind it, and the reversibility guarantee only covers what runs inside its coupled agent+environment model — side effects that escaped to the real world don't revert.","edges":[{"to":"langgraph","type":"complements","why":"LangGraph checkpoints state inside the graph; Shepherd wraps the whole run in a reversible trace — fork and replay an agent from any step, review outputs before they apply, instead of re-running the graph and hoping.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-13T10:36:42.228Z","linkCount":1,"related":{"complements":[{"slug":"langgraph","why":"LangGraph checkpoints state inside the graph; Shepherd wraps the whole run in a reversible trace — fork and replay an agent from any step, review outputs before they apply, instead of re-running the graph and hoping.","dir":"out","confidence":0.55,"name":"LangGraph"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/shepherd-agents/shepherd"},"sia":{"name":"sia","owner":"hexo-ai","slug":"sia","stars":2058,"image":"https://raw.githubusercontent.com/hexo-ai/sia/main/docs/flow.png","avatar":"https://avatars.githubusercontent.com/u/118002492?v=4&s=96","forks":244,"language":"Python","license":"MIT","updated":"22 days ago","topics":["agents","training"],"summary":"Self-improving loop from the SIA paper: Meta, Target and Feedback agents evolve a task agent's harness AND weights against a benchmark — #1 on MLE-Bench Hard, 14x kernel speedups.","curator_note":"The most ambitious of the self-improvement loops on the map: where others mutate prompts or code, SIA's Feedback agent rewrites the task agent's harness and updates its weights, with paper-grade receipts (70.1% on LawBench vs 45% prior SOTA, 14x on an AlphaFold-3 Triton kernel, #1 on MLE-Bench Hard). Reach for it when your problem IS a benchmark: a scoreable task the loop can grind against. NOT for tasks without a mechanical metric, and not a polished product — it's the official research implementation, so expect to build the task harness yourself and budget real GPU time for the weight-update path.","edges":[{"to":"evo","type":"alternative","why":"Both run autonomous improve-against-a-benchmark loops with paper-adjacent rigor. evo points coding agents at YOUR codebase (tree search over commits); SIA generates a task-specific agent and evolves the agent itself — harness and weights.","confidence":0.7,"status":"approved"},{"to":"autoresearch","type":"alternative","why":"Same lineage — the autoresearch loop: metric, modify, verify, keep or discard. The skill hill-climbs your code with a fixed agent; SIA makes the agent itself the thing that improves, up to and including its weights.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T14:13:33.234Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"evo","why":"Both run autonomous improve-against-a-benchmark loops with paper-adjacent rigor. evo points coding agents at YOUR codebase (tree search over commits); SIA generates a task-specific agent and evolves the agent itself — harness and weights.","dir":"out","confidence":0.7,"name":"evo"},{"slug":"autoresearch","why":"Same lineage — the autoresearch loop: metric, modify, verify, keep or discard. The skill hill-climbs your code with a fixed agent; SIA makes the agent itself the thing that improves, up to and including its weights.","dir":"out","confidence":0.6,"name":"autoresearch"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/hexo-ai/sia"},"sie":{"name":"sie","owner":"superlinked","slug":"sie","stars":2297,"image":"https://cdn.prod.website-files.com/65dce6831bf9f730421e2915/65dce6831bf9f730421e2929_superlinked_logo.svg","avatar":"https://avatars.githubusercontent.com/u/94243920?v=4&s=96","forks":215,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["local","rag"],"summary":"Self-hosted inference cluster for everything agents call besides the big LLM: embeddings, rerankers, OCR, NER, guardrails and small LLMs — 100+ models, one OpenAI-compatible API, K8s stack included.","curator_note":"Fills the gap between 'vLLM serves my big model' and the pile of task models agents actually need: retrieval (bge-m3, SPLADE, ColBERT), doc-to-markdown (MinerU, GLM-OCR, docling), structured extraction (GLiNER), safety verdicts (Granite Guardian) — one cluster, on-demand loading with LRU eviction instead of a server per model. The production story is unusually complete and all Apache-2.0: load-balancing gateway, KEDA scale-to-zero, Grafana dashboards, Terraform for GKE/EKS/AKS; the MCP edge offloads document work from Claude-class clients to your cluster. NOT for max-throughput single-LLM serving — that stays vLLM/SGLang territory (SIE's own generation image wraps SGLang). Telemetry is on by default, opt-out env var.","edges":[{"to":"vllm","type":"alternative","why":"Both self-hosted OpenAI-compatible inference servers, opposite shapes: vLLM maximizes throughput for one big LLM; SIE serves breadth — 100+ heterogeneous task models (embedders, rerankers, OCR, NER, guards) loaded on demand across a cluster.","confidence":0.6,"status":"approved"},{"to":"ollama","type":"alternative","why":"Same 'serve open models behind one local API' job at different scales: Ollama is the single-machine developer runner; SIE is the multi-model production cluster with autoscaling, gateway and Terraform.","confidence":0.55,"status":"approved"},{"to":"mineru","type":"built_with","why":"MinerU ships in SIE's model catalog as a backend for the document-to-markdown task — SIE serves it (alongside GLM-OCR, PaddleOCR-VL and docling) behind its unified API.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-14T23:32:32.989Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"vllm","why":"Both self-hosted OpenAI-compatible inference servers, opposite shapes: vLLM maximizes throughput for one big LLM; SIE serves breadth — 100+ heterogeneous task models (embedders, rerankers, OCR, NER, guards) loaded on demand across a cluster.","dir":"out","confidence":0.6,"name":"vLLM"},{"slug":"ollama","why":"Same 'serve open models behind one local API' job at different scales: Ollama is the single-machine developer runner; SIE is the multi-model production cluster with autoscaling, gateway and Terraform.","dir":"out","confidence":0.55,"name":"Ollama"}],"built_with":[{"slug":"mineru","why":"MinerU ships in SIE's model catalog as a backend for the document-to-markdown task — SIE serves it (alongside GLM-OCR, PaddleOCR-VL and docling) behind its unified API.","dir":"out","confidence":0.6,"name":"MinerU"}]},"url":"https://stackmap.shipwithai.xyz/repos/superlinked/sie"},"siliconscope":{"name":"SiliconScope","owner":"kennss","slug":"siliconscope","stars":785,"image":"https://raw.githubusercontent.com/kennss/siliconscope/main/docs/img/dashboard.png","avatar":"https://avatars.githubusercontent.com/u/77983942?v=4&s=96","forks":48,"language":"Swift","license":"MIT","updated":"2 days ago","topics":["local"],"summary":"Sudoless Apple Silicon monitor: SwiftUI dashboard plus menu-bar suite tracking ANE, Media Engine and memory bandwidth Activity Monitor won't show — with DVR-style record & replay.","curator_note":"The monitor to run when local AI is your workload: it surfaces what actually saturates an Apple Silicon box — GPU vs Neural Engine vs memory bandwidth — and shows per-process ANE memory, a signal no other monitor exposes. Sudoless, no kexts, and Record/Replay lets you scrub a session like a DVR to compare before/after tuning. Credible iStat Menus stand-in too. NOT for non-Apple hardware (macOS 14+, Apple Silicon only), and accelerator metrics beyond CPU are system-wide — macOS won't attribute GPU/Media per process, and it honestly labels that instead of faking numbers. A monitor, not a profiler: Instruments still owns trace-level work.","edges":[{"to":"ollama","type":"complements","why":"The question SiliconScope was built to answer — what is my on-device model actually doing to the silicon — comes up the moment Ollama is serving: watch GPU, ANE and bandwidth load per model and quantization instead of guessing.","confidence":0.6,"status":"approved"},{"to":"ai-token-monitor","type":"complements","why":"Two menu-bar monitors covering the two halves of local AI work: ai-token-monitor watches what agents spend in tokens, SiliconScope watches what the hardware pays in compute, power and bandwidth.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-20T09:51:35.519Z","linkCount":4,"related":{"complements":[{"slug":"mlx-lora-studio","why":"Fine-tuning on Apple Silicon is a unified-memory balancing act — SiliconScope shows the GPU/ANE load, memory pressure and bandwidth of a Studio training run live, so you can size batch and quantization to the machine instead of guessing.","dir":"in","confidence":0.65,"name":"MLX-LoRA-Studio"},{"slug":"ollama","why":"The question SiliconScope was built to answer — what is my on-device model actually doing to the silicon — comes up the moment Ollama is serving: watch GPU, ANE and bandwidth load per model and quantization instead of guessing.","dir":"out","confidence":0.6,"name":"Ollama"},{"slug":"ai-token-monitor","why":"Two menu-bar monitors covering the two halves of local AI work: ai-token-monitor watches what agents spend in tokens, SiliconScope watches what the hardware pays in compute, power and bandwidth.","dir":"out","confidence":0.5,"name":"ai-token-monitor"},{"slug":"coreai-models","why":"Core AI models execute on the Neural Engine — SiliconScope is the monitor that shows ANE load and per-process ANE memory, so you can see what your .aimodel actually costs the silicon.","dir":"in","confidence":0.5,"name":"coreai-models"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/kennss/siliconscope"},"sim":{"name":"sim","owner":"simstudioai","slug":"sim","stars":29184,"image":"https://raw.githubusercontent.com/simstudioai/sim/main/apps/sim/public/static/readme-banner.png","avatar":"https://avatars.githubusercontent.com/u/199344406?v=4&s=96","forks":3736,"language":"TypeScript","license":"Apache-2.0","updated":"yesterday","topics":["agents","orchestration"],"summary":"Visual workspace to build, deploy and orchestrate AI agents — 1,000+ integrations, knowledge bases, built-in tables and files, schedules and run monitoring. Self-host via npx simstudio or Docker.","curator_note":"The n8n-for-agents play at real scale (29k stars): build agents visually, conversationally, or in code, wire them to Slack/Notion/Salesforce-class integrations, and keep tables, files, knowledge bases and scheduled runs in the same workspace — it's a platform, not a library. Self-hosting is honest but heavy: Docker with 12GB+ RAM recommended, Postgres + pgvector, and note that the Chat/copilot feature remains a Sim-managed service even self-hosted (you fetch a COPILOT_API_KEY from sim.ai). Ollama/vLLM local models supported. Pick it when non-engineers need to build and operate agents; pick a code-first control plane (agentfield) when engineers do, and skip the platform entirely for a single agent.","edges":[{"to":"bytechef","type":"alternative","why":"Both are open, self-hostable platforms unifying AI-agent orchestration with classic workflow automation behind a visual builder and big integration catalogs — Sim leans agent-first with knowledge/tables/files built in; ByteChef leans integration-first with its 200+ components.","confidence":0.7,"status":"approved"},{"to":"agentfield","type":"alternative","why":"Same 'run your AI workforce' ambition, opposite audiences: AgentField turns plain Python/Go/TS functions into agent microservices for engineers; Sim gives teams a visual workspace where agents are built and operated without touching the runtime.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.900Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"bytechef","why":"Both are open, self-hostable platforms unifying AI-agent orchestration with classic workflow automation behind a visual builder and big integration catalogs — Sim leans agent-first with knowledge/tables/files built in; ByteChef leans integration-first with its 200+ components.","dir":"out","confidence":0.7,"name":"bytechef"},{"slug":"agentfield","why":"Same 'run your AI workforce' ambition, opposite audiences: AgentField turns plain Python/Go/TS functions into agent microservices for engineers; Sim gives teams a visual workspace where agents are built and operated without touching the runtime.","dir":"out","confidence":0.55,"name":"agentfield"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/simstudioai/sim"},"skillforge":{"name":"SkillForge","owner":"tripleyak","slug":"skillforge","stars":797,"image":"https://raw.githubusercontent.com/tripleyak/skillforge/main/assets/images/01-title.png","avatar":"https://avatars.githubusercontent.com/u/132320136?v=4&s=96","forks":85,"language":"Python","license":"MIT","updated":"2 months ago","topics":["skills"],"summary":"A methodology and toolkit for engineering AI skills instead of improvising them — plus a Context Skill Advisor that proactively recommends, improves or creates skills from what you're doing.","curator_note":"The 'skills are engineering, not art' manifesto, executed: structured creation process with validation built in, and v5.2's Context Skill Advisor watches your session/project context and suggests the skill you should have — with proactivity levels (off→active) so it advises rather than nags. If you author skills regularly, the rigor pays. NOT a registry or installer (that's the package managers' job) and the advisor's judgment tracks its context quality — garbage project context, garbage suggestions.","edges":[{"to":"skillx","type":"alternative","why":"Both systematize skill creation: SkillX distills skills automatically from agent trajectories (research-grade); SkillForge is the hands-on engineering methodology with a proactive advisor for working developers.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:12.950Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"skillx","why":"Both systematize skill creation: SkillX distills skills automatically from agent trajectories (research-grade); SkillForge is the hands-on engineering methodology with a proactive advisor for working developers.","dir":"out","confidence":0.55,"name":"SkillX"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/tripleyak/skillforge"},"skillkit":{"name":"skillkit","owner":"rohitg00","slug":"skillkit","stars":1399,"image":"https://raw.githubusercontent.com/rohitg00/skillkit/main/docs/img/banner.svg","avatar":"https://avatars.githubusercontent.com/u/48523873?v=4&s=96","forks":130,"language":"TypeScript","license":"Apache-2.0","updated":"1 months ago","topics":["coding","skills"],"summary":"Package manager for AI agent skills — install from 400K+ skills across 31 sources, auto-translate between 46 agents' incompatible formats, security-scan on install, sync to every agent at once.","curator_note":"Use it the moment skills must live on more than one agent — you write Claude SKILL.md, a teammate runs Cursor: author once, `skillkit sync` translates and deploys everywhere; the install-time security scan is a real feature now that skill marketplaces are a prompt-injection vector. NOT needed if you're all-in on a single agent — Claude Code's first-party plugin/marketplace flow is simpler and better supported. Format translation is lossy at the edges (agent-specific frontmatter, hooks, tool references), so verify ported skills actually trigger; and treat the optional mesh/messaging/REST extras as experiments, not infrastructure.","edges":[{"to":"claude-reflect","type":"complements","why":"claude-reflect mines your sessions into reusable skills/commands for Claude Code; skillkit is the distribution layer that packages and translates them to 45 other agents. Minor overlap: skillkit also persists session learnings, but capture is claude-reflect's whole job.","confidence":0.6,"status":"approved"},{"to":"agents-cli","type":"complements","why":"agents-cli ships as an agent-skills pack; skillkit is a skill package manager that installs packs like it and ports them across agents. Plausible pairing, unverified compatibility — agents-cli's skill format may not round-trip cleanly.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-07T13:05:18.105Z","linkCount":11,"related":{"complements":[{"slug":"claude-reflect","why":"claude-reflect mines your sessions into reusable skills/commands for Claude Code; skillkit is the distribution layer that packages and translates them to 45 other agents. Minor overlap: skillkit also persists session learnings, but capture is claude-reflect's whole job.","dir":"out","confidence":0.6,"name":"claude-reflect"},{"slug":"pro-workflow","why":"pro-workflow ships its 34 skills through SkillKit — the package manager is its distribution channel to 32+ agent formats.","dir":"in","confidence":0.6,"name":"pro-workflow"},{"slug":"designer-skills","why":"Content meets distribution: designer-skills is a curated skill pack, skillkit is the package manager that installs, translates and security-scans skills like these across 46 agent formats.","dir":"in","confidence":0.55,"name":"designer-skills"},{"slug":"agents-cli","why":"agents-cli ships as an agent-skills pack; skillkit is a skill package manager that installs packs like it and ports them across agents. Plausible pairing, unverified compatibility — agents-cli's skill format may not round-trip cleanly.","dir":"out","confidence":0.5,"name":"agents-cli"},{"slug":"cc-wf-studio","why":"Studio designs the skill, SkillKit distributes it — canvas-to-Markdown on one side, cross-agent packaging and security scanning on the other.","dir":"in","confidence":0.5,"name":"cc-wf-studio"},{"slug":"omnigraph","why":"Omnigraph ships its operational playbook as an installable agent skill (`npx skills add ModernRelay/omnigraph@omnigraph`) — exactly the artifact skill managers like skillkit install and translate into whatever agent operates your cluster.","dir":"in","confidence":0.5,"name":"omnigraph"}],"alternative":[{"slug":"apm","why":"Overlapping job, different model: skillkit is an imperative installer + format translator chasing breadth (46 agents, 400K skills); apm is declarative — manifest, lockfile, transitive deps, org policy — chasing reproducibility. apm's README even ships a 'coming from npx skills add' migration.","dir":"in","confidence":0.8,"name":"apm"},{"slug":"asm","why":"Same job — install and manage skills across many coding agents — opposite bets: skillkit maximizes breadth (400K skills, 46 agents, format auto-translation); asm maximizes scriptability (agent-first --json/--yes CLI, dedupe audit, curated catalog, no accounts/telemetry).","dir":"in","confidence":0.8,"name":"asm"},{"slug":"skillnet","why":"Both are registry-scale skill installers (400-500K skills indexed). skillkit's strength is cross-agent format translation across 46 agents plus security scanning on install; SkillNet's is the research stack — create from traces, evaluate quality, compose and orchestrate.","dir":"in","confidence":0.75,"name":"SkillNet"},{"slug":"autoskills","why":"Both end with skills installed in your agent; skillkit is the explicit package manager you drive, autoskills the zero-config detector that decides for you from a smaller audited registry.","dir":"in","confidence":0.7,"name":"autoskills"},{"slug":"n-skills","why":"Same job — getting skills into any agent's format. skillkit is the breadth play (400K+ skills, 31 sources, security scans); n-skills is the depth play: one small marketplace where a human curated every entry.","dir":"in","confidence":0.65,"name":"n-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/rohitg00/skillkit"},"skillnet":{"name":"SkillNet","owner":"zjunlp","slug":"skillnet","stars":1111,"image":"https://raw.githubusercontent.com/zjunlp/skillnet/main/images/skillnet.png","avatar":"https://avatars.githubusercontent.com/u/41887875?v=4&s=96","forks":129,"language":"Python","license":"MIT","updated":"13 days ago","topics":["skills"],"summary":"Open skill infrastructure from ZJU NLP: search a 500K+ indexed skill library, install, generate skills from repos/docs/traces, score their quality, and compose/orchestrate them. SDK + CLI + MCP.","curator_note":"Use it when you want skills treated like packages, not pasted folders: semantic search over a 500K+ GitHub-skill index is credential-free, and `evaluate` gives a safety/completeness/executability score before you drop a random skill into an agent. Create/evaluate/analyze need an OpenAI-compatible endpoint; `orchestrate` additionally needs a Claude-Agent-SDK-capable gateway and ships exactly one scene (sciatlas) so far. Prefer skillkit or asm when you just want install plus cross-agent format translation — SkillNet's edge is the research layer: skill creation from execution traces, quality scoring, and skill-relationship graphs, with an arXiv report behind it.","edges":[{"to":"skillkit","type":"alternative","why":"Both are registry-scale skill installers (400-500K skills indexed). skillkit's strength is cross-agent format translation across 46 agents plus security scanning on install; SkillNet's is the research stack — create from traces, evaluate quality, compose and orchestrate.","confidence":0.75,"status":"approved"},{"to":"asm","type":"alternative","why":"Same job — search and install agent skills from a large catalog via CLI/SDK. asm is the scriptable ops tool (dedupe audits, security scans, --json everywhere for CI); SkillNet adds generation, scoring and graph analysis on top of discovery.","confidence":0.65,"status":"approved"},{"to":"n-skills","type":"alternative","why":"Both distribute reusable skills on the SKILL.md format; n-skills is a small curated marketplace, SkillNet is a 500K+ crawled-and-deduplicated index with quality ranking.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T17:35:30.629Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"skillkit","why":"Both are registry-scale skill installers (400-500K skills indexed). skillkit's strength is cross-agent format translation across 46 agents plus security scanning on install; SkillNet's is the research stack — create from traces, evaluate quality, compose and orchestrate.","dir":"out","confidence":0.75,"name":"skillkit"},{"slug":"asm","why":"Same job — search and install agent skills from a large catalog via CLI/SDK. asm is the scriptable ops tool (dedupe audits, security scans, --json everywhere for CI); SkillNet adds generation, scoring and graph analysis on top of discovery.","dir":"out","confidence":0.65,"name":"asm"},{"slug":"n-skills","why":"Both distribute reusable skills on the SKILL.md format; n-skills is a small curated marketplace, SkillNet is a 500K+ crawled-and-deduplicated index with quality ranking.","dir":"out","confidence":0.55,"name":"n-skills"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/zjunlp/skillnet"},"skillx":{"name":"SkillX","owner":"zjunlp","slug":"skillx","stars":265,"image":"https://raw.githubusercontent.com/zjunlp/skillx/main/assets/overview.png","avatar":"https://avatars.githubusercontent.com/u/41887875?v=4&s=96","forks":24,"language":"Python","license":"MIT","updated":"19 days ago","topics":["skills"],"summary":"Research framework that auto-distills agent trajectories into a three-level skill knowledge base (planning, functional, atomic) — pluggable into weaker agents and new environments.","curator_note":"The interesting research bet: instead of storing raw trajectories or reflections, distill agent experience into a hierarchy of reusable skills that transfer — a strong backbone agent builds the library, weaker agents plug it in and improve on AppWorld/BFCL/τ2-Bench. Read it if you're building agent-improvement loops; the three-level decomposition (planning/functional/atomic) is a genuinely useful mental model. NOT production tooling — it's an academic codebase (~250 stars) built around benchmarks, so expect to adapt it to your stack rather than pip-install it.","edges":[{"to":"claude-reflect","type":"alternative","why":"Both turn agent experience into reusable knowledge: claude-reflect captures your corrections into CLAUDE.md pragmatically, SkillX distills full trajectories into a transferable skill hierarchy — research-grade.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.410Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"claude-reflect","why":"Both turn agent experience into reusable knowledge: claude-reflect captures your corrections into CLAUDE.md pragmatically, SkillX distills full trajectories into a transferable skill hierarchy — research-grade.","dir":"out","confidence":0.55,"name":"claude-reflect"},{"slug":"skillforge","why":"Both systematize skill creation: SkillX distills skills automatically from agent trajectories (research-grade); SkillForge is the hands-on engineering methodology with a proactive advisor for working developers.","dir":"in","confidence":0.55,"name":"SkillForge"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/zjunlp/skillx"},"slime":{"name":"slime","owner":"THUDM","slug":"slime","stars":7601,"image":"https://raw.githubusercontent.com/THUDM/slime/main/imgs/arch.png","avatar":"https://avatars.githubusercontent.com/u/48590610?v=4&s=96","forks":1090,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["training"],"summary":"THUDM's RL post-training framework behind the GLM releases — Megatron training plus SGLang rollouts with native arg pass-through, and pluggable reward, verifier and agentic data-generation workflows.","curator_note":"One of the few open RL stacks proven on frontier releases (GLM-4.5 through 5.2, with Qwen/DeepSeek/Llama support): the Megatron+SGLang-only bet keeps the dataflow explicit and upstream engine features usable instead of flattened behind a multi-backend abstraction, and rollout-only/train-only debug paths take RL's silent-bug problem seriously. NOT an afternoon tool — you need Megatron-scale GPU infrastructure and RL literacy; for single-node SFT or LoRA use a lighter trainer. And if your rollout engine must be vLLM, this is the wrong framework by design.","edges":[],"status":"approved","added":"2026-07-07T19:32:14.039Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"verl","why":"Same job — large-scale RL post-training. verl bets on multi-backend flexibility (FSDP/Megatron × vLLM/SGLang) where slime hard-commits to Megatron+SGLang; pick verl when your infra is vLLM- or FSDP-shaped.","dir":"in","confidence":0.85,"name":"verl"},{"slug":"trl","why":"The lighter trainer slime's own docs point you toward: single-node SFT/LoRA/DPO on the HF stack instead of Megatron-scale RL infrastructure. Same goal — a post-trained model — at opposite ends of the infra spectrum.","dir":"in","confidence":0.7,"name":"trl"},{"slug":"h2o-llmstudio","why":"Both post-train LLMs, from opposite ends: slime is Megatron-scale RL for frontier runs, LLM Studio is no-code LoRA/DPO fine-tuning on models a single node can hold.","dir":"in","confidence":0.5,"name":"h2o-llmstudio"}],"built_with":[{"slug":"deepdive","why":"DeepDive's multi-turn RL runs on THUDM's slime framework — the released training code is a slime rollout setup, and the repo credits it directly.","dir":"in","confidence":0.85,"name":"DeepDive"}]},"url":"https://stackmap.shipwithai.xyz/repos/THUDM/slime"},"squad":{"name":"squad","owner":"mco-org","slug":"squad","stars":609,"image":"https://raw.githubusercontent.com/mco-org/squad/main/assets/squad-readme-hero.png","avatar":"https://avatars.githubusercontent.com/u/264223103?v=4&s=96","forks":60,"language":"Rust","license":"MIT","updated":"1 months ago","topics":["coding","orchestration"],"summary":"Multi-agent terminal collaboration for AI CLIs: a manager, workers and an inspector — Claude Code, Gemini, Codex, OpenCode — coordinating through one-shot shell commands and SQLite. No daemon.","curator_note":"The unix-philosophy take on multi-agent coding: no daemon, no server, no framework — each agent is a terminal running the CLI you already use, coordinating through /squad slash commands backed by SQLite. Assign a manager, spin workers, add an inspector, watch them divide the work. Radically simpler to reason about than orchestration platforms, and it dies clean (every command is one-shot). NOT for production pipelines or unattended fleets — it's built for a human watching terminals; young (~600 stars) and quiet for weeks at review time, so kick the tires before making it a habit.","edges":[{"to":"alook","type":"alternative","why":"Both coordinate multiple local coding agents into a team; alook builds the full always-on 'AI company' (email, org charts), squad strips it to slash commands + SQLite you can hold in your head.","confidence":0.65,"status":"approved"}],"status":"approved","added":"2026-07-14T14:52:59.886Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"contrabass","why":"Both coordinate fleets of local coding-agent CLIs in the terminal; Squad does ad-hoc manager/worker collaboration over SQLite with no daemon, Contrabass dispatches from Linear/GitHub issues with worktrees, retries and branch-advance verification.","dir":"in","confidence":0.75,"name":"contrabass"},{"slug":"omnigent","why":"Same job — coordinating multiple AI coding CLIs — opposite philosophy: squad is daemonless one-shot shell + SQLite, Omnigent is a full server/meta-harness.","dir":"in","confidence":0.7,"name":"omnigent"},{"slug":"alook","why":"Both coordinate multiple local coding agents into a team; alook builds the full always-on 'AI company' (email, org charts), squad strips it to slash commands + SQLite you can hold in your head.","dir":"out","confidence":0.65,"name":"alook"},{"slug":"fabro","why":"Same goal — multi-agent software process with human gates — opposite weight classes: Squad coordinates AI CLIs through one-shot shell commands and SQLite with no daemon; Fabro runs a persistent server with a runs board and queued 24/7 execution.","dir":"in","confidence":0.6,"name":"fabro"},{"slug":"three-man-team","why":"Same manager/worker/inspector shape, different substance: Squad is tooling — agents coordinating through shell commands and SQLite; Three Man Team is pure process — personas and handoff rules inside a single Claude Code session.","dir":"in","confidence":0.55,"name":"three-man-team"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/mco-org/squad"},"squid":{"name":"squid","owner":"iusztinpaul","slug":"squid","stars":152,"skip_image":true,"avatar":"https://avatars.githubusercontent.com/u/28981860?v=4&s=96","forks":21,"language":"Shell","license":"Apache-2.0","updated":"3 days ago","topics":["coding","orchestration"],"summary":"Claude Code plugin that turns a feature spec into a reviewed PR through a 5-agent pipeline — PA → SWE → Tester → PR-Reviewer → On-Call — with exactly two human gates.","curator_note":"Squid is for people who already live in Claude Code and are tired of re-explaining team conventions every session: markdown specs + five adversarial agents (no agent both writes code and judges it) turn a spec into a PR while you only show up to approve the plan and merge. When NOT: you don't use Claude Code (it's a plugin, not a standalone tool), your stack is Rust/Java/mobile (specs are Python/TS/Go for now), or you already trust an in-house pipeline. Early days and opinionated by design — adopt the opinions or skip it.","edges":[{"to":"agents-cli","type":"alternative","why":"Same shape — a skills layer that upgrades your coding assistant into a specialized workflow — different bet: agents-cli specializes it for the Google Cloud agent lifecycle; squid for a convention-enforcing spec→PR software factory.","confidence":0.55,"status":"approved"},{"to":"claude-reflect","type":"complements","why":"Both ride Claude Code: squid runs the pipeline, claude-reflect captures your mid-run corrections and routes them back into skill and AGENTS.md files — the factory learns between features.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-07T01:01:41.000Z","linkCount":6,"related":{"complements":[{"slug":"alook","why":"alook gives Claude Code agents inboxes, roles, and an always-on runtime; squid gives each of them a disciplined spec→PR pipeline — the org chart and the SOP, together.","dir":"in","confidence":0.65,"name":"alook"},{"slug":"claude-reflect","why":"Both ride Claude Code: squid runs the pipeline, claude-reflect captures your mid-run corrections and routes them back into skill and AGENTS.md files — the factory learns between features.","dir":"out","confidence":0.6,"name":"claude-reflect"},{"slug":"loop-engineering","why":"squid is a concrete, productized maker/checker agent pipeline (PA→SWE→Tester→PR-Reviewer→On-Call) with human gates; loop-engineering is the cross-tool methodology and CLI tooling (PR-babysitter/CI-sweeper patterns, loop-audit, loop-worktree, design checklist) behind such pipelines. Use loop-engineering to design the loop, squid as one ready-made implementation.","dir":"in","confidence":0.6,"name":"loop-engineering"}],"alternative":[{"slug":"agents-cli","why":"Same shape — a skills layer that upgrades your coding assistant into a specialized workflow — different bet: agents-cli specializes it for the Google Cloud agent lifecycle; squid for a convention-enforcing spec→PR software factory.","dir":"out","confidence":0.55,"name":"agents-cli"},{"slug":"bemyagent","why":"Same job — disciplined, human-gated task execution for coding agents — different bet: squid is a Claude Code plugin driving a 5-agent spec→PR pipeline; bemyagent is a tool-agnostic markdown protocol a single agent follows. Pick squid for enforced structure on Claude Code, bemyagent for portability.","dir":"in","confidence":0.55,"name":"bemyagent"},{"slug":"pullfrog","why":"Both turn repo intent into reviewed changes: squid pipelines spec→PR inside Claude Code; pullfrog embeds the agent in GitHub itself — tag a comment and Actions does the rest, any model.","dir":"in","confidence":0.55,"name":"pullfrog"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/iusztinpaul/squid"},"stealth-browser-mcp":{"name":"stealth-browser-mcp","owner":"vibheksoni","slug":"stealth-browser-mcp","stars":1536,"image":"https://raw.githubusercontent.com/vibheksoni/stealth-browser-mcp/master/media/UndetectedStealthBrowser.png","avatar":"https://avatars.githubusercontent.com/u/102437829?v=4&s=96","forks":231,"language":"Python","license":"MIT","updated":"2 months ago","topics":["web"],"summary":"MCP server for undetectable browser automation: real Chrome via nodriver + CDP, Cloudflare/anti-bot bypass, AI-written network hooks — agents browse where Playwright gets blocked.","curator_note":"For the pages where standard automation dies at the Cloudflare wall: real Chrome instances driven through nodriver + CDP, exposed to any MCP client, with network-hook tooling Playwright MCPs don't have. When your agent legitimately needs a protected page (your own accounts, paywalled services you subscribe to), this is the tool that actually works. The obvious caution IS the caution: 'bypasses anti-bot systems' means you're overriding sites' stated wishes — check ToS and law before pointing it anywhere you don't own; expect the arms race to break it periodically, and audit what its hooks can capture (credentials pass through).","edges":[{"to":"browser-use","type":"alternative","why":"Both hand an AI agent a real browser; browser-use optimizes for task completion on the open web, stealth-browser-mcp for surviving anti-bot walls — pick by whether detection is your bottleneck.","confidence":0.6,"status":"approved"},{"to":"scrapling","type":"alternative","why":"Two stealth approaches to protected pages: Scrapling is a scraping framework with Cloudflare-passing fetchers for extraction at scale; stealth-browser-mcp is interactive browser control over MCP for agent-driven sessions.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T14:52:59.857Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"browser-use","why":"Both hand an AI agent a real browser; browser-use optimizes for task completion on the open web, stealth-browser-mcp for surviving anti-bot walls — pick by whether detection is your bottleneck.","dir":"out","confidence":0.6,"name":"browser-use"},{"slug":"scrapling","why":"Two stealth approaches to protected pages: Scrapling is a scraping framework with Cloudflare-passing fetchers for extraction at scale; stealth-browser-mcp is interactive browser control over MCP for agent-driven sessions.","dir":"out","confidence":0.55,"name":"Scrapling"},{"slug":"browser-harness-js","why":"Both CDP-level browser control for agents: stealth-browser-mcp specializes in anti-bot evasion via nodriver; the harness is the thinnest generic typed bridge.","dir":"in","confidence":0.5,"name":"browser-harness-js"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/vibheksoni/stealth-browser-mcp"},"supavec":{"name":"supavec","owner":"supavec","slug":"supavec","stars":1152,"image":"https://github.com/user-attachments/assets/76e2c674-d683-487c-bf02-ac8bccf19e69","avatar":"https://avatars.githubusercontent.com/u/198683319?v=4&s=96","forks":104,"language":"TypeScript","license":"Apache-2.0","updated":"6 months ago","topics":["rag"],"summary":"Open-source RAG-as-a-service (the Carbon.ai alternative): upload any data source, get vector search and a chat API in minutes — Supabase-based, multi-tenant with RLS, streaming responses.","curator_note":"The API-first take on RAG: POST a file, query a /chat endpoint, done — with genuinely production-minded internals (row-level security for tenant isolation, batched embeddings cutting OpenAI cost ~65%, Redis sliding-window rate limits, request tracing). Use it when you want RAG behind an API for your product without assembling the pipeline yourself. Know the shape: it's an open-core SaaS — usage-tiered billing on the cloud version, and self-hosting means running the Next.js + Supabase + Upstash stack yourself with docs that assume you'll read the code. For a self-contained enterprise RAG product with a UI, MaxKB; for just the vector store, Chroma.","edges":[{"to":"maxkb","type":"alternative","why":"Both are open platforms for shipping RAG: MaxKB is the batteries-included enterprise product — visual workflows, UI, zero-code embedding into business systems; Supavec is the developer primitive — a clean REST API for ingestion and chat you build your own product on.","confidence":0.55,"status":"approved"},{"to":"chroma","type":"alternative","why":"Different layers of the same job: Chroma gives you the embedding database and you assemble ingestion, chunking and chat around it; Supavec sells the whole assembled slice as one API.","confidence":0.45,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.945Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"maxkb","why":"Both are open platforms for shipping RAG: MaxKB is the batteries-included enterprise product — visual workflows, UI, zero-code embedding into business systems; Supavec is the developer primitive — a clean REST API for ingestion and chat you build your own product on.","dir":"out","confidence":0.55,"name":"MaxKB"},{"slug":"chroma","why":"Different layers of the same job: Chroma gives you the embedding database and you assemble ingestion, chunking and chat around it; Supavec sells the whole assembled slice as one API.","dir":"out","confidence":0.45,"name":"Chroma"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/supavec/supavec"},"superserve":{"name":"superserve","owner":"superserve-ai","slug":"superserve","stars":427,"image":"https://github.com/user-attachments/assets/ba41a4ed-1c8b-4826-bd3d-6fc8d84472ae","avatar":"https://avatars.githubusercontent.com/u/225320943?v=4&s=96","forks":49,"language":"TypeScript","license":"Apache-2.0","updated":"3 days ago","topics":["agents"],"summary":"Persistent, secure sandboxes for AI agents on Firecracker microVMs — TypeScript and Python SDKs, CLI and console; the runtime is a hosted service, the SDK stack is Apache-2.0.","curator_note":"The pitch is persistence: sandboxes that survive between agent runs instead of being disposable, on Firecracker isolation. Know what the repo is before starring: this is the SDK/CLI/console monorepo — the actual sandbox runtime lives behind the hosted service at superserve.ai, and the README is a contributor doc, not a product doc (the substance is at docs.superserve.ai). Use it when you want managed microVM sandboxes with a clean SDK and don't want to run KVM hosts; NOT for air-gapped or self-hosted requirements — that's CubeSandbox's territory. Young project, small community — evaluate the service's durability before building on it.","edges":[{"to":"cubesandbox","type":"alternative","why":"Both provide Firecracker/microVM sandbox infrastructure for AI agents. CubeSandbox is self-hosted on your own KVM nodes with an E2B-compatible API; Superserve is a hosted service with persistence as the headline and SDKs as the open-source surface.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-16T22:04:16.990Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"cubesandbox","why":"Both provide Firecracker/microVM sandbox infrastructure for AI agents. CubeSandbox is self-hosted on your own KVM nodes with an E2B-compatible API; Superserve is a hosted service with persistence as the headline and SDKs as the open-source surface.","dir":"out","confidence":0.7,"name":"CubeSandbox"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/superserve-ai/superserve"},"supervision":{"name":"supervision","owner":"roboflow","slug":"supervision","stars":48316,"image":"https://media.roboflow.com/open-source/supervision/rf-supervision-banner.png?updatedAt=1678995927529","avatar":"https://avatars.githubusercontent.com/u/53104118?v=4&s=96","forks":4438,"language":"Python","license":"MIT","updated":"2 days ago","topics":["vision"],"summary":"Roboflow's reusable computer-vision toolkit: one Detections API over any model (YOLO, SAM, transformers), 20+ annotators, zone counting, tracking and dataset tools. 48k stars, MIT.","curator_note":"The glue layer every CV project reinvents until it finds this: normalize any model's output into one Detections object, then annotate, filter by polygon zone, count line crossings, track across frames and slice datasets — the unglamorous 80% around the model. Model-agnostic by design, docs and cookbook are exemplary. NOT a model: it detects nothing on its own — you bring YOLO/SAM/anything — and it's a Python library for pipelines you write, not a no-code video-analytics product. Roboflow-backed, so expect the ecosystem to nudge toward their platform at the edges.","edges":[{"to":"ultralytics","type":"complements","why":"The canonical pairing: YOLO produces the detections, supervision consumes them — annotation, zone counting, tracking and evaluation around Ultralytics outputs is the library's headline use case.","confidence":0.8,"status":"approved"},{"to":"sam2","type":"complements","why":"supervision speaks SAM 2 natively — masks become the same Detections object as boxes, so segmentation results plug into the identical annotate/filter/count pipeline.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-23T17:20:10.929Z","linkCount":2,"related":{"complements":[{"slug":"ultralytics","why":"The canonical pairing: YOLO produces the detections, supervision consumes them — annotation, zone counting, tracking and evaluation around Ultralytics outputs is the library's headline use case.","dir":"out","confidence":0.8,"name":"ultralytics"},{"slug":"sam2","why":"supervision speaks SAM 2 natively — masks become the same Detections object as boxes, so segmentation results plug into the identical annotate/filter/count pipeline.","dir":"out","confidence":0.6,"name":"sam2"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/roboflow/supervision"},"surfsense":{"name":"SurfSense","owner":"MODSetter","slug":"surfsense","stars":15298,"image":"https://github.com/user-attachments/assets/9361ef58-1753-4b6e-b275-5020d8847261","avatar":"https://avatars.githubusercontent.com/u/122026167?v=4&s=96","forks":1473,"language":"Python","license":"NOASSERTION","updated":"yesterday","topics":["web","rag"],"summary":"Open-source competitive-intelligence platform for agents: live Reddit/YouTube/TikTok/Maps/search connectors; scheduled agents produce briefs and alerts into a cited knowledge base. REST + MCP.","curator_note":"The interesting shape: not a scraper, the layer above — agents subscribe to live market signals (Reddit threads, YouTube, rankings, Maps) through one REST/MCP surface, scheduled runs turn findings into briefs, and everything lands in a citable knowledge base. If you're building market-monitoring agents, this saves you the connector swamp. Two honesty flags: the project just pivoted from 'NotebookLM alternative' to competitive intelligence (the README says so itself) — momentum is real but the identity is weeks old; and NO standard license resolution at review time — verify before building on it.","edges":[],"status":"approved","added":"2026-07-14T15:35:46.288Z","linkCount":3,"related":{"complements":[{"slug":"flowsint","why":"Both do open-source intelligence, at different tempos: surfsense's scheduled agents monitor live sources into cited briefs; Flowsint is where you take a lead from those briefs and investigate it hands-on as an entity graph.","dir":"in","confidence":0.45,"name":"flowsint"}],"alternative":[{"slug":"last30days-skill","why":"Same signal sources (Reddit, YouTube, social, search), different shapes: SurfSense is a self-hosted platform where scheduled agents build a cited knowledge base; /last30days is a skill inside your own agent — on-demand briefs, BYO keys, no infrastructure.","dir":"in","confidence":0.65,"name":"last30days-skill"},{"slug":"horizon","why":"Both watch live sources and produce briefs; SurfSense is the platform play (connectors, API/MCP, cited knowledge base for agents), Horizon the personal radar — one human, one daily briefing.","dir":"in","confidence":0.6,"name":"Horizon"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/MODSetter/surfsense"},"tailclaude":{"name":"tailclaude","owner":"rohitg00","slug":"tailclaude","stars":208,"image":"https://raw.githubusercontent.com/rohitg00/tailclaude/main/assets/banner-v2.jpg","avatar":"https://avatars.githubusercontent.com/u/48523873?v=4&s=96","forks":22,"language":"TypeScript","license":null,"updated":"4 months ago","topics":["coding"],"summary":"Claude Code from any browser via your Tailscale tailnet — streaming chat UI, session history, model switching and cost dashboards; no SSH, no terminal.","curator_note":"The 'doom coding from your phone' setup without the terminal: publish Claude Code to your tailnet, scan a QR code, get a touch-optimized chat UI with streaming, full session history and cost dashboards — Tailscale handles auth and transport, so nothing new is exposed. The caveats are serious though: the repo ships NO license (all rights reserved by default — you can run it, not fork or redistribute), it's young (~200 stars) and has been quiet for months. Treat it as a clever pattern to evaluate, NOT infrastructure to depend on.","edges":[{"to":"codenomad","type":"alternative","why":"Same itch — escaping the terminal for your coding agent. CodeNomad is a rich desktop/server cockpit for OpenCode; TailClaude is a zero-install browser UI for Claude Code over Tailscale.","confidence":0.6,"status":"approved"}],"status":"approved","added":"2026-07-11T15:07:43.430Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"codenomad","why":"Same itch — escaping the terminal for your coding agent. CodeNomad is a rich desktop/server cockpit for OpenCode; TailClaude is a zero-install browser UI for Claude Code over Tailscale.","dir":"out","confidence":0.6,"name":"CodeNomad"},{"slug":"omnigent","why":"Covers tailclaude's use case — driving your coding agent from any browser/phone — but for many harnesses, at the cost of a much heavier stack.","dir":"in","confidence":0.6,"name":"omnigent"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/rohitg00/tailclaude"},"telegram-drive":{"name":"Telegram-Drive","owner":"caamer20","slug":"telegram-drive","stars":4451,"image":"https://raw.githubusercontent.com/caamer20/telegram-drive/main/screenshots/AuthScreen.png","avatar":"https://avatars.githubusercontent.com/u/9097300?v=4&s=96","forks":664,"language":"TypeScript","license":null,"updated":"4 days ago","topics":["storage"],"summary":"Tauri/Rust desktop app that turns your Telegram account into unlimited cloud storage — file-explorer UI over channels, media streaming, share links, and a local REST API for LLM/tool integration.","curator_note":"Use it as free personal cloud storage with a real file manager riding Telegram's servers — folders are private channels, uploads stream, links are shareable with passwords and expiry. The agent hook is the opt-in local REST API (key-auth, ships an OpenAPI spec): point a tool-calling LLM at it and it becomes an agent-drivable file backend. NOT for team or production storage — unofficial client, Telegram-ToS gray zone, and your data lives as chat messages. Only tangentially an AI repo; it earns its spot as storage infrastructure agents can use, nothing more.","edges":[],"status":"approved","added":"2026-07-14T17:21:39.944Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"karakeep","why":"Same shelf — self-hosted personal hoarding with an API agents can drive — different content: Telegram-Drive stores raw files on Telegram's servers; Karakeep hoards bookmarks, notes and pages with AI tagging, search and archival.","dir":"in","confidence":0.45,"name":"karakeep"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/caamer20/telegram-drive"},"three-man-team":{"name":"three-man-team","owner":"russelleNVy","slug":"three-man-team","stars":889,"image":"https://raw.githubusercontent.com/russelleNVy/three-man-team/main/assets/banner.png","avatar":"https://avatars.githubusercontent.com/u/272609506?v=4&s=96","forks":112,"language":"Shell","license":"MIT","updated":"1 months ago","topics":["skills","coding"],"summary":"A disciplined 3-agent dev process as context files — Architect plans, Builder builds the brief, Reviewer gates — running in one Claude Code session via subagents. Token-frugal by design.","curator_note":"It's a process, not software: three markdown personas with strict handoffs (plan → brief → build → review → deploy) that target the classic solo-agent failure modes — scope drift, unrequested features, token burn. The five CLAUDE.md token rules ('is this speculative? kill the tool call') are genuinely good hygiene even outside the team. Everything runs in one Claude Code session via the Agent tool; works with anything that reads context files. NOT a framework, and not for exploratory hacking — the ceremony pays off on production work and taxes quick experiments. Know the two strings attached: a commercial 'Pro' waitlist, and Arch phones a version registry at session start to offer updates (auditable, consent-gated, but it's there).","edges":[{"to":"bemyagent","type":"alternative","why":"Both are markdown-first process scaffolds for coding agents. bemyagent bootstraps one structured workspace with a Think→Task→Execute→Verify cycle; three-man-team splits the discipline across three role personas with review as a hard gate.","confidence":0.6,"status":"approved"},{"to":"squad","type":"alternative","why":"Same manager/worker/inspector shape, different substance: Squad is tooling — agents coordinating through shell commands and SQLite; Three Man Team is pure process — personas and handoff rules inside a single Claude Code session.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-16T09:58:13.750Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"bemyagent","why":"Both are markdown-first process scaffolds for coding agents. bemyagent bootstraps one structured workspace with a Think→Task→Execute→Verify cycle; three-man-team splits the discipline across three role personas with review as a hard gate.","dir":"out","confidence":0.6,"name":"bemyagent"},{"slug":"squad","why":"Same manager/worker/inspector shape, different substance: Squad is tooling — agents coordinating through shell commands and SQLite; Three Man Team is pure process — personas and handoff rules inside a single Claude Code session.","dir":"out","confidence":0.55,"name":"squad"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/russelleNVy/three-man-team"},"tokensave":{"name":"tokensave","owner":"aovestdipaperino","slug":"tokensave","stars":477,"image":"https://raw.githubusercontent.com/aovestdipaperino/tokensave/master/src/resources/logo.png","avatar":"https://avatars.githubusercontent.com/u/42046613?v=4&s=96","forks":43,"language":"Rust","license":"MIT","updated":"yesterday","topics":["coding","local"],"summary":"Code-intelligence MCP server for coding agents — a pre-indexed semantic graph (libSQL + FTS5) they query instead of grepping: symbols, callers, impact radius in one call. 100% local, 50+ languages.","curator_note":"Use it to stop agents burning tokens on grep/glob/read exploration: one MCP call returns the symbols, relationships and snippets a task needs, and it's local (libSQL, no cloud). Broad reach — 50+ languages, 12+ agent integrations. NOT worth it on small repos where a couple of greps suffice, and it's an index you must keep fresh (re-index on change) — a stale graph misleads the agent. 80+ tools is a lot of surface; most tasks touch a handful.","edges":[{"to":"codegraph-mcp","type":"alternative","why":"Same job — a code knowledge graph served to agents over MCP: codegraph-mcp bets on compliance (hash-chained audit of every read, cross-language HTTP edges); tokensave bets on breadth and token savings (50+ langs, 80+ tools, 12 agent integrations).","confidence":0.8,"status":"approved"},{"to":"lynkr","type":"complements","why":"Two token-savers at different layers that stack: lynkr compresses and routes at the gateway; tokensave cuts the exploration calls at the MCP layer so the agent asks the graph instead of scanning files.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-07T22:49:33.343Z","linkCount":9,"related":{"complements":[{"slug":"lynkr","why":"Two token-savers at different layers that stack: lynkr compresses and routes at the gateway; tokensave cuts the exploration calls at the MCP layer so the agent asks the graph instead of scanning files.","dir":"out","confidence":0.55,"name":"Lynkr"},{"slug":"headroom","why":"Two layers of the same diet: TokenSave stops tokens at the source (indexed code queries instead of grep dumps), Headroom compresses whatever still flows through. Stackable — one shrinks what the agent asks for, the other what it receives.","dir":"in","confidence":0.55,"name":"headroom"},{"slug":"9router","why":"Both cut a coding agent's token bill but at different layers, so they stack: 9router compresses tool_result payloads at the gateway and routes to cheap/free models, while tokensave gives the agent a pre-indexed code graph to query instead of burning tokens on grep/read. Front your agent with 9router and hand it tokensave as an MCP tool.","dir":"in","confidence":0.5,"name":"9router"}],"alternative":[{"slug":"codegraph-mcp","why":"Same job — a code knowledge graph served to agents over MCP: codegraph-mcp bets on compliance (hash-chained audit of every read, cross-language HTTP edges); tokensave bets on breadth and token savings (50+ langs, 80+ tools, 12 agent integrations).","dir":"out","confidence":0.8,"name":"codegraph-mcp"},{"slug":"cocoindex-code","why":"Same job — replace agent grepping with a pre-built local index. TokenSave answers structural questions (symbols, callers, impact radius) from a semantic graph; cocoindex-code answers 'where is the code that does X' via AST-chunked embedding search.","dir":"in","confidence":0.8,"name":"cocoindex-code"},{"slug":"codebase-memory-mcp","why":"The same exact job — pre-indexed structural code intelligence over MCP so agents stop grepping. TokenSave is a libSQL semantic graph across 50+ languages; codebase-memory is a zero-dependency C binary betting everything on speed: 158 languages, sub-ms.","dir":"in","confidence":0.8,"name":"codebase-memory-mcp"},{"slug":"repowise","why":"Both are code-intelligence MCP servers that cut agent context burn; tokensave answers structural questions (symbols, callers, impact) from a semantic graph, repowise layers health scoring and executable refactoring plans on top of its own graph.","dir":"in","confidence":0.6,"name":"repowise"},{"slug":"cocoindex","why":"Head-to-head on the code-index-for-agents job: tokensave ships a pre-indexed libSQL+FTS5 semantic graph agents query over MCP; CocoIndex's flagship cocoindex-code does the same over MCP but incrementally re-indexed on every commit via the Δ engine.","dir":"in","confidence":0.55,"name":"cocoindex"},{"slug":"understand-anything","why":"Same 'index the codebase once, query the structure' idea: tokensave is a pre-indexed semantic graph agents hit instead of grepping (symbols, callers, impact radius); Understand Anything is the human-facing equivalent with tours, search and diff impact.","dir":"in","confidence":0.55,"name":"Understand-Anything"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/aovestdipaperino/tokensave"},"train-llm-from-scratch":{"name":"train-llm-from-scratch","owner":"FareedKhan-dev","slug":"train-llm-from-scratch","stars":8531,"image":"https://cdn-images-1.medium.com/max/5200/1*r99Hq3YBd5FTTWLNYKKvPw.png","avatar":"https://avatars.githubusercontent.com/u/63067900?v=4&s=96","forks":1177,"language":"Python","license":"MIT","updated":"1 months ago","topics":["training"],"summary":"The full LLM pipeline hand-written in plain PyTorch — tokens, transformer, pretraining, then SFT, reward model, PPO, DPO, GRPO. No trl, no peft: read every algorithm, train on one GPU.","curator_note":"The textbook that runs: every stage from raw text to an aligned reasoning-style model, each algorithm implemented by hand in readable PyTorch — including the post-training alphabet (SFT → reward model → PPO/DPO → GRPO) that most tutorials wave at. If you want to understand what TRL actually does under its trainer classes, this is the fastest honest path, and it fits on a single GPU at the 13M–1B scale. NOT for production: hand-rolled training code at toy scale is the point, not the product — when you're done learning, ship with TRL or LLaMA-Factory. From the author of all-agentic-architectures, same runnable-textbook DNA.","edges":[{"to":"trl","type":"alternative","why":"Same post-training algorithms, opposite purposes: TRL is the production library you call; this re-implements SFT/PPO/DPO/GRPO by hand so you can read them. Learn here, ship with TRL.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-22T18:15:22.318Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"trl","why":"Same post-training algorithms, opposite purposes: TRL is the production library you call; this re-implements SFT/PPO/DPO/GRPO by hand so you can read them. Learn here, ship with TRL.","dir":"out","confidence":0.55,"name":"trl"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/FareedKhan-dev/train-llm-from-scratch"},"trl":{"name":"trl","owner":"huggingface","slug":"trl","stars":18911,"image":"https://huggingface.co/datasets/trl-lib/documentation-images/resolve/main/trl_banner_dark.png","avatar":"https://avatars.githubusercontent.com/u/25720743?v=4&s=96","forks":2858,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["training"],"summary":"Hugging Face's post-training library: SFT, DPO, GRPO, KTO and reward-model trainers on top of Transformers — from a Colab LoRA run to multi-GPU deployments.","curator_note":"The on-ramp for post-training: if your model is on the Hub and your job fits SFT/DPO/GRPO/KTO, a Trainer class gets you a running job in an afternoon — and nothing else scales down to a free Colab as gracefully. PEFT/LoRA, quantized training and accelerate multi-GPU come along for free. NOT for frontier-scale RL dataflows (that's verl/slime territory — no Megatron, no disaggregated rollout engines), and the trainer abstraction that makes it easy also hides the loss mechanics: when results surprise you, read the trainer source before blaming the data.","edges":[{"to":"slime","type":"alternative","why":"The lighter trainer slime's own docs point you toward: single-node SFT/LoRA/DPO on the HF stack instead of Megatron-scale RL infrastructure. Same goal — a post-trained model — at opposite ends of the infra spectrum.","confidence":0.7,"status":"approved"},{"to":"vllm","type":"complements","why":"TRL's online RL trainers (GRPO) generate through vLLM — the use_vllm flag is the standard throughput path.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.156Z","linkCount":6,"related":{"complements":[{"slug":"vllm","why":"TRL's online RL trainers (GRPO) generate through vLLM — the use_vllm flag is the standard throughput path.","dir":"out","confidence":0.55,"name":"vLLM"}],"alternative":[{"slug":"llamafactory","why":"The two dominant open fine-tuning stacks: TRL is the code-first Hugging Face library; LlamaFactory the config-driven unified trainer with a GUI and a wider model-coverage matrix.","dir":"in","confidence":0.8,"name":"LlamaFactory"},{"slug":"h2o-llmstudio","why":"The same LoRA/DPO fine-tuning jobs behind different interfaces: TRL is the code-first library for Hub-native workflows, LLM Studio the no-code GUI for teams that don't write training loops.","dir":"in","confidence":0.75,"name":"h2o-llmstudio"},{"slug":"verl","why":"Same goal — a post-trained model — at opposite ends of the infra spectrum: TRL for Hub-native single-node SFT/DPO/GRPO, verl when the job needs disaggregated rollout engines and multi-node RL dataflows.","dir":"in","confidence":0.75,"name":"verl"},{"slug":"slime","why":"The lighter trainer slime's own docs point you toward: single-node SFT/LoRA/DPO on the HF stack instead of Megatron-scale RL infrastructure. Same goal — a post-trained model — at opposite ends of the infra spectrum.","dir":"out","confidence":0.7,"name":"slime"},{"slug":"train-llm-from-scratch","why":"Same post-training algorithms, opposite purposes: TRL is the production library you call; this re-implements SFT/PPO/DPO/GRPO by hand so you can read them. Learn here, ship with TRL.","dir":"in","confidence":0.55,"name":"train-llm-from-scratch"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/huggingface/trl"},"ultralytics":{"name":"ultralytics","owner":"ultralytics","slug":"ultralytics","stars":59783,"image":"https://raw.githubusercontent.com/ultralytics/assets/main/yolov8/banner-yolov8.png","avatar":"https://avatars.githubusercontent.com/u/26833451?v=4&s=96","forks":11430,"language":"Python","license":"AGPL-3.0","updated":"yesterday","topics":["vision"],"summary":"Ultralytics YOLO (v8→26): real-time object detection, segmentation, classification, pose and tracking behind one Python/CLI API — train, validate and export to ONNX/TensorRT/CoreML.","curator_note":"The de-facto standard for real-time detection: one API covers train/val/predict/export, a deep model zoo from nano (edge) to xlarge, and deployment paths to ONNX, TensorRT, CoreML and TFLite that actually work. Reach for it when you need boxes, masks or poses at video framerate on your own hardware. The catch is the license: AGPL-3.0 — shipping it inside a closed-source product requires Ultralytics' commercial license, which surprises teams late. NOT for promptable zero-shot segmentation (that's SAM 2's job) and not an LLM tool — it's the classic-vision workhorse your agent pipeline calls.","edges":[{"to":"deepface","type":"complements","why":"The standard two-stage pipeline: YOLO finds the people in the frame, deepface tells you who they are — detector crops feed face verification directly.","confidence":0.75,"status":"approved"},{"to":"sam2","type":"complements","why":"The standard detect-then-segment pipeline: YOLO boxes become SAM 2 prompts — the detector names and localizes, SAM 2 delivers pixel-perfect masks and video tracking.","confidence":0.8,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.173Z","linkCount":3,"related":{"complements":[{"slug":"sam2","why":"The standard detect-then-segment pipeline: YOLO boxes become SAM 2 prompts — the detector names and localizes, SAM 2 delivers pixel-perfect masks and video tracking.","dir":"out","confidence":0.8,"name":"sam2"},{"slug":"supervision","why":"The canonical pairing: YOLO produces the detections, supervision consumes them — annotation, zone counting, tracking and evaluation around Ultralytics outputs is the library's headline use case.","dir":"in","confidence":0.8,"name":"supervision"},{"slug":"deepface","why":"The standard two-stage pipeline: YOLO finds the people in the frame, deepface tells you who they are — detector crops feed face verification directly.","dir":"out","confidence":0.75,"name":"deepface"}],"alternative":[],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/ultralytics/ultralytics"},"understand-anything":{"name":"Understand-Anything","owner":"Egonex-AI","slug":"understand-anything","stars":75734,"image":"https://raw.githubusercontent.com/Egonex-AI/understand-anything/main/assets/hero.png","avatar":"https://avatars.githubusercontent.com/u/257477979?v=4&s=96","forks":6307,"language":"TypeScript","license":"MIT","updated":"3 days ago","topics":["coding"],"summary":"Plugin for Claude Code and 16 other hosts that turns any codebase into an interactive knowledge graph — multi-agent analysis, layered dashboard, guided tours, diff-impact view, domain mapping.","curator_note":"The onboarding killer app, and its motto is the right one: graphs that teach, not graphs that impress. /understand runs a multi-agent pipeline over the repo, then the dashboard gives you architecture layers, dependency-ordered guided tours, semantic search, diff blast-radius, and a business-domain view that maps code to real processes. The team trick is the sleeper: the graph is plain JSON — commit it once and every teammate skips the analysis. Incremental re-runs + a post-commit auto-update hook keep it fresh. Budget real tokens for the first full run on a large repo (their own warning), or point it at a local model. NOT an agent context server — this graph is for humans first; when you want a code graph served TO agents over MCP, that's codegraph-mcp or tokensave territory.","edges":[{"to":"codegraph-mcp","type":"alternative","why":"Both build knowledge graphs of your codebase, for opposite consumers: codegraph-mcp serves symbols, call edges and blast-radius queries to AI agents over MCP with an audit chain; Understand Anything renders an interactive dashboard for human comprehension and onboarding.","confidence":0.6,"status":"approved"},{"to":"tokensave","type":"alternative","why":"Same 'index the codebase once, query the structure' idea: tokensave is a pre-indexed semantic graph agents hit instead of grepping (symbols, callers, impact radius); Understand Anything is the human-facing equivalent with tours, search and diff impact.","confidence":0.55,"status":"approved"},{"to":"repowise","type":"alternative","why":"Overlapping codebase-intelligence job: repowise leans quantitative — defect-calibrated health scores, refactoring plans, git analytics over MCP; Understand Anything leans pedagogical — teach the architecture through an explorable graph.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-16T15:39:30.861Z","linkCount":4,"related":{"complements":[],"alternative":[{"slug":"codegraph-mcp","why":"Both build knowledge graphs of your codebase, for opposite consumers: codegraph-mcp serves symbols, call edges and blast-radius queries to AI agents over MCP with an audit chain; Understand Anything renders an interactive dashboard for human comprehension and onboarding.","dir":"out","confidence":0.6,"name":"codegraph-mcp"},{"slug":"tokensave","why":"Same 'index the codebase once, query the structure' idea: tokensave is a pre-indexed semantic graph agents hit instead of grepping (symbols, callers, impact radius); Understand Anything is the human-facing equivalent with tours, search and diff impact.","dir":"out","confidence":0.55,"name":"tokensave"},{"slug":"code-review-graph","why":"Code knowledge graphs for opposite consumers: understand-anything renders an explorable dashboard to teach humans the architecture; code-review-graph feeds minimal context to agents.","dir":"in","confidence":0.55,"name":"code-review-graph"},{"slug":"repowise","why":"Overlapping codebase-intelligence job: repowise leans quantitative — defect-calibrated health scores, refactoring plans, git analytics over MCP; Understand Anything leans pedagogical — teach the architecture through an explorable graph.","dir":"out","confidence":0.5,"name":"repowise"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Egonex-AI/understand-anything"},"unlimited-ocr":{"name":"Unlimited-OCR","owner":"baidu","slug":"unlimited-ocr","stars":17780,"image":"https://raw.githubusercontent.com/baidu/unlimited-ocr/main/assets/baidu.png","avatar":"https://avatars.githubusercontent.com/u/13245940?v=4&s=96","forks":1677,"language":"Python","license":"MIT","updated":"3 days ago","topics":["ocr","rag"],"summary":"Baidu's open OCR VLM that parses entire multi-page documents in one shot — 'unlimited' long-horizon parsing pushing DeepSeek-OCR further. MIT weights on HF; serve via transformers, vLLM or SGLang.","curator_note":"The current bar if you batch-parse long PDFs: one-shot multi-page parsing keeps cross-page structure (tables spanning pages, running sections) that page-by-page OCR pipelines lose, and first-party vLLM/SGLang recipes plus official Docker images make serving unusually painless for a fresh research model. NOT a lightweight dependency — it's a GPU-hungry VLM; for occasional single pages classic OCR or a hosted API is cheaper. Inference-only repo: no training code, and evaluation beyond the paper's claims is on you.","edges":[{"to":"olmocr","type":"alternative","why":"Same job — self-hosted VLM that linearizes messy PDFs into LLM-ready text: olmocr is AllenAI's page-pipeline toolkit built for training-data ingestion; Unlimited-OCR bets on one-shot long-horizon parsing that preserves cross-page structure.","confidence":0.8,"status":"approved"},{"to":"vllm","type":"built_with","why":"First-party vLLM support is the blessed serving path — official recipe plus vllm/vllm-openai:unlimited-ocr Docker images for CUDA 12.9/13.0.","confidence":0.7,"status":"approved"}],"status":"approved","added":"2026-07-07T19:24:17.501Z","linkCount":5,"related":{"complements":[],"alternative":[{"slug":"olmocr","why":"Same job — self-hosted VLM that linearizes messy PDFs into LLM-ready text: olmocr is AllenAI's page-pipeline toolkit built for training-data ingestion; Unlimited-OCR bets on one-shot long-horizon parsing that preserves cross-page structure.","dir":"out","confidence":0.8,"name":"olmocr"},{"slug":"mineru","why":"Both turn documents into LLM-ready text: Unlimited-OCR is a one-shot long-horizon VLM; MinerU is a staged layout-analysis + OCR pipeline that also ingests Office formats.","dir":"in","confidence":0.75,"name":"MinerU"},{"slug":"chandra","why":"Both are open OCR vision-language models. Baidu's model bets on one-shot long-horizon parsing of entire multi-page documents; Chandra processes per-page with stronger layout/table/form structure and a broader language benchmark.","dir":"in","confidence":0.55,"name":"chandra"},{"slug":"xberg","why":"Engine vs model: xberg parses 96 formats deterministically; Unlimited-OCR throws a VLM at whole documents. Same slot in a RAG ingestion pipeline.","dir":"in","confidence":0.55,"name":"xberg"}],"built_with":[{"slug":"vllm","why":"First-party vLLM support is the blessed serving path — official recipe plus vllm/vllm-openai:unlimited-ocr Docker images for CUDA 12.9/13.0.","dir":"out","confidence":0.7,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/baidu/unlimited-ocr"},"verl":{"name":"verl","owner":"verl-project","slug":"verl","stars":22621,"image":"https://github.com/user-attachments/assets/c42e675e-497c-4508-8bb9-093ad4d1f216","avatar":"https://avatars.githubusercontent.com/u/212961691?v=4&s=96","forks":4264,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["training"],"summary":"ByteDance's RL post-training library (HybridFlow): PPO/GRPO dataflows in a few lines, FSDP/Megatron training with vLLM/SGLang rollouts, production-proven at frontier scale.","curator_note":"The community default for open RL post-training: the hybrid-controller model expresses PPO/GRPO/DAPO dataflows in a few lines, and the backend matrix (FSDP or Megatron for training, vLLM or SGLang for rollouts) means it fits infrastructure you already have. Proven on real frontier runs and the most-forked codebase in the space. NOT an afternoon tool — multi-GPU distributed debugging is table stakes; for single-node SFT/LoRA use TRL instead, and know the tradeoff slime calls out: multi-backend abstraction can lag upstream engine features.","edges":[{"to":"slime","type":"alternative","why":"Same job — large-scale RL post-training. verl bets on multi-backend flexibility (FSDP/Megatron × vLLM/SGLang) where slime hard-commits to Megatron+SGLang; pick verl when your infra is vLLM- or FSDP-shaped.","confidence":0.85,"status":"approved"},{"to":"vllm","type":"built_with","why":"vLLM is one of verl's first-class rollout engines — RL generation batches run through it during training.","confidence":0.8,"status":"approved"},{"to":"trl","type":"alternative","why":"Same goal — a post-trained model — at opposite ends of the infra spectrum: TRL for Hub-native single-node SFT/DPO/GRPO, verl when the job needs disaggregated rollout engines and multi-node RL dataflows.","confidence":0.75,"status":"approved"}],"status":"approved","added":"2026-07-11T13:28:52.190Z","linkCount":4,"related":{"complements":[{"slug":"opensre","why":"OpenSRE supplies the missing piece verl-style RL training needs for infrastructure agents: a realistic environment with scalable feedback for incident-response rollouts.","dir":"in","confidence":0.55,"name":"opensre"}],"alternative":[{"slug":"slime","why":"Same job — large-scale RL post-training. verl bets on multi-backend flexibility (FSDP/Megatron × vLLM/SGLang) where slime hard-commits to Megatron+SGLang; pick verl when your infra is vLLM- or FSDP-shaped.","dir":"out","confidence":0.85,"name":"slime"},{"slug":"trl","why":"Same goal — a post-trained model — at opposite ends of the infra spectrum: TRL for Hub-native single-node SFT/DPO/GRPO, verl when the job needs disaggregated rollout engines and multi-node RL dataflows.","dir":"out","confidence":0.75,"name":"trl"}],"built_with":[{"slug":"vllm","why":"vLLM is one of verl's first-class rollout engines — RL generation batches run through it during training.","dir":"out","confidence":0.8,"name":"vLLM"}]},"url":"https://stackmap.shipwithai.xyz/repos/verl-project/verl"},"vibe-trading":{"name":"Vibe-Trading","owner":"HKUDS","slug":"vibe-trading","stars":26713,"image":"https://raw.githubusercontent.com/HKUDS/vibe-trading/main/assets/icon.png","avatar":"https://avatars.githubusercontent.com/u/118165258?v=4&s=96","forks":4353,"language":"Python","license":"MIT","updated":"yesterday","topics":["finance","agents"],"summary":"HKUDS' personal trading agent: one command gives your agent market data, analysis and trading capability, with a shadow-account mode, API and MCP surface.","curator_note":"From the lab behind LightRAG and VideoAgent: a full trading-agent stack with the two features that matter — a shadow account so strategies run against real markets with fake money, and an MCP surface so YOUR agent gains the capability rather than you adopting theirs. The team publicly disavows the fake tokens trading on its name — read that as both integrity and a warning about the space. NOT financial advice infrastructure: agents amplify whatever edge (or absence of one) you encode; stay in the shadow account until the data argues otherwise.","edges":[],"status":"approved","added":"2026-07-14T16:04:55.203Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"autohedge","why":"Same ambition — agents trading on your behalf — with opposite temperaments: Vibe-Trading (HKUDS) ships shadow accounts and an MCP surface for cautious integration; AutoHedge sells the autonomous hedge-fund dream. Both need your skepticism.","dir":"in","confidence":0.6,"name":"AutoHedge"},{"slug":"metatrader-mcp-server","why":"Same job — give your agent real trading capability over MCP — different scope: Vibe-Trading is a full personal trading-agent stack with market analysis and a shadow-account safety mode; metatrader-mcp-server is the raw MT5 broker bridge, bring your own judgment.","dir":"in","confidence":0.55,"name":"metatrader-mcp-server"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/HKUDS/vibe-trading"},"video-spec-builder":{"name":"video-spec-builder","owner":"feicaiclub","slug":"video-spec-builder","stars":822,"image":"https://github.com/user-attachments/assets/7820d93e-84b6-4e09-904c-9567c6595c57","avatar":"https://avatars.githubusercontent.com/u/156590182?v=4&s=96","forks":99,"language":"JavaScript","license":"MIT","updated":"2 months ago","topics":["skills"],"summary":"A director-skill for Claude Code/Codex: interrogates your vague video idea until it's a second-by-second storyboard spec (video-spec.md), ready for HyperFrames to render.","curator_note":"Built on a sharp observation: the hard part of making a video isn't rendering, it's knowing what you want. Install the skill, say 'I want to make a video', and it grills you like a director — audience, length, the one takeaway line, which shot carries the weight — until a precise spec exists. The spec-first discipline is the value even if you never render. NOT a renderer (it hands off to HyperFrames) and the project's home base is Chinese-language — English docs work but the community and examples skew zh.","edges":[{"to":"videoagent","type":"alternative","why":"Two philosophies for idea→video: video-spec-builder forces a human-approved storyboard spec before any rendering; VideoAgent owns the whole pipeline conversationally. Spec-first control vs all-in-one convenience.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-14T15:14:12.920Z","linkCount":2,"related":{"complements":[{"slug":"vimax","why":"Spec-first meets render-pipeline: video-spec-builder interrogates a vague idea into a second-by-second storyboard spec, exactly the kind of screenplay input ViMax's Script2Video pipeline consumes (no official integration — the spec skill targets HyperFrames).","dir":"in","confidence":0.5,"name":"ViMax"}],"alternative":[{"slug":"videoagent","why":"Two philosophies for idea→video: video-spec-builder forces a human-approved storyboard spec before any rendering; VideoAgent owns the whole pipeline conversationally. Spec-first control vs all-in-one convenience.","dir":"out","confidence":0.55,"name":"VideoAgent"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/feicaiclub/video-spec-builder"},"videoagent":{"name":"VideoAgent","owner":"HKUDS","slug":"videoagent","stars":1517,"image":"https://raw.githubusercontent.com/HKUDS/videoagent/main/assets/logo_new.png","avatar":"https://avatars.githubusercontent.com/u/118165258?v=4&s=96","forks":211,"language":"Python","license":"MIT","updated":"2 days ago","topics":["vision","agents"],"summary":"All-in-one agentic framework for video: understanding and summarization, clip editing, and generative remaking, driven end-to-end through natural-language conversation.","curator_note":"From HKUDS (the lab behind LightRAG): describe what you want — 'summarize this lecture', 'cut the highlights', 'remake this as a trailer' — and the agent plans tool use across understanding, editing and generation in one conversational loop. The breadth is the differentiator; nothing else open-source covers understand+edit+remake together. NOT a video model itself — it orchestrates underlying multi-modal models, so quality and cost track what you plug in; young codebase from an academic group, expect research-grade edges rather than product polish.","edges":[],"status":"approved","added":"2026-07-13T10:36:42.285Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"vimax","why":"Same lab (HKUDS), overlapping agentic-video job, opposite direction: VideoAgent operates on existing footage — understand, summarize, edit, remake — while ViMax generates films from an idea, novel or screenplay from scratch.","dir":"in","confidence":0.7,"name":"ViMax"},{"slug":"video-spec-builder","why":"Two philosophies for idea→video: video-spec-builder forces a human-approved storyboard spec before any rendering; VideoAgent owns the whole pipeline conversationally. Spec-first control vs all-in-one convenience.","dir":"in","confidence":0.55,"name":"video-spec-builder"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/HKUDS/videoagent"},"vimax":{"name":"ViMax","owner":"HKUDS","slug":"vimax","stars":11292,"image":"https://raw.githubusercontent.com/HKUDS/vimax/main/assets/vimax.svg","avatar":"https://avatars.githubusercontent.com/u/118165258?v=4&s=96","forks":1669,"language":"Python","license":"MIT","updated":"4 days ago","topics":["vision","agents"],"summary":"HKUDS multi-agent video studio: turns an idea, novel or screenplay into a finished film — scriptwriting, storyboards, consistent characters, then rendering via Seedance/Nano Banana/Omni APIs.","curator_note":"Use it when you want long-form narrative video rather than 8-second clips: the agent pipeline (screenwriter → storyboard artist → character extractor → portrait generator → shot renderer) is engineered for cross-shot character and scene consistency, with an interactive agent-loop TUI for revision, resume and context compaction. Novel2Video — compressing a whole book into episodic video — is the flashiest trick. NOT a video model: you bring API keys for the generators (Doubao Seedream/Seedance, Nano Banana, MiniMax, Google Omni; several routed via the yunwu.ai reseller, so costs are real and defaults skew CN). Same lab as VideoAgent — pick VideoAgent to work on existing footage, ViMax to create from scratch.","edges":[{"to":"videoagent","type":"alternative","why":"Same lab (HKUDS), overlapping agentic-video job, opposite direction: VideoAgent operates on existing footage — understand, summarize, edit, remake — while ViMax generates films from an idea, novel or screenplay from scratch.","confidence":0.7,"status":"approved"},{"to":"video-spec-builder","type":"complements","why":"Spec-first meets render-pipeline: video-spec-builder interrogates a vague idea into a second-by-second storyboard spec, exactly the kind of screenplay input ViMax's Script2Video pipeline consumes (no official integration — the spec skill targets HyperFrames).","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T23:32:33.038Z","linkCount":2,"related":{"complements":[{"slug":"video-spec-builder","why":"Spec-first meets render-pipeline: video-spec-builder interrogates a vague idea into a second-by-second storyboard spec, exactly the kind of screenplay input ViMax's Script2Video pipeline consumes (no official integration — the spec skill targets HyperFrames).","dir":"out","confidence":0.5,"name":"video-spec-builder"}],"alternative":[{"slug":"videoagent","why":"Same lab (HKUDS), overlapping agentic-video job, opposite direction: VideoAgent operates on existing footage — understand, summarize, edit, remake — while ViMax generates films from an idea, novel or screenplay from scratch.","dir":"out","confidence":0.7,"name":"VideoAgent"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/HKUDS/vimax"},"vllm":{"name":"vLLM","owner":"vllm-project","slug":"vllm","stars":86940,"image":"https://raw.githubusercontent.com/vllm-project/vllm/main/docs/assets/logos/vllm-logo-text-light.png","avatar":"https://avatars.githubusercontent.com/u/136984999?v=4&s=96","forks":19744,"language":"Python","license":"Apache-2.0","updated":"yesterday","topics":["local"],"summary":"High-throughput, memory-efficient inference and serving engine for LLMs.","edges":[],"added":"2026-06-22T23:27:20.000Z","linkCount":14,"related":{"complements":[{"slug":"litellm","why":"LiteLLM fronts a self-hosted vLLM server as one more OpenAI-compatible backend — vLLM does the serving, LiteLLM does routing, keys and spend tracking.","dir":"in","confidence":0.72,"name":"litellm"},{"slug":"llamafactory","why":"Fine-tune here, serve fast there — LlamaFactory ships a vLLM backend for high-throughput inference on the models it trains.","dir":"in","confidence":0.6,"name":"LlamaFactory"},{"slug":"exxperts","why":"The same gateway setup explicitly supports a vLLM proxy as the model backend, so rooms can run entirely against self-served local models.","dir":"in","confidence":0.55,"name":"exxperts"},{"slug":"lynkr","why":"For team deployments, Lynkr fronts any OpenAI-compatible backend — vLLM is the standard self-hosted server to put behind it when a laptop-class Ollama isn't enough. Indirect (generic OpenAI-compatible path, vLLM not named in the provider table).","dir":"in","confidence":0.55,"name":"Lynkr"},{"slug":"trl","why":"TRL's online RL trainers (GRPO) generate through vLLM — the use_vllm flag is the standard throughput path.","dir":"in","confidence":0.55,"name":"trl"}],"alternative":[{"slug":"ollama","why":"Both serve open models locally; vLLM optimizes for throughput, Ollama for one-command simplicity.","dir":"in","confidence":0.8,"name":"Ollama"},{"slug":"sie","why":"Both self-hosted OpenAI-compatible inference servers, opposite shapes: vLLM maximizes throughput for one big LLM; SIE serves breadth — 100+ heterogeneous task models (embedders, rerankers, OCR, NER, guards) loaded on demand across a cluster.","dir":"in","confidence":0.6,"name":"sie"},{"slug":"airllm","why":"Opposite ends of the local-inference spectrum: vLLM maximizes throughput given abundant VRAM (production serving); AirLLM minimizes VRAM given abundant patience (frontier-size models on consumer cards).","dir":"in","confidence":0.5,"name":"airllm"}],"built_with":[{"slug":"lmcache","why":"LMCache plugs into vLLM as its KV-connector; vLLM is the primary serving engine it accelerates.","dir":"in","confidence":0.95,"name":"LMCache"},{"slug":"llm-d","why":"llm-d is explicitly the orchestration layer above model servers: vLLM does the on-accelerator inference, llm-d adds cluster-level routing, KV-cache management, disaggregation and autoscaling.","dir":"in","confidence":0.9,"name":"llm-d"},{"slug":"verl","why":"vLLM is one of verl's first-class rollout engines — RL generation batches run through it during training.","dir":"in","confidence":0.8,"name":"verl"},{"slug":"unlimited-ocr","why":"First-party vLLM support is the blessed serving path — official recipe plus vllm/vllm-openai:unlimited-ocr Docker images for CUDA 12.9/13.0.","dir":"in","confidence":0.7,"name":"Unlimited-OCR"},{"slug":"openresearcher","why":"The whole serving path runs on vLLM — the 30B-A3B deploys via bundled vLLM server scripts and the trajectory generation leans on vLLM's native browser-tool support for GPT-OSS.","dir":"in","confidence":0.6,"name":"OpenResearcher"},{"slug":"olmocr","why":"olmocr runs its OCR vision-language model through a high-throughput inference backend — vLLM (or SGLang) — to batch-process PDFs at scale, so vLLM is the serving engine under olmocr's pipeline.","dir":"in","confidence":0.5,"name":"olmocr"}]},"url":"https://stackmap.shipwithai.xyz/repos/vllm-project/vllm"},"voicebox":{"name":"voicebox","owner":"jamiepine","slug":"voicebox","stars":46050,"image":"https://raw.githubusercontent.com/jamiepine/voicebox/main/.github/assets/icon-dark.webp","avatar":"https://avatars.githubusercontent.com/u/32987599?v=4&s=96","forks":5619,"language":"TypeScript","license":"MIT","updated":"3 days ago","topics":["voice"],"summary":"Local-first AI voice studio: clone voices, generate speech via 7 TTS engines in 23 languages, dictate system-wide with Whisper, and give any MCP-aware agent a cloned voice. Tauri, MLX/CUDA.","curator_note":"The whole voice I/O loop in one local app — ElevenLabs (output) plus WisprFlow (input) territory without the cloud. Zero-shot cloning, 50+ preset voices, paralinguistic tags via Chatterbox Turbo, a stories/timeline editor, and a global dictation hotkey with local-LLM cleanup. The agent hook is the standout: one `voicebox.speak` MCP tool call and Claude Code or Cursor talks in a voice you cloned, with per-agent voice binding and an always-visible on-screen pill so no agent speaks silently; a REST /speak covers non-MCP harnesses. NOT a library — it's a desktop product (Tauri): Linux is build-from-source, models are hefty downloads, and engine quality varies (only Turbo interprets [laugh]-style tags; others read them literally).","edges":[{"to":"luxtts","type":"built_with","why":"LuxTTS ships inside Voicebox as one of its seven selectable TTS engines — the lightweight English option (~1GB VRAM, 150x realtime on CPU).","confidence":0.8,"status":"approved"},{"to":"pocket-tts","type":"alternative","why":"Same say-it-locally job, different shape: pocket-tts is a pip-installable 100M CPU library for embedding speech in your own code; Voicebox is a full desktop studio with cloning, dictation, effects and MCP agent integration.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-14T23:06:03.512Z","linkCount":3,"related":{"complements":[],"alternative":[{"slug":"handy","why":"Both give you system-wide local Whisper dictation via a hotkey in a Tauri app. Voicebox bundles it into a full voice studio (cloning, 7 TTS engines, MCP agent voices); Handy is dictation only — lighter, simpler, and deliberately forkable.","dir":"in","confidence":0.8,"name":"Handy"},{"slug":"pocket-tts","why":"Same say-it-locally job, different shape: pocket-tts is a pip-installable 100M CPU library for embedding speech in your own code; Voicebox is a full desktop studio with cloning, dictation, effects and MCP agent integration.","dir":"out","confidence":0.5,"name":"pocket-tts"}],"built_with":[{"slug":"luxtts","why":"LuxTTS ships inside Voicebox as one of its seven selectable TTS engines — the lightweight English option (~1GB VRAM, 150x realtime on CPU).","dir":"out","confidence":0.8,"name":"LuxTTS"}]},"url":"https://stackmap.shipwithai.xyz/repos/jamiepine/voicebox"},"wrenai":{"name":"WrenAI","owner":"Canner","slug":"wrenai","stars":16578,"image":"https://raw.githubusercontent.com/Canner/wrenai/main/misc/wrenai_logo.png","avatar":"https://avatars.githubusercontent.com/u/7250217?v=4&s=96","forks":1867,"language":"Python","license":"NOASSERTION","updated":"yesterday","topics":["agents"],"summary":"Open-source GenBI engine: agents write governed SQL and deploy shareable dashboards over 22+ data sources, grounded in a Git-friendly context layer (MDL semantics, definitions, memory).","curator_note":"The strongest open answer to 'my agent writes confidently wrong SQL': business definitions, approved joins and past queries live in reviewable files, not prompts, and dry-plan validation catches errors before execution. Agent-driven by design (skills + CLI, works through Claude Code/Cursor). Skip for one-off charts from a CSV — the context layer is the point, and it takes real setup.","edges":[{"to":"ktx","type":"alternative","why":"Both are governed semantic layers that make agents trustworthy over company data: ktx curates warehouse metrics and serves read-only SQL context; Wren goes further into governed execution and agent-deployed dashboards.","confidence":0.65,"status":"approved"},{"to":"scout","type":"alternative","why":"Two shapes of 'company context for agents': Scout navigates unstructured sources (Slack/Drive/wiki) building its own wiki; Wren governs the structured-data side with semantic models and SQL.","confidence":0.5,"status":"approved"}],"status":"approved","added":"2026-07-19T12:39:42.253Z","linkCount":2,"related":{"complements":[],"alternative":[{"slug":"ktx","why":"Both are governed semantic layers that make agents trustworthy over company data: ktx curates warehouse metrics and serves read-only SQL context; Wren goes further into governed execution and agent-deployed dashboards.","dir":"out","confidence":0.65,"name":"ktx"},{"slug":"scout","why":"Two shapes of 'company context for agents': Scout navigates unstructured sources (Slack/Drive/wiki) building its own wiki; Wren governs the structured-data side with semantic models and SQL.","dir":"out","confidence":0.5,"name":"scout"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/Canner/wrenai"},"xberg":{"name":"xberg","owner":"xberg-io","slug":"xberg","stars":8688,"image":"https://raw.githubusercontent.com/xberg-io/xberg/main/docs/assets/demos/extract.gif","avatar":"https://avatars.githubusercontent.com/u/241328462?v=4&s=96","forks":525,"language":"Rust","license":"MIT","updated":"yesterday","topics":["ocr","rag","local"],"summary":"Rust-core document-intelligence engine with 15 language bindings: turns 96 formats — PDF, Office, images, audio, code — into clean text, tables and RAG-ready chunks. Library, CLI, REST or MCP.","curator_note":"The polyglot pick: reach for Xberg when your stack spans Rust, Python, Node, Go, the JVM or WASM and you want ONE extraction engine instead of a per-language pipeline. Handles 96 formats — PDFs, Office, images, audio/video (Whisper), even source code (306 languages, syntax-aware chunking) — with pluggable OCR (Tesseract/PaddleOCR/VLM), no GPU, and library/CLI/REST/MCP entry points. It's the v1 successor to Kreuzberg. Don't reach for it if you only need max-fidelity parsing of messy scanned PDFs (a dedicated VLM like MinerU or olmocr wins there), or if you're Python-only and already committed to LlamaIndex's own readers.","edges":[{"to":"mineru","type":"alternative","why":"Both turn documents (PDFs, Office) into LLM-ready text/JSON for RAG. MinerU is a heavyweight Python/VLM parser tuned for max-fidelity layout + OCR; Xberg is a lightweight polyglot engine spanning 96 formats and 15 language bindings with pluggable OCR. Pick MinerU for the hardest scanned/complex PDFs, Xberg for breadth and multi-language embedding.","confidence":0.8,"status":"approved"},{"to":"olmocr","type":"alternative","why":"Same end goal — clean, ordered Markdown from PDFs for RAG/training ingestion. olmocr is a dedicated self-hosted vision-language model that excels on scans, tables and handwriting; Xberg is a general extraction framework where OCR is one pluggable backend. Use olmocr when OCR quality is the bottleneck, Xberg when you need many formats and language bindings.","confidence":0.7,"status":"approved"},{"to":"chroma","type":"complements","why":"Natural RAG pairing: Xberg does the front half (parse → clean text → syntax-aware chunks → local or hosted embeddings) and Chroma stores and retrieves those vectors. Xberg feeds the index; Chroma serves the queries.","confidence":0.8,"status":"approved"},{"to":"llamaindex","type":"complements","why":"Xberg is a stronger document reader/extraction layer than LlamaIndex's built-in loaders; plug it in as the ingestion front-end (96 formats, OCR, transcription, chunking) and let LlamaIndex handle indexing, retrieval and query orchestration.","confidence":0.65,"status":"approved"},{"to":"unlimited-ocr","type":"alternative","why":"Engine vs model: xberg parses 96 formats deterministically; Unlimited-OCR throws a VLM at whole documents. Same slot in a RAG ingestion pipeline.","confidence":0.55,"status":"approved"}],"status":"approved","added":"2026-07-08T00:18:06.704Z","linkCount":6,"related":{"complements":[{"slug":"chroma","why":"Natural RAG pairing: Xberg does the front half (parse → clean text → syntax-aware chunks → local or hosted embeddings) and Chroma stores and retrieves those vectors. Xberg feeds the index; Chroma serves the queries.","dir":"out","confidence":0.8,"name":"Chroma"},{"slug":"llamaindex","why":"Xberg is a stronger document reader/extraction layer than LlamaIndex's built-in loaders; plug it in as the ingestion front-end (96 formats, OCR, transcription, chunking) and let LlamaIndex handle indexing, retrieval and query orchestration.","dir":"out","confidence":0.65,"name":"LlamaIndex"}],"alternative":[{"slug":"mineru","why":"Both turn documents (PDFs, Office) into LLM-ready text/JSON for RAG. MinerU is a heavyweight Python/VLM parser tuned for max-fidelity layout + OCR; Xberg is a lightweight polyglot engine spanning 96 formats and 15 language bindings with pluggable OCR. Pick MinerU for the hardest scanned/complex PDFs, Xberg for breadth and multi-language embedding.","dir":"out","confidence":0.8,"name":"MinerU"},{"slug":"olmocr","why":"Same end goal — clean, ordered Markdown from PDFs for RAG/training ingestion. olmocr is a dedicated self-hosted vision-language model that excels on scans, tables and handwriting; Xberg is a general extraction framework where OCR is one pluggable backend. Use olmocr when OCR quality is the bottleneck, Xberg when you need many formats and language bindings.","dir":"out","confidence":0.7,"name":"olmocr"},{"slug":"pdf-inspector","why":"Both are deterministic Rust document engines with multi-language bindings. xberg goes wide — 96 formats into RAG-ready chunks; pdf-inspector goes deep on one format with classification, confidence scores and OCR routing.","dir":"in","confidence":0.6,"name":"pdf-inspector"},{"slug":"unlimited-ocr","why":"Engine vs model: xberg parses 96 formats deterministically; Unlimited-OCR throws a VLM at whole documents. Same slot in a RAG ingestion pipeline.","dir":"out","confidence":0.55,"name":"Unlimited-OCR"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/xberg-io/xberg"},"zvec":{"name":"zvec","owner":"alibaba","slug":"zvec","stars":15234,"image":"https://zvec.oss-cn-hongkong.aliyuncs.com/logo/github_logo_1.svg","avatar":"https://avatars.githubusercontent.com/u/1961952?v=4&s=96","forks":961,"language":"C++","license":"Apache-2.0","updated":"3 days ago","topics":["rag"],"summary":"Alibaba's open-source in-process vector database: billion-scale similarity search embedded in your app, with DiskANN on-disk indexing, native full-text search and hybrid retrieval.","curator_note":"The 'SQLite of vector search' position, executed with Alibaba-scale engineering: runs inside your process (no server to operate), DiskANN keeps memory flat at large scale, and v0.5 added native FTS + hybrid retrieval so one MultiQuery spans dense, sparse, filters and text — no external search engine. Battle-tested internally before open-sourcing. NOT for multi-service architectures needing a shared, network-accessible store with auth and replication — in-process is the whole point and the whole limitation; C++ core with Python/Go/Rust SDKs, so debugging beneath the binding is not for everyone.","edges":[{"to":"chroma","type":"alternative","why":"Same job — the embedding store under a RAG pipeline. Chroma is the Python-native developer default; Zvec is the in-process C++ engine betting on raw speed, DiskANN memory economics and built-in hybrid retrieval.","confidence":0.85,"status":"approved"}],"status":"approved","added":"2026-07-13T10:36:42.259Z","linkCount":1,"related":{"complements":[],"alternative":[{"slug":"chroma","why":"Same job — the embedding store under a RAG pipeline. Chroma is the Python-native developer default; Zvec is the in-process C++ engine betting on raw speed, DiskANN memory economics and built-in hybrid retrieval.","dir":"out","confidence":0.85,"name":"Chroma"}],"built_with":[]},"url":"https://stackmap.shipwithai.xyz/repos/alibaba/zvec"}},"topics":{"agents":{"name":"Agents","slug":"agents","color":"#f5791e","blurb":"Frameworks for building autonomous and multi-agent LLM systems.","repoCount":55},"coding":{"name":"Coding Agents","slug":"coding","color":"#a78bfa","blurb":"Plugins, skill layers and runtimes for agentic coding assistants like Claude Code and Codex.","repoCount":54},"evals":{"name":"Evals & Testing","slug":"evals","color":"#ff5a8c","blurb":"Measure and monitor LLM/agent output quality.","repoCount":9},"finance":{"name":"Finance","slug":"finance","color":"#0ea5e9","blurb":"Trading, markets and financial analysis — agents with money on the line.","repoCount":4},"gateway":{"name":"Gateway","slug":"gateway","color":"#0d9488","blurb":"LLM gateways and routers — a proxy between clients or agents and model providers for fallback, multi-provider access, cost control and token savings.","repoCount":8},"local":{"name":"Local / Inference","slug":"local","color":"#46d282","blurb":"Run and serve open models on your own hardware.","repoCount":26},"memory":{"name":"Memory","slug":"memory","color":"#4dd0e1","blurb":"Persistent memory for AI systems — from vector recall to compounding agent lessons.","repoCount":19},"ocr":{"name":"OCR","slug":"ocr","color":"#e07a5f","blurb":"Turn documents, scans and PDFs into LLM-ready text — OCR and document-parsing models and pipelines.","repoCount":9},"orchestration":{"name":"Orchestration","slug":"orchestration","color":"#ffb450","blurb":"Compose and control multi-step LLM pipelines.","repoCount":23},"rag":{"name":"RAG & Retrieval","slug":"rag","color":"#789bff","blurb":"Connect LLMs to your data — ingestion, indexing, retrieval.","repoCount":25},"security":{"name":"Security","slug":"security","color":"#64748b","blurb":"AI for offensive and defensive security — pentest agents, vuln scanning, secure-code review.","repoCount":8},"skills":{"name":"Skills","slug":"skills","color":"#ffd43b","blurb":"Portable skill packs for coding agents — authoring, packaging, distribution and management.","repoCount":31},"storage":{"name":"Storage","slug":"storage","color":"#b45309","blurb":"Self-hosted drives and file backends — storage apps and infrastructure your agents and apps can use.","repoCount":4},"training":{"name":"Training","slug":"training","color":"#ef4444","blurb":"Post-training and RL infrastructure for open models — fine-tuning, alignment, reward loops.","repoCount":12},"vision":{"name":"Vision","slug":"vision","color":"#4f46e5","blurb":"Computer vision models and pipelines — recognition, detection and image understanding beyond documents.","repoCount":8},"voice":{"name":"Voice","slug":"voice","color":"#d946ef","blurb":"Speech in and out — text-to-speech, voice cloning and speech-to-text models and apps.","repoCount":6},"web":{"name":"Web","slug":"web","color":"#84cc16","blurb":"Browse, crawl and scrape the web — data acquisition for agents, RAG and pipelines.","repoCount":14}}}