StackMap
Subscribe
Explore / topics

Evals & Testing

Measure and monitor LLM/agent output quality.

agents-cli logo
5
google avatargoogle 6k · 9 days ago
agents-cli

CLI + agent-skills layer that turns your coding assistant into a Google Cloud agent-lifecycle expert: scaffold ADK projects, run and evaluate them, then deploy and publish to Gemini Enterprise.

codingevalsskills
→ pairs well with claude-reflect, skillkit
CopilotKit avatar
4
CopilotKit avatarCopilotKit 949 · 7 days ago
aimock

Mock everything an AI app talks to: 13 providers across 15 API surfaces plus MCP, A2A, AG-UI, vector DBs, search and rerank on one local port — with record-and-replay fixtures.

evals
→ pairs well with deepeval, litellm
AutoGen logo
9
microsoft avatarmicrosoft 61.2k · 5 months ago
AutoGen

Multi-agent conversation framework for building LLM applications with cooperating agents.

agentsevals
→ pairs well with Ragas, CubeSandbox
deepeval logo
8
confident-ai avatarconfident-ai 18.5k · 3 days ago
deepeval

Pytest for LLM apps: 40+ research-backed metrics — G-Eval, RAG suite, agent task completion, hallucination — as unit tests you run in CI, judged by any LLM including local ones.

evals
→ pairs well with deepteam, aimock
deepteam logo
3
confident-ai avatarconfident-ai 3k · 10 days ago
deepteam

Open-source red-teaming framework for LLM systems: 50+ vulnerabilities, jailbreak/injection/multi-turn attacks against agents, RAG pipelines and chatbots — plus guardrails. Runs locally.

securityevals
→ pairs well with deepeval, garak
future-agi logo
3
future-agi avatarfuture-agi 2.1k · 3 days ago
future-agi

Self-hostable platform for the whole agent-quality loop: tracing, evals, simulations, datasets, guardrails and an LLM gateway — one feedback loop from prototype to production. Apache 2.0.

evals
→ alternative to LangSmith, Tracely-ai
giskard-oss logo
4
Giskard-AI avatarGiskard-AI 5.8k · 3 days ago
giskard-oss

Giskard v3: modular Python evals and red-teaming for agentic systems — scenario-based checks with LLM-as-judge, plus an automatic vulnerability scanner across OWASP LLM Top-10 categories.

evalssecurity
→ alternative to deepteam, Ragas
cmu-l3 avatar
5
cmu-l3 avatarcmu-l3 288 · 5 days ago
gym-anything

CMU framework that turns real software — browsers, IDEs, EMRs, CAD — into standardized agent environments: start the app, hand the agent a task, score it with automatic verifiers.

evalsagents
→ pairs well with browser-use, labs-molt
L
5
langchain-ai 0 · 1 week ago
LangSmith

Trace, test and monitor LLM apps in production.

evals
→ pairs well with LangGraph, deepagents
opensre logo
3
Tracer-Cloud avatarTracer-Cloud 11.2k · 3 days ago
opensre

Open-source framework for AI SRE agents plus the RL training and evaluation environment they need — connect 60+ tools you already run and investigate incidents on your own infra.

agentsevals
→ pairs well with verl, superlog
OrcaReplay logo
5
Continuum-AI-Corp avatarContinuum-AI-Corp 269 · 3 days ago
OrcaReplay

Time travel for coding agents: a local proxy records any run (Claude Code, Codex, SDKs, even TLS-only harnesses), replays it byte-for-byte offline, and forks from any step onto a different model.

evalscoding
→ pairs well with aimock, agent-flow
Ragas logo
8
vibrantlabsai avatarvibrantlabsai 15.9k · 7 months ago
Ragas

Evaluation toolkit for your RAG and agent pipelines — faithfulness, relevance, and more.

evalsrag
→ pairs well with LlamaIndex, AutoGen
icedreamc avatar
6
icedreamc avataricedreamc 22 · 4 months ago
STALE

Benchmark plus memory pipeline for stale memories: STALE probes whether agents notice stored facts stopped being true; CUP-Mem adds conflict-aware writes, invalidation and premise verification.

evalsmemory
→ pairs well with metronix-memory, LightMem
supercov logo
4
supercorp-ai avatarsupercorp-ai 133 · 4 days ago
supercov

Code quality and coverage for coding agents: a Rust CLI that scores files with Jev, runs your existing test command, and turns uncovered paths into small actionable targets.

codingevals
→ pairs well with ripwire, open-code-review
Tracely-ai logo
4
Jwuthri avatarJwuthri 1.4k · 3 days ago
Tracely-ai

Trace-native CI/CD for agents: OTLP traces are graded on arrival, failures cluster into issues, and one click freezes a failing run into a hermetic regression case that blocks the PR.

evalsagents
→ pairs well with superlog, deepeval