StackMap
Subscribe
Explore / topics

Evals & Testing

Measure and monitor LLM/agent output quality.

agents-cli logo
5
google avatargoogle 5.7k · 2 days ago
agents-cli

CLI + agent-skills layer that turns your coding assistant into a Google Cloud agent-lifecycle expert: scaffold ADK projects, run and evaluate them, then deploy and publish to Gemini Enterprise.

codingevalsskills
pairs well with claude-reflect, skillkit
AutoGen logo
9
microsoft avatarmicrosoft 60.5k · 4 months ago
AutoGen

Multi-agent conversation framework for building LLM applications with cooperating agents.

agentsevals
pairs well with Ragas, CubeSandbox
deepeval logo
4
confident-ai avatarconfident-ai 17.7k · today
deepeval

Pytest for LLM apps: 40+ research-backed metrics — G-Eval, RAG suite, agent task completion, hallucination — as unit tests you run in CI, judged by any LLM including local ones.

evals
pairs well with deepteam, Ragas
deepteam logo
3
confident-ai avatarconfident-ai 2.6k · 8 days ago
deepteam

Open-source red-teaming framework for LLM systems: 50+ vulnerabilities, jailbreak/injection/multi-turn attacks against agents, RAG pipelines and chatbots — plus guardrails. Runs locally.

securityevals
pairs well with deepeval, garak
future-agi logo
2
future-agi avatarfuture-agi 1.7k · today
future-agi

Self-hostable platform for the whole agent-quality loop: tracing, evals, simulations, datasets, guardrails and an LLM gateway — one feedback loop from prototype to production. Apache 2.0.

evals
alternative to LangSmith, deepeval
giskard-oss logo
4
Giskard-AI avatarGiskard-AI 5.8k · yesterday
giskard-oss

Giskard v3: modular Python evals and red-teaming for agentic systems — scenario-based checks with LLM-as-judge, plus an automatic vulnerability scanner across OWASP LLM Top-10 categories.

evalssecurity
alternative to deepteam, Ragas
2
cmu-l3 avatarcmu-l3 276 · 3 days ago
gym-anything

CMU framework that turns real software — browsers, IDEs, EMRs, CAD — into standardized agent environments: start the app, hand the agent a task, score it with automatic verifiers.

evalsagents
pairs well with browser-use, OpenSandbox
3
langchain-ai 0 · 1 week ago
LangSmith

Trace, test and monitor LLM apps in production.

evals
pairs well with LangGraph, deepagents
opensre logo
2
Tracer-Cloud avatarTracer-Cloud 10.7k · today
opensre

Open-source framework for AI SRE agents plus the RL training and evaluation environment they need — connect 60+ tools you already run and investigate incidents on your own infra.

agentsevals
pairs well with verl, allama
Ragas logo
7
explodinggradients avatarexplodinggradients 15.4k · 5 months ago
Ragas

Evaluation toolkit for your RAG and agent pipelines — faithfulness, relevance, and more.

evalsrag
pairs well with LlamaIndex, AutoGen