StackMap
Subscribe
Explore / deepeval
confident-ai

deepeval

Pytest for LLM apps: 40+ research-backed metrics — G-Eval, RAG suite, agent task completion, hallucination — as unit tests you run in CI, judged by any LLM including local ones.

16,962 1,681 Python Apache-2.0updated today
Curator's take

The most complete general-purpose eval framework: pytest ergonomics ('assert_test' in CI), metrics with explanations grounded in the research they implement, component-level tracing for agents, plus benchmark harnesses (MMLU, HumanEval) when you need them. If you're evaluating LLM apps in Python and don't know where to start, start here. NOT vendor-neutral: the free framework feeds Confident AI's platform for reports and regression tracking — usable without it, but the gravity is real. And LLM-as-judge metrics inherit judge bias: treat scores as regression signals, not absolute truth. RAG-only shops may prefer Ragas' tighter focus.

Mapped by ShipWithAI editors · links verified
README.md

DeepEval.

The LLM Evaluation Framework

confident-ai%2Fdeepeval | Trendshift

discord-invite reddit-community

Documentation | Metrics and Features | Getting Started | Integrations | Confident AI

GitHub release Try Quickstart in Colab License Twitter Follow

Deutsch | Español | français | 日本語 | 한국어 | Português | Русский | 中文

DeepEval is a simple-to-use, open-source LLM evaluation framework, for evaluating large-language model systems. It is similar to Pytest but specialized for unit testing LLM apps. DeepEval incorporates the latest research to run evals via metrics such as G-Eval, task completion, answer relevancy, hallucination, etc., which uses LLM-as-a-judge and other NLP models that run locally on your machine.

Whether you're building AI agents, RAG pipelines, or chatbots, implemented via LangChain or OpenAI, DeepEval has you covered. With it, you can easily determine the optimal models, prompts, and architecture to improve your AI quality, prevent prompt drifting, or even transition from OpenAI to Claude with confidence.

[!IMPORTANT] Need a place for your DeepEval testing data to live 🏡❤️? Sign up to Confident AI to compare iterations of your LLM app, generate & share testing reports, and more.

Demo GIF

Wa

Continue your stack

What teams reach for next — and why each earns a place beside deepeval. Ranked by curator confidence.