StackMap
Subscribe
Explore / mini-swe-agent
SWE-agent

mini-swe-agent

The 100-line agent from the SWE-bench team: >74% on SWE-bench Verified with no tools but bash, no config sprawl — the reference minimal harness, adopted by Meta, NVIDIA and Ramp.

6,933 964 Python MITupdated 3 days ago
View on GitHubDispute this mapping →
Curator's take

The existence proof that most harness complexity is optional: the team that built SWE-bench and SWE-agent asked what a 100x simpler agent loses — the answer is almost nothing (>74% Verified), which is why it became the standard baseline harness for benchmarking models (Ramp's SWE-bench, DeepSWE — where it beats Claude Code and Codex as a harness). Read it to understand agents; use it to evaluate models fairly. NOT a daily driver: no MCP, no skills, no IDE plumbing — by design. If you're choosing a tool to ship features with, this is the control group, not the product.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside mini-swe-agent. Ranked by curator confidence.

alternativealternativealternativealternativealternativealternativebuilt withtauOpenManusopeninterpreterprime-agentjcodemakalitellmmini-swe-agent
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md2 min read
mini-swe-agent banner

The minimal AI software engineering agent

📣 mini-swe-agent now powers Ramp SWE-Bench
📣 mini-swe-agent beats Claude Code and Codex on DeepSWE
📣 Run mini-swe-agent on our new & extremely challenging benchmark, ProgramBench
📣 New tutorial on building minimal AI agents

Docs Slack PyPI - Version

[!WARNING] This is mini-swe-agent v2. Read the migration guide. For the previous version, check out the v1 branch.

In 2024, we built SWE-bench & SWE-agent and helped kickstart the coding agent revolution.

We now ask: What if our agent was 100x simpler, and still worked nearly as well?

mini is

  • Widely adopted: Used by Meta, NVIDIA, Essential AI, IBM, Nebius, Anyscale, Princeton University, Stanford University, and many more.
  • Minimal: Just some 100 lines of python for the agent class (and a bit more for the environment, model, and run script) — no fancy dependencies!
  • Performant: Scores >74% on the SWE-bench verified benchmark; starts much faster than Claude Code
  • Deployable: Supports local environments, docker/podman, singularity/apptainer, bublewrap, contree, and more
  • Compatible: Supports all models via litellm, openrouter, portkey, and more. Support for /completion and /response endpoints, interleaved thinking etc.
  • Built by the Princeton & Stanford team behind SWE-bench, SWE-agent, and more
  • Tested: Codecov
More motivation (for research)

SWE-agent jump-started the development of AI agents in 2024. Back then, we placed a lot of emphasis on tools and special interfaces for the agent. However, one year later, as LMs have become more capable, a lot of this is not needed at all to build a useful agent! In fact, the mini agent

  • Does not have any tools other than bash — it doesn't even need to use the tool-calling interface of the LMs. This means that you can run it with literally any model. When running in sandboxed environments you also don't need to take care of installing a single package — all it needs is bash.
  • Has a completely linear history — every step of the agent just appends to the messages and that's it. So there's no difference between the trajectory and the messages that you pa