cua vs gym-anything
Computer-use infrastructure for agents: sandboxed Linux, macOS and Windows desktops locally or in the cloud, Cua Driver for native apps (CLI/MCP/SDK), Lume VMs, CUA-S1 models, Cua Bench. — versus — CMU framework that turns real software — browsers, IDEs, EMRs, CAD — into standardized agent environments: start the app, hand the agent a task, score it with automatic verifiers.
Both turn real software into graded agent tasks; gym-anything standardizes environments behind automatic verifiers, Cua Bench creates tasks on Cua desktops and exports trajectories for training.
| cua | gym-anything | |
|---|---|---|
| Stars | 28k | 288 |
| Forks | 1.9k | 40 |
| Language | Rust | Shell |
| License | MIT | MIT |
| Last activity | today | 5 days ago |
| Topics | agents, sandboxes | evals, agents |
| Curated connections | 4 | 5 |
cua — the curator's take
The broadest open stack for computer-use agents: one sandbox API that runs a desktop locally or on their cloud, Cua Driver that operates native apps on macOS, Windows and Linux in the background without taking your pointer (CLI, MCP or typed SDKs), Lume for macOS and Linux VMs on Apple Silicon, Cua Bench for tasks and trajectories, and CUA-S1, small decision models for forms. Use it when your agent has to touch software that has no API. The breadth is the cost: many components, a cloud product (Fleets) promoted throughout, and the installer signs you in. If your agent only needs the web, browser-use is far less; if it only needs to run code, a plain sandbox is enough.
gym-anything — the curator's take
The missing middle layer for computer-use agents: a Core runtime (environment lifecycle, actions, observations, verifiers), a benchmark collection wrapping real applications like Moodle, and reference agents (Claude, Gemini, Qwen, Kimi) — three parts connected by contracts, each independently replaceable, one CLI to run it all. The `doctor` setup command and environment caching signal real operational care. Use it to evaluate your CUA agent beyond browser-only benchmarks, or to gym-ify internal software by adding a task folder with a setup script and a checker. NOT an agent framework — the agents are references, bring your own. Young (arXiv 2026, CMU L3): environment coverage is the current bottleneck.