StackMap
Subscribe

cua vs gym-anything

Computer-use infrastructure for agents: sandboxed Linux, macOS and Windows desktops locally or in the cloud, Cua Driver for native apps (CLI/MCP/SDK), Lume VMs, CUA-S1 models, Cua Bench. — versus — CMU framework that turns real software — browsers, IDEs, EMRs, CAD — into standardized agent environments: start the app, hand the agent a task, score it with automatic verifiers.

The curated verdict

Both turn real software into graded agent tasks; gym-anything standardizes environments behind automatic verifiers, Cua Bench creates tasks on Cua desktops and exports trajectories for training.

cuagym-anything
Stars28k288
Forks1.9k40
LanguageRustShell
LicenseMITMIT
Last activitytoday5 days ago
Topicsagents, sandboxesevals, agents
Curated connections45

cua — the curator's take

The broadest open stack for computer-use agents: one sandbox API that runs a desktop locally or on their cloud, Cua Driver that operates native apps on macOS, Windows and Linux in the background without taking your pointer (CLI, MCP or typed SDKs), Lume for macOS and Linux VMs on Apple Silicon, Cua Bench for tasks and trajectories, and CUA-S1, small decision models for forms. Use it when your agent has to touch software that has no API. The breadth is the cost: many components, a cloud product (Fleets) promoted throughout, and the installer signs you in. If your agent only needs the web, browser-use is far less; if it only needs to run code, a plain sandbox is enough.

gym-anything — the curator's take

The missing middle layer for computer-use agents: a Core runtime (environment lifecycle, actions, observations, verifiers), a benchmark collection wrapping real applications like Moodle, and reference agents (Claude, Gemini, Qwen, Kimi) — three parts connected by contracts, each independently replaceable, one CLI to run it all. The `doctor` setup command and environment caching signal real operational care. Use it to evaluate your CUA agent beyond browser-only benchmarks, or to gym-ify internal software by adding a task folder with a setup script and a checker. NOT an agent framework — the agents are references, bring your own. Young (arXiv 2026, CMU L3): environment coverage is the current bottleneck.