Benchmarks for agents in enterprise workflows.
AI Agents today are measured on capability. READY Benchmarks are the first to measure deployment readiness: how an agent performs in enterprise environments, at what level of human oversight, and at what cost.
- 2
- Seed benchmarks
- 5,824
- Clinical cases
- 15+
- Models on a live leaderboard
From workflow to deployment profile
Showing how an agent delivers reliable work, not just measuring capability.
Inputs
What you hand the benchmark
- A real enterprise workflowthe job you want an agent to own
- Task instances + environmentreal records, tools, and data to act on
- Success criteria + reliability targetwith cost and risk limits
Testbed
What the benchmark does with it
- Each agent works every caseend-to-end, on its own
- Add a human-oversight policyaccept · escalate · clarify · correct · take over
- Optimize the cheapest policythat still hits the reliability target
- Freeze it and qualify on held-out cases
Deployment profile
What you get back
- Status✓ Qualified
- Reliability76%
- Autonomous coverage78%
- Human-oversight burden22%
- Operating cost≈ $5.50 / case
The cost of reliable AI
Total cost of task completion = model tokens + human review to hit your reliability target
Two agents with the same task accuracy can need very different amounts of human review. Explore the trade-off between the reliability you need and what it costs to get there.
READY Benchmarks
The first two benchmarks below cover healthcare workflows. The READY framework is industry-agnostic, with more benchmarks to come soon.
CliniCARE-Bench
A deployment-oriented benchmark for selective autonomy in retrospective clinical audit, end-to-end, auditable agent reasoning over real MIMIC-IV records.
PSEBench
A controllable, verifiable benchmark for policy-grounded patient-safety event triage against Minnesota's 29 Reportable Adverse Health Events (MN29).
Open for contributions
The framework is designed for extension: a new benchmark is a single structured record. We are looking for domain experts to author realistic enterprise tasks and researchers to build evaluation environments.