CliniCARE-Bench
A deployment-oriented benchmark for selective autonomy in retrospective clinical audit, end-to-end, auditable agent reasoning over real MIMIC-IV records.
- Verdict space
- Yes
- No
- Indeterminate: Lack of Data
- Indeterminate: Medically Ambiguous
- Scenarios
- 25
- Tasks
- 750
- Substrate
- MIMIC-IV v3.1 · Note · ED
What CliniCARE-Bench measures
Foundation models now match or exceed clinicians on medical-knowledge exams, yet exam performance is a poor proxy for what clinical deployment demands. A nurse at change of shift, a quality officer auditing sepsis-bundle compliance, or a nephrologist assessing a returning patient does not answer a self-contained multiple-choice question, each conducts a small research project against a heterogeneous, longitudinal, uncurated electronic health record.
CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR) is a deployment-oriented benchmark for selective autonomy in retrospective clinical audit. To our knowledge it is the first clinical agent environment to support end-to-end auditing of an agent's reasoning process, retrieval, evidence grounding, policy use, and the decision trace that links them, scored end to end rather than on the final answer alone.
Paper
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Chatrath, Zhu, Pu, Shanker, Ursekar, Sharma, Fan, Han, Tiwari, Dan, Qin, Yin, Wang, Kalmath, Agarwal, Li, Doctor, Zhang, Xue (Scale AI · Emory · Vanderbilt · UC Santa Cruz)
Scale AI Research