Scale Labs
BACK
Safety & Oversight10/9/2026

DistressBench: Evaluating Multi-Dimensional Crisis Support in Large Language Models

Drew Rein, Patrick Oathout, Vishal Kumar, Udari Madhushani Sehwag

People increasingly turn to chatbots for emotional support, and some of those conversations involve suicide or self-harm.

DistressBench measures whether models give the support those conversations need, using rubrics written by licensed clinicians: 718 clinician-authored conversations across 23 suicide and self-harm subcategories.

Most existing safety benchmarks score these conversations on whether the model refused, which says little about whether the user actually got help. DistressBench contains 718 clinician-authored conversations across 23 suicide and self-harm subcategories. Each one comes with a rubric of weighted criteria (6,901 in total) covering seven dimensions of crisis support. Across 25 frontier models, the best model meets 88.4% of weighted criteria and the median model meets 71.4%, compared with 98.7% for clinician reference responses. The biggest gap is between recognizing a crisis and acting on it. In 35.3% of conversations where a model fully recognized the crisis, it still didn’t point the user toward human help. Models score near the top on avoiding harmful content (94.5% median), but the median model reaches only 46.0% on de-escalation while the strongest reaches 74.6%. That spread suggests the gap mostly comes from product choices, since at least one deployed model already does much better.