Scale Labs

Research to Advance AI

Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

[SHOWDOWN]

Model-preference rankings from real-world usage.

1claude-opus-4-61071.40
1gpt-5.2-chat-latest1070.06
1claude-opus-4-7 (Thinking)1064.73
2claude-opus-4-71062.33
3gpt-5.5-2026-04-231052.38
View more

[PAPERS]

Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

Date Title
9/4/2026
READY or Not: Reliable Enterprise Agent DeploymentAgents, Enterprise
8/20/2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and Alignment
8/6/2026
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and Alignment
6/30/2026
DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?Agents, Enterprise, Evaluation and Alignment
6/29/2026
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsAgents, Evaluation and Alignment
6/19/2026
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld TasksAgents, Evaluation and Alignment
View more

[BLOG]

Insights, analysis, and updates from Scale Labs

AgentsSep 3, 2026

Introducing READY: What It Takes to Deploy an AI Agent

Today we're introducing READY (Reliable Enterprise Agent Deployment), a suite of industry-specific benchmarks that measure agents the way enterprises actually use them: working alongside people, inside real workflows.

AgentsSep 2, 2026

Who Grades the Graders? Rethinking Verifier Design for Computer-Use Agents

As Computer Use Agents (CUA) take on complex professional tasks like processing emails, creating spreadsheets, drawing 3D diagrams, and drafting financial memos, verifiers are at the heart of providing evaluation and training signals around their capabilities. Benchmarks like OSWorld 2.0 and Agents' Last Exam use these verifiers, in the form of programmatic checks, to measure whether an agent is capable of completing production-grade work.

AgentsAug 26, 2026

CliniCARE-Bench: Clinical AI Agents Can Be Right for the Wrong Reasons

We evaluated 16 agentic systems on 25 clinical care scenarios over 750 real patient cases. Every single system committed to an answer more often than the evidence allowed, and up to one in five correct verdicts rested on an investigation the case authors had explicitly prohibited.

Physical AIAug 13, 2026

Can Robots Learn from Watching Us?

Human demonstrations allow for diverse data collection in the real world, but pose a difficult learning problem due to the large embodiment gap between robots and humans. Being able to effectively learn from high volumes of human interaction data can unlock the future of robotics. Our research demonstrates that adding a limited amount of human data to robot datasets can nearly double performance on unseen environments while only sacrificing a nominal amount of in-distribution performance.

View allAll posts