Research to Advance AI
Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.
[LEADERBOARDS]
Benchmarks for frontier, agentic, and safety capabilities
[SHOWDOWN]
Model-preference rankings from real-world usage.
[PAPERS]
Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.






READY or Not: Reliable Enterprise Agent Deployment
[BLOG]
Insights, analysis, and updates from Scale Labs
Introducing READY: What It Takes to Deploy an AI Agent
Today we're introducing READY (Reliable Enterprise Agent Deployment), a suite of industry-specific benchmarks that measure agents the way enterprises actually use them: working alongside people, inside real workflows.
Who Grades the Graders? Rethinking Verifier Design for Computer-Use Agents
As Computer Use Agents (CUA) take on complex professional tasks like processing emails, creating spreadsheets, drawing 3D diagrams, and drafting financial memos, verifiers are at the heart of providing evaluation and training signals around their capabilities. Benchmarks like OSWorld 2.0 and Agents' Last Exam use these verifiers, in the form of programmatic checks, to measure whether an agent is capable of completing production-grade work.
CliniCARE-Bench: Clinical AI Agents Can Be Right for the Wrong Reasons
We evaluated 16 agentic systems on 25 clinical care scenarios over 750 real patient cases. Every single system committed to an answer more often than the evidence allowed, and up to one in five correct verdicts rested on an investigation the case authors had explicitly prohibited.
Can Robots Learn from Watching Us?
Human demonstrations allow for diverse data collection in the real world, but pose a difficult learning problem due to the large embodiment gap between robots and humans. Being able to effectively learn from high volumes of human interaction data can unlock the future of robotics. Our research demonstrates that adding a limited amount of human data to robot datasets can nearly double performance on unseen environments while only sacrificing a nominal amount of in-distribution performance.
View allAll posts