Scale Labs

Humanity's Sixth Sense

What HSS Measures

Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels.

Humanity's Sixth Sense (HSS), in partnership with Elorian, is a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance.

Release Artifacts

1. HSS benchmark dataset via huggingface: including 522 human crafted samples spans four main domains and eleven subdomains.

2. Evaluation harness: evaluation code repo including model registry, evaluation prompts, grading prompts for reproduce and new models' evaluation.

All confirmatory metrics are evaluated on a frozen, versioned release containing 522 tasks.

Key Takeaways

  • The top-performing model at launch, GPT-6-astra (maximum reasoning effort), achieves only a 53.6% pass rate, compared to 93.1% for human annotators. Most models perform below 40%, with a median pass rate of 30.9%.
  • Across all 25 models, the average reasoning token usage is 4,000 per task: even on tasks intuitive to humans, models generate substantial reasoning chains before answering, and frequently overthink without arriving at the correct answer.
  • Video tasks prove more challenging than static images for 23 of the 25 models, with an average drop of 7.3 percentage points.
  • Social understanding is the weakest domain across vendors; it is the lowest-performing domain for 21 of the 25 models, averaging 24.4% accuracy compared to 34.1% across the remaining three domains.

Key Stats

  • 522 tasks across 4 domains and 11 subdomains
  • 288 image-based and 234 video-based tasks; 17.6 hours of video in total (median 76 s, longest 28 min)
  • 723 rubric criteria
  • 3,466 tasks authored, 522 admitted after three independent review rounds: a 15.1% acceptance rate
  • 25 models from eight vendors on the leaderboard
  • Human baseline from 20 annotators answering under the same free-form conditions as the models
  • Scoring is pass@1 averaged over three attempts per task, with 95% bootstrap intervals over tasks

How to Read the Leaderboard

Pass@1 is the primary metric. Pass@1 is the fraction of attempts that are correct, averaged over tasks. Each model is attempted three times per task; the score is an average over sampled attempts rather than a single deterministic run. A response that never commits to an answer, including one that spends its entire output budget on reasoning is scored as failing every criterion.

Rubric score is retained as a diagnostic: it awards partial credit and so separates a systematic near-miss from a total miss. The two coincide on the 71.5% of tasks that carry a single criterion.

The ± value is the half-width of a 95% cluster bootstrap interval that resamples tasks and keeps their attempts together, so the interval reflects which tasks the benchmark happens to contain rather than rerun noise.

Measuring Settings

All models are evaluated with high reasoning effort and the maximum allowed output length. We keep temperature and top-p at their per-model defaults.

Scoring and Judging

Model outputs are graded by Claude-Opus-5 as an automated LLM judge. The judge sees the question, the reference answer, the rubric, and the candidate answer, and returns a per-criterion verdict with a one-sentence justification.

Limitations

Visual channel only. Video items are supplied without their audio track and without a transcript, so items whose resolution would depend on speech or off-screen sound are outside the benchmark's scope by construction.

Judge dependence. Scores are mediated by an LLM judge, which is known to carry biases including self-preference; cross-judge agreement is reported to bound the effect.

Contamination. Media sourced from the public web may appear in training data. The questions are newly authored against that media rather than collected with it, but prior exposure to an individual image or clip cannot be fully excluded. Theory of mind results are read as measurements of agreement with annotator judgement, not as ground truth about the people depicted.

Resources:

Performance Comparison

1

GPT-6-Astra (max)

53.60±4.10

2

GPT-6.1-Sol (max)

46.60±3.80

3

Claude-Opus-5.5 (xhigh)

44.60±3.80

4

Gemini-3.8-Flash (high)

41.60±3.70

5

Claude-Fable-5.1

40.80±3.60

6

Gemini-3.7-Flash (high)

39.80±3.70

7

Muse Spark 1.3 (max)

37.40±3.50

8

Claude-Fable-5 (max)

34.50±3.70

9

Qwen3.8-Max (max)

33.00±3.60

10

Gemini-3.5-Flash (high)

32.60±3.40

11

Gemini-3.6-Flash (high)

31.90±3.40

12

GPT-6-Sol (max)

31.20±3.40

13

Qwen3.8-Flash (max)

30.90±3.30

14

Claude-Opus-5 (max)

30.50±3.50

15

GPT-5.6-Sol (max)

30.00±3.30

16

GLM-5.3-Flash (max)

26.80±3.00

17

Kimi-K3 (max)

25.50±3.10

18

MiMo-v2.6-Flash (high)

25.40±2.90

19

MiMo-v2.6-Pro (high)

25.20±2.90

20

Claude-Sonnet-5 (max)

24.80±3.20

21

Muse Spark 1.2 (xhigh)

24.50±3.10

22

Claude-Opus-4.8 (max)

23.50±3.20

23

GPT-5.6-Terra (max)

22.70±3.20

24

GPT-5.6-Luna (max)

21.60±3.20

25

GPT-6-Luna (max)

21.00±3.00