BENCHMARK EVIDENCE / ACTIVE

Humanity's Last Exam

multidisciplinary reasoning. Metric: accuracy. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.

2observations
Anthropicevaluator
activestatus
2026-10-04verified

METHOD

Read the score with its configuration.

Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.

vendor-reported

anthropic-fable-5-1-launch-table

1 observation
ModelScoreReasoningToolsSource
Claude Fable 5.1 60.9% not published no tools Evidence ↗
vendor-reported

anthropic-fable-5-1-launch-table-tools

1 observation
ModelScoreReasoningToolsSource
Claude Fable 5.1 65.0% not published with tools Evidence ↗