BENCHMARK EVIDENCE / ACTIVE
Terminal-Bench
v4.0
agentic coding. Metric: task-success. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.
1observations
Anthropicevaluator
activestatus
2026-10-04verified
METHOD
Read the score with its configuration.
Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.
- Provider-reported launch evaluation; configuration differs from independent evaluator runs and remains a separate comparable group.
vendor-reported
1 observationanthropic-fable-5-1-launch-table
| Model | Score | Reasoning | Tools | Source |
|---|---|---|---|---|
| Claude Fable 5.1 | 55.8% | not published | provider-defined | Evidence ↗ |