BENCHMARK EVIDENCE / ACTIVE
AA-Briefcase
v1.1
agentic knowledge work. Metric: elo. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.
3observations
Artificial Analysisevaluator
activestatus
2026-10-04verified
METHOD
Read the score with its configuration.
Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.
- Agentic knowledge-work evaluation using realistic business workflows and deliverables.
independent
3 observationsartificial-analysis-v4.3.2
| Model | Score | Reasoning | Tools | Source |
|---|---|---|---|---|
| Claude Opus 5.5 | 1,808 Elo | max | benchmark-defined | Evidence ↗ |
| GPT-6 Astra | 1,507 Elo | high | benchmark-defined | Evidence ↗ |
| GPT-6 Sol | 1,268 Elo | high | benchmark-defined | Evidence ↗ |