Table
Ranks only the selected comparable group by the benchmark's declared metric direction.
EVALUATION EXPLORER V1
Filter by benchmark, evaluator, provider and comparable group. Switch between ranked evidence, cost-vs-score and context-vs-score views without mixing incompatible runs.
INTERACTIVE ANALYSIS
The explorer only ranks observations inside the selected benchmark and comparable group. Missing coverage stays missing; under-review benchmarks carry a visible warning.
| Model | Score | Context | Workload cost | Reasoning | Evidence |
|---|---|---|---|---|---|
| Claude Opus 5.5Anthropic | 58 | 1,000,000tokens | $0.6100K input + 10K output | max | Evidence ↗ |
| Claude Fable 5.1Anthropic | 53 | 1,000,000tokens | $1.5100K input + 10K output | max | Evidence ↗ |
| GPT-6 AstraOpenAI | 51 | 1,050,000tokens | $1.5100K input + 10K output | high | Evidence ↗ |
| Grok 4.7xAI | 46 | 500,000tokens | $0.26100K input + 10K output | xhigh | Evidence ↗ |
| GPT-6 SolOpenAI | 42 | 1,050,000tokens | $0.3100K input + 10K output | high | Evidence ↗ |
| Gemini 3.8 FlashGoogle | 41 | 1,048,576tokens | $0.1125100K input + 10K output | high | Evidence ↗ |
READING THE VIEWS
A model is Pareto-efficient when no other visible model is both better on the selected benchmark and better on the second axis. Cost uses the SXF Standard workload assumption of 100K uncached input + 10K output tokens; context uses the provider-published context window.
Ranks only the selected comparable group by the benchmark's declared metric direction.
Highlights models that offer a non-dominated tradeoff between direct token cost and benchmark score.
Shows whether a larger context window trades off against the selected evaluation result.
Latency and throughput stay out until SXF has a methodology-pinned speed dataset with matching model configurations.