EVALUATION EXPLORER V1

Benchmark evidence.
Decision views, not one score.

Filter by benchmark, evaluator, provider and comparable group. Switch between ranked evidence, cost-vs-score and context-vs-score views without mixing incompatible runs.

12benchmarks
34observations
6models evaluated
2026-10-04verified

INTERACTIVE ANALYSIS

Choose one measurement frame.

The explorer only ranks observations inside the selected benchmark and comparable group. Missing coverage stays missing; under-review benchmarks carry a visible warning.

Raw evidence ↗
ModelScoreContextWorkload costReasoningEvidence
Claude Opus 5.5Anthropic 58 1,000,000tokens $0.6100K input + 10K output max Evidence ↗
Claude Fable 5.1Anthropic 53 1,000,000tokens $1.5100K input + 10K output max Evidence ↗
GPT-6 AstraOpenAI 51 1,050,000tokens $1.5100K input + 10K output high Evidence ↗
Grok 4.7xAI 46 500,000tokens $0.26100K input + 10K output xhigh Evidence ↗
GPT-6 SolOpenAI 42 1,050,000tokens $0.3100K input + 10K output high Evidence ↗
Gemini 3.8 FlashGoogle 41 1,048,576tokens $0.1125100K input + 10K output high Evidence ↗

READING THE VIEWS

Pareto means efficient, not universally best.

A model is Pareto-efficient when no other visible model is both better on the selected benchmark and better on the second axis. Cost uses the SXF Standard workload assumption of 100K uncached input + 10K output tokens; context uses the provider-published context window.

01

Table

Ranks only the selected comparable group by the benchmark's declared metric direction.

02

Cost vs score

Highlights models that offer a non-dominated tradeoff between direct token cost and benchmark score.

03

Context vs score

Shows whether a larger context window trades off against the selected evaluation result.

04

No speed view yet

Latency and throughput stay out until SXF has a methodology-pinned speed dataset with matching model configurations.