EVALUATION INTELLIGENCE

AI benchmarks.
Configuration attached.

Benchmark scores are evidence, not universal truth. SXF stores who ran the evaluation, which version, which model configuration and which results are actually comparable.

12benchmark definitions
30independent observations
4vendor-reported
2026-10-04verified
INTERACTIVE EXPLORERCost vs score. Context vs score. Comparable groups only.Analyze the current evidence without collapsing benchmarks into one universal ranking.
Open Explorer ↗

COMPARABILITY RULE

Same name does not mean same measurement.

SXF only treats results as numerically comparable when benchmark identity, version, evaluator methodology and comparable group line up. Tool use, fallback behavior and reasoning effort remain part of the observation.

01

Independent first

Third-party evaluator results are separated from provider launch claims.

02

Version pinned

Live leaderboards can change methodology. Every snapshot stores its observed date and benchmark version.

03

Configuration pinned

Reasoning effort, tools and fallback behavior stay attached to the score.

04

No synthetic IQ

Different benchmarks are not collapsed into a homemade universal intelligence score.

BENCHMARK REGISTRY

What SXF tracks.

Open dataset ↗
independentactive

Artificial Analysis Intelligence Index 4.3.2

composite intelligence · score · higher-is-better

6 observations Open benchmark evidence ↗
independentactive

AutomationBench-AA

business automation · task-success · higher-is-better

3 observations Open benchmark evidence ↗
independentactive

Humanity's Last Exam

multidisciplinary reasoning · accuracy · higher-is-better

3 observations Open benchmark evidence ↗
vendor-reportedactive

AutomationBench

professional workflows · task-success · higher-is-better

1 observations Open benchmark evidence ↗
vendor-reportedactive

Terminal-Bench 4.0

agentic coding · task-success · higher-is-better

1 observations Open benchmark evidence ↗
vendor-reportedactive

Humanity's Last Exam

multidisciplinary reasoning · accuracy · higher-is-better

2 observations Open benchmark evidence ↗

MODEL COVERAGE

Models with evaluation evidence.

Model database ↗