BENCHMARK EVIDENCE / ACTIVE

GDPval-AA
v2.1

professional knowledge work. Metric: elo. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.

3observations
Artificial Analysisevaluator
activestatus
2026-10-04verified

METHOD

Read the score with its configuration.

Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.

independent

artificial-analysis-v4.3.2

3 observations
ModelScoreReasoningToolsSource
Claude Opus 5.5 1,867 Elo max benchmark-defined Evidence ↗
GPT-6 Astra 1,485 Elo high benchmark-defined Evidence ↗
GPT-6 Sol 1,396 Elo high benchmark-defined Evidence ↗