BENCHMARK EVIDENCE / ACTIVE
Artificial Analysis Intelligence Index
v4.3.2
composite intelligence. Metric: score. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.
6observations
Artificial Analysisevaluator
activestatus
2026-10-04verified
METHOD
Read the score with its configuration.
Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.
- Composite index spanning reasoning, coding, science, knowledge work, long context and agentic evaluations.
- Scores are configuration-specific; reasoning effort and fallback behavior must stay attached to each observation.
independent
6 observationsartificial-analysis-v4.3.2
| Model | Score | Reasoning | Tools | Source |
|---|---|---|---|---|
| Claude Opus 5.5 | 58 | max | benchmark-defined | Evidence ↗ |
| Claude Fable 5.1 | 53 | max | benchmark-defined | Evidence ↗ |
| GPT-6 Astra | 51 | high | benchmark-defined | Evidence ↗ |
| Grok 4.7 | 46 | xhigh | benchmark-defined | Evidence ↗ |
| GPT-6 Sol | 42 | high | benchmark-defined | Evidence ↗ |
| Gemini 3.8 Flash | 41 | high | benchmark-defined | Evidence ↗ |