BENCHMARK EVIDENCE / UNDER-REVIEW
SciCode
scientific coding. Metric: score. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.
3observations
Artificial Analysisevaluator
under-reviewstatus
2026-10-04verified
METHOD
Read the score with its configuration.
Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.
- Artificial Analysis states this evaluation is under review after an independent audit flagged possible dataset errors.
- SXF displays the observation with a warning and does not use SciCode alone to declare a model superior.
independent
3 observationsartificial-analysis-v4.3.2
| Model | Score | Reasoning | Tools | Source |
|---|---|---|---|---|
| Claude Opus 5.5 | 67.0% | max | benchmark-defined | Evidence ↗ |
| GPT-6 Astra | 55.0% | high | benchmark-defined | Evidence ↗ |
| GPT-6 Sol | 55.0% | high | benchmark-defined | Evidence ↗ |