BENCHMARK EVIDENCE / ACTIVE

AutomationBench

professional workflows. Metric: task-success. Direction: higher-is-better. Results below stay separated by evaluator evidence type and comparable group.

1observations
OpenAIevaluator
activestatus
2026-10-04verified

METHOD

Read the score with its configuration.

Scores from different comparable groups are intentionally not merged into one ranking. Reasoning effort, tools, fallback behavior, evaluator and benchmark version stay attached to every observation.

vendor-reported

openai-astra-launch-table

1 observation
ModelScoreReasoningToolsSource
GPT-6 Astra 41.4% not published provider-defined Evidence ↗