GPT-6 Luna
- Context
- 1,050,000
- Max output
- 128,000
- Input / MTok
- $0.10
- Cached / MTok
- $0.01
- Output / MTok
- $0.50
- Input types
- text + image
HIGH-VOLUME ECONOMICS / VERIFIED 2026-10-04
Compare lower-cost OpenAI and Google models on Standard token rates, context capacity, output limits and multimodal input.
LIVE VERIFIED FACTS
DECISION FACTORS
GPT-6 Luna publishes the larger context window: 1,050,000 vs 1,048,576 tokens (1.0×).
Gemini 3.7 Flash additionally lists PDF, audio, video.
For 100K uncached input + 10K output tokens, GPT-6 Luna is lower at $0.015 vs $0.1125 (7.5× difference).
GPT-6 Luna: 128,000. Gemini 3.7 Flash: 65,536.
HOW TO DECIDE
Context, modalities and direct token economics are comparable from official sources. Coding quality, latency, reliability and agent success should be measured on your own acceptance tests before production routing.
Use representative prompts, files, tools and expected outputs from the workload you plan to ship.
Include retries, cached tokens, long-context rules and tool charges—not only headline input price.
Record hallucinations, tool errors, timeout behavior and human corrections alongside pass rate.
A portfolio can outperform a one-model policy when different task classes have different cost and capability needs.
INDEPENDENT EVALUATIONS
SXF does not infer quality from specifications or compare benchmark scores across different evaluator versions/configurations.
WHAT CHANGED
No post-baseline factual changes recorded for GPT-6 Luna, Gemini 3.7 Flash since 2026-09-27.
QUICK ANSWERS
GPT-6 Luna publishes the larger context window: 1,050,000 vs 1,048,576 tokens (1.0×).
For 100K uncached input + 10K output tokens, GPT-6 Luna is lower at $0.015 vs $0.1125 (7.5× difference).
No. This page compares source-backed specifications and economics. Quality, latency and task success require workload-specific evaluation evidence.