Not published
Only provider-documented reasoning controls are shown; missing controls remain unpublished rather than inferred.
MODEL REFERENCE / GOOGLE
Google's flagship production text-to-speech model for high-fidelity, expressive long-form and multi-speaker audio generation.
PRIMARY-SOURCE VERIFIED
Google documents a 8,192-token context window. Input modalities: text. Output: audio. Pricing status: Official paid API.
CAPABILITY CONTRACT
Only provider-documented reasoning controls are shown; missing controls remain unpublished rather than inferred.
Input modalities come from the cited official model documentation.
Output modality and output-limit claims remain separate so an undocumented token cap is not guessed.
PRICING STATUS
$0.50 input · $0.125 cached · $9.00 output per 1 million tokens.
SXF preserves the provider-published native billing basis. Only calculator-eligible generative token models enter token-cost arithmetic; specialist units and input-only embeddings remain visible without forced conversion.
CAVEATS
Google lists Gemini 3.8 Flash TTS as generally available as of September 22, 2026.
Standard pricing uses text input tokens and audio output tokens; audio output tokens correspond to 25 tokens per second.
The model is excluded from the generic generative token calculator because its output-token dimension represents audio rather than text.
EXPLORE
OFFICIAL SOURCES
WHAT CHANGED
No post-baseline factual changes recorded for Gemini 3.8 Flash TTS since 2026-10-06.
VERIFIED CHANGE HISTORY
Baseline verified: 8,192 context · 16,384 max output · $1.00 input / $0.25 cached / $18.00 output per 1 million tokens
Official evidence ↗Append-only SXF ledger · 1 verified event · chain head a81376cb9fe2…