Not published
Only provider-documented reasoning controls are shown; missing controls remain unpublished rather than inferred.
MODEL REFERENCE / GOOGLE
Google's generally available speech-to-text model for high-accuracy non-streaming transcription with language detection, diarization, word-level timestamps, and vocabulary biasing.
PRIMARY-SOURCE VERIFIED
Google does not publish a token context window for this service. Input modalities: audio. Output: text. Pricing status: Official paid API.
CAPABILITY CONTRACT
Only provider-documented reasoning controls are shown; missing controls remain unpublished rather than inferred.
Input modalities come from the cited official model documentation.
Output modality and output-limit claims remain separate so an undocumented token cap is not guessed.
PRICING STATUS
$2.00 input · $12.00 output per 1 million tokens.
SXF preserves the provider-published native billing basis. Only calculator-eligible generative token models enter token-cost arithmetic; specialist units and input-only embeddings remain visible without forced conversion.
CAVEATS
Google announced Gemini 3.5 Transcribe as generally available on August 26, 2026, with no shutdown date currently announced.
The model accepts audio and returns text/word annotations; Google documents up to one hour of audio per request, or up to 30 minutes when diarization or word-level timestamps are enabled.
Standard pricing is $2 per 1M audio input tokens and $12 per 1M text output tokens. The model is excluded from the generic text-token calculator because the input token dimension is audio-specific.
EXPLORE
OFFICIAL SOURCES
WHAT CHANGED
No post-baseline factual changes recorded for Gemini 3.5 Transcribe since 2026-10-06.
VERIFIED CHANGE HISTORY
Baseline verified: context not published · Not published max output · $2.00 input / $12.00 output per 1 million tokens
Official evidence ↗Append-only SXF ledger · 1 verified event · chain head a81376cb9fe2…