Gemini API Changelog·· 2026-08-26精选AI 评分60
Google 发布 Gemini 3.5 Transcribe 与 Transcribe Live 语音转文字模型
August 26, 2026
AI 导读
Google 将 Gemini 3.5 Transcribe 系列语音转文字模型正式开放(GA),包含两款基于 Gemini 音频理解能力的专用模型。
推荐理由
Gemini 3.5 Transcribe 两款语音转文字模型正式开放,可据此了解其语言覆盖与流式能力边界。
正文 · 原文
Gemini 3.5 Transcribe generally available (GA): Released two dedicated speech-to-text models based on Gemini's audio understanding:
- Gemini 3.5 Transcribe (
gemini-3.5-transcribe): High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps, and custom vocabulary biasing (up to 1,000 terms). - Gemini 3.5 Transcribe Live (
gemini-3.5-transcribe-live): Low-latency, bidirectional streaming speech-to-text over WebSockets using the Live API, supporting interim and finalized transcription events, Smart transcription mode, and multiple Voice Activity Detection (VAD) strategies.
To get started, see the Audio transcription guide, the Live transcription guide, and the Gemini 3.5 Transcribe model page.
- Gemini 3.5 Transcribe (
来源:Gemini API Changelog · ai.google.dev