Google DeepMind launches Gemini 3.5 Live Translate

On June 9, 2026, Google DeepMind published the model card for Gemini 3.5 Live Translate (Gemini 3.5 Audio), a streaming speech-to-speech model that processes continuous audio streams to deliver immediate, human-like spoken responses rather than waiting for a sentence to finish. The model card lists an audio input context window of up to 128K tokens and audio-plus-text output of up to 64K tokens, and notes the system was evaluated across translation quality, latency, and speech naturalness.

The model was distributed across Google Translate, Google AI Studio, and the Gemini API, putting real-time voice-to-voice translation into both developer and consumer surfaces on the same day.

This marks a structural shift from turn-based translation, where systems waited for a full utterance before processing, toward translation generated concurrently with speech.

Sources

Last verified June 29, 2026