Google DeepMind released Gemini 3.5 Live Translate on June 9, 2026, launching an audio model that provides near real-time speech-to-speech translation in over 70 languages.
The model automatically detects spoken languages and retains the speaker's intonation, pacing, and pitch in the generated audio. Unlike turn-by-turn translation tools that wait for a speaker to finish, Gemini 3.5 Live Translate processes audio as a continuous stream, staying just a few seconds behind the original speech.
API and Enterprise Access
Google DeepMind made the model available today in public preview for developers through the Gemini Live API and Google AI Studio. Integration partners including Agora, Fishjam, LiveKit, Pipecat, and Vision Agents are using the API to support voice applications. Ride-hailing service Grab is testing the model for drivers and passengers on its platform, where users make over 10 million voice calls each month. Entertainment company CJ ENM is also testing the software.
Meet and Mobile Features
Google Meet will integrate the model in a private preview starting this month for select Google Workspace business accounts before a broader rollout later this year. The integration expands Meet's real-time translation support from five languages to over 70, allowing more than 2,000 language pairs in a single meeting.
Mobile devices running the Google Translate app on Android and iOS can use the feature globally when connected to headphones. Android devices are also gaining a listening mode that streams translated speech through the earpiece, allowing individuals to hear live translations during events such as guided tours without using public speakers or external headphones.
Audio outputs generated by the model contain SynthID watermarks woven directly into the signal to keep synthetic speech detectable. The release marks 20 years since Google launched its first machine learning translation experiments, which now process more than one trillion words per month.
