Change language to
0:00

Google has launched Gemini 3.8 Flash TTS, a new text-to-speech model that can design voices from plain-language prompts and recreate a speaker’s voice from a 30-second sample. The release also includes the cheaper Gemini 3.8 Flash-Lite TTS for high-volume dubbing and voice agents. Both are rolling out through the Gemini API and Google AI Studio, according to Google’s announcement. The Next Web first highlighted the 30-second cloning workflow, while Crypto Briefing also reported the launch.

Hume AI benchmark chart comparing Gemini 3.8 Flash TTS models

Subscribe to our Newsletters for more Tech Stories

Gemini 3.8 Flash TTS turns prompts into voices

Gemini 3.8 Flash TTS is aimed at games, podcasts, audiobooks and interactive media. Developers can describe a character’s role, accent and vocal characteristics in more than 100 languages and dialects, then save the resulting voice for reuse. Google also says the model provides access to more than 2,000 production-ready voices, including regional varieties such as Quebec French and Scots English.

The model is designed to behave more like a vocal studio than a preset menu. Scripts can direct delivery line by line, with cues for pacing, emotional tone, laughs, sighs and other small performance details. It also supports two-speaker scenes from a single script and long-form audio generation.

The official announcement includes a Hume AI text-to-speech quality chart comparing Gemini 3.8 Flash TTS and Flash-Lite with models from ElevenLabs, Cartesia, OpenAI and Inworld. Google says the full Flash model is built for creative direction, while Flash-Lite is optimised for cheaper, higher-volume workloads such as dubbing and voice agents.

Gemini 3.8 Flash TTS can replicate a voice

The more consequential feature is voice replication. Gemini 3.8 Flash TTS can recreate a vocal profile from a 30-second recording, but Google says the process requires a verbal consent recording from the voice owner. The system checks that the consent recording matches the reference speaker before allowing the voice to be created.

That safeguard does not make voice cloning risk-free, but it is a more concrete barrier than simply asking users to promise they own the audio. Every clip generated by Google’s Gemini Audio models carries a SynthID watermark, while replicated voices receive C2PA content credentials. Google says those measures are intended to help identify synthetic speech and preserve information about how it was made.

The models are available now through the Gemini API and AI Studio, with further integrations listed for products including Gemini Notebook and Google Vids. Google previously launched Gemini 3.5 Transcribe for cleaning up voice notes, while Nvidia’s PersonaPlex takes a different route by handling live, two-way voice interaction.

That leaves Google with a broader audio stack: custom voices without a recording session, voice replication with consent checks, and a lower-cost model for businesses that need to generate a great deal of speech. The robots are getting better direction.

What is Gemini 3.8 Flash TTS?

It is Google’s text-to-speech model for designing expressive voices, directing delivery and generating long-form or multi-speaker audio.

Can Gemini 3.8 Flash TTS clone a voice?

Yes. Google says it can replicate a vocal profile from a 30-second sample when the voice owner provides matching verbal consent.

Where can developers use Gemini 3.8 Flash TTS?

The models are rolling out through the Gemini API and Google AI Studio, with additional integrations listed for Gemini Notebook and Google Vids.