ElevenLabs has launched v4 and v4 Turbo, two text-to-speech models that add more expressive voice direction and support more than 90 languages. Turbo is designed for lower-latency voice agents, while both models can clone a voice from a 10-second sample, according to the company.
TechCrunch reports that the new models lift language support from 70 in v3 to more than 90, with the largest quality improvements reported in Japanese, Brazilian Portuguese, Mandarin and Cantonese. ElevenLabs says v4 can also use surrounding text to shape delivery, rather than treating each spoken line as an isolated recording.
ElevenLabs v4 gives generated speech more direction
The company says v4 can follow natural-language instructions about tone and pacing, as well as inline tags for laughter, anger, accents and sound effects. Users can stack tags and direct how particular phrases should be delivered. The model also aims to keep a speaker’s identity consistent across longer passages and multi-speaker scenes, which could help with audiobooks, game dialogue, dubbing and branded voice work.
The launch is part of a wider contest to make synthetic speech sound less like a clean recording of words and more like a performance. ElevenLabs’ own blind comparison chart shows v4 winning between 65% and 81% of preference judgements against four named competitors, including models from Google, Cartesia and Inworld. The company says listeners judged expressiveness and naturalness; these are vendor-run comparisons, not an independent test.

For another recent example of that competition, Google’s Gemini 3.8 Flash TTS can clone a voice from a 30-second sample, with a matching verbal-consent recording. That makes the 10-second capture claim a useful point of comparison, but sample length alone does not establish permission to reproduce a real person’s voice.
| Model | Intended use | Company-reported detail |
|---|---|---|
| Eleven v4 | Narration, dubbing and expressive dialogue | More than 90 languages; voice cloning from 10 seconds |
| Eleven v4 Turbo | Low-latency conversational agents | More than 90 languages; about 150ms median time to first speech in ElevenLabs’ test |
ElevenLabs v4 Turbo is built for voice agents
Turbo is the option for products where waiting for a response can make a conversation feel awkward. ElevenLabs reports a median time to first audible speech of about 150 milliseconds in its September test, measured over WebSocket streaming. The company also says v4 can begin generating audio as the language model starts producing its answer, instead of waiting for the full response.
Those figures describe the speech model’s benchmark, not the complete time a customer will wait for an AI agent. The language model, network, application and any checks around the conversation can all add delay. ElevenLabs has been expanding its enterprise calling business, and TechCrunch reported that more than 55% of the company’s business now comes from large companies. A quicker, more expressive agent model fits that push, though businesses will still need to test it with their own scripts and callers.
The company says v4 Turbo is designed to work with its ElevenAgents platform, while both new models are available through ElevenLabs’ creative tools and API. That gives developers the choice of using a single vendor stack or integrating speech with their own agent systems.
More than 90 languages, plus voice cloning from 10 seconds
ElevenLabs v4 and v4 Turbo support more than 90 languages, up from 70 in v3, according to TechCrunch. ElevenLabs says the update improves not just pronunciation but also rhythm, emotion and accent consistency when a voice speaks another language. That could matter for dubbing and localisation, where a voice that changes character halfway through a long recording is not particularly useful.
ElevenLabs says its Instant Voice Clones can now capture a voice from just 10 seconds of audio. Its official v4 announcement also describes improved speaker similarity and consistency across generated passages. The easier capture process does not, by itself, answer the rights question: anyone cloning an identifiable person’s voice still needs to make sure they have permission to use it.
For developers comparing the wider voice-agent stack, Nvidia’s open-source PersonaPlex takes a different route: it listens and speaks at the same time rather than focusing only on text-to-speech output. It is a different architecture, not a like-for-like replacement for Eleven v4 Turbo.
What UAE creators should check about Arabic support
ElevenLabs already offers an Arabic text-to-speech service, but its v4 launch post does not publish a complete language list or name Arabic specifically. The company’s broader claim of more than 90 languages is not enough to confirm which Arabic varieties work with each new model. UAE teams preparing Arabic dubbing, customer support or voice content should check the model’s current language and accent options in the product or API before building a workflow around it.
No UAE-specific pricing or release terms were announced in the sources reviewed. The models are available now through ElevenLabs’ own products and API, but the company has not published v4-specific regional pricing in its launch announcement.
Does the 150ms figure mean an AI agent completes its response in 150ms?
No. It measures time to first audible speech in ElevenLabs’ test, not the full answer. The language model, network, application and any checks around the conversation can add delay.
Which model should a team use for a long-form voiceover?
The standard Eleven v4 is the better fit to test first for expressive narration, dubbing and longer dialogue. Turbo is aimed at interactive agents where fast turn-taking is the priority; teams needing both should test their own scripts in each model.
Does a 10-second voice sample give permission to clone someone?
No. Ten seconds is a technical capture threshold, not proof of consent or rights. Get permission from the speaker and check applicable law and platform rules before using an identifiable voice clone.
Can UAE teams assume Eleven v4 supports Arabic?
No. The launch announces more than 90 languages but does not name Arabic or identify which Arabic varieties the new models support. Check the current model documentation or test the target accent before committing to a production workflow.


















