ElevenLabs released Eleven v4 on 2026-09-28, describing it as “our most emotive text-to-speech model yet,” together with a low-latency variant, Eleven v4 Turbo. The post, by co-founders Mati Staniszewski and Piotr Dabkowski, says v4 is “built on an entirely new architecture,” ranked first by Artificial Analysis and “preferred by ~75% of listeners in blind head-to-head tests over competing models.” Both models are available in ElevenAgents, ElevenCreative and the API.
The pitch is direction rather than plain reading. Users can describe how a line should be delivered in natural language or use inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing], and ElevenLabs says v4 follows these tags more accurately than prior models. Multi-speaker scenes are generated with awareness of the whole conversation, so speakers respond to each other instead of being stitched together line by line. Both models support more than 90 languages, a voice recorded in one language can speak others with a native accent, and Instant Voice Clones need about 10 seconds of audio.
Eleven v4 Turbo is aimed at real-time agents: ElevenLabs gives a median inference latency of about 100 ms and a median time to first speech of about 150 ms, and says it was optimized together with its ElevenAgents platform as one system.
Why it matters: expressive control and low latency had been a trade-off in voice agents, and this release claims to close it, which matters for customer-service, gaming and healthcare voice deployments. Ten-second cloning across 90+ languages also lowers the bar for voice impersonation. What it does not show: the ranking and preference figures come from ElevenLabs’ own post (citing a third-party arena and its own blind tests), and the post does not detail the new architecture or any new safeguards against misuse of cloned voices.