ElevenLabs releases v4 speech models with support for 90 languages

ElevenLabs expands speech platform with v4 and v4 Turbo
ElevenLabs has launched ElevenLabs v4 and v4 Turbo, two speech models that add support for more than 90 languages, greater expression control and lower latency for voice-agent applications. The release follows the company’s v3 model, introduced last year, which supported 70 languages.
The company said the v4 generation uses a new architecture designed to improve control and accelerate voice cloning. Users can clone a voice from 10 seconds of audio, ElevenLabs said. The model is also intended to preserve a speaker’s identity more consistently across longer passages of text.
ElevenLabs said it saw its largest quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The broader language coverage and those stated gains matter for companies building multilingual speech products rather than deploying a separate experience for each market.
More detailed control over delivery
The new models retain context from the text as they generate speech, allowing delivery to change with the meaning of a passage. ElevenLabs introduced inline expression tags with v3; v4 expands that mechanism by allowing multiple tags to be stacked and followed in sequence.
That feature is aimed at creative production as well as operational dialogue, where tone can change within a single response. It gives developers another layer of instruction over generated speech, alongside the voice identity and the text submitted to the model.
Voice agents are a central use case
ElevenLabs positioned v4 for voice agents because of lower latency intended to make conversations feel more fluid. The model can begin generating audio when the LLM behind an agent starts producing its answer, rather than waiting for the full response to be completed.
The company also said v4 can treat confrontations, escalations and holds differently to support issue resolution. That focus arrives as the company’s enterprise calling business has grown: more than 55% of ElevenLabs’ business now comes from large companies.
The release comes as ElevenLabs annualized revenue run rate places the company’s annualized revenue run rate above $600 million, following rapid growth from roughly $330 million at the start of the year. ElevenLabs has also expanded its workforce to more than 800 people and hired across India, Europe and Brazil.
Competitive speech market
Speech-model competition has intensified, with Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs developing expressive voice models, while Google and OpenAI have improved their own voice offerings. ElevenLabs raised $500 million in a Sequoia-led round earlier this year at an $11 billion valuation.
For businesses, the practical next step is to evaluate v4 and v4 Turbo on representative dialogue, long-form narration and priority languages, paying particular attention to latency, expression-tag consistency and the quality of cloned voices in real workflows.

