Pradhyuman Yadav

ElevenLabs’ new v4 speech model supports more expression control and 90 languages

By Pradhyuman,

ElevenLabs launched two speech models, named v4 and v4 Turbo. The models support over 90 languages, up from 70 in the previous version. ElevenLabs reported that the largest quality improvements occurred in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. The v4 model architecture enables voice cloning using 10 seconds of audio and lets users stack inline tags to control vocal expressions in sequence.

ElevenLabs stated that lower latency in v4 helps voice agents handle fluid conversations. Large companies generate over 55% of the company's enterprise calling business. The startup reported that its annualized revenue run rate grew from about $330 million at the start of the year to over $600 million, while total headcount passed 800 employees.

Sequoia led a $500 million investment round earlier this year that valued ElevenLabs at $11 billion. Co-founder and CEO Mati Staniszewski told TechCrunch that the company aims for an initial public offering in the next years.

Maintained by Pradhyuman.

Filed under: Funding & Business, Models & Research

Related articles