On June 1, 2025, ElevenLabs introduced Eleven v3 (alpha), which the company describes as its most expressive Text to Speech model to date. Where earlier generations of ElevenLabs models optimized primarily for fidelity and speed, Eleven v3 is architected around emotional depth, conversational nuance, and fine-grained linguistic accuracy. The model supports more than 74 languages, a substantial expansion over prior releases, and is aimed at long-form and creative use cases such as videos, audiobooks, and media production tools rather than sub-100ms real-time conversation.
The headline capability of Eleven v3 is its support for inline audio tags. Users can direct the vocal performance by embedding descriptive prompts directly into the script, using tags such as [excited], [whispers], [sighs], [laughing], [happily], and [shouts]. For example, a script line written as "[whispers] Something's coming... [sighs]" or "[happily][shouts] We did it! [laughs]" causes the model to generate the corresponding non-verbal cues and emotional inflections. This gives creators a lightweight, text-native way to choreograph a performance without needing separate voice direction or post-production editing.
Eleven v3 also introduces a dialogue mode designed for multi-speaker conversations. To support this, ElevenLabs shipped a dedicated Text to Dialogue API endpoint that accepts structured speaker turns, allowing developers to generate natural back-and-forth exchanges between distinct voices in a single request. This complements the existing Text to Speech endpoint, which remains available for single-voice synthesis. The model can be invoked via the API by specifying the model ID eleven_v3.
Beyond expressiveness, the model brings substantial improvements in text normalization. ElevenLabs reports a significantly reduced error rate when vocalizing specialized notation such as chemical formulas, phone numbers, and mathematical expressions, areas where earlier TTS systems frequently mispronounced or skipped tokens. This makes the model more reliable for technical, educational, and enterprise content where accurate rendering of structured text matters.
At launch, ElevenLabs offered aggressive introductory pricing. Self-serve users accessing Eleven v3 through the UI received an 85% discount, making it roughly five times cheaper than standard pricing, with an equivalent 85% reduction applied to business plan pricing for enterprise UI users. The promotion was scheduled to run through the end of June 2025, encouraging creators to experiment with the alpha during its initial rollout window.
ElevenLabs was explicit about where Eleven v3 fits in its model lineup. Because the alpha prioritizes expressiveness over ultra-low latency, the company recommended that developers building real-time and conversational applications continue to use Turbo v2.5 or Flash for those workloads, reserving Eleven v3 for scenarios where richness of performance is the priority. At launch the model was available through both the ElevenLabs website and the API, with Studio support noted as coming soon.
The company later updated the announcement to reflect that Eleven v3 graduated out of alpha and became generally available, with the Text to Dialogue API endpoint opened to all users. The general availability milestone confirmed the model's position as ElevenLabs' flagship expressive TTS offering, combining audio-tag-driven emotional control, multi-speaker dialogue generation, 74+ language coverage, and improved normalization into a single production-ready system.