On December 16, 2024, ElevenLabs introduced Flash, a text-to-speech model engineered from the ground up for ultra-low latency. Flash generates speech in just 80 milliseconds, excluding application and network latency, making it the fastest speech synthesis model in the ElevenLabs lineup. The model is purpose-built for real-time use cases such as conversational AI, voice agents, and chatbots, where even small delays between a user finishing a sentence and the agent responding can break the natural flow of dialogue.

The release comes in two variants. Flash v2 supports English only, while Flash v2.5 extends coverage to 34 languages, enabling deployment across global markets from a single model family. Developers integrate the models through the API using the model IDs eleven_flash_v2 and eleven_flash_v2_5. Both share the same latency-first design philosophy, trading a small amount of quality for a dramatic reduction in time-to-first-audio compared to ElevenLabs' higher-fidelity Turbo models.

On pricing, Flash is notably economical: it costs 1 credit for every 2 characters of generated speech, roughly half the credit consumption of ElevenLabs' standard models. This pricing, combined with the low compute cost of the model, makes Flash the cost-effective choice for high-volume, latency-sensitive applications where large amounts of speech are generated in real time across many concurrent sessions.

ElevenLabs is transparent about the trade-offs involved. The company notes that Flash has slightly lower quality and emotional depth compared to its Turbo models, which remain the recommendation for content where expressiveness and nuance are paramount. However, for the specific demands of conversational agents, the reduction in latency more than compensates for the modest quality difference, since responsiveness is the dominant factor in perceived conversation quality.

To validate the model, ElevenLabs benchmarked Flash against comparable ultra-low-latency speech synthesis models. In the company's testing, Eleven Flash consistently outscored other ultra-low-latency models, establishing it as a leader in the sub-100ms tier. This positions Flash as both the fastest and, according to ElevenLabs' evaluations, the highest-quality option among models designed specifically for real-time synthesis.

Flash is presented as ElevenLabs' recommended model for low-latency, conversational voice agents and is available directly through the Conversational AI platform. This tight integration means developers building voice agents on ElevenLabs can adopt Flash as the speech layer without additional wiring, pairing the 75ms synthesis latency with the platform's turn-taking, telephony, and orchestration features. The launch sits within a broader wave of ElevenLabs conversational releases in late 2024, alongside the introduction of Conversational AI Agents on December 1, 2024, the AI Engineer Pack on December 9, 2024, and building on Eleven Turbo v2.5 from July 17, 2024, and the company's earlier work on real-time dubbing dating back to October 31, 2023.