Text to Speech API
Ultra-realistic and low latency speech generation
Build with high-quality, controllable TTS for real-time and bulk applications. Models optimized for latency, fidelity, and long-form consistency.
- Lovable
- Synthesia
- Stripe
- Perplexity
- Twilio
Built on the most powerful Voice AI models
Choose the right model for your use case: from ultra-low latency agents to expressive, long-form narration.

Eleven v4
Our most emotive, high quality model.
- Exceptional voice cloning capabilities
- 90+ languages supported
- 10,000 character limit
- Multi-speaker dialogue
- ~$0.08 per minute

Eleven v4 Turbo
Our most expressive model for realtime speech.
- Ultra-low latency (~100ms)
- Exceptional voice cloning capabilities
- 90+ languages supported
- Audio tags for fine-grained control
- ~$0.04 per minute
Everything you need to build production-ready speech
Generate expressive, controllable speech with models built for real-time, long-form, and production use.
Control emotion and delivery
Create controllable, expressive speech, layered with emotion, audio events, and immersive soundscapes.

Access 11,000+ voices
Explore an ever-growing collection of expressive, lifelike voices for any use case.

Voice design & cloning
Create in over 30 languages with natural voices, expressive accents, and localized audio tailored to your audience.

Multi-speaker dialogue
Create natural multi-speaker conversations across 70+ languages with expressive, controllable voices.

Audio events and direction
Control delivery with audio tags, timing cues, and narrative direction built into the speech.

Pronunciation dictionaries
Define custom pronunciations to ensure consistent, accurate speech for names and terminology.



