How to choose the right model

This guide shows you how to choose the right ElevenLabs model for your use case.

ElevenLabs offers a range of models optimised for different requirements. The right choice depends on your use case, latency requirements, and quality expectations. Refer to the models reference for full specifications.

By requirement

Quality

Use eleven_v4

The flagship model with the highest fidelity, richest emotional expression, and broadest language support.

Low-latency

Use eleven_v4_turbo

Optimised for real-time applications with ~100ms latency.

By use case

Content creation

Use eleven_v4

Ideal for professional content, audiobooks, and video narration.

Conversational agents

Use eleven_v4_turbo for the most expressive delivery, or eleven_flash_v2_5 and eleven_flash_v2 for the lowest latency.

Use the 2.5 model for language support outside of English.

Optimised for real-time conversational applications.

Transcription

Use scribe_v2 for batch transcription, scribe_v2_medical for medical and clinical audio, or scribe_v2_realtime for real-time transcription.

State-of-the-art accuracy across 90+ languages with speaker diarisation and word-level timestamps.

Voice changer

Use eleven_multilingual_sts_v2

Specialised for Speech-to-Speech conversion.

Next steps