Introducing Eleven v4Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12

Skip to content

What are Audio Tags? The complete guide to emotional TTS

Written by
Jack Limebear
Published
Last updated

ListenListen to this article

Audio Tags are bracketed, natural-language cues, such as [laughs], [gasps], [excited], and [worried], that Eleven v3 reads as performance direction rather than words to speak. When you write a script for a text to speech model, inserting these Audio Tags gives you granular control over the emotional delivery of your writing. 

Eleven v3 became publicly available in March of 2026. Before v3, getting a specific emotional performance out of a Text to Speech model was a process of regenerating content and hoping for the best. Audio Tags strip away that guesswork, giving you full control over an emotional text to speech performance.

This guide covers what Audio Tags are, how they work, every tag category available, and how they compare to SSML (Speech Synthesis Markup Language). We’ll also cover common use cases, each linking to its own dedicated guide.

Summary

  • Audio Tags are bracketed cues like [shouts] or [excited] that direct emotion, pacing, delivery, and tone within TTS in Eleven v3.
  • Eleven v3 doesn’t support SSML but replaces it with Audio Tags for a more natural writing experience.
  • Tags range across a few categories, such as emotions, delivery and pacing, human reactions, accents, and sound effects.
  • Punctuation and capitalization in your text allow you to shape the pacing of an AI voice performance.
  • You can access Audio Tags on ElevenLabs via the UI or through the API.

How do Audio Tags work?

Audio Tags sit inline with your transcript and are wrapped in square brackets. Whenever you want the delivery to shift, let’s say from normal to [worried], you simply need to add an audio tag.

You can also use more than one tag in a single sentence and can combine tags for a more layered performance, such as [tired] It’s been a long day… [upset] How many more days can I take?

The architecture behind Eleven v3 allows the model to read context at a deeper level than earlier models, letting it to follow emotional cues, tone shifts, speaker transitions, and context clues without needing a separate parameter or setting. Simply tell the model what the moment calls for, and it’ll perform the line exactly as you planned.

Eleven v3 also handles multi-speaker dialogue that feels spontaneous, including interruptions, mood shifts, and the natural back-and-forth of real conversation, all from tags and script structure alone. 

Audio Tags vs. SSML

If you’ve worked with other TTS systems, you’ve likely come across Speech Synthesis Markup Language. Platforms like Amazon Polly and Microsoft Azure Speech use SSML tags like <break> or <emphasis> to give a degree of additional context to a TTS script. 

Eleven v3 doesn’t support SSML break tags or the rest of the SSML tag set. Instead, Audio Tags, punctuation, and the very fabric of a text’s structure take over the role, all while offering granular control over delivery.

You can use the table below to map out common SSML tags to the Eleven v3 equivalent.

What it does
<break time=”1s”>
Inserts a pause
<prosody rate=”slow”>
Slows down speech
<prosody rate=”fast”>
Speeds speech up
<emphasis>
Stresses a particular word
<prosody pitch=”high”>
Shifts pitch higher
<prosody pitch=”low”>
Shifts pitch lower
Eleven v3 equivalent
<break time=”1s”>
[pause], an ellipsis in your text, or a line break
<prosody rate=”slow”>
[drawn out] or [slowly]
<prosody rate=”fast”>
[rushed]
<emphasis>
[shouts] or capitalizing a word
<prosody pitch=”high”>
Emotional tags like [excited]
<prosody pitch=”low”>
Emotional TTS tags like [softly] or [sorrowful]

From a writer’s perspective, adding [whispers] to a line takes much less effort than configuring a prosody rate.

The categories of Audio Tags in v3: Directing performance with Audio Tags

You can place Audio Tags anywhere in your script to shape delivery in real time. You can also use combinations of tags within a script or even in a sentence.

Tags fall into the following core categories.

Emotions

These tags can help you set the emotional tone of the voice, whether it's somber, intense, or upbeat. For example, you could use one or a combination of [sad], [angry], [happily], and [sorrowful].

[sorrowful] I couldn't sleep that night. The air was too still, and the moonlight kept sliding through the blinds like it was trying to tell me something.
[quietly] And suddenly,
that's when I saw it.
0:00
Okay, you are not going to believe this. You know how I've been totally stuck on that short story, like staring at the screen for hours, just nothing? [sighs] I was seriously about to just trash the whole thing, start over, give up probably.
But then [chuckles] last night, I was just doodling, not even thinking about it, right? And this one little phrase popped into my head, just completely out of the blue. And it wasn't even for the story initially, but then I typed it out just to see, and it was like the floodgates opened. Suddenly, I knew exactly where the character needed to go, what the ending had to be. It all just clicked. [sighs] I stayed up till like 3:00 a.m. just typing like a maniac. Didn't even stop for coffee. [chuckles] And it's, it's good, like really good. It feels so complete now, you know? Like it finally has a soul. [sighs] I am so incredibly pumped to finish editing it now. It went from feeling like a chore to feeling like
magic. Seriously, I'm still buzzing.
0:00

Delivery and pacing

These are more about the tone and performance. You can use these tags to adjust volume and energy for scenes that need restraint or force. Examples include tags such as [whispers], [shouts], and [softly].

Could you switch my accent in the old model? [dismissive] Didn't think so, [cheeky] but you can now, so check this out. In just a sec, I'm going to speak with a different accent. And just between you and me, [whispers] I don't really know how, but okay. First, let's change it up [australian accent] so that I can fit in with the locals in Melbourne when I visit next month. [laughing] Whoa. Yeah, man, this is sick. Okay, let's try a different one, see if you can guess. [french accent] My love is like a red, red rose.
0:00
So, I was thinking we could-
[jumping in] Test our new timing features.
[surprised] Exactly! How did you-
[overlapping] Know what you were thinking? Lucky guess. Sorry, go ahead.
[cautiously] Okay. So if we both try to talk at the same time-
We'll probably crash the system?
[panicking] Wait, are we crashing? I can't tell if this is a feature or a-
[interrupting] Bug. Did I just cut you off again?
[sighing] Yes. But honestly, this is kind of fun.
0:00

Human reactions

True natural speech includes reactions. For example, you can use this to add realism by embedding natural, unscripted moments into speech. Tags like [laughs], [clears throat], and [sighs] all help create a TTS model capable of delivering human emotion.

Other examples of different audio tag categories include:

  • Accents and character voices: Accent tags shift the voice into a specific region or persona without switching models, such as [French accent], [British accent], or [pirate voice]. Use them to turn a single voice into a flexible cast rather than one fixed performance.
  • Sound effects: Sound effect tags add non-speech audio directly into the generation, such as [gunshot], [explosion], and [clapping]. These work well for immersive audio, games, and dramatized narration where the scene needs more than a voice. Alternatively, use the ElevenCreative AI sound effect generator.

While tags offer you a high degree of control, you can also use punctuation to shape rhythm and emphasis. For example, add an ellipsis for a natural pause or use commas to create natural breathing patterns. Try writing a word in all caps to add emphasis: “I want it right NOW.”

We're off under the lights here for this semifinal clash, the stadium buzzing with anticipation. Eleven Labs united in their iconic black and white shirts, pushing forward with intent straight from the opening whistle. [excited] The ball is zipped out wide, early attack here. Driving down the wing, pace to burn, [shouting] he skids past one, skips past two. Oh, this is beautiful. One-on-one with the fullback, cuts inside. Oh, that's a lovely bit of footwork. PURE MAGIC on the pitch. ElevenLabs on top form tonight.
0:00
Oh my God. [laughing] You guys, like no joke, I just tried this TTS thing and it was, like, weirdly emotional. Like, it literally said hi, and I was, like, on the verge of tears. [laughing] I don't even cry, okay? I'm a Capricorn.
0:00

What can you do with Audio Tags?

Audio Tags unlock the emotional complexity of scripts in an intuitive way. Write naturally, and add tags as you go.

Here is a quick glimpse of some use cases of emotional TTS with Audio Tags.

Situational awareness in AI audio

Tags such as [whisper] or [shouting] let Eleven v3 react to the moment. From softening a warning to holding an elongated pause for suspense, you can build dynamic situational awareness into your script.

For a full breakdown, take a look at our situational awareness in AI audio guide.

Directing character performance in speech

When building out a script with multiple characters, you want to clearly demonstrate how each voice is different. From tinkering with their personality using [sarcastically] to defining how they speak with [British accent], you can capture a full range of emotions with our TTS engine.

Discover more detail about directing character performance in AI speech.

Expressing emotional context in speech

Audio Tags allow you to shape the emotional delivery of a line as it progresses. You can build emotions throughout a speech, reaching a crescendo the moment you envision your character at their full potential. Tags like [awe], [booming], and [big laugh] help build an entire emotional range into your AI voice performance.

Read more about using emotional context in AI speech. 

Bringing multi-character dialogue to life

When building multi-character dialogue performances, you can begin each line with an audio tag to build out rich character discussions. 

Take the following as an example: Tom: [coughing] [beginning to speak] sorry I’m a little bit ill today— Emma: [interrupting] —Ill? You know how I hate coughing, Tom! Tom: [surprised] It’s just a little cough, it’s nothing serious. How would you feel if— Emma: [overlapping] [annoyed] —I think you need to go home. 

See what else you can do with multi-character dialogue on Eleven v3.

Precision delivery control for AI speech

You can control the pacing of speech with Eleven v3 with Audio Tags or through punctuation. Tags like [pause], [rushed], or [drawn out] let you actively shape the line with precision. Alternatively, add ellipses, full stops, and exclamation marks to build emotion naturally into your TTS performance.

We’ve written an in-depth guide about precision delivery control on Eleven v3.

Emotional TTS is available on the API

Audio Tags work the same way through the Text to Speech API as they do within the ElevenLabs UI. Write the tag inline with the script text, and the API will return the tagged performance in the generated audio, from audiobook pipelines to ad voiceovers. 

Learn more about Eleven v3 or contact sales to get started today.

Audio Tags and emotional TTS FAQ

Similar articles

Create with the highest quality AI Audio