All guides
BreezeBlueTTSAI VoiceVoiceoverContent12 min read

Design any character voice — and then direct it

BreezeBlue, Breeze TTS 2

From Adam

Hey, it's Adam. The soccer commentator, the gravelly documentary narrator, the tiny chirpy bird — none of those voices existed until I typed them. But designing the voice isn't actually the impressive part. Directing it is: same voice, same line, completely different performance, because you write notes like you would for an actor instead of stapling on an emotion tag. This guide is the prompt library — the frameworks BreezeBlue's own team uses, plus a stack of copy-paste examples for both jobs. BreezeBlue partners with me on this content; the prompts below are the real ones. You commented VOICE, so here's everything.

the idea

Two prompts, two completely different jobs 🎭

This is the thing almost everyone gets wrong, and it's why most AI voiceover still sounds like a robot reading a script. There are two separate prompts in this workflow, and they answer different questions:

Casting versus directing. A brilliant voice with no direction reads flat; great direction on a badly cast voice fights itself. You need both, and they're written differently.

The unlock in Breeze TTS 2 is that the second one is written in plain natural language — you describe the situation the way you'd brief an actor, rather than being limited to bracket-style emotion cues.

before you start

What you need ✅

Try BreezeBlue — free tier available

Try BreezeBlue →
part one

PART ONE — designing the voice 🧬

The five things every good voice description contains

BreezeBlue's own framework. You don't have to hit all five, but the good ones do:

  1. 1Persona & context — who is speaking, and where? A tired customer support agent. A late-night radio host.
  2. 2Basic profile — age, gender, register, accent. Middle-aged female alto. Bright synthetic soprano.
  3. 3Timbre & texture — physical, audible adjectives, not personality traits. Breathy, gravelly, velvety, nasal, polished.
  4. 4Pacing & delivery — speed, pauses, articulation. Slow and deliberate. Rapid-fire. Dramatic pauses.
  5. 5Emotion & fit — the underlying feeling, and what kind of script this voice is built for.

Write it as natural-language paragraphs, not a bulleted list. Three to five sentences: persona first, then texture, then pacing and emotion.

🔑 The rule that matters most: pair the description with a matching script. A generic description ("a warm male voice") with a generic script ("have a wonderful day") proves nothing. The script has to be a line that only this character would say — that's what tests whether the voice actually works.

The prompt library — voice design

Every one is a description + its matching script. Swap the details for your own character.

📋 1. Late-night radio host (BreezeBlue's own prime example)
Voice Description: Create the impression of a late-night radio host speaking to someone who cannot sleep. He has a deep, velvety baritone with a soft magnetic rasp. His pacing is slow and intimate, with gentle pauses after comforting phrases. The tone should feel reassuring, reflective, and emotionally supportive, like a private comfort shared at 2 AM.Script: You are not alone tonight. Let your shoulders drop, take one slow breath with me, and let the noise of the day fade out for a little while.
📋 2. Soccer commentator at the winning goal
Voice Description: A live football commentator in the final seconds of a decisive match. Male, mid-forties, bright and forward-placed tenor that thins and strains at the top of his range when he pushes volume. His pacing accelerates as play builds, words crowding together, then breaks open into long held vowels. The tone is euphoric and slightly hoarse, like a man who has been shouting for ninety minutes and has one more shout left.Script: He's through, he's through — nobody's tracking him — and OH, he's done it, he's actually done it, right at the death!
📋 3. Gravelly old documentary narrator
Voice Description: An elderly documentary narrator with decades of broadcasting behind him. Male, seventies, low and gravelled bass with a dry, papery rasp at the edges of long vowels. His pacing is unhurried and weighted, with deliberate pauses before key facts, and he lands consonants softly rather than crisply. The tone is authoritative but warm — a man telling you something he finds genuinely remarkable.Script: For eleven months of the year, this valley holds nothing at all. And then, almost overnight, it holds everything.
📋 4. Tiny chirpy bird
Voice Description: A very small bird character with an outsized personality. Extremely high, bright, thin soprano with a light nasal chirp and quick flutters at the ends of words. Pacing is rapid and bouncy with almost no pauses, words tumbling out in excitable bursts. The tone is relentlessly cheerful and slightly nosy, built for animated character dialogue.Script: Oh! Oh oh oh — you're new here, aren't you? I've seen everyone, and I have definitely not seen you!
📋 5. Product explainer for a creator channel
Voice Description: A confident creator explaining a tool they genuinely use. Late twenties, neutral accent, clear mid-range voice with a slightly dry, matter-of-fact texture and no broadcast polish. Pacing is brisk and conversational, with short pauses before the point rather than after it. The tone is direct and unimpressed by hype — built for short-form voiceover where the viewer decides in two seconds whether to stay.Script: Most tools stop at the voice. This one lets you direct it, which turns out to be the entire difference.
📋 6. Warm customer support agent
Voice Description: A support agent who has genuinely heard this problem before and is not annoyed by it. Female, thirties, warm mid alto with a soft rounded timbre and a slight smile audible in the vowels. Pacing is calm and even, slowing deliberately when giving instructions, with clean articulation on numbers and names. The tone is patient and competent — reassuring without being saccharine, built for voice agents and phone support.Script: I can see exactly what happened here, and it's a quick fix. Let's sort it out together — it'll take about two minutes.
📋 7. Villain / antagonist for a game
Voice Description: An antagonist who never needs to raise his voice. Male, fifties, low resonant baritone with a smooth polished surface and a cold edge underneath. Pacing is slow and unhurried with long confident pauses, every word fully finished. Articulation is precise to the point of being clipped. The tone is amused and utterly unthreatened, built for game dialogue where menace comes from calm.Script: You made it further than the others. That's not a compliment — it just means this takes slightly longer.
📋 8. Meditation and sleep-content narrator
Voice Description: A guide leading a body-scan meditation. Gender neutral, soft breathy mid-range voice with a lot of air in the tone and almost no hard consonants. Pacing is very slow with long, unhurried pauses between phrases and audible settled breath. The tone is tranquil and unhurried, built for long-form sleep and relaxation audio where the listener should stop tracking the words.Script: Let the weight of your hands go. There is nothing to do here and nowhere else you need to be.

Test: read your description back and ask four questions — can you picture who's speaking? Are there real auditory textures, not just "nice" or "gentle"? Are pacing and pauses explicit? Does the script only make sense in this character's mouth? Any no, rewrite that part.

part two

PART TWO — directing the performance 🎬

The golden rule: direct the action, never the adjective

This is where most AI voiceover dies. Compare:

The first is a label. The second is a direction — situation, physical delivery, pacing, breath. You're not telling it what emotion to have; you're telling it what to physically do.

The five performance dimensions

Hit two or three per instruction — not all five:

  1. 1Context — where are they and what do they want? Hiding from a patrol. Delivering a live news report.
  2. 2Emotion & intensity — be specific. Restrained panic. Muted frustration. Intimate grief. (Not "sad.")
  3. 3Pacing — speed, pauses, breath. Fast and clipped. Long pauses. Almost no room to breathe.
  4. 4Emphasis — which exact words to hit or soften. Push the commands together. Drop the volume on the final warning.
  5. 5Physical / vocal actions — audible cues: sighing, shaking, laughing, choking back tears, coughing.

Always read the transcript first, find the conflict or the keyword, and write the direction around it.

The prompt library — direction

Each is a transcript plus its instruction. The transcript stays the same; change the instruction and you change the performance.

📋 The whisper / hiding delivery
Transcript: We have to stay completely hidden behind these crates until the guard patrol passes the checkpoint, or we risk getting caught immediately.Instruction: Play this as someone hiding close to danger. Keep it barely above a whisper, tense and controlled, with careful breath.
📋 The time-pressure / crisis delivery
Transcript: You need to pull the red lever while pushing the green button simultaneously or else the entire system is going to overheat and shut down in three seconds.Instruction: Deliver it as emergency instructions under time pressure. Keep it fast, precise, and clear.
📋 The tired / resigned delivery
Transcript: [sigh] I suppose we have to restart the engine diagnostics from scratch since the logs were corrupted during the last solar flare event.Instruction: Let the sigh carry tired resignation. Speak slowly and flatly, with muted frustration.
📋 Same line, controlled anger
Transcript: I asked you for that file three times this week.Instruction: Anger held tightly under the surface, not shouted. Keep the volume low and even, slow the pace down, and put a hard stop after "three times." Let the control itself sound like the threat.
📋 Same line, hurt instead
Transcript: I asked you for that file three times this week.Instruction: Play it as disappointment rather than anger. Softer and slightly breathy, pace trailing off toward the end as if they've already given up on the answer. Let "three" carry the weight.
📋 The confident open (short-form hook)
Transcript: Most people are using this completely wrong.Instruction: Deliver it as an opening line that has to stop someone scrolling. Land it flat and certain with no upward inflection, a small pause before "wrong," and full confidence — the tone of someone stating a fact, not making an argument.
📋 The calm reveal (mid-video turn)
Transcript: And that's the part nobody tells you about.Instruction: Drop the energy deliberately from the previous line. Quieter, slower, more intimate — as if leaning in. Small pause before "nobody." Let the drop in volume do the work.
📋 The live news report
Transcript: Crews are still working at the site and officials say the road will remain closed through the morning.Instruction: Deliver as a live on-scene broadcast report. Measured and clear with even pacing, professional detachment, and clean emphasis on "closed" and "morning." No warmth, no drama.
📋 The de-escalating support agent
Transcript: I completely understand why that's frustrating, and I'm going to fix it right now.Instruction: The caller is upset. Drop all brightness. Slow down, lower the pitch slightly, and lead with sincere apology in the tone rather than the words. Steady and unhurried — the sound of someone taking control of a problem calmly.

Test: run the same transcript twice with two different instructions. If the two takes sound genuinely different, your instructions are doing real work. If they sound the same, you wrote adjectives instead of directions.

the dial

guidance_scale — and the setting most people get wrong 🎚

One parameter controls how strictly the output obeys your prompt. It runs 1.0 to 10.0, higher being more creative and expressive. Leave it blank and a random value gets applied — which is why your results feel inconsistent for no reason.

Here's the part almost nobody notices: the two jobs want different settings.

Set it explicitly, every time.

shortcuts

Enhance, the library, and reference input ✨

the rules

🔑 What separates good output from robot output

  1. 1Two prompts, two jobs. Design the voice, then direct the line.
  2. 2Physical adjectives, never personality labels. "Gravelly," not "nice."
  3. 3The script must be a line only that character would say.
  4. 4Direct the action, not the emotion. "Tense low whisper, careful breath" beats "sound scared."
  5. 5Two or three performance dimensions per instruction. Not five.
  6. 6Set guidance_scale explicitly — 2.0–4.0 to design, 1.0–3.0 to direct.
  7. 7Keep both fields under 500 characters. Precision beats padding.
straight talk

The honest bits ⚠️