What Is TTS? Text to Speech Explained
TTS is short for text to speech — technology that turns written text into spoken audio. Here's how it works, where it's used, and how to try it free.
If you've seen "TTS" in a Discord server, a TikTok comment section, or a Twitch stream and weren't sure what it meant, this is the short version: TTS is short for text to speech. It's technology that reads written text out loud in a synthetic voice.
You've almost certainly heard it already — in navigation apps, screen readers, airport announcements, and a large share of the faceless narration videos on your feed.
What TTS actually does
A text-to-speech system takes a string of text and produces audio of someone speaking it. That sounds simple, but the interesting part is everything the model has to figure out on its own:
- Pronunciation — "read" in "I read a book yesterday" versus "I like to read"
- Rhythm and stress — which words to lean on so a sentence sounds natural instead of flat
- Intonation — the rise at the end of a question, the fall at the end of a statement
- Pauses — where to breathe, and how long to hold on a comma versus a period
Early systems handled this with hand-written rules and stitched-together recordings, which is why old TTS sounds so robotic. Modern neural TTS learns all of it from data, which is why current voices can be genuinely hard to distinguish from a human read.
How modern TTS works
Today's systems generally run in two stages. First, a model converts your text into an intermediate acoustic representation — essentially a detailed map of how the audio should sound over time. Then a second model, called a vocoder, turns that representation into an actual waveform you can play.
What changed in the last few years is scale and training data. Models like Kokoro-82M pack surprisingly natural output into a small package, while larger commercial systems add emotional range, voice cloning, and fine-grained control over delivery.
The practical result: quality that used to require a recording studio and a voice actor now takes a few seconds and costs a fraction of a cent.
To put actual numbers on it, here is what the engines behind this site cost to run, per 1,000 characters of text — roughly one minute of speech:
| Engine | Type | Approx. cost per 1,000 characters |
|---|---|---|
| Kokoro-82M | Open weights | ~$0.0006 |
| Chatterbox | Open weights | ~$0.001 |
| Fish Audio S1 | Commercial | ~$0.015 |
| MiniMax speech-2.8 | Commercial | ~$0.06 |
| ElevenLabs Multilingual v2 | Commercial | ~$0.10 |
The spread is roughly 160×, and it buys you emotional range, voice cloning and consistency across long passages — not basic intelligibility. Open-weight models cleared the "sounds like a person reading" bar some time ago. That is why a free tier can be genuinely usable rather than a crippled demo.
Where people use TTS
Content creation. Faceless YouTube channels, TikTok narration, and short-form explainers lean heavily on TTS. It lets one person produce a consistent voice across hundreds of videos without recording a thing.
Gaming and streaming. Discord has a built-in /tts command, and Twitch streamers use donation TTS so viewer messages get read aloud on stream. Both are so common that "TTS" is basically gamer vocabulary at this point.
Accessibility. Screen readers have used TTS for decades. For people with visual impairments or reading disabilities like dyslexia, it's not a novelty — it's how the web is usable at all.
Learning. Language learners use TTS to hear correct pronunciation on demand, and students convert readings into audio to review while commuting.
Business. Phone systems, e-learning modules, product demos, and internal training videos all use synthetic voice to avoid re-recording every time the script changes.
What TTS still gets wrong
Synthetic voice has closed most of the gap, but four failure modes show up reliably enough to plan around:
Homographs it cannot disambiguate. "Lead" the metal and "lead" the verb are spelled identically, and no model gets this right every time from context alone. Same for "bass", "tear", "wound", "live". If a word is critical, listen before you publish.
Proper nouns and jargon. Product names, surnames, place names and technical terms are where TTS most often embarrasses itself, because they follow no consistent pronunciation rules. Anything invented after the training data was collected is a coin flip.
Numbers and abbreviations. "1998" can be read as "nineteen ninety-eight" or "one thousand nine hundred ninety-eight", and both are defensible. Dates, version numbers, currency and units are all ambiguous in writing and unambiguous in speech, so the model has to guess.
Sustained emotional arc. A single sentence can be delivered convincingly. A ten-minute emotional narration is still where a human performer wins — models drift toward a neutral average over long passages.
The practical workaround for the first three is the same: spell it out phonetically in the input text. Writing "Kokoro (koh-koh-roh)" costs you nothing and removes the guesswork.
How to judge whether a TTS voice is good enough
Demos are chosen to flatter the model. Test with your own text instead, and listen for these specifically:
- Sentence endings. Weak models flatten every sentence to the same falling tone. Good ones distinguish a question from a statement from a list item.
- Comma pauses. Listen to whether a comma produces an actual beat or the model runs straight through it.
- The second minute. Many voices are convincing for twenty seconds and start sounding mechanical once you notice the rhythm repeating. Generate at least a minute.
- Your own vocabulary. Feed it the product names and terms you will actually use, not the sample text.
If a voice clears all four on your material, the engine tier is not your bottleneck — your script is.
What "TTS" means in different contexts
Within technology the abbreviation is stable, but expectations shift by platform — on Discord it usually means the /tts command, on Twitch it usually means donation TTS, and on TikTok it usually means the in-app narrator voice.
Step outside tech and the letters stop meaning speech at all: in sneaker and clothing threads, TTS means true to size. We wrote a separate guide to what TTS stands for in each context for anyone who arrived at the abbreviation cold.
Free vs. paid TTS
Most tools split along a predictable line. Free tiers typically give you solid neural voices with a daily character limit — plenty for videos, study material, and everyday narration. Paid tiers unlock higher-end engines with more emotional range, voice cloning, longer inputs, and commercial guarantees.
One thing worth checking before you commit to any tool: whether the free output is actually licensed for commercial use. Plenty of "free" generators quietly prohibit monetized use, which is a problem if you're publishing to a monetized channel.
Try it yourself
The fastest way to understand TTS is to hear it. You can generate speech free on this site — no sign-up needed, download as MP3, and commercial use is included on every tier including free.
If you're publishing to a specific platform, these are set up for it:
- Discord TTS — audio for soundboards and voice channels
- TikTok voice generator — voiceovers for short-form video
- Twitch TTS — alert audio and stream segments
- Kokoro TTS online — try the open-weight model in your browser
Frequently asked
Is TTS the same as AI voice? Mostly, in casual use. "TTS" is the older, broader term for any text-to-audio system; "AI voice" usually implies a modern neural model. Today they overwhelmingly refer to the same thing.
Is TTS the opposite of STT? Yes — speech to text (STT) transcribes audio into written text, going the other direction.
Can TTS sound like a real person? Current models get close enough that most listeners can't reliably tell in short clips. Longer passages and emotionally complex delivery are still where human performers hold an edge.
Does YouTube demonetise TTS narration? No. YouTube's policy targets content with no original contribution, not synthetic voice specifically. A narrated video with an original script, original edit and genuine commentary is fine; a channel that reads scraped articles over stock footage is at risk whether the voice is synthetic or human.
How many characters is a minute of speech? Around 900 characters, or roughly 150 spoken words, at a normal narration pace. That is a useful number for estimating how far a character allowance will actually go — 10,000 characters is about eleven minutes of audio.