Guide2026/09/21

TTS for YouTube: Rules, Voices and Script Length

Using TTS for YouTube voiceovers is allowed and can be monetized. What gets faceless channels demonetized, when to disclose, and how many words fill 8 minutes.

TTS for YouTube means generating a video's narration from a written script with a text-to-speech voice instead of recording it yourself. YouTube allows it, and videos narrated this way can earn: its monetization rules judge what a video is — templated, repetitive, or an AI posing as an expert — not whether the voice is synthetic. What decides whether a faceless channel earns is the script, the pacing and the pattern across uploads, and all three can be planned before you generate a second of audio.

The script-length figures in this guide are not a rule of thumb. They come from timing seven of the free voices on this site at default speed.

Is TTS allowed on monetized YouTube channels?

Yes. The confusion dates from July 2025, when YouTube renamed its "repetitious content" policy to "inauthentic content". A wave of videos announced that AI voices were being banned. They were not: the policy was aimed at mass-produced uploads, and a synthetic narrator was never the test.

In July 2026 YouTube spelled the policy out in more detail. Content that cannot earn now falls into three groups:

  1. Generic or repetitive content. Videos that look like they were made from a template, or that feel interchangeable when you watch several from the same channel in a row. YouTube's examples include image slideshows, templated storylines, characters put in the same situation with the same outcome, and AI content built on generic templates without the creator's own insight.
  2. Unsatisfying or off-putting content. Disturbing themes repeated without a real story, emotionally manipulative formats, unrelated AI clips stitched together for shock, and deceptive visuals such as a fake celebrity death.
  3. AI personas on sensitive topics. Content that presents itself as a human expert giving advice on health, legal questions, money or politics. Channels that upload it are not allowed to monetize.

Notice what none of the three mentions: the voice. A TTS narrator reading an original script about something you researched is not on the list. A TTS narrator reading the fortieth reworded copy of the same top-ten template is — and a human voice reading it would be too.

Two practical consequences for a faceless channel:

  • Reviewers look at the channel, not one video. Twenty uploads with the same opening, the same structure and the same stock footage read as a template even if every script is new. Vary the format deliberately.
  • The third group is the one TTS channels walk into by accident. A finance or health channel with a synthetic "host" who talks like a credentialed adviser fits the description exactly. Present the narrator as a narrator, put your sources on screen, and never invent credentials.

Do you have to disclose an AI voice?

YouTube asks creators to disclose realistic AI content that could mislead viewers about what happened. Its examples are photorealistic: a real person appearing to do something they did not, a realistic scene that never took place. Non-realistic content and minor edits do not need disclosure, and its list of things you do not need to disclose includes cloning your own voice for voiceovers or dubs, and repairing voice audio.

A stock synthetic narrator over your own footage sits between those examples. The line that settles most cases is whether a viewer could be misled. A voice presented as a specific real person other than you should be disclosed. A clearly narrated explainer is a judgment call.

When in doubt, disclose. YouTube's help page says disclosing "won't limit a video's audience or impact its eligibility to earn money". The toggle sits in the upload flow in YouTube Studio.

Since May 2026 the labels are harder to miss. For photorealistic AI content the label can appear on the player itself rather than only in the expanded description, and YouTube may apply one automatically when a file carries C2PA content credentials or other signals it recognises. Creators who repeatedly fail to disclose risk having labels applied for them, videos removed, or suspension from the Partner Program.

The monetization bar changes on February 1, 2027

Today a channel can join the YouTube Partner Program for ad revenue with 1,000 subscribers plus either 4,000 public watch hours in the last 12 months or 10 million Shorts views in the last 90 days.

From February 1, 2027, new applicants need 8,000 watch hours in the last 365 days, or 20 million Shorts views in the last 90 days — double the current figures. Channels already in the program are not affected by the new entry bar, and the thresholds for fan funding and shopping stay the same. Separately, from the same date, sharing in Shorts ad revenue requires 10 million Shorts views in the last 90 days even for existing members; channels below that keep earning from long-form videos.

Nothing in the announcement targets synthetic voices. What it changes for a TTS channel starting now is the arithmetic. 8,000 hours is 480,000 minutes. At an average view duration of three minutes, that is 160,000 views in a year. Watch time per view is the lever you control, and it is mostly decided by the script.

How many words fill a YouTube video?

This is where most guides say "about 150 words per minute" and move on. We measured instead: seven of our free English voices read a fixed test passage at default speed through the production engines, and the audio files were timed. The spread between the slowest and fastest voice is more than two to one.

VoiceEngineMeasured pace60-second Short8-minute video10-minute video
HarperZonos84 wpm84 words672 words840 words
MilesZonos125 wpm125 words1,000 words1,250 words
BellaKokoro152 wpm152 words1,216 words1,520 words
TaraOrpheus157 wpm157 words1,256 words1,570 words
EmmaKokoro181 wpm181 words1,448 words1,810 words
AdamKokoro187 wpm187 words1,496 words1,870 words

What the table means in practice:

  • Pick the voice before you finish the script. The same 1,200-word script runs 6 minutes 25 seconds with Adam and 14 minutes 17 seconds with Harper.
  • Eight minutes is a real threshold. YouTube lets you turn on mid-roll ads only on videos of 8 minutes or longer. With a brisk voice that takes about 1,500 words; with a deliberate one, under 700. Do not pad a script to reach it — choose a pace that suits the material and write to that.
  • Speed control moves the numbers. Most voices here run from 0.75× to 1.5×, which shifts these figures roughly in proportion; the Orpheus voices, Tara and Dan, speak at a fixed speed. The table is the 1× baseline.
  • Characters, not words, are what TTS allowances count. Our 137-word test passage averaged 5.55 characters per word, spaces and punctuation included, so an 8-minute script for the Kokoro and Orpheus voices is roughly 6,700 to 8,300 characters.
  • Write your ad breaks into the script. YouTube says ad slots at natural breakpoints, such as a pause in the audio, are more likely to serve ads. End a section cleanly where you want a mid-roll, and if the generated audio runs straight through, add a second of silence in your editor.

To time a finished script against a particular voice, paste it into the words to minutes calculator, which uses the same measured speeds.

Choosing a voice for a faceless channel

Pace matters more than timbre, because pace sets how much you can say before a viewer drifts. By format:

  • Explainers and documentary-style videos: a warm, mid-paced read. Bella (152 wpm) suits this well; Emma (181 wpm) if you want a British voice.
  • Lists, tech and news round-ups: a brisk, confident voice such as Adam (187 wpm), which fits more information into each minute.
  • Stories, mysteries and horror: slower is better. Miles (125 wpm) and Harper (84 wpm) leave room for tension, and a slow read carries a short script further.
  • Shorts: Shorts can now run up to three minutes, but the hook still has to land in the first sentence. A faster voice fits more into a 60-second cut.

If your channel is not in English, use a voice built for the language rather than an English voice reading translated text — the accent gives it away within a sentence. Free native voices exist for several languages here: four Hindi voices on the Hindi text to speech page, including Priya, and two each on the Spanish and Portuguese pages. The measured paces above are for the English voices, so time a non-English script by generating a paragraph and checking its length.

Whatever you choose, keep it. A consistent narrator is part of a faceless channel's identity: viewers learn the voice before they learn the name. Switching voices halfway through a series makes it sound like a different channel.

A workflow that stays out of the template trap

  1. Own the script. Draft it however you like, but the facts, the angle and the structure have to be yours. That is where YouTube's test for original, authentic value is decided.
  2. Time it before you generate. Check the length against your voice's measured pace, then cut or expand the script — not the audio.
  3. Generate in sections. Listen at normal speed, fix mispronounced names by spelling them phonetically, and regenerate only the section that broke.
  4. Export WAV for editing. WAV is lossless, so it survives trimming, levelling and re-encoding better than MP3. Keep MP3 for quick drafts.
  5. Vary the format between uploads. Different openings, your own footage or diagrams, on-screen evidence. The July 2026 guidance judges the pattern across a channel.
  6. Disclose when the result could mislead. It costs nothing.

Free or paid voices?

The free voices here run on open-weight models, and a commercial license comes with every tier, the free one included — the audio can go into monetized YouTube videos. The main free engine, Kokoro, is Apache-2.0 licensed; the Kokoro guide covers what it does well and where it falls short, and if you are new to the technology, what TTS is explains how these models work.

The free allowance is 5,000 characters a day with no account, about five to six minutes of narration at Kokoro's pace. Signing in raises it to 10,000 characters, which covers one 8-minute video a day, and past that the free voices continue at one credit per 1,000 characters if you opt in. A script longer than one request is split into parts and joined into a single download automatically. You can start on the Kokoro TTS page without signing up.

Paid voices buy one thing reliably: emotional range. For explainers and lists the free voices are enough. For story channels, where the performance is the product, compare the premium engines on the pricing page before committing to a voice you will use for a hundred episodes.

Frequently asked

Can you monetize YouTube videos that use TTS? Yes. YouTube's monetization rules target templated, repetitive, off-putting or fake-expert content, not synthetic voices. A TTS voice reading an original, researched script is eligible like any other video.

Do I have to disclose a TTS voice on YouTube? Only when the result is realistic enough to mislead, for example a voice presented as a real person. YouTube says disclosing does not affect reach or earnings, so when unsure, disclose.

How many words do I need for an 8-minute YouTube video? Between about 670 and 1,500 words, depending on the voice: 1,216 words at 152 words per minute, 1,496 at 187. Eight minutes is the length at which you can turn on mid-roll ads.

What is the best TTS voice for YouTube? The one that fits the format: a calm mid-paced voice for explainers, a brisk one for lists and news, a slower one for stories. Keep the same voice across a series so the channel sounds consistent.

Will the 2027 Partner Program changes affect TTS channels? Only through the thresholds. From February 1, 2027, new applicants need 8,000 watch hours in 365 days or 20 million Shorts views in 90 days. Nothing in the announcement targets synthetic voices.

Can I use TTS for YouTube Shorts? Yes. Shorts can run up to three minutes, and a 60-second Short holds 125 to 187 words with most of our voices, roughly 700 to 1,050 characters.

Is free TTS audio safe to use on a monetized channel? It depends on what the service grants, not on the technology. Here, free generations carry the same commercial license as paid ones. Some services restrict free-tier output to personal use or say nothing at all, so check before you publish.