Open-weight model

Zonos TTS Online

Speech studio

34 voices · 18 languages

5,000 free characters/day
155 / 400
Pick a voicePlay any of them
Free voices
Natural voices3 cr / 1k chars
Premium voices10 cr / 1k chars
Ultra voices20 cr / 1k chars

Try Zyphra’s Zonos-v0.1 in your browser — no CUDA, no Python, no weights to download. Pick one of the preset voices, type your text, and download 44kHz audio.

Apache-2.0 model44kHz outputNo install

How it works

Three steps, about a minute

  1. 1

    Paste or write your text

    Drop in a script, a paragraph, or a whole article. Longer text is split into parts automatically and joined back into one file.

  2. 2

    Pick a voice and preview it

    Almost every voice has a preview you can play before spending anything, paid ones included. Filter by language, then hear two or three before committing.

  3. 3

    Generate and download

    Audio comes back in seconds as MP3 or WAV, with a commercial licence on every tier — including the free one.

Studio headphones on a stand beside an audio interface

Examples

Hear it on real scripts

Every clip below was generated here on the free tier — no account, no credits. Play one to hear the voice on that kind of writing, or load the script into the generator and make your own version of it.

Phone greeting

IVR or voicemail line — the classic use for a consistent voice.

Thank you for calling Northgate Supply. Our team is available Monday to Friday, eight until six. For an existing order, press one. For everything else, stay on the line and someone will be with you shortly.

Read by Emma · British, calm

Product ad

Thirty seconds of ad copy, the hardest thing to read naturally.

You already own the shoes. That's the annoying part. They're at the back of the cupboard, they still fit, and the only thing standing between you and Saturday morning is the fifteen minutes it takes to find them. We can't help with that. But everything after it, we can.

Read by Poppy · British, bright

Article to audio

Turn something written to be read into something worth listening to.

Reading aloud is the oldest editing tool there is. The eye forgives a sentence that has no breath in it; the ear refuses. Which is why writers who read their own drafts out loud tend to produce shorter sentences, and why anyone about to publish something long should hear it first.

Read by Harper · American, clear

Included

What you get

Free without an account
Thousands of characters a day as a guest, no sign-up and nothing to install. An account raises the limit; it is not a gate on the tool.
Commercial use included
Monetised videos, client work and ads are all fine on every tier. You own what you generate.
MP3 and WAV export
MP3 for editors and uploads, WAV where you need uncompressed audio. Speed is adjustable from 0.75× to 1.5× on engines that support it.
Long text, one file
Text past the per-request limit is split at sentence boundaries, generated in order, and stitched back into a single download.
Voices across 18 languages
Several languages are read by voices trained on that language rather than an English voice reading foreign words — the difference is audible immediately.
Credits that never expire
Premium and ultra voices are billed in credits. One-time packs never expire; monthly plans refill every billing cycle instead.

What makes Zonos interesting

Most open TTS models converge on the same recipe: a transformer backbone generating audio tokens at 24kHz. Zonos-v0.1 broke from that twice. Its hybrid variant swapped much of the transformer for an SSM (Mamba-style) backbone, and both variants output at 44kHz — closer to CD quality than to the 24kHz most of the field settled on.

Zyphra released both variants under Apache-2.0 with weights on Hugging Face, which made Zonos a fixture in open-TTS comparisons. The trade-off is the usual one: running it yourself requires a CUDA GPU, and the setup is involved enough that most people never actually hear the model before deciding whether to invest in it.

Hear it before you set it up

This page exists to remove that gap. The generator above calls the hybrid variant on cloud GPUs — the same deployment that serves this site’s free tier — so the audio you get is the model’s real output, not a cherry-picked demo reel.

Three preset voices are available: two American (female and male) and one British female. Run the same sentence through each, then through Kokoro or Chatterbox from the voice picker, and you will have a better comparison than any benchmark table gives you.

Try it here vs. self-host

The short version:

  • Use this page to evaluate the model, produce short clips, or check whether 44kHz output matters for your use case
  • Self-host when you want Zonos’s voice cloning, need volume beyond short clips, or want to fine-tune
  • Each clip here is capped at 400 characters by the hosted deployment — for long text, Kokoro on the same free tier goes to 5,000

Questions

Frequently asked