Open-weight model
Zonos TTS Online
Speech studio
34 voices · 18 languages
Try Zyphra’s Zonos-v0.1 in your browser — no CUDA, no Python, no weights to download. Pick one of the preset voices, type your text, and download 44kHz audio.
How it works
Three steps, about a minute
- 1
Paste or write your text
Drop in a script, a paragraph, or a whole article. Longer text is split into parts automatically and joined back into one file.
- 2
Pick a voice and preview it
Almost every voice has a preview you can play before spending anything, paid ones included. Filter by language, then hear two or three before committing.
- 3
Generate and download
Audio comes back in seconds as MP3 or WAV, with a commercial licence on every tier — including the free one.

Examples
Hear it on real scripts
Every clip below was generated here on the free tier — no account, no credits. Play one to hear the voice on that kind of writing, or load the script into the generator and make your own version of it.
Phone greeting
IVR or voicemail line — the classic use for a consistent voice.
“Thank you for calling Northgate Supply. Our team is available Monday to Friday, eight until six. For an existing order, press one. For everything else, stay on the line and someone will be with you shortly.”
Product ad
Thirty seconds of ad copy, the hardest thing to read naturally.
“You already own the shoes. That's the annoying part. They're at the back of the cupboard, they still fit, and the only thing standing between you and Saturday morning is the fifteen minutes it takes to find them. We can't help with that. But everything after it, we can.”
Article to audio
Turn something written to be read into something worth listening to.
“Reading aloud is the oldest editing tool there is. The eye forgives a sentence that has no breath in it; the ear refuses. Which is why writers who read their own drafts out loud tend to produce shorter sentences, and why anyone about to publish something long should hear it first.”
Included
What you get
- Free without an account
- Thousands of characters a day as a guest, no sign-up and nothing to install. An account raises the limit; it is not a gate on the tool.
- Commercial use included
- Monetised videos, client work and ads are all fine on every tier. You own what you generate.
- MP3 and WAV export
- MP3 for editors and uploads, WAV where you need uncompressed audio. Speed is adjustable from 0.75× to 1.5× on engines that support it.
- Long text, one file
- Text past the per-request limit is split at sentence boundaries, generated in order, and stitched back into a single download.
- Voices across 18 languages
- Several languages are read by voices trained on that language rather than an English voice reading foreign words — the difference is audible immediately.
- Credits that never expire
- Premium and ultra voices are billed in credits. One-time packs never expire; monthly plans refill every billing cycle instead.
What makes Zonos interesting
Most open TTS models converge on the same recipe: a transformer backbone generating audio tokens at 24kHz. Zonos-v0.1 broke from that twice. Its hybrid variant swapped much of the transformer for an SSM (Mamba-style) backbone, and both variants output at 44kHz — closer to CD quality than to the 24kHz most of the field settled on.
Zyphra released both variants under Apache-2.0 with weights on Hugging Face, which made Zonos a fixture in open-TTS comparisons. The trade-off is the usual one: running it yourself requires a CUDA GPU, and the setup is involved enough that most people never actually hear the model before deciding whether to invest in it.
Hear it before you set it up
This page exists to remove that gap. The generator above calls the hybrid variant on cloud GPUs — the same deployment that serves this site’s free tier — so the audio you get is the model’s real output, not a cherry-picked demo reel.
Three preset voices are available: two American (female and male) and one British female. Run the same sentence through each, then through Kokoro or Chatterbox from the voice picker, and you will have a better comparison than any benchmark table gives you.
Try it here vs. self-host
The short version:
- Use this page to evaluate the model, produce short clips, or check whether 44kHz output matters for your use case
- Self-host when you want Zonos’s voice cloning, need volume beyond short clips, or want to fine-tune
- Each clip here is capped at 400 characters by the hosted deployment — for long text, Kokoro on the same free tier goes to 5,000
Questions
Frequently asked
Keep going