Voice Engines & Pricing

The AI speech engines available on Speech Generation, what each costs in credits, and what's included free.

2026/07/27

Overview

Speech Generation runs several text-to-speech engines side by side. Free engines cover everyday narration at no cost; paid engines are there when you need studio-grade delivery, more emotional range, or broader language support.

You pick an engine by choosing a voice — every voice belongs to one tier, and the credit cost is shown before you generate.

Free tier

Free voices cost 0 credits. They're limited by a daily character allowance rather than by credits.

GuestsFree account
Characters per day5,00010,000
Characters per generation1,0003,000
MP3 / WAV downloadYesYes
Commercial licenseYesYes

Engines: Kokoro-82M and Chatterbox — both open-weight models with permissive licenses. Kokoro is English-strong and remarkably natural for its size; Chatterbox covers more languages and has more expressive delivery. You can try Kokoro directly on the Kokoro TTS online page.

Paid engines are billed in credits at the per-1,000-character rates below, charged in proportion to what you actually send — a 200-character line costs about a fifth of a 1,000-character one, with a 1-credit minimum. Short experiments stay cheap. Credits come from one-time packs and never expire — there's no subscription and nothing resets monthly.

TierEngineCredits per 1,000 charactersBest for
NaturalFish Audio s13Everyday narration with a noticeable quality lift over free
PremiumMiniMax speech-2.8-turbo10Client work, ads, polished video voiceover — and the only paid tier with WAV and speed control
UltraElevenLabs Multilingual v220Maximum realism, emotional range, multilingual delivery

A 1,000-character script is roughly 60–80 seconds of speech, so 3 credits ≈ a minute of Natural-tier audio.

What's included on every tier

  • Commercial license. Audio you generate is yours to use commercially — including on monetized YouTube, TikTok, and client projects — on the free tier as well as paid.
  • MP3 export, plus WAV on every tier except Ultra — ElevenLabs' endpoint only returns MP3, so the WAV option is unavailable there rather than silently ignored.
  • 18 languages, with dedicated pages for the most-requested ones.
  • No watermarks and no forced attribution.

Choosing an engine

A practical approach: draft with the free tier, and only upgrade the takes that need it.

  • Free is genuinely good enough for study material, drafts, internal videos, and a lot of published short-form content.
  • Natural is the sweet spot for regular publishing — better prosody, still cheap.
  • Premium / Ultra earn their cost when the voice is the product: ads, client deliverables, long-form narration, or anything where a flat read would undercut the work.

Every engine is available in the generator on the home page, so you can run the same script through several and compare before spending anything.

Failed generations

If a generation fails, the credits are refunded automatically and the failure is recorded in your history with the reason. Free-tier character allowance is returned the same way.

Coming soon

API access is in progress — see the TTS API page for status and to join the waitlist. Voice cloning is planned for a later release.