Answers
How much does AI text-to-speech cost per script?
Short answer
One cent per started 200 characters. A 1,000-character script — roughly 150 spoken words — costs $0.05, and a 60-second voice-over usually lands under $0.10. Thirty voices are available, single speaker or a two-person dialogue, returned as a 24 kHz WAV in the same response. You are charged only if the audio is produced.
- Price
- $0.01 per started 200 characters
- 1,000-character script
- $0.05
- Model
- Gemini 3.1 Flash TTS
- Voices
- 30, single speaker or exactly two for dialogue
- Output
- WAV, 24 kHz mono, returned synchronously
- Failed request
- Not charged
Facts on this page were checked against the live platform on .
How does the per-character billing round?
Up, per started block of 200 characters. A 210-character line is billed as two blocks — $0.02 — the same as a 400-character one. For anything longer than a sentence this rounding is noise, but it does mean that splitting a script into many tiny calls is more expensive than sending it in one piece.
Characters are counted on the text you send, including spaces and punctuation. Style directions written into the text count too, which is a fair trade because they are what makes the delivery sound intentional rather than read aloud.
What do you get for a cent?
Roughly 30 spoken words per block, so a cent buys about twelve seconds of speech. A typical 30-second social voice-over is $0.02–$0.03; a two-minute explainer narration is around $0.10.
The output arrives in the same response — there is no job to poll, unlike images and video. That makes speech the easiest thing to script: send text, receive a WAV.
Dialogue is one call, not two. Assign two speakers with their own voices and the model performs the exchange with the timing between the lines intact, which is what makes it sound like a conversation rather than two takes edited together.
How does it fit a video workflow?
It is the cheap half of a UGC-style ad. Veo 3.1 with native audio roughly doubles the price of a clip; a silent clip plus a generated voice-over is usually cheaper and gives you a script you control word for word. Gemini Omni Flash is the exception — its audio is always included in the per-second rate and lip-syncs, so for talking heads it is worth generating in-model.
Voice-over is also the part you will iterate most. At a cent per 200 characters, ten takes of a 30-second read cost less than a single Veo generation.
Where this is documented
Longer reads on the same thing
More on this question
Can I control tone and pacing?
Yes — through style instructions in the text and a language code. There is no separate emotion slider; the direction is written in.
How many speakers can one call have?
One, or exactly two for a dialogue. Three or more voices means separate calls stitched in an editor.
Is speech available over MCP?
Yes, as a paid tool that answers synchronously with a signed WAV URL. It is also one of the three tools available over x402 for agents paying per call.
Does speech show up in my generation history?
Speech made in the web studio does. TTS requested through MCP is logged against the API key rather than the visual history.
Questions next door
How much does one AI-generated video actually cost?
Between $0.10 and $4.40 per clip, depending on model, resolution, length and whether the audio track is generated. A 4-second silent 720p clip on Veo 3.1 Lite is $0.10; an 8-second 4K clip with audio on Veo 3.1 is $4.40. Gemini Omni Flash bills per second — $0.10 a second at 720p, so $0.30 to $1.00 per clip. You pay per generation, not per month.
Read the answerIs there a free trial, and what can I do with it?
Every new account starts with $0.20 of real balance, no card and no trial clock. That is six images on Nano Banana 2 Lite, or two 4-second silent clips on Veo 3.1 Lite, or a mix. It is ordinary balance rather than trial credit: it does not expire, and whatever you generate with it is yours to use.
Read the answerCan an AI agent generate images without an account or API key?
Yes, over x402. An agent holding a wallet posts a request, receives HTTP 402 with the exact price for those arguments, pays in USDC on Base and repeats the request with a payment header. No sign-up, no API key, no prepaid balance. Images and speech settle only after the file exists, so a failure costs nothing.
Read the answer