Answers

How much does AI text-to-speech cost per script?

By BananaBanana TeamPublished Last checked

Short answer

One cent per started 200 characters. A 1,000-character script — roughly 150 spoken words — costs $0.05, and a 60-second voice-over usually lands under $0.10. Thirty voices are available, single speaker or a two-person dialogue, returned as a 24 kHz WAV in the same response. You are charged only if the audio is produced.

Price
$0.01 per started 200 characters
1,000-character script
$0.05
Model
Gemini 3.1 Flash TTS
Voices
30, single speaker or exactly two for dialogue
Output
WAV, 24 kHz mono, returned synchronously
Failed request
Not charged
Billing counts started blocks, not characters, so 201 characters cost the same as 400. Nothing is charged when the audio is not produced.

Facts on this page were checked against the live platform on .

01In detail

How does the per-character billing round?

Up, per started block of 200 characters. A 210-character line is billed as two blocks — $0.02 — the same as a 400-character one. For anything longer than a sentence this rounding is noise, but it does mean that splitting a script into many tiny calls is more expensive than sending it in one piece.

Characters are counted on the text you send, including spaces and punctuation. Style directions written into the text count too, which is a fair trade because they are what makes the delivery sound intentional rather than read aloud.

What do you get for a cent?

Roughly 30 spoken words per block, so a cent buys about twelve seconds of speech. A typical 30-second social voice-over is $0.02–$0.03; a two-minute explainer narration is around $0.10.

The output arrives in the same response — there is no job to poll, unlike images and video. That makes speech the easiest thing to script: send text, receive a WAV.

Dialogue is one call, not two. Assign two speakers with their own voices and the model performs the exchange with the timing between the lines intact, which is what makes it sound like a conversation rather than two takes edited together.

How does it fit a video workflow?

It is the cheap half of a UGC-style ad. Veo 3.1 with native audio roughly doubles the price of a clip; a silent clip plus a generated voice-over is usually cheaper and gives you a script you control word for word. Gemini Omni Flash is the exception — its audio is always included in the per-second rate and lip-syncs, so for talking heads it is worth generating in-model.

Voice-over is also the part you will iterate most. At a cent per 200 characters, ten takes of a 30-second read cost less than a single Veo generation.

02Go deeper

Where this is documented

03Guides

Longer reads on the same thing

04Also asked

More on this question

Can I control tone and pacing?

Yes — through style instructions in the text and a language code. There is no separate emotion slider; the direction is written in.

How many speakers can one call have?

One, or exactly two for a dialogue. Three or more voices means separate calls stitched in an editor.

Is speech available over MCP?

Yes, as a paid tool that answers synchronously with a signed WAV URL. It is also one of the three tools available over x402 for agents paying per call.

Does speech show up in my generation history?

Speech made in the web studio does. TTS requested through MCP is logged against the API key rather than the visual history.

05Related

Questions next door

Try it on your own prompt

New accounts start with $0.20 of balance — no card, nothing expires.