Narration experimentSep 24, 2026

What does it cost to read a rabbit aloud?

by Aliyah Daleen

Reading aloud is an unusually unforgiving place to save money. A voice that sounds convincing for a greeting can become tiring across a chapter. A misplaced pause can change the sense of a sentence. Still, narration is something a reader might use for hours, so a small difference in the price of each paragraph eventually matters. We wanted something more useful than a list of prices: the same words, spoken by the models we are considering, with the bill beside the play button.

We used the third paragraph of Chapter I of Lewis Carroll’s Alice’s Adventures in Wonderland. It gives the voices a parenthetical aside, an exclamation, a long sentence and a rabbit in a hurry. Every request contains the same passage below. These are single takes, with their settings exposed, so you can listen for the differences without taking our word for how any of them sounds.

There was nothing so very remarkable in that; nor did Alice think it so very much out of the way to hear the Rabbit say to itself, “Oh dear! Oh dear! I shall be late!” (when she thought it over afterwards, it occurred to her that she ought to have wondered at this, but at the time it all seemed quite natural); but when the Rabbit actually took a watch out of its waistcoat-pocket, and looked at it, and then hurried on, Alice started to her feet, for it flashed across her mind that she had never before seen a rabbit with either a waistcoat-pocket, or a watch to take out of it, and burning with curiosity, she ran across the field after it, and fortunately was just in time to see it pop down a large rabbit-hole under the hedge.
Lewis Carroll, Alice’s Adventures in Wonderland, Chapter I
141 words · 733 characters · Public-domain source text

The reference is Green Room’s current ElevenLabs Flash v2.5 configuration, using Jessa. Alongside it are SpeechifyAI Simba 3.2, Inworld Realtime TTS-2 Flash, Murf Falcon 2, and Google’s Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS. The Google pair both use Gacrux: identical input, voice and PCM output settings, with only the model changed and no extra delivery instruction. That makes the pair a closer comparison than the other voices, although it is still just one generation each.

ElevenLabsCurrent app

Flash v2.5

Voice: Jessa

Sample cost
$0.0367
Estimated usage
Per audio hour
$3.02
Projected from this take
Audio duration
43.70 s
Request time
1.92 s
Settings, timing & cost basis

Estimate: published ElevenAPI Flash rate $0.05/1k input characters (2026-09-24). Not an invoice charge; character-cost response header retained separately.

Alignment: Character timing captured

Playback RMS: -23.10 dBFS · Sample peak: -2.01 dBFS

{
  "model_id": "eleven_flash_v2_5",
  "voice_id": "yj30vwTGJxSHezdAGsv9",
  "output_format": "mp3_44100_128",
  "voice_settings": "provider defaults, exactly as current app"
}
SpeechifyAIComparison take

Simba 3.2

Voice: Imogen

Sample cost
$0.0073
Estimated usage
Per audio hour
$0.78
Projected from this take
Audio duration
33.88 s
Request time
6.15 s
Settings, timing & cost basis

Estimate at Starter $10/M billable characters; returned billable count. Excludes subscription/free allowance.

Alignment: Speech marks captured

Playback RMS: -23.10 dBFS · Sample peak: -7.15 dBFS

{
  "voice_id": "imogen_32",
  "model": "simba-3.2",
  "output_format": "wav_48000"
}
InworldComparison take

Realtime TTS-2 Flash

Voice: Cordelia

Sample cost
$0.0110
Estimated usage
Per audio hour
$0.87
Projected from this take
Audio duration
45.48 s
Request time
6.97 s
Settings, timing & cost basis

Estimate at on-demand $15/M using returned billed count. No volume discount.

Alignment: Word timestamps captured

Playback RMS: -23.10 dBFS · Sample peak: -5.88 dBFS

{
  "voiceId": "Cordelia",
  "modelId": "inworld-tts-2-flash",
  "audioConfig": {
    "audioEncoding": "WAV",
    "sampleRateHertz": 48000,
    "speakingRate": 1
  },
  "temperature": 1,
  "enhanceGeneration": true,
  "timestampType": "WORD"
}
GoogleComparison take

Gemini 3.8 Flash-Lite TTS

Voice: Gacrux

Sample cost
$0.0081
Measured billing
Per audio hour
$0.70
Projected from this take
Audio duration
41.60 s
Request time
11.75 s
Settings, timing & cost basis

Measured: OpenRouter generation receipt total_cost; excludes credit-purchase fees/tax.

Alignment: Not captured in this run

Playback RMS: -23.10 dBFS · Sample peak: -6.64 dBFS

{
  "model": "google/gemini-3.8-flash-lite-tts",
  "voice": "Gacrux",
  "response_format": "pcm"
}
GoogleComparison take

Gemini 3.8 Flash TTS

Voice: Gacrux

Sample cost
$0.0125
Measured billing
Per audio hour
$1.04
Projected from this take
Audio duration
43.12 s
Request time
14.07 s
Settings, timing & cost basis

Measured: OpenRouter generation receipt total_cost; excludes credit-purchase fees/tax.

Alignment: Not captured in this run

Playback RMS: -23.10 dBFS · Sample peak: -5.57 dBFS

{
  "model": "google/gemini-3.8-flash-tts",
  "voice": "Gacrux",
  "response_format": "pcm"
}
MurfComparison take

Falcon 2

Voice: Theo

Sample cost
$0.0073
Estimated usage
Per audio hour
$0.58
Projected from this take
Audio duration
45.64 s
Request time
3.48 s
Settings, timing & cost basis

Estimate at published $10/M input characters; no returned billing receipt.

Alignment: Not captured in this run

Playback RMS: -23.10 dBFS · Sample peak: -4.08 dBFS

{
  "voiceId": "Theo",
  "model": "falcon-2",
  "locale": "en-UK",
  "style": "Narration",
  "format": "WAV",
  "sampleRate": 24000,
  "channelType": "MONO",
  "rate": 0,
  "pitch": 0
}

Starting a sample pauses the others. All published samples use the same listening format: mono MP3 at 96 kbps and 24 kHz, converted from the providers’ original outputs. We matched whole-recording RMS volume, including pauses, to -23.10 dBFS using constant gain. The decoded MP3s are within 0.02 dB of that target, with at least 1 dB of sample-peak headroom. There is no dynamic compression or limiting; RMS matching does not guarantee identical perceived loudness. Request time is the full request and download, not the delay until a listener could first hear streamed speech.

On this paragraph, Murf Falcon 2 has the lowest recorded projection at $0.58 per audio hour, compared with $3.02 for our current ElevenLabs setup. That is 81% lower on the cost bases shown below. This is a cost observation about these takes; it does not establish a listening-quality winner.

One paragraph, six models · USD · 2026-09-24
Model Sample Audio hour Cost basis
ElevenLabs
Flash v2.5
$0.0367 $3.02 Estimated usage
SpeechifyAI
Simba 3.2
$0.0073 $0.78 Estimated usage
Inworld
Realtime TTS-2 Flash
$0.0110 $0.87 Estimated usage
Google
Gemini 3.8 Flash-Lite TTS
$0.0081 $0.70 Measured billing
Google
Gemini 3.8 Flash TTS
$0.0125 $1.04 Measured billing
Murf
Falcon 2
$0.0073 $0.58 Estimated usage

The two cost labels matter. “Measured billing” means a returned OpenRouter generation receipt. “Estimated usage” means a published provider rate applied to the recorded character count, with the exact assumption in the card. An estimate is not an invoice: subscriptions, included allowances and negotiated rates can change what someone pays. We have not treated a missing price as zero, or counted temporary allowances as the lasting cost of running narration.

The hourly column uses each recording’s actual duration: sample cost multiplied by 3,600, divided by its length in seconds. This is useful for thinking about listening, but pace is a confounding factor. A slower reading produces more audio minutes from the same text and can therefore look cheaper per hour without being cheaper per book. The sample-price column holds the text constant; the hourly column tells us what that particular performance would cost at its particular pace.

There is another requirement that a play button cannot show. Green Room needs to know which words are being spoken so the reading highlight can follow along. Audio alone leaves that connection missing. Each card records the alignment evidence returned by its request. A good-sounding model without usable timestamps would need an additional alignment step, and its cost and reliability would belong in the comparison too. A provider advertising timestamps somewhere in its API is not enough; the narration route we actually use must supply them.

We have kept the app’s narration model unchanged while making this comparison. The next decision needs longer passages, dialogue, difficult names and checks that no words disappear or repeat. Voice preference also deserves room: an economical take that you would stop listening to after ten minutes has missed the point. For now, the recordings and their costs give us a concrete place to start choosing.

Captured 2026-09-24. All prices are USD. Settings and provider sources are linked with each recording. Source passage SHA-256: fe13915de3945aff9a34fd49a1b093e065ba6943f138fc5af530094e27865843.

← All devlog posts

The cast has been waiting for you.

Join the waitlist — beta invites go out in order, and the devlog keeps you honest company until then.