Reading aloud is an unusually unforgiving place to save money. A voice that sounds
convincing for a greeting can become tiring across a chapter. A misplaced pause can
change the sense of a sentence. Still, narration is something a reader might use for
hours, so a small difference in the price of each paragraph eventually matters. We
wanted something more useful than a list of prices: the same words, spoken by the models
we are considering, with the bill beside the play button.
We used the third paragraph of Chapter I of Lewis Carroll’s
Alice’s Adventures in Wonderland. It gives the voices a parenthetical aside, an
exclamation, a long sentence and a rabbit in a hurry. Every request contains the same
passage below. These are single takes, with their settings exposed, so you can listen
for the differences without taking our word for how any of them sounds.
There was nothing so very remarkable in that; nor did Alice think it so very much out
of the way to hear the Rabbit say to itself, “Oh dear! Oh dear! I shall be late!”
(when she thought it over afterwards, it occurred to her that she ought to have
wondered at this, but at the time it all seemed quite natural); but when the Rabbit
actually took a watch out of its waistcoat-pocket, and looked at it, and then hurried
on, Alice started to her feet, for it flashed across her mind that she had never
before seen a rabbit with either a waistcoat-pocket, or a watch to take out of it, and
burning with curiosity, she ran across the field after it, and fortunately was just in
time to see it pop down a large rabbit-hole under the hedge.
The reference is Green Room’s current ElevenLabs Flash v2.5 configuration, using Jessa.
Alongside it are SpeechifyAI Simba 3.2, Inworld Realtime TTS-2 Flash, Murf Falcon 2, and
Google’s Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS. The Google pair both use
Gacrux: identical input, voice and PCM output settings, with only the model changed and
no extra delivery instruction. That makes the pair a closer comparison than the other
voices, although it is still just one generation each.
ElevenLabsCurrent app
Flash v2.5
Voice: Jessa
Sample cost
$0.0367
Estimated usage
Per audio hour
$3.02
Projected from this take
Audio duration
43.70 s
Request time
1.92 s
Settings, timing & cost basis
Estimate: published ElevenAPI Flash rate $0.05/1k input characters (2026-09-24).
Not an invoice charge; character-cost response header retained separately.
Starting a sample pauses the others. All published samples use the same listening
format: mono MP3 at 96 kbps and 24 kHz, converted from the providers’ original outputs.
We matched whole-recording RMS volume, including pauses, to -23.10 dBFS using constant
gain. The decoded MP3s are within 0.02 dB of that target, with at least 1 dB of
sample-peak headroom. There is no dynamic compression or limiting; RMS matching does not
guarantee identical perceived loudness. Request time is the full request and download,
not the delay until a listener could first hear streamed speech.
On this paragraph, Murf Falcon 2 has the lowest recorded projection at $0.58 per audio
hour, compared with $3.02 for our current ElevenLabs setup. That is 81% lower on the
cost bases shown below. This is a cost observation about these takes; it does not
establish a listening-quality winner.
One paragraph, six models · USD · 2026-09-24
Model
Sample
Audio hour
Cost basis
ElevenLabs Flash v2.5
$0.0367
$3.02
Estimated usage
SpeechifyAI Simba 3.2
$0.0073
$0.78
Estimated usage
Inworld Realtime TTS-2 Flash
$0.0110
$0.87
Estimated usage
Google Gemini 3.8 Flash-Lite TTS
$0.0081
$0.70
Measured billing
Google Gemini 3.8 Flash TTS
$0.0125
$1.04
Measured billing
Murf Falcon 2
$0.0073
$0.58
Estimated usage
The two cost labels matter. “Measured billing” means a returned OpenRouter generation
receipt. “Estimated usage” means a published provider rate applied to the recorded
character count, with the exact assumption in the card. An estimate is not an invoice:
subscriptions, included allowances and negotiated rates can change what someone pays. We
have not treated a missing price as zero, or counted temporary allowances as the lasting
cost of running narration.
The hourly column uses each recording’s actual duration: sample cost multiplied by
3,600, divided by its length in seconds. This is useful for thinking about listening,
but pace is a confounding factor. A slower reading produces more audio minutes from the
same text and can therefore look cheaper per hour without being cheaper per book. The
sample-price column holds the text constant; the hourly column tells us what that
particular performance would cost at its particular pace.
There is another requirement that a play button cannot show. Green Room needs to know
which words are being spoken so the reading highlight can follow along. Audio alone
leaves that connection missing. Each card records the alignment evidence returned by its
request. A good-sounding model without usable timestamps would need an additional
alignment step, and its cost and reliability would belong in the comparison too. A
provider advertising timestamps somewhere in its API is not enough; the narration route
we actually use must supply them.
We have kept the app’s narration model unchanged while making this comparison. The next
decision needs longer passages, dialogue, difficult names and checks that no words
disappear or repeat. Voice preference also deserves room: an economical take that you
would stop listening to after ten minutes has missed the point. For now, the recordings
and their costs give us a concrete place to start choosing.
Captured 2026-09-24. All prices are USD. Settings and provider sources are linked with
each recording. Source passage SHA-256:
fe13915de3945aff9a34fd49a1b093e065ba6943f138fc5af530094e27865843.