· 5 min read

Google Shipped Gemini 3.8 TTS With 2,000+ Voices. Here's What It Actually Costs to Add Voice to a Solo Product.

Google shipped two new text-to-speech models on September 23: Gemini 3.8 Flash TTS and Flash-Lite TTS. They clone a voice from a 30-second sample, take line-by-line direction like a voice actor, stage two-speaker conversations, and currently sit at number one on Hume AI's Voice Design Benchmark. They're available right now through the Gemini API and Google AI Studio. What isn't available yet is a pricing page, and for a solo builder deciding whether to add voice to a product, that gap matters more than the benchmark score.

What the models actually do

The headline capability is voice creation from a natural-language prompt, plus cloning from a real sample with consent verification built in. Beyond that, the models take direction the way you'd brief a voice actor: pacing, acting cues, laughter, sighs, the small interjections that make a synthetic voice sound like it's actually listening rather than reading. They handle dual-speaker scenes, which is the piece that's historically forced people into stitching together two separate TTS calls and hoping the timing lines up. Language coverage runs past 100, including regional variants like Mexican Spanish and Quebec French, with particularly strong benchmark performance in Japanese, Brazilian Portuguese, Vietnamese, Arabic, and Hindi.

On the benchmark side, the models hold the top overall spot on Hume AI's Voice Design Benchmark at 71.4, and lead the Voice Arena leaderboard in multiple languages. I'd treat any single benchmark number with some skepticism, benchmarks get gamed and Hume's own scoring methodology isn't something I've independently audited, but the combination of a top Voice Design Benchmark score and separate Voice Arena wins across languages is a stronger signal than either alone.

The part that actually matters for a solo build

Here's the calculation every indie builder runs the moment "add voice" comes up: ElevenLabs for the voice, a queue or worker to manage generation, storage for the audio files, and a separate bill to track against usage. That's not a huge stack, but it's another vendor, another API key, another line item, and another thing that breaks at 2am. If Gemini 3.8 TTS is genuinely competitive on quality, folding voice into the same API you're likely already calling for text generation removes an entire vendor relationship, not just a feature gap.

But Google hasn't published pricing for these models yet, and that's not a small omission. Gemini's existing per-token pricing on text models has moved before, and audio generation pricing across the industry runs anywhere from cheap per-character rates to expensive per-second billing that punishes long-form content. Until there's a number, you genuinely cannot model whether this replaces a $22-a-month ElevenLabs starter plan or costs more than what you're trying to route around. I'd hold off building a cost projection into a pricing page or an internal budget until Google actually publishes a rate card, no matter how good the demo audio sounds.

There's also the enterprise-first rollout pattern to watch. The consumer-facing pieces, Gemini Notebook and Google Vids, are live now. The enterprise API is "coming soon." That ordering, developer API first, enterprise API later, is normal, but it also means the free-tier and rate-limit story for a solo builder testing this in a side project isn't settled yet either.

Where I'd actually apply this

If you're currently paying for a dedicated TTS vendor and your Gemini usage is already meaningful, this is worth a real evaluation the moment pricing lands, not before. Voice cloning and multi-speaker staging are exactly the features that turn a decent podcast tool, an audiobook generator, or a customer-facing voice assistant from "clearly synthetic" into something a listener will actually sit through. If you've been holding off on a voice feature because juggling a second AI vendor felt like more infrastructure than the feature was worth, this is the update that might change that math, once the number exists to do the math with.

The honest counter-take: benchmark-topping launches from major labs come with marketing incentives to look best on the specific tests they cite, and Google choosing to lead with Hume's benchmark rather than a broader, independently-run comparison is worth noticing. I'd want to hear real-world voice quality reports from builders who've shipped with it before treating "number one on Voice Design Benchmark" as the whole story.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts