Case Study Deskcasestudydesk.com
Text-to-Speech AISourced

Case Study: Speechify cuts TTS cost per million characters 44% on Baseten

Speechify Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Speechify
Industry
Text-to-Speech AI
Challenge
Self-managed GPU fleet slowed model launches and raised latency
Headline result
44% lower cost and 4.5x faster cold starts

Key results

44%
Lower cost per million characters
30-50%
Lower p99 inference latency
across SIMBA TTS family
4.5x
Faster cold starts
940
GPUs retired
across 18 zones

The challenge

Speechify ran its text-to-speech models on a self-managed stack that peaked at roughly 1,500 GPUs across 18 zones in 7 regions, serving more than 60 million users and synthesizing over 161 billion characters a month. Infrastructure complexity meant a new model could take days of platform-team work before it shipped, and cold starts ran around 9-11 minutes.

The solution

Speechify migrated to Baseten, adopting rolling deploys with fully automated CI/CD, traffic-based autoscaling, multi-region inference, and the open-source Truss tool for self-service model deployment. Baseten now hosts more than 10 production model deployments across the SIMBA TTS family and related models.

The reason we came to Baseten in the first place was the complexity of managing our own infrastructure. Our priority is continuing to deliver the best TTS platform for our 60M+ users, and we didn't want our inference infrastructure to stand in the way of that.

KK
Kai Krause
VP of Engineering and AI, Speechify

The results, in context

Cost per million characters dropped 44% while platform traffic grew 7% over the same period, and the new SIMBA 3.0 vLLM models ran at 76 ms p50 time-to-first-byte. Across the TTS family, p99 inference latency fell 30-50% and replica startup became 4.5x faster, while Speechify retired roughly 940 self-managed GPUs across 18 zones.

Products used

Baseten TrussBaseten Rolling deploysBaseten Multi-region inferenceBaseten Traffic-based autoscalingBaseten Baseten Inference Stack