Case Study: Speechify cuts TTS cost per million characters 44% on Baseten
Key results
The challenge
Speechify ran its text-to-speech models on a self-managed stack that peaked at roughly 1,500 GPUs across 18 zones in 7 regions, serving more than 60 million users and synthesizing over 161 billion characters a month. Infrastructure complexity meant a new model could take days of platform-team work before it shipped, and cold starts ran around 9-11 minutes.
The solution
Speechify migrated to Baseten, adopting rolling deploys with fully automated CI/CD, traffic-based autoscaling, multi-region inference, and the open-source Truss tool for self-service model deployment. Baseten now hosts more than 10 production model deployments across the SIMBA TTS family and related models.
“The reason we came to Baseten in the first place was the complexity of managing our own infrastructure. Our priority is continuing to deliver the best TTS platform for our 60M+ users, and we didn't want our inference infrastructure to stand in the way of that.”
KKKai KrauseVP of Engineering and AI, Speechify
The results, in context
Cost per million characters dropped 44% while platform traffic grew 7% over the same period, and the new SIMBA 3.0 vLLM models ran at 76 ms p50 time-to-first-byte. Across the TTS family, p99 inference latency fell 30-50% and replica startup became 4.5x faster, while Speechify retired roughly 940 self-managed GPUs across 18 zones.