Case Study: Gamma scales to 70M+ users with up to 80% faster image generation on Baseten
Key results
The challenge
Gamma's AI platform generates presentations, websites, and documents, where image-generation speed directly shapes the user experience. Its early closed-source models were slow, with some taking 10+ seconds for a single high-resolution image, and Gamma wanted the performance and cost efficiency of open-source models without building an internal ML team.
The solution
Gamma partnered with Baseten's forward deployed engineers to benchmark and optimize open-source image models such as SDXL, Flux, and Qwen for production, tuning them for ultra-low latency and higher cost-efficiency so fewer replicas were needed. Baseten handled the model optimization and infrastructure management.
“Baseten's FDE team has effectively been our team of in-house ML inference specialists. By partnering with Baseten, we've been able to scale to over 70 million users and billions of requests. We've never had a need to scale our AI or infrastructure teams, and we haven't made a single hire for either.”
JNJon NoronhaCo-founder and CPO, Gamma
The results, in context
Gamma achieved 30%-80% faster image generation per model and 20% improved efficiency through reduced replica counts, producing more than 3 million images per day. The company scaled to over 70 million users and $100+ million ARR with a 50-person team and no machine learning or infrastructure hires.