Case Study Deskcasestudydesk.com
Document AISourced

Case Study: Reducto cuts P90 latency 3x by moving 30+ models to Modal

Reducto Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Reducto
Industry
Document AI
Challenge
Reducto's Kubernetes setup couldn't scale 30+ models independently under bursty document loads.
Headline result
Reducto cut P90 latency 3x and cold boots 83% after migrating from Kubernetes

Key results

3x
P90 latency reduction
83%
Cold boot time reduction
~70s to ~12s
1,000+
GPUs scaled in under an hour
load testing
2 lines
Code to deploy an endpoint
from ~150 lines plus config

The challenge

Reducto ran 30+ production models on Kubernetes that had to scale together despite varied usage patterns, producing unpredictable P90 latency during large bursty uploads of millions of pages. Traffic spikes strained infrastructure and threatened customer SLAs.

The solution

Reducto migrated to Modal, using GPU memory snapshotting, independent per-model scaling, and parameterized functions for customer-specific autoscaling pools. Granular regional compute control gave the team more placement flexibility.

Modal gives us a lot of flexibility to do pretty complex stuff that we wouldn't get with an LLM inference service.

RC
Raunak Chowdhuri
Founder, Reducto

The results, in context

Reducto reported a 3x reduction in P90 latency and an 83% reduction in cold boot times, from roughly 70 seconds to about 12 seconds. It scaled to 1,000+ GPUs in under an hour during load testing and reduced endpoint deployment from about 150 lines plus configuration to 2 lines of code.

Products used

Modal Modal GPU computeModal Serverless inference