Case Study: Notion cut AI response latency from ~2 seconds to 350ms with Fireworks AI
Key results
The challenge
Notion set out to deliver AI features to a very large user base without the response times undermining the experience. Baseline model latency of roughly two seconds was too slow for interactive use.
The solution
Notion fine-tuned models served on Fireworks AI to reduce response latency while preserving quality, enabling AI features to be rolled out across its user base.
“By fine-tuning models, we reduced latency from about 2 seconds to 350 milliseconds, significantly improving performance and enabling us to launch AI features at scale.”
SSSarah SachsHead of AI Engineering, Notion
The results, in context
After fine-tuning, Notion reduced AI response latency from about 2 seconds to 350 milliseconds, roughly a 4x improvement, which the company says let it launch AI features at scale. Notion states it serves more than 100 million users.