Case Study Deskcasestudydesk.com
SoftwareSourced

Case Study: Notion deploys new frontier models in under 24 hours with evals on Braintrust

Notion Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Notion
Industry
Software
Challenge
Ship the latest frontier models within hours while keeping ~70 engineers aligned on eval quality.
Headline result
Notion reported deploying new frontier models in under 24 hours of release while keeping about 70 engineers aligned on evals in Braintrust.

Key results

<24hrs
To deploy a new frontier model
70
Engineers aligned on evals
80%
Of AI team work based on evaluating feedback and traces

The challenge

Notion set out to give customers access to the latest frontier models as quickly as possible, ideally within hours of release, which required rigorous evaluation at scale. The AI team needed to keep roughly 70 engineers aligned on a shared evaluation practice and to surface narrow, needle-in-a-haystack failures affecting specific segments such as multilingual workspaces across very large LLM traces.

The solution

Notion adopted Braintrust as its evaluation framework, pairing regression evals that catch breakage with frontier evals that measure improvements from new models. The team deployed custom evaluation code for Notion-specific data structures and used Brainstore for performant search across large trace volumes, enabling systematic iteration from evaluation to production.

I first started working with Braintrust on my first day at Notion. I sat down in Braintrust and looked at some of the worst experiences our customers had and tried to understand how we can be better.

SS
Sarah Sachs
AI Modeling Lead, Notion

The results, in context

Notion reported deploying new frontier models in under 24 hours of release while keeping about 70 engineers aligned on evals. The company stated that roughly 80% of what its AI team does is based on evaluating feedback and traces in Braintrust.

Products used

Braintrust Braintrust (evals, logging)Braintrust Brainstore