Case Study Deskcasestudydesk.com
Enterprise AISourced

Case Study: Writer serves custom 70B LLMs with 60% higher throughput on Baseten

Writer Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Writer
Industry
Enterprise AI
Challenge
Serving custom 70B domain-specific LLMs in production
Headline result
60% higher TPS and 35% lower cost per million tokens

Key results

60%
Higher tokens per second
FP16 on four A100 GPUs
23%
Lower time to first token
35%
Lower cost per million tokens
70B
Parameter custom LLMs served

The challenge

Writer builds domain-specific Palmyra LLMs for enterprises, including 70-billion-parameter models such as Palmyra-Med-70B and Palmyra-Fin-70B for compliance-heavy industries. Serving 70B models requires multiple high-end A100 or H100 GPUs and cutting-edge optimization, detail-oriented inference work that pulled focus from the team's core competence of model training.

The solution

Writer worked with Baseten's model performance and forward-deployed engineers to build TensorRT-LLM model-specific engines for each LLM, compiling specialized CUDA instructions tuned to real-world sequence shapes and batch sizes and using in-flight batching for production serving.

Inference for custom-built LLMs could be a major headache. Thanks to Baseten, we're getting cost-effective high-performance model serving without any extra burden on our internal engineering teams. Instead, we get to focus our expertise on creating the best possible domain-specific LLMs for our customers.

WA
Waseem Alshikh
CTO and Co-Founder, Writer

The results, in context

In a benchmark running the LLMs in FP16 on four NVIDIA A100 GPUs, Writer saw 60% higher tokens per second, 23% lower time to first token, and 35% lower cost per million tokens. Writer surpassed its performance requirements ahead of launching the new models.

Products used

Baseten TensorRT-LLMBaseten Custom model deploymentBaseten In-flight batching