Case Study Deskcasestudydesk.com
Contact Center AISourced

Case Study: Cresta reported 100x lower cost per inference vs GPT-4 with Fireworks AI

Cresta Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Cresta
Industry
Contact Center AI
Challenge
High cost of serving many custom models in real time
Headline result
Multi-LoRA serving on a shared base model cut inference cost 100x

Key results

100x
Lower cost per inference
vs GPT-4, via Multi-LoRA

The challenge

Cresta runs real-time contact-center AI where low latency is critical, and needed to deploy many fine-tuned model variants on private enterprise data without incurring the cost of serving each independently.

The solution

Cresta used Fireworks AI's Multi-LoRA capability to serve multiple LoRA adapters through a single base model, fine-tuning cutting-edge base models for RAG-powered tasks.

Fireworks' Multi-LoRA capabilities align with Cresta's strategy to deploy custom AI through fine-tuning cutting-edge base models. It helps unleash the potential of AI on private enterprise data.

TS
Tim Shi
Co-Founder and CTO, Cresta

The results, in context

Cresta reported a 100x cost reduction per inference unit compared with GPT-4 by serving multiple LoRA adapters through one base model, and said its fine-tuned variants consistently outperform GPT-4 on RAG-powered tasks. The company cites low-latency, high-throughput serving as key for its real-time applications.

Products used

Fireworks AI Fireworks AIFireworks AI Multi-LoRAFireworks AI Fine-tuning