Case Study: Cresta reported 100x lower cost per inference vs GPT-4 with Fireworks AI
Key results
The challenge
Cresta runs real-time contact-center AI where low latency is critical, and needed to deploy many fine-tuned model variants on private enterprise data without incurring the cost of serving each independently.
The solution
Cresta used Fireworks AI's Multi-LoRA capability to serve multiple LoRA adapters through a single base model, fine-tuning cutting-edge base models for RAG-powered tasks.
“Fireworks' Multi-LoRA capabilities align with Cresta's strategy to deploy custom AI through fine-tuning cutting-edge base models. It helps unleash the potential of AI on private enterprise data.”
TSTim ShiCo-Founder and CTO, Cresta
The results, in context
Cresta reported a 100x cost reduction per inference unit compared with GPT-4 by serving multiple LoRA adapters through one base model, and said its fine-tuned variants consistently outperform GPT-4 on RAG-powered tasks. The company cites low-latency, high-throughput serving as key for its real-time applications.