Case Study: Trilogy ran billion-token workloads at ~1/5 the cost with Fireworks AI
Key results
The challenge
Trilogy wanted to validate whether open-weight models could match proprietary systems for enterprise agentic workloads while significantly lowering inference cost.
The solution
Trilogy used Fireworks AI as a unified inference layer to run billion-token-scale internal and agentic workloads, including via its OpenSymphony system.
“Open models in the last few months have now reached parity with proprietary models at an order of magnitude less pricing.”
LGLeonardo GonzalezVP, Trilogy AI Center of Excellence
The results, in context
Trilogy reported open-weight models delivering inference at roughly one-fifth the cost of proprietary systems for comparable workloads, while sustaining billion-token-scale agentic workflows. In production it observed a 93.6% prompt cache hit rate, 150 tokens per second throughput, and 75,000 tokens per request.