Case Study Deskcasestudydesk.com
Enterprise SoftwareSourced

Case Study: Trilogy ran billion-token workloads at ~1/5 the cost with Fireworks AI

Trilogy Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Trilogy
Industry
Enterprise Software
Challenge
Validate open-weight models for enterprise-scale workloads
Headline result
Open-weight models hit enterprise scale with a 93.6% prompt cache hit rate

Key results

1/5
Inference cost
vs proprietary systems
93.6%
Prompt cache hit rate
in production
150
Tokens/sec throughput
observed
75,000
Tokens per request
production

The challenge

Trilogy wanted to validate whether open-weight models could match proprietary systems for enterprise agentic workloads while significantly lowering inference cost.

The solution

Trilogy used Fireworks AI as a unified inference layer to run billion-token-scale internal and agentic workloads, including via its OpenSymphony system.

Open models in the last few months have now reached parity with proprietary models at an order of magnitude less pricing.

LG
Leonardo Gonzalez
VP, Trilogy AI Center of Excellence

The results, in context

Trilogy reported open-weight models delivering inference at roughly one-fifth the cost of proprietary systems for comparable workloads, while sustaining billion-token-scale agentic workflows. In production it observed a 93.6% prompt cache hit rate, 150 tokens per second throughput, and 75,000 tokens per request.

Products used

Fireworks AI Fireworks AIFireworks AI Serverless inferenceFireworks AI Prompt caching