Case Study Deskcasestudydesk.com
MediaSourced

Case Study: Hedra cut infrastructure costs 60% and tripled inference speed on Together AI

Hedra Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
Hedra
Industry
Media
Challenge
Massive-scale video generation strained cost and autoscaling
Headline result
Scaling viral AI video generation with lower cost and faster inference

Key results

60%
total cost savings
vs. prior infrastructure
faster inference
on NVIDIA Blackwell
300×
compute growth supported
GPU usage

The challenge

Hedra generates millions of AI videos monthly, a workload with volatile traffic spikes and heavy GPU demand. Managing its own autoscaling and unit economics at that scale was operationally costly.

The solution

Hedra migrated its video models to Together AI, running on NVIDIA Blackwell architecture with managed autoscaling and a committed base plus incremental flex capacity.

The results, in context

Together AI's page states the migration cut cloud costs by 60%, tripled inference speed on Blackwell hardware, and supported 300× growth in the company's compute usage while maintaining consistent performance and cost.

Products used

Together AI Dedicated EndpointsTogether AI NVIDIA Blackwell GPUsTogether AI Autoscaling