Case Study Deskcasestudydesk.com
Digital MediaSourced

Case Study: SimilarWeb cut ELK downtime from one crash a month to zero with Logz.io

SimilarWeb Case StudySourced & dated by Case Study Desk
Key facts · TL;DR
Company
SimilarWeb
Industry
Digital Media
Challenge
SimilarWeb's self-managed ELK stack crashed about once a month, causing multi-hour downtime and diverting engineering time to maintenance.
Headline result
Monthly ELK downtime reduced from one crash to zero

Key results

1 to 0
Monthly ELK downtime incidents
Eliminated recurring Elasticsearch crashes
3-4 hrs
Downtime per incident eliminated
Previous recovery time per Elasticsearch failure

The challenge

SimilarWeb, a digital-media company that analyzes website traffic, ran a self-managed ELK stack that experienced at least one crash per month. Each failure could take three to four hours to recover and could delay log ingestion by around three hours, which was significant when those gaps overlapped with incidents under investigation. Maintaining ELK diverted resources from the broader engineering and development organization.

The solution

SimilarWeb migrated to Logz.io's managed log-management service to centralize logs for debugging and troubleshooting before issues reached end users. The team also uses anomaly detection, drop filters, and sub-accounts to distribute data ownership across its R&D groups.

Logz.io frees a lot of our time to actually maintain and develop the system instead of just monitoring.

OT
Or Tzabary
VP R&D, Production Engineering, SimilarWeb

The results, in context

After migrating, SimilarWeb reduced ELK downtime incidents from at least one per month to zero, eliminating the recurring three-to-four-hour recovery windows. Freed from repeatedly fixing the same infrastructure problems, the team redirected time to delivering product impact.

Products used

Logz.io Log Management