Nobody monitors the monitor: Chaos engineering for the security team's own infrastructure

No ratings

Presented at BSides Tallinn 2026 by

We ask a lot of questions about other people's infrastructure resilience. Turns out we'd never asked them about our own. This is the story of what happened when security team finally did. We sat down to write a Business Continuity Plan for our security monitoring platform — a SIEM ingesting 60+ log sources, running detections across infrastructure serving millions of users across 50 countries — and discovered that "untested" covers a lot more ground than we'd assumed. Some failure modes were obvious. The interesting ones weren't: the silent degradation scenario where the platform stays technically up but detection rules quietly stop running and no alert fires; the compounding case where a routine outage overlaps with an active incident and your tolerable downtime drops from days to hours; the log volume spike that leaves you triaging a live attack with an increasingly incomplete picture — and no indication that the picture is incomplete. Attendees will leave with: - Why your SIEM needs a BIA, not just an SLA. The difference between "it should recover in 4 hours" and "here's what breaks if it doesn't." - The failure scenarios that don't look like failures. Silent degradation and partial log loss are harder to detect - and more dangerous - than a clean outage. - A chaos test list for security infrastructure. What to test and what we found when we actually ran pieces of it.