DevOps Interview DropExpert TierScenario+100 XP on read

Describe the last Sev-1 incident you handled.

Core Summary

This question isn't really about the outage, it's about how you behave during one. A good answer shows metric-driven diagnosis, mitigation running in parallel with investigation, real numbers, and a lesson that changed a system rather than blamed a person. Vague heroics score badly; a specific timeline scores well.

Hints

Hint 1: Structure: what broke, how fast you detected it, what you did, how long it took

Hint 2: Specific numbers matter: error rate went from 0.1% to 8%, latency went from 150ms to 2s

Hint 3: Show that you prioritized mitigation while investigation was still ongoing

Reported in interviews at Google, Amazon, Meta, Netflix