Describe a major outage you handled. How did you manage it?
Tell it as a timeline, and separate the technical work from the coordination work — the second is what distinguishes senior operations people.
- Detection and initial assessment. How you found out, and how quickly you established scope: who is affected, how badly, and is it getting worse.
- Communication, started early. Notifying stakeholders and setting an update cadence before you have answers. Silence during an outage causes more damage than the outage.
- Restore service before finding the cause. Failing over, rolling back, restarting, or activating a workaround. Root cause analysis is for afterwards.
- The diagnosis and fix, and how you verified recovery rather than assuming it.
- The post-incident review. What the root cause was, what the contributing factors were, and the specific actions taken so it cannot recur.
Note: Naming an incident command structure — someone owning the technical fix, someone owning communication — is a strong signal. So is describing a blameless review, because it shows a culture where people report problems early rather than hiding them.





