Sixty Automated Jobs Kept Running While the System Was Down
A storage controller failed on a Tuesday afternoon and took our order system down for five hours. The...
Tag archive
A storage controller failed on a Tuesday afternoon and took our order system down for five hours. The...
Our order capture service began failing intermittently at twelve minutes past eight on a Tuesday. It...
Your incident comms are worse than your incident During an incident, engineers focus on...
A nightclub ejection moved a conflict to a parking lot with no cameras, no staff, no protocol. Here's the ops failure and what a real perimeter system looks like.
A 4-field post-incident review template small IT teams can finish in 15 minutes — with one guardrail shipped per incident.
The 5-Minute Rule That Kills Paging Burnout Every escalation tree fails the same way:...

Originally published at norvik.tech Introduction Explore the incident-first strategy...

Per-seat pricing in on-call and incident management tools tracks headcount, not cost or usage. Why vendors use it and what it does to your rotation.
Nine rows, three tiers, named owners, and a one-sentence "up means" per tier-1 service — the page that incident severity, status components, and capacity effort all secretly depend on.
When an unruly-passenger report becomes federal evidence overnight, your logging schema matters. Here's what transit security ops need to capture—and why.
Post-mortems don't fail in the meeting. They fail in the two weeks after, when the action list...
A Columbia officer died on a domestic call tied to a restraining order. Here's the information-chain failure every security ops builder needs to understand.