Architecture & Incident Reports

Applied Systems Engineering Blog

Technical post-mortems, benchmark telemetry, and incident runbooks authored by our principal consulting engineers.

Featured Architecture Post8 min read

Anatomy of a Production Kubernetes Cascading Failure: Post-Mortem & Fix

A step-by-step breakdown of how a single misconfigured liveness probe triggered an avalanche of node evictions and a 38-minute outage across three availability zones.

By Jordan Reeves (Principal Cloud Architect)
August 12, 2026
Live Incident Takeaways
  • Decouple liveness probes from downstream state checks
  • Enforce PodDisruptionBudgets on multi-AZ deployments
  • Add exponential backoff jitter on reconnection storms
Read Full Post-Mortem
Knowledge Vault

All Engineering Runbooks & Articles

Filter by technical discipline or search specific error scenarios and architecture patterns.