With the recent Telstra outage last week—reportedly lasting about 12
hours and traced back to a single faulty NTP server that effectively
sent parts of the network back 20 years—it was a painful reminder that
even highly redundant networks can still harbor significant single
points of failure. Ouch... especially for anything relying on security
certificates.
Anyone who has managed large-scale, mission-critical infrastructure
knows that troubleshooting these types of outages is anything but easy.
Redundancy doesn't eliminate complexity. When your NOC is fielding
thousands of customer calls, executives are demanding status updates,
and engineering teams are peeling back layers of overlays, routing
protocols, and dependencies, diagnosing the root cause becomes
incredibly intense—especially during a customer-facing outage with real
revenue impact.
It got me thinking: how many of us know there are significant single
points of failure lurking in our own environments?
I can think of two in a network I support today (which will remain
nameless). So why do organizations knowingly leave these risks in
place?
The reality is that you can spend millions chasing the holy grail of
five nines (99.999%) availability, but very few organizations have an
unlimited budget. At some point, every engineering team has to make
trade-offs based on cost, complexity, operational risk, and return on
investment. Perfect networks don't exist.
So I'm curious:
What factors influence the decision to leave known single points of
failure unaddressed?How do you determine when the cost and complexity
of eliminating a risk outweigh the likelihood and impact of failure?How
do you validate that your architecture is "good enough"?Does your NOC
provide meaningful feedback that influences future network designs—such
as recurring outages, convergence issues, scalability concerns, or
operational pain points?
I've always believed one of the biggest contributors to operational
stability is standardization (aka deterministic). The fewer "snowflake"
designs we build, the easier networks become to operate, troubleshoot,
and evolve.
Interested to hear how others approach these decisions.
/vrode
outages anarchist & masochist