With the recent Telstra outage last week—reportedly lasting about 12 hours and traced back to a single faulty NTP server that effectively sent parts of the network back 20 years—it was a painful reminder that even highly redundant networks can still harbor significant single points of failure. Ouch... especially for anything relying on security certificates. Anyone who has managed large-scale, mission-critical infrastructure knows that troubleshooting these types of outages is anything but easy. Redundancy doesn't eliminate complexity. When your NOC is fielding thousands of customer calls, executives are demanding status updates, and engineering teams are peeling back layers of overlays, routing protocols, and dependencies, diagnosing the root cause becomes incredibly intense—especially during a customer-facing outage with real revenue impact. It got me thinking: how many of us know there are significant single points of failure lurking in our own environments? I can think of two in a network I support today (which will remain nameless). So why do organizations knowingly leave these risks in place? The reality is that you can spend millions chasing the holy grail of five nines (99.999%) availability, but very few organizations have an unlimited budget. At some point, every engineering team has to make trade-offs based on cost, complexity, operational risk, and return on investment. Perfect networks don't exist. So I'm curious: What factors influence the decision to leave known single points of failure unaddressed?How do you determine when the cost and complexity of eliminating a risk outweigh the likelihood and impact of failure?How do you validate that your architecture is "good enough"?Does your NOC provide meaningful feedback that influences future network designs—such as recurring outages, convergence issues, scalability concerns, or operational pain points? I've always believed one of the biggest contributors to operational stability is standardization (aka deterministic). The fewer "snowflake" designs we build, the easier networks become to operate, troubleshoot, and evolve. Interested to hear how others approach these decisions. /vrode outages anarchist & masochist