| ▲ | orf an hour ago | |
> Outages happen as a result of the system's design being shoddy in the first place, and an attitude where bugs shipped are acceptable because it's only a single component, rest of the system shouldn't fail. This is correct for some classes of failure but not others. If your authorization system goes down, should you just let any request pass? No. You also need to fail correctly. If we extend the civil engineering analogy, then that would mean fail safely. Two trains derailed in the UK last week. Should we all abandon trains? No? Why do plane crashes happen? That’s the most safety conscious area for software and hardware. Because failures happen. Everyone expects them. Bridges don’t last forever, they know they will fail after a specific time frame and so they design around that. Planes have triplicate systems everywhere. And yet, despite all this, failures still happen. And the fact they know that bridges have specific failure modes and lifespans is the result of a long, long history of bridges failing. | ||