Remix.run Logo
legulere an hour ago

One advantage of percentages is that you can multiply them if the outages aren't related. If you rely on 10 systems with a 99% uptime each you end up with roughly 90% uptime in total.

hinkley 22 minutes ago | parent [-]

After a merger we ended up with 2 operations systems and one of them thought the other one were clowns. But they thought everyone were clowns so it was hard for the rest of us to tell. They never actually apologized for being assholes, but their tune changed when they found they had bitten off more than they could chew later and they had to come hat in hand to a bunch of groups to delegate responsibilities back to them.

So a production outage happens, and we go to look at our runbooks to figure out what to do. The Wiki is also down, so no runbooks. What the actual fuck?

So it turns out the alleged clowns put all of our internal infrastructure onto the same SAN system in different partitions. They got a lot less judgy and the other Ops team got a lot more autonomy after that event.

Anyone with an Operational IQ above room temperature knows that you don't put offline and online resources onto the same hardware. Not only do they have separate duty cycles but also offline services can end up taking out the online services in unexpected ways, in part because they are assumed to be a bit sloppy and so they get less operational scrutiny.

I sped up the average runtime of a batch job by quite a humongous amount by throttling the request rate we made to a service that was used by user-facing services. If several unrelated batch jobs kicked off during the wrong time of day, we would start getting circuit breakers opening everywhere, including occasionally production.

I throttled us to use something like 8 or 10% of the official capacity of the service and no more. But I did it as limiting the number of in-flight requests, so that created back pressure if the servers were under heavy load and sending responses slowly, or let us go faster if the servers were returning results faster than usual. That peak shaving saved the other team a load of grief and let them push some capacity planning work down the roadmap to work on other things. And it dropped our failure rate by more than a factor of six. Low enough that we no longer had to babysit that process, which meant we started using it much, much more.