| ▲ | arccy 2 hours ago |
| Thundering herd / cascading outages. You take out a large enough portion of your fleet, and the remaining load overloads your remaining nodes one by one as they restart, so you can never have enough healthy nodes. |
|
| ▲ | erulabs 2 hours ago | parent | next [-] |
| SRE team debates correctness versus availability for the 540th time this year You're both correct, of course! |
| |
| ▲ | atmosx 2 hours ago | parent [-] | | What this guy said :point_up: My personal take-away is this: whatever you choose, make sure it's consistent across services (not serviceA behaves like X and serviceB like Y) and make sure eng teams know _how_ these are configured and what can go wrong. They'll figure out the rest. |
|
|
| ▲ | jaggederest 2 hours ago | parent | prev [-] |
| That's a problem for circuitbreakers on these kinds of actions, not lying on health checks. Something like healthcheck fails -> restart -> healthcheck fails -> restart -> healthcheck fails -> circuit breaker trip, alarm raised, give up until manual intervention or X minutes have passed |
| |
| ▲ | deathanatos an hour ago | parent [-] | | That circuitbreaker exists, by default. It is "CrashloopBackoff", here, and TFA covers it. (& it's an "until X minutes have passed" kind, by default.) | | |
| ▲ | dilyevsky 41 minutes ago | parent [-] | | backoff is only applied to individual pods/containers not across pods. the point is at scale it's easy to get into a situation where it's not possible to recover without (usually manual) full service drain |
|
|