Remix.run Logo
proxysna 4 hours ago

Wider audience started to use status pages because the service unreliability became so much more noticeable than before and not the other way around. I never had to use a status page for bear blog or protonmail because i never had and issue with it or just i never noticed.

I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.

Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.

Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".

Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.

SoftTalker an hour ago | parent | next [-]

> 1% is 3.5 days

I think that illustrates the author's point quite well. 99% uptime sounds good, but when you think about a 3+ day outage that doesn't sound very good. Imagine Facebook or TikTok being down for 3 days.

Of course most of the time it's not all one outage, but a bunch of short ones. Still, it might communicate the impact better, especially depending on the argument you're trying to win.

rdmuser 4 hours ago | parent | prev | next [-]

Yeah I'm at a point where I have a bookmark folder of status pages mostly for very large orgs because I've been using those pages relatively regularly. This is not something I felt the need for historically.

fmbb 2 hours ago | parent | next [-]

Nothing beats downdetector.com anyway. Always quicker!

hinkley an hour ago | parent [-]

Downdetector also doesn't have a motivation to lie.

I've yet to find a status page that wasn't lying about the actual status.

Also 97% up is bullshit for the 3% of people who are offline.

Saucelabs was doubly bad for this because I'm absolutely certain based on traces that they had some sort of demux bug where they would send events from their tunnel to the wrong job. I could see it in the logs that a test timeout was often the cause of an event firing that was looking for something that never happened, because the event immediately preceding it in the script was never fired. Which meant it was either dropped or went somewhere it shouldn't.

Then it stopped one day and there was nothing in their release notes about it. Lies compounded by further lies.

That's just the most memorable example I have. Stuff like this happens all the time and with many services it plays out the same. There's a perverse incentive not to be transparent about problems with the service, so the status pages play down the intensity of the situation.

gplk 2 hours ago | parent | prev [-]

Same ! To such a point I ended up installing a menubar app (like vitalsbar). Never felt the need to have a realtime overview to be able to work...

fbd_0100 4 hours ago | parent | prev | next [-]

Just two weeks ago protonmail had a massive outage related to data center cooling failure.

nwallin 2 hours ago | parent | prev [-]

Microsoft has absolutely gone to shit in the past ~year. Github, Teams, Windows, Azure, Exchange, it's all been fucking trash. Github was running at 80% uptime for a few months, and if they're telling me they're at 98% or whatever now, then they're cooking the books, period. There's no way they're even that reliable.

What's going on at Microsoft? Are they just copy-pasting their github issue reports into copilot and hitting send it without doing code reviews?

fingerlocks 2 hours ago | parent | next [-]

Everyone got laid off. Massive cost cutting across all orgs. Engineers are now evaluated on AI usage and pull request frequency, and not bugs fixed or performance improvements.

VirusNewbie 2 hours ago | parent | prev [-]

About four years ago I did interview loops at GCP, Netflix, and Azure at the same time. The latter was a “hiring event” so all my interviewers were from different teams, either managers or TLs.

It was the interview equivalent of the multi-headed dragon meme, where the last one looks absolutely stupid. The contrast was insane, microsoft was an absolute shit show compared to the other two companies in terms of talent, personality, organization and more.