Remix.run Logo
alamsterdam 5 hours ago

It's been hours :(

I have sympathy for the on-call team trying to resolve it, most of us have been there done that.

But seems something is systematically going wrong at GH

delduca 4 hours ago | parent | next [-]

> But seems something is systematically going wrong at GH

Yes, we call it: Microslop.

yashap 5 hours ago | parent | prev | next [-]

It's honestly insane how terrible their reliability is. Over the past 4 weeks, 8 days with GitHub Actions outages, many of them multi-hour outages: >3 hours on each of July 9th, July 20th and today, and 1.5 hrs on July 23rd.

Outages happen, but this many outages so close together, and so many of them so major/long lasting, something is systematically wrong for sure. It's been seriously hamstringing our ability to ship code at my company.

alamsterdam 5 hours ago | parent [-]

Mon and Dad (MS) are (generally) pretty solid with uptime.

What is happening at GH?

Rate of change trying to keep up with new challengers? Over-reliance on AI? Engineers trying to debug slop?

yashap 5 hours ago | parent | next [-]

Yeah who knows, would be interesting to hear an inside take if any readers here are also GH devs!

They're at 93.91% uptime over the past 90 days, according to https://mrshu.github.io/github-statuses/ , and that doesn't even include today's outage yet.

A glorious one nine of reliability.

marcprux 4 hours ago | parent | next [-]

I see two nines in there…

warmwaffles 4 hours ago | parent | prev [-]

Don't throw shade, there are two nines in that percentage for now. Miles a part though.

alamsterdam 4 hours ago | parent [-]

911 has one! :)

simoncion 4 hours ago | parent | prev [-]

> What is happening at GH?

In large part, the move from AWS to Azure. Azure's just bad.

PsylentKnight 3 hours ago | parent | next [-]

Based on previous posts I've seen about this, IIRC the timing seems to imply that it has more to do with them being hammered with AI slop than it does with the Azure transition. Who really knows though

alamsterdam 3 hours ago | parent | prev [-]

wasn't that years ago? or was that Skype? hahaha

niwtsol 2 hours ago | parent | prev | next [-]

"most of us have been there done that" - so true. That feeling in your gut when you realize something you just did caused an outage is pretty unique.

tempaccount420 4 hours ago | parent | prev | next [-]

They're one rewrite in Rust away from fixing everything. (jk)

hinkley 3 hours ago | parent | prev [-]

Time to touch some grass. Better use of my time and energy than twisting the remaining things on my todo list today to make more progress on them than I have managed. I should have taken a long lunch but I rebased the hell out of a PR instead.