| ▲ | ieie3366 2 hours ago |
| To everyone who is angry: calm down. Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside. If a work tool going down triggers you enough mentally to start angrily ranting online, it’s a sign you need to chill out and focus more on your health |
|
| ▲ | askonomm 2 hours ago | parent | next [-] |
| That's an overly optimistic view of things. For a lot of us GitHub is critical infrastructure, which if it goes down loses us and our customers money. |
| |
| ▲ | datsci_est_2015 2 hours ago | parent | next [-] | | It is interesting, how much money is being lost during this outage? My significant other was just let go from their job as a scapegoat for an organizational error: 3 layers of failure - IC, manager, director, and the IC was let go. The error caused a 7 figure loss for the company that has 10 figures of revenue per year. The manager and director may not see any consequences, though the director will probably be forced out by end-of-year due to incompetence. The new executive has taken to firing employees much more eagerly than their predecessor, like some sort of Jack Welch acolyte. Their firing has put a lot of things into perspective for me. Mostly, fuck "at-will" employment and its negative effect on the American social contract. But also this "angry ranting" online that the original poster was referencing. Not everyone has the privilege to calmly respond to things that directly impact their livelihood. | | |
| ▲ | theappsecguy an hour ago | parent [-] | | Any engineered system where an individual can accidentally cause a 7 figure outage is poorly designed. And engineering leadership that decider to terminate an individual due to such failure (as long as there was no malicious actions) is completely clueless. | | |
| |
| ▲ | Waterluvian 2 hours ago | parent | prev | next [-] | | I'd be curious what the statistics might actually be for people who are directly affected because their business is suffering vs. people affected because their employer's business is suffering. | | |
| ▲ | 6LLvveMx2koXfwn 2 hours ago | parent | next [-] | | Most people's employment prospects are directly correlated with their employers ability to make money. | | |
| ▲ | john_strinlai 2 hours ago | parent [-] | | most people's employers arent firing people over a few hours of github outage. not to mention that any business which could potentially lose enough money that they would need to let go of developers from a github outage should probably already have some business continuity plans in place. |
| |
| ▲ | polycaster 2 hours ago | parent | prev [-] | | For Germany it's roughly 1:11 if you go by the self-employment rate (~8% of the workforce). |
| |
| ▲ | nullify88 an hour ago | parent | prev | next [-] | | GitHub Enterprise Cloud has been chugging along with no issues. I hope your critical infrastructure isn't dependent on a free tier / service. And that you have a business continuity process in place. | | | |
| ▲ | antaviana 2 hours ago | parent | prev | next [-] | | "It's just money. It's made up. Pieces of paper with pictures on it so we don't have to kill each other just to get something to eat". Jeremy Irons in movie Margin Call | | | |
| ▲ | parthdesai 2 hours ago | parent | prev | next [-] | | How is github being down losing you guys money? | | |
| ▲ | dewey 2 hours ago | parent | next [-] | | If you pay developers x money / day and one of their core tools is down for n hours during the day and they spend their money on HN instead that's pretty straight forward to calculate. | | | |
| ▲ | aurareturn 2 hours ago | parent | prev | next [-] | | I was in the middle of a hot fix. Our pipeline goes through Github. | | |
| ▲ | parthdesai 2 hours ago | parent [-] | | That's understandable, but surely there's a way to bypass it? You need to have a breakglass procedures lol | | |
| ▲ | penultimatename an hour ago | parent [-] | | “lol just use a workaround” doesn’t work in an environment with hundreds or thousands of employees coupled with audit, security, and other legal requirements to ship software. | | |
| ▲ | parthdesai 43 minutes ago | parent [-] | | If you're actually bleeding money, you better believe you'll get permissions for a workaround, if you know what you're doing. It all ties back to the OP, where the issue you've might not be as bad as you think. I have been in situations where we have dropped all procedures to push a hot fix because we were actively bleeding money, and in situations where you know there is an issue, and you let it be. |
|
|
| |
| ▲ | phoronixrly 2 hours ago | parent | prev [-] | | Imagine Github being critical infrastructure for you... The ineptitude... | | |
| ▲ | parthdesai 2 hours ago | parent [-] | | We use github quite a bit as well. Github being down currently is not actively losing us money, i.e. having customer impact. |
|
| |
| ▲ | kakacik 2 hours ago | parent | prev | next [-] | | If its so critical why relying on it, and not having ie some mirror or some other way to handle any sort of outage like this. its not like Microsoft is your friend or good business partner, ever. With every single of these enterprise 'cloud' offerings you are giving (almost) complete power over your business/project to somebody else who couldn't care less about your success or failure, you are simply irrelevant for them. I see it at work too, every time critical external systems go down whole bank stops still, just because few bucks were saved yearly on some on-prem servers. Look at it this way, you are learning some important lesson today and finding great area of improvement for resiliency from now on. | |
| ▲ | AlexandrB 2 hours ago | parent | prev | next [-] | | Boomer opinion but trusting third parties to be critical infrastructure, especially with no SLA in sight, will always end in tears. "The cloud" is very convenient, but its providers will never care about your infrastructure or your customers as much as you will. | | |
| ▲ | Maxion 2 hours ago | parent | next [-] | | Older millennial here and I agree. If you rely on Github you should at least be have your processes such that you can work around it. | | | |
| ▲ | otterley 2 hours ago | parent | prev | next [-] | | 3P-maintained infrastructure is what makes civilizations work efficiently. We're not all digging our own wells, generating our own electricity, and burning or burying our own garbage. | | |
| ▲ | kalaksi 2 hours ago | parent [-] | | You are still supposed to be prepared for outages. | | |
| ▲ | otterley an hour ago | parent [-] | | Agreed. So is the issue really then that people were inadequately prepared with backup plans and now they're suffering the consequences? It's not all that different from, say, an AWS region having a service impact. People would rather complain about AWS than prepare and utilize a well-tested recovery plan to shift to a standby region. Oftentimes there's no fallback plan because the business already considered it and decided it was too costly relative to the benefit, but when the incident happens, they still can't help but complain. Humans being humans. |
|
| |
| ▲ | aenis 2 hours ago | parent | prev | next [-] | | Boomer here as well, but I'd add that trusting your own org for critical infra usually also ends in tears. Most everything in IT involves failure, including in well designed systems designed by great engineers. I worked for a few years in an exceedingly well capitalised place which ran everything in their own data centers, money no object, with a truck parked somewhere, ready to go, with a smaller version of our critical infra. We had a serious business-stopping outage once every 18 months or so, every time for fringe reasons one only learns about when trying to run a large data center. Its convenient to blame the cloud and pretend that self-hosting in private sector was so, so great with six nines. | | |
| ▲ | 6031769 2 hours ago | parent [-] | | At this point github is barely managing one nine. | | |
| ▲ | aenis 2 hours ago | parent [-] | | True. I'd be very annoyed if I had anything mission critical on github. While we have everything in the cloud, we do actually self-host that part. |
|
| |
| ▲ | swills an hour ago | parent | prev | next [-] | | Agree. But also, it's affecting everyone equally, whether they have a free personal account or are part of an enterprise account with SLA. Understanding the practical value of an SLA is an interesting problem. | |
| ▲ | raffael_de 2 hours ago | parent | prev | next [-] | | +1 (as a millenial) ... especially given that setting up a git server for non-OSS company code isn't too much of a challenge really. also, no need to self-denigrate this reasonable opinion in preemptive obedience. | | |
| ▲ | otterley an hour ago | parent [-] | | GitHub is way more than just a git repo host. It manages code reviews, merge (pull) requests, and has an entire CI/CD workflow engine in it. Replicating all that is a challenge that most orgs are not up to. |
| |
| ▲ | poisonborz 2 hours ago | parent | prev | next [-] | | How did this become a boomer opinion? It is proved truth thousand times a day. Not that you shouldn't use third parties - but in this industry you can shrink this exposure to the minimum, and have plan B for anything else. | |
| ▲ | ethagnawl 2 hours ago | parent | prev | next [-] | | You're right. However, I'm also of the boomer opinion that you should get what you pay for. "Ranting online" about a service (you pay for) being unavailable is a reasonable reaction. It's not like they have a call center you can dial into for support ... | |
| ▲ | SideburnsOfDoom 2 hours ago | parent | prev | next [-] | | People are already saying things like "We need a plan B in case we urgently need to deploy a fix to production, and GitHub Actions is unavailable again". But in general, it's not feasible to do everything in house. And I don't think that GitHub is devoid of SLA: https://github.com/customer-terms/github-online-services-sla The issue is that they're not achieving two nines uptime in practice. | |
| ▲ | TZubiri 2 hours ago | parent | prev | next [-] | | Boomer who has never had a business? Depending on third party suppliers is business as usual. | | | |
| ▲ | BryanBeshore 2 hours ago | parent | prev [-] | | Truly a boomer opinion. Respect! |
| |
| ▲ | TZubiri 2 hours ago | parent | prev [-] | | And it's not just the user's fault, GH spent a considerable effort in marketing to position themselves as such, see Github Actions and similar junk. |
|
|
| ▲ | jjice 2 hours ago | parent | prev | next [-] |
| > Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code I sympathized with them when they said this a handful of months ago, but then I saw this [0] page that shows how it's been shot for years prior (which tracks with my memory). I feel for the GH engineers that have to deal with this, especially the SREs. I also don't hate the downtime right now, as I'll make a cup of coffee and do something else. I will say though, I did have a hotfix a week or so ago during the Actions outage, which really was a pain. You're right that getting angry and ranting isn't the right reaction here, but I do no give them the LLM load excuse. I don't give them an out for having awful uptime during work hours for a product we pay quite a bit for as an org. [0] https://damrnelson.github.io/github-historical-uptime/ |
| |
| ▲ | twistedpair 2 hours ago | parent | next [-] | | They _are_ the LLM coding agent vendor, and are _owned_ by MS, the ~majority~ biggest shareholder of OpenAI. How can you NOT consider that 20x+ scaling in your capacity roadmap projections, where you are trying to get everyone to use these agents as part of your core OKRs? Or, maybe your #1 IT priority was moving everything to Azure instead ;) | | |
| ▲ | rafram 2 hours ago | parent [-] | | > the majority owner of OpenAI Largest shareholder (27%), not majority. | | |
| |
| ▲ | debarshri 2 hours ago | parent | prev | next [-] | | It looks like right after Microsoft acquisition this started going off. This is the outcome of violating single responsibility principle in business. | |
| ▲ | Topfi 2 hours ago | parent | prev | next [-] | | Github being under the CoreAI division probably also doesn't help the engineers prioritize addressing infrastructure issues and makes using LLM load as an excuse feel self-inflicted. Akin to feeling sorry when a pyromaniacs house burns down... | |
| ▲ | chrisjj 2 hours ago | parent | prev [-] | | They've had a couple of months and still have not brought in sufficient extra capacity?? What's their excuse? |
|
|
| ▲ | eddieroger 2 hours ago | parent | prev | next [-] |
| For the amount of money enterprises are paying to GitHub, there is a reasonable (and contractual) expectation of uptime. I don't think my boss would find it to fun if I bailed work to hit the gym just because GitHub was down. |
| |
| ▲ | iamcoder18 2 hours ago | parent | next [-] | | Yes, but GitHub enterprise has different uptime (https://us.githubstatus.com/posts/dashboard) and actually does have an SLA. Edit: it seems GHE is also affected by the outage and isn't much better. | | |
| ▲ | cyansmoker 2 hours ago | parent | next [-] | | As an enterprise user I can tell you that: - we are getting no benefit beyond getting excuses replies to our emails
- right or wrong, but we depend heavily on GitHub actions so this has a very real impact on us | |
| ▲ | twistedpair 2 hours ago | parent | prev | next [-] | | Since upgrading to Cloud Enterprise... why didn't my GH uptime get any better?
Guess on-prem GHE server is the only way to do that. | | |
| ▲ | joeyhage 2 hours ago | parent | next [-] | | My understanding is the us.githubstatus.com is only if you are using GitHub enterprise with a custom subdomain and data residency in the US because then *some* of the infrastructure is separate from GitHub.com. | |
| ▲ | mh- 2 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | eddieroger an hour ago | parent | prev | next [-] | | Yeah, I know, and it's down. | |
| ▲ | tedivm 2 hours ago | parent | prev [-] | | This, like all of their dashboards, is complete bullshit. As a GitHub enterprise user i am absolutely affected by the outage today, and have been affected by all of the GitHub Cloud outages. There is no separate enterprise platform unless you buy their self hosted server product. |
| |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | leishman 2 hours ago | parent | prev | next [-] |
| > Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code Which means they have a pricing problem. Rate limit non-paying users. |
| |
| ▲ | xnorswap 2 hours ago | parent | next [-] | | Just take away free Actions and it would surely solve a lot of their scale issues. It's absurd that I have dozens of repos, many with GH Actions that run CI, test and then package and push to prod/package managers, and haven't paid GH anything. | |
| ▲ | anoplus 2 hours ago | parent | prev | next [-] | | Right. It is not an excuse and they were notorious for outages before LLM era. Also, LLM can serve Github engineers as well, so we are all playing in the same field. | |
| ▲ | cluckindan 2 hours ago | parent | prev | next [-] | | They already do | | |
| ▲ | rzzzt 2 hours ago | parent [-] | | I can't check the commit history of some OSS projects without hitting a rate limit if I'm not signed in. My request rate is one request per (arbitrary time interval) at that point. |
| |
| ▲ | chrisjj 2 hours ago | parent | prev [-] | | No. It means they have a capacity problem. |
|
|
| ▲ | denysvitali 2 hours ago | parent | prev | next [-] |
| Stop blaming LLMs:
https://damrnelson.github.io/github-historical-uptime/ |
|
| ▲ | sausagefeet 2 hours ago | parent | prev | next [-] |
| Actually, I do have an important fix for a deal we are trying to close that does need to go out. And I am paying GitHub to host this. Once or twice, I can see, but GitHub's SLA is getting worse than just hosting it myself, and that's the entire reason I pay GitHub. |
|
| ▲ | world2vec 2 hours ago | parent | prev | next [-] |
| Geez, sorry that the paid service my company paid for is down and I can't do my work. |
|
| ▲ | sarjann 2 hours ago | parent | prev | next [-] |
| This has to be rage bait, this is a critical piece of infrastructure for many people. Stuff going down can lead to deployments failing and as you mentioned in many cases emergency hotfix's. The idea that we should be fine with this unreliability is just amazing. It's not a mental health issue to have problems when important infra fails. |
| |
| ▲ | john_strinlai 2 hours ago | parent [-] | | >The idea that we should be fine with this unreliability is just amazing. the comment doesnt say you should be "fine" with the unreliability. they are saying people shouldnt get so emotionally worked up over it. which, while i wouldn't phrase it in the way the parent did, i agree with the direction of their point. | | |
| ▲ | sarjann 2 hours ago | parent [-] | | Except he tries justifying it. "Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code" "Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside."
with assumptions like above. If it was, "hey they f*ed up but there's no point having an overly emotional reaction" that would be fine. But he seems to be justifying this. Especially with Github's record up to now of unreliability I think it's completely fair to be annoyed. |
|
|
|
| ▲ | madeforhnyo 2 hours ago | parent | prev | next [-] |
| > due to LLM sloppers pushing large amounts of trash code Which they encouraged by pushing Copilot down everyone's throat |
|
| ▲ | constant_flux 2 hours ago | parent | prev | next [-] |
| GH shouldn't have promoted LLM usage the way they have if they didn't have the infrastructure to support it. Regardless of how you feel about the code quality (which you have zero evidence of), if GH has made commitments to supporting broad LLM usage, they need to back that up with the proper hardware and without whatever fragile SDLC processes they have. |
| |
| ▲ | pamcake 2 hours ago | parent [-] | | > GH shouldn't have promoted LLM usage the way they have No qualifiers necessary. But try arguing with a lawnmower... |
|
|
| ▲ | bhouston 2 hours ago | parent | prev | next [-] |
| > To everyone who is angry: calm down. Github’s servers are constantly on fire Is this rage bait? Isn't Google, Amazon, and all the other services in the world similarly impacted by LLM's? Isn't Claude, OpenAI, etc.? Why can they handle the load but not Github? |
|
| ▲ | s_dev 2 hours ago | parent | prev | next [-] |
| What if we're not angry and just perplexed? Do we calm down further or can we still ask questions? |
| |
|
| ▲ | armSixtyFour 35 minutes ago | parent | prev | next [-] |
| Hard disagree, people are paying GitHub a lot of money. If they are facing an influx of commits from bots maybe they need to start metering commits from those accounts and charge them money. This is a self imposed problem. |
|
| ▲ | malfist 2 hours ago | parent | prev | next [-] |
| Why do you get to tell me how I should feel? |
|
| ▲ | AgentK20 2 hours ago | parent | prev | next [-] |
| what if I do have an emergency hotfix for a production outage though ;-; |
|
| ▲ | _joel 2 hours ago | parent | prev | next [-] |
| It's every other day now. We pay for a service, we expect service. |
|
| ▲ | DharmaPolice 2 hours ago | parent | prev | next [-] |
| This might be naïve but wouldn't the appropriate response be to reduce access/rate limit free/new accounts in order for service to be maintained for everyone else? |
|
| ▲ | stetrain 2 hours ago | parent | prev | next [-] |
| > as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code Github and their parent company are active participants in pushing LLM-driven coding. |
|
| ▲ | phtrivier 2 hours ago | parent | prev | next [-] |
| I understand the need to level-head the discussion and calming everyone. However, your "(you don't)" comment is not going to calm down all the people, who, you know, DO have a hotfix to push now, and DO have an angry customer that could not care less for which part of our infrastructure is breaking _their_ workflow. The only things would calm everyone down is guarantee that Microsoft would be paying for _our_ SLA breach compensation. But they don't. And I don't think anyone is paying their GH bill with a prorata of the number of time the platform was actually available. (I'm also aware that the wording of the contract probably clearly says that you should not use GitHub for anything critical, that Microsoft is only a small startup in their garage, that you can't credibly expect 90% uptime anyway, and that it's all the fault of LLM slop ! Bad LLM slop ! Also, please buy our LLMs to generate more slop, please.) |
| |
|
| ▲ | anonymousab 2 hours ago | parent | prev | next [-] |
| Allowing and accepting the LLM load is a willful choice they made and are making at the expense of their users, including their paying enterprise users. They could easily tighten things up in that regard, and make a choice that is right for their main users at the expense of The MS corpo mandate/mission. It is a choice to do otherwise. |
| |
| ▲ | pamcake 2 hours ago | parent [-] | | Allowing and accepting? More like encouraging and instigating. |
|
|
| ▲ | kilroy123 2 hours ago | parent | prev | next [-] |
| As a _paying_ customer, I disagree. |
|
| ▲ | robofanatic 2 hours ago | parent | prev | next [-] |
| > constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code I find it hard to believe that github APIs don't have rate limits to handle traffic spikes. |
|
| ▲ | AustinDev 2 hours ago | parent | prev | next [-] |
| Alternatively, don't use GH for critical infrastructure they are no longer reliable, simple as. |
|
| ▲ | slowin 2 hours ago | parent | prev | next [-] |
| Well, they’re a critical piece of infrastructure with terrible stability. I’m in the process of migrating us off GitHub now. I’m not a Meta fan but it’s interesting that they manage to keep their systems up with an order of magnitude more traffic. GitHub’s uptime is inexcusable. |
|
| ▲ | raincole 2 hours ago | parent | prev | next [-] |
| Or, you know, it's perfectly reasonable and natural to feel angry when a service you paid for gets worse over time. > Unless you have a emergency hotfix (you don’t) Oh, we do. Given the sheer number of users, it's almost guaranteed someone is on fire ever time GitHub is down. Statistics is a very charming branch of reality. |
|
| ▲ | javier2 2 hours ago | parent | prev | next [-] |
| While I agree in principle, Github with their CI is very critical to many parties. |
|
| ▲ | semiquaver 2 hours ago | parent | prev | next [-] |
| Why would you take what they say at face value? I believe that this is basically a lie covering up deeper rot. |
|
| ▲ | Topfi 2 hours ago | parent | prev | next [-] |
| No one ever has to commit anything time sensitive. |
|
| ▲ | delduca 2 hours ago | parent | prev | next [-] |
| > increased something like 50x due to LLM This should not be an excuse since they probably can use AI to fix it /partially sarcasm |
|
| ▲ | bradhe 2 hours ago | parent | prev | next [-] |
| > due to LLM sloppers pushing large amounts of trash code hey bud your bias is showing |
|
| ▲ | thejosh 2 hours ago | parent | prev | next [-] |
| it's been trash for years, their status page just doesn't reflect that. |
|
| ▲ | OtomotO 2 hours ago | parent | prev | next [-] |
| You know, I am not employed. I have little time in my day to spend with my daughter. But I also have responsibilities. And one client has decided to bet on github. It's the only client I've ever worked with who has their code on github. EVERY other client has hosted their own gitforge or used bitbucket. So now I am not angry because some critical piece of US-american infrastructure is down all the time. I am angry because instead of spending time with my daughter I have to work on this later, because there are due dates and "Well, fucking GitHub was down" ain't gonna cut it. |
|
| ▲ | 2 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | nerdix 2 hours ago | parent | prev | next [-] |
| Wouldn't be surprised if we lost free private repos because of the slopocalypse. Though most don't seem to be shy about sharing their slop with the world so I'm not sure if that will actually increase stability. |
|
| ▲ | c2403 2 hours ago | parent | prev | next [-] |
| In this case, perhaps they shouldn't allow AI slop or promote usage of AI agents... |
|
| ▲ | cesarb 2 hours ago | parent | prev | next [-] |
| > go hit the gym or walk outside. My mind immediately went to this classic xkcd: https://xkcd.com/303/ |
|
| ▲ | catigula 2 hours ago | parent | prev | next [-] |
| FYI this is an argument to not use Github, not an argument to use it. |
|
| ▲ | sharts 2 hours ago | parent | prev | next [-] |
| It has nothing to do with load. |
| |
|
| ▲ | dimgl 2 hours ago | parent | prev | next [-] |
| What if you, you know, have a job? |
|
| ▲ | TZubiri 2 hours ago | parent | prev | next [-] |
| is the infrastructure of the free tier shared with that of paid users? That might be the issue. Otherwise if usage of paid users scales, then it would be fine. |
| |
| ▲ | john_strinlai 2 hours ago | parent [-] | | >is the infrastructure of the free tier shared with that of paid users? github enterprise is fully operational at the moment | | |
|
|
| ▲ | 2 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | neo_doom 2 hours ago | parent | prev [-] |
| Agreed. This impatience culture is so toxic. Been down less than 30 mins. |
| |
| ▲ | anonymars 2 hours ago | parent | next [-] | | That's cool, I mean if you'd gotten here earlier you could have said "been down less than 5 minutes" So with the "elapsed time" out of the way (now over an hour, incidentally), how about "how long will it be down?" | |
| ▲ | cb33 7 minutes ago | parent | prev [-] | | 2 hours later - still down. Is this "impatience culture" or incompetence culture at this point? |
|