| ▲ | orf 7 hours ago | |||||||||||||||||||||||||
> First of all tons of more outages than ever, remember that React useEffect fkup [0]? Complete insane that this would happen at an infra company that runs a third or so of the web. Bugs happen all the time. They roughly increase with scale, not decrease. There’s an argument to be made about better testing, but this specific bug seems like a perfect one to slip through: multiple services, hard to spot at code review, involves JS/frontend, invisible at low traffic (test/UT envs). So IMO it’s not completely insane. Is the implication that Cloudflare should have no bugs whatsoever? | ||||||||||||||||||||||||||
| ▲ | minraws 6 hours ago | parent [-] | |||||||||||||||||||||||||
Bugs happen all the time yes, outages don't. If outages are increasing with scale you get 1 or maybe 2 free passes. After that you either have in-ept Engineering or just in-ept leadership. I used to be all in on CF a while ago, now I am moving off them almost entirely. Same issue with GitHub, I can understand if you can't build for the scale when you couldn't predict it but if after over 12-18 months things don't seem to be improving what are you even doing? I honestly think all of these companies are deluded if they think people will stick around with all these weekly outage events. I have a homelab server I have had 2 outages in 1 year because my shitty ISP went down. Still at 99.9% uptime, I have done nothing special. I now have backup internet as well. Is it big? Nope but it doesn't need to be cf scale. And scale is the reason to use these services why would I use cloudflare if a homelab would have been enough? If they aren't designing and scaling their systems to handle this scale they might as well close shop, someone else might do it better. As a infra/dev person who does his own thing on the side, I might be the most impacted by these outages, so I might be coming off as harsh. But they cost me both time/money and headache in extra development work. Imagine prod deploys are down for 2 days why? Because GitHub actions keep failing... Oh serving new OTA updates broke? Why? dig into the code.. go oncall with users instead of doing work, realize it's a CF outage and the writes failed. (Feel the tears streaming down your face). If I have to waste dev time, with AI and me together we could self host it with higher reliability with significantly cheaper costs at this point even at fairly decent scale. I think any infra company that has more than 1 outage a year is already not worth investing in. But more than 3 and you might be better off self hosting, even in this ram apocalypse. If all people in SF are this unserious about reliability (which hasn't been my experience but HN seems especially open to break the prod if you have to) Then well software companies really do deserve to be replaced by AI. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||