| ▲ | Onavo 6 hours ago |
| People like S3 because they have hard engineering guarantees around bit rot and work well as a high level abstraction of a network filesystem with all of the low level failure recovery built-in. You don't have to worry about doing your own RAID configs. The bigger question is whether non-AWS services can offer the same level of guarantees. I have heard horror stories for example when it comes to downtime on Hetzner's S3 object store. I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing. |
|
| ▲ | someonebaggy 2 hours ago | parent | next [-] |
| I find that Ceph is pretty annoying to operate but if you do, it works fine. Keeps redundant copies of data on multiple disks, across multiple racks if your diversity is that wide. Supports erasure coding for effective redundancy less than 2x. Don't go for the "object gateway" compatibility layer - just use raw Ceph if you're writing your own app. |
|
| ▲ | sgt 5 hours ago | parent | prev [-] |
| Meanwhile, just using disks and regular servers is still reliable and more so than ever, especially with some redundancy. Consider your cloud costs. |
| |
| ▲ | shye 3 hours ago | parent [-] | | If you there's no SLAs to engineer for, just set the target at zero, and drive your cost to zero as well. So you have to design for _some_ defined value of reliability. Backblaze's famous reports put HD failure rate is ~1.39%, so for the 11 9s you get as a guarantee from S3. Assuming nothing else fails, you'd need at least 6 independent copies to get that, plus all the effort to engineer recovery, and constant upkeep. Suddenly, S3, even when considering bandwidth costs, seems like a steal. | | |
| ▲ | someonebaggy 2 hours ago | parent [-] | | The usual SLA is not expressed numerically but as a feeling. And a normal server suffices to deliver that feeling. I'd bet 99% of apps are used by under a hundred people and if they have to take a day off due to a head crash, it's annoying but not catastrophic. If you have two sets of hardware the server can run on (cold standby), and RAID, and are competent at physical IT work, you can have faulty hardware replaced in an hour. Drive fails - replace it. Anything else fails - swap the drives to the other machine, boot it up and then troubleshoot the original. Most likely you don't even need that. If the app server runs on a standard platform like Windows you can shuffle it over to some spare tower PC. You hear horror stories of dusty towers that nobody knows what they're doing - the horror there isn't from running server software on a tower, but from unmaintained servers no matter the form factor |
|
|