| ▲ | shye an hour ago | |
If you there's no SLAs to engineer for, just set the target at zero, and drive your cost to zero as well. So you have to design for _some_ defined value of reliability. Backblaze's famous reports put HD failure rate is ~1.39%, so for the 11 9s you get as a guarantee from S3. Assuming nothing else fails, you'd need at least 6 independent copies to get that, plus all the effort to engineer recovery, and constant upkeep. Suddenly, S3, even when considering bandwidth costs, seems like a steal. | ||
| ▲ | someonebaggy 42 minutes ago | parent [-] | |
The usual SLA is not expressed numerically but as a feeling. And a normal server suffices to deliver that feeling. I'd bet 99% of apps are used by under a hundred people and if they have to take a day off due to a head crash, it's annoying but not catastrophic. If you have two sets of hardware the server can run on (cold standby), and RAID, and are competent at physical IT work, you can have faulty hardware replaced in an hour. Drive fails - replace it. Anything else fails - swap the drives to the other machine, boot it up and then troubleshoot the original. Most likely you don't even need that. If the app server runs on a standard platform like Windows you can shuffle it over to some spare tower PC. You hear horror stories of dusty towers that nobody knows what they're doing - the horror there isn't from running server software on a tower, but from unmaintained servers no matter the form factor | ||