Remix.run Logo
▲ pitched an hour ago

FTA, this is a fault-tolerant server where all hardware is redundant and can be hot swapped while running. The servers you’re thinking about are redundant at the software level so rebooting one server won’t cause the service to go down.

What is remarkable is that apparently that VOS thing has never crashed. I wonder if they also do software redundancy under the hood to keep uptime going during reboots. If the CPU is hotswappable, it must have something.

▲wildzzz 18 minutes ago | parent | next [-]

Probably uses formally verified code along with plenty of housekeeping processes to keep any failures from shitting the whole bed. In critical system design, you build in redundancies that work in parallel such that any one failure will not interrupt the system.

The main computer system in the Space Shuttle is an excellent example of this. It had 5 identical IBM System/4 Pi machines. Three of them ran identical code and handled the same work. The fourth ran a completely different codebase to handle the same work, preventing a bug in the main code from killing the whole system. A fifth computer handled other tasks but could be swapped over to the critical role if needed. You could lose 2/5 computers and still have insurance against a cosmic ray flipping a bit.

▲serf 14 minutes ago | parent | prev [-]

VOS is a parallel lockstep OS. You drop nodes and replace them to keep the whole operational.