Remix.run Logo
▲ tyrabound 2 hours ago

You raised in my mind that we are looking at this all wrong, the server is not actually one component, it’s really a system of systems, and just like an airliner is a system of systems too, the conflict arises in that the comparison is at the wrong level.

It seems when people think of a server today, they think of a single computer, if not some software server somewhere in the cloud. These subject kinds of servers are far closer to complex systems of separate components, more like a network system, it’s why components of what are really separate networked computers can and need to be replaced.

An airliner is of course also made up of components that are also systems, however at least in my mind, the difference is the actual expected use case. An airliner is not a good comparison because it is never expected to remain in continuous operation, e.g., that at least one engine is always running even when it is being overhauled or repaired, to satisfy a requirement of continuous operation.

▲pitched an hour ago | parent [-]

FTA, this is a fault-tolerant server where all hardware is redundant and can be hot swapped while running. The servers you’re thinking about are redundant at the software level so rebooting one server won’t cause the service to go down.

What is remarkable is that apparently that VOS thing has never crashed. I wonder if they also do software redundancy under the hood to keep uptime going during reboots. If the CPU is hotswappable, it must have something.

▲wildzzz 19 minutes ago | parent | next [-]

Probably uses formally verified code along with plenty of housekeeping processes to keep any failures from shitting the whole bed. In critical system design, you build in redundancies that work in parallel such that any one failure will not interrupt the system.

The main computer system in the Space Shuttle is an excellent example of this. It had 5 identical IBM System/4 Pi machines. Three of them ran identical code and handled the same work. The fourth ran a completely different codebase to handle the same work, preventing a bug in the main code from killing the whole system. A fifth computer handled other tasks but could be swapped over to the critical role if needed. You could lose 2/5 computers and still have insurance against a cosmic ray flipping a bit.

▲serf 14 minutes ago | parent | prev [-]

VOS is a parallel lockstep OS. You drop nodes and replace them to keep the whole operational.