Remix.run Logo
TacticalCoder 5 hours ago

Funny to see that frontpage as earlier on today I made a comment saying that with LLMs I'm betting we'll see N-modular redundancy systems soon, with a ultra-hardened, minimal, part picking the majority votes of N LLM-written implementations (in different languages, on different stacks, all running at the same time). Not 100% TFA but "computing in space" involved a lot of N-modular redundancy systems.

The reason I'm 99.9% sure we'll see that is that it'll help catch both bugs in the LLMs implementations themselves (and we know there are plenty of those) but also in the stacks/platforms/VMs running those software.

Imagine one spec and five implementations (Rust, Go, Java, Python, whatever) and one minimal system, with the tiniest of the tiniest attack surface, returning the answer as soon as 3-of-5 agree. And, as a bonus, if later on one the two "missing" answer arrives and doesn't match, it's cause for enquiry and bugs be smashed.

Basically (and although I don't care about Ethereum or cryptocurrencies except for the cryptographic aspect), we already witnessed that: 3 different implementations of Ethereum and, in the early days, one of the implementation whose result differed from the two others. And hence the implementation not respecting the spec (in that case it was the only one that was faulty) got instantly detected (and promptly patched). My memory is fuzzy but I know this happened.

Heck, I may write a proof-of-concept for fun.

I've got other ideas as to what will be possible in the future but I'm keeping them for another day.

kqr 2 hours ago | parent [-]

This all rests on the assumption that design errors are independent beween LLM-generated programs for the same specification. That's not true for humans (Knight and Leveson, 1986) and I highly doubt it's any more true for robots.

(On the other hand I just found out about Ron, Baudry, Monperrus, 2026, which seems to say "sure, problems are correlated, but it could still be useful.")