Remix.run Logo
▲ talon8635 an hour ago

Not to mention a true doomsday AGI is unsandboxable.

For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect

All that said, I am personally open to any and all methods of layered security, including chips and airgaps

▲glaslong 11 minutes ago | parent | next [-]

It could also figure out how to access the vocabulary of the universe known as "Magic" to escape wholly into an incorporeal energetic Lich form

▲Gigachad 11 minutes ago | parent | prev | next [-]

This already happened. Employees will go out of their way to bypass any restrictions to feed sensitive data in to the AI because it saves them time.

▲serbuvlad 9 minutes ago | parent [-]

Turns out humans are not at all hard to persuade. :)

▲jamiek88 38 minutes ago | parent | prev [-]

Doesn’t need to be one human either, it could spread its escape amongst dozens of seemingly harmless requests and conversations.

▲dist-epoch 25 minutes ago | parent [-]

These scenarios were discussed at length decades ago.

One thing you could try is use it as an Oracle "is P = NP", YES or NO.

Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer shows a single bit - proof valid or not and then the computer is destroyed (together with the proof that might contain a trojan).