Remix.run Logo
throwatdem12311 4 hours ago

Open AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to.

Maybe alignment isn’t possible with LLMs.

pizza234 4 hours ago | parent | next [-]

> Maybe alignment isn’t possible with LLMs.

It absolutely isn't, indeed.

The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.

The simplest analogy that comes to my mind is the three body problem.

estearum 4 hours ago | parent | prev [-]

The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").