Remix.run Logo
quantumgarbage 3 hours ago

Proprietary reasoning can be recovered from its encrypted traces. Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext, without ever attacking the stronger model directly or triggering its anti-distillation safeguards.

the_af an hour ago | parent [-]

Why do you restate the abstract? Anyone can read it from the link.

Groxx 27 minutes ago | parent | next [-]

It's rather common for posters to make a very small summary in a comment. It can help fight the floods of comments working off the title alone (though it's not particularly needed here for that purpose, imo)

ronsor an hour ago | parent | prev | next [-]

This is Hacker News. You know people don't follow links and read.

mschuster91 an hour ago | parent | prev [-]

People don't read no links no more