| ▲ | quantumgarbage 3 hours ago | |||||||||||||||||||
Proprietary reasoning can be recovered from its encrypted traces. Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext, without ever attacking the stronger model directly or triggering its anti-distillation safeguards. | ||||||||||||||||||||
| ▲ | the_af an hour ago | parent [-] | |||||||||||||||||||
Why do you restate the abstract? Anyone can read it from the link. | ||||||||||||||||||||
| ||||||||||||||||||||