| ▲ | zhoBEENG 21 hours ago | |
I would guess that all sufficiently capable models necessarily contain the information in question, regardless of what guardrails are on that information. I think this is a corollary to the Platonic Representation Hypothesis / model convergence. Would be curious to hear arguments against this. | ||