| ▲ | ben_w 2 hours ago | |
At the present time, that's like being handed a random human and asking if there's any proof/evidence of their brains including signs of being a secret double agent. What we can monitor for is output, and very crude probes into their inner states. We can't de-compile them; this is extremely unfortunate, given we do know it's possible to insert backdoors: https://www.anthropic.com/news/sleeper-agents-training-decep... We can say e.g. R1 apparently writes less secure code if you say you're Uyghur etc., we don't know if this specific example is deliberate influence from the top, or if it's like how western models misbehave all over the place in all kinds of other ways. The latter is basically guaranteed as a general problem for all models, while the former is merely a case where we know they have motive, means, and opportunity. | ||