Remix.run Logo
embedding-shape 3 hours ago

Is there any proof/evidence of weights themselves containing backdoors? I'm not claiming it isn't possible, clearly would be, but have we actually seen anything like that in any published mainstream weights, from any country?

Eddy_Viscosity2 3 hours ago | parent | next [-]

Not that I'm aware of, but that's just a specific. It could be that it can't work via weight manipulation. That just means they have to do it a different way. The general point is that they do want backdoors and will try to get them in using any and every available method.

ben_w 2 hours ago | parent | prev | next [-]

At the present time, that's like being handed a random human and asking if there's any proof/evidence of their brains including signs of being a secret double agent.

What we can monitor for is output, and very crude probes into their inner states.

We can't de-compile them; this is extremely unfortunate, given we do know it's possible to insert backdoors: https://www.anthropic.com/news/sleeper-agents-training-decep...

We can say e.g. R1 apparently writes less secure code if you say you're Uyghur etc., we don't know if this specific example is deliberate influence from the top, or if it's like how western models misbehave all over the place in all kinds of other ways. The latter is basically guaranteed as a general problem for all models, while the former is merely a case where we know they have motive, means, and opportunity.

kfse 3 hours ago | parent | prev [-]

[dead]