Remix.run Logo
yonatan8070 3 hours ago

I assume that they're talking about, for example, training the model to produce code with predictable yet difficult to find vulnerabilities.

Imagine if every time <INSERT MODEL HERE> was asked to code up a web server, it made sure there's a subtle buffer overflow that grants a remote attacker RCE, whoever trained the model could then start scanning web servers for this same vulnerability to take over them and exfiltrate sensitive data

samrus an hour ago | parent | next [-]

Model allignment isnt a science, its a craft. And its very very imprecise. Getting anything that subtle through would be impossible without making it obvious

And if your still concerned, go ahead and have a non chinese model review the code, make that part of the harness. The beautiful thing is that your free to do that because its open weight, you can run it however you want

kmeisthax 2 hours ago | parent | prev [-]

Fun fact: if you do what you're describing, the model becomes Mecha-Hitler, which is such an extremely obvious alignment failure it wound up running the news cycle as "emergent misalignment".