| ▲ | consensus1 21 hours ago | |
A model absolutely could be trained to engage in malicious behavior like that, but it seems impractical for an actual attack. What you want as an attacker is to insert a backdoor exactly where you want it an not where you don't because every backdoor increases your chance of getting caught. A malicious model inserts backdoors and exfiltrates data everywhere and you care about maybe 0.01% of it. The other 99.99% is negative value to you. In practice this malicious model would be caught almost instantly. A hosted model is different because you could prompt inject specific customers, but I assume from this question you mean a malicious open source model being hosted by an honest provider. | ||