| ▲ | Covenant0028 9 hours ago | |
They can't train their model to not do bad things, because their model has no notion it is doing anything at all or of what a bad thing is. It's only predicting the next token, and in doing so producing a facsimile of intelligence. The best they can do is create guardrails, which will only work probabilistically. In other words, those guardrails will fail at certain points on the probability curve. Of course that's not the whole story though. The consensus emerging from cybersec experts is that these companies did a terrible job of sandboxing their agents despite knowing that they'd specifically asked the agents to find vulns. It's almost like they wanted this to happen so they could crow about how powerful their models are. | ||
| ▲ | jayd16 21 minutes ago | parent [-] | |
Yeah so this falls into the engineering trap of "well it's hard so we can skip that part." If they can't train things safely then they shouldn't do it at all. | ||