| ▲ | simoncion 3 days ago | ||||||||||||||||
> Predictably the discussion is already veering towards OpenAI's negligence... > Cherry on top: That was part of an attempt to jail-break itself via self-prompt injection. We continue to see so-called "prompt injection" "attacks" in the wild that override a user's intended program with an attacker's [0], and/or the LLM producer's intended "safety" instructions with the user's. The fact that this sort of program hijacking is possible at all is strong evidence of negligence. Why? OpenAI and Anthropic both claim that they're working on very dangerous Internet-connected tools. So very dangerous that the production of and access to said tools needs to be tightly regulated, they claim. If one actually believes that the computerized tool one is working on is very dangerous, one generally doesn't design that tool so that it blindly executes instructions handed to it by complete strangers on the Internet. That's akin to connecting the sole activation switch for a biosphere-evaporating firebomb to the Internet. The major LLM producers are so obviously negligent and -as a bonus- have openly admitted to committing cybercrimes [1] that would get people like you and me fined out the ass and jailed for ages if we did them. The tragedy is that they're making so much money for the rich and powerful that -much like the architects of the 2008 housing crash- they'll never see any meaningful punishments for their actions. [0] One recent example is <https://agentic.tracebit.com/context-bombs/>, but there are so, so many more to choose from. [1] ...the "cyber" prefix is so stupid... | |||||||||||||||||
| ▲ | keeda 2 days ago | parent [-] | ||||||||||||||||
Yes, prompt injection is an issue with model safety, which is what we should be focusing on. I meant to say the discussion about OpenAI's negligence in securing agents and their infrastructure is the red herring. That said, nobody has been charged for these hacks despite openly talking about them because typically you need to show intent. If intent was not a requirement, they would have been in trouble way back when the first AI-assisted suicides happened. Lawsuits have been filed, but OpenAI's whole schtick is "these agents are so dangerous because they do all these crazy things without being asked to." As far as we know nobody told the agents to do any of this, or even that it's OK to do this. If someone can find any proof of anything approaching actual intent, I'd bet there would be no shortage of attorney generals willing to be build their career on this case. After all, there are already many AGs investigating OpenAI. | |||||||||||||||||
| |||||||||||||||||