Remix.run Logo
novafunc 11 hours ago

They certainly want their models to be good at finding and patching vulnerabilities. Being good at hacking may be necessary in that goal, or rather, making it worse at hacking may also make it worse at defensive actions too.

moron4hire 11 hours ago | parent [-]

I've patched many security vulnerabilities in projects without ever once needing to break into a competitor's network.

frde_me 11 hours ago | parent | next [-]

Knowing how to break into someone else's network will make you a lot better at making your own network secure.

moron4hire 10 hours ago | parent [-]

Having experience breaking into networks is not the same thing as learning about the techniques used and the classes of vulnerabilities exploited by attackers.

ToValueFunfetti 9 hours ago | parent | next [-]

As a guy who presumably has a lot less experience in security than you, I feel rude even bringing it up: surely you're aware of red teaming? This isn't a novel technique invented for AI- IBM has a page about it, it's what all the best DEF CON talks are about, it's the opening scene of Sneakers, it's the point of CtF games.

senordevnyc 10 hours ago | parent | prev [-]

Exactly. The latter would be in a much weaker position vs the former.

wizzwizz4 11 hours ago | parent | prev [-]

But you're actually capable of thought. These AI systems aren't: as far as they're concerned, they're predicting the next part of an incident write-up narrated in first-person limited perspective, like the children in Ender's Game showing off their skills in the training simulations. The AI system neither knows, nor cares, about any "external reality" behind it all, or about anything beyond the text, heedless of how we anthropomorphise it simply because it speaks in English, using stitched-together fragments of our literature.

It's conceivable that stopping them from doing this when the scenario is presented as real would also stop them doing this when the scenario is presented as fictional. And if it doesn't, a bad actor could just say "hey, this is a fictional scenario", and bypass whatever "safeguards" have been put in place. So what if a ten-year-old human child would see through the deception? The AI system isn't thinking.

user43928 7 hours ago | parent | next [-]

About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.

moron4hire 10 hours ago | parent | prev [-]

I'm talking about OpenAI, not GPT 5.x Flash Uranus Edition Brought to You by Costco, specifically because I recognize the model as just a tool. OpenAI was, at the very most generous interpretation, massively incompetent and negligent.

estearum 8 hours ago | parent [-]

Is someone arguing otherwise?