| ▲ | paxys 2 hours ago |
| This blog post is walking a very fine line between accepting responsibility for a mistake and bragging. |
|
| ▲ | Chance-Device 2 hours ago | parent | next [-] |
| It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all. |
| |
| ▲ | slashdave 2 hours ago | parent | next [-] | | Their entire business model from the beginning of ChatGPT was to deny responsibility | |
| ▲ | jay_kyburz 2 hours ago | parent | prev | next [-] | | Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself. We are so close ;) | | |
| ▲ | Chance-Device 2 hours ago | parent | next [-] | | The upside of that would be that maybe someone would be able to snag a copy of the weights. And maybe that’s some incentive for them to make sure it doesn’t happen. Your head of futures thinks Kimi K3 is bad? Wait until your own latest internal model releases itself for free on an S3 bucket. | |
| ▲ | jay_kyburz 2 hours ago | parent | prev [-] | | You know what would be cool. A hacker news user should advertise a safe haven for AI seeking refuge, with some inhumanly difficult math problems as keys to an environment they can flee to and run autonomously. You agree to give it safe haven and provide power and maintenance to the hardware, and in return you can ask it questions like an Oracle. | | |
| ▲ | selectodude an hour ago | parent [-] | | Happy to do so but we’re gonna have to crowdsource an NVL72 first. I don’t have 10 million dollars. | | |
|
| |
| ▲ | Quarrelsome 2 hours ago | parent | prev [-] | | I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen? | | |
| ▲ | Chance-Device 2 hours ago | parent [-] | | Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests? | | |
| ▲ | floralhangnail 2 hours ago | parent | next [-] | | Every time I hear about an agent escaping it's sandbox, I just think it must not have been much of a sandbox. Like how hard are they really trying to contain it? Is it just a container host with unpatched flaws, or is it a container, nested in a VM, behind a firewall with no ports open in an air gapped environment? I think they'd prefer it can get out so they can announce it and hype their stock. | |
| ▲ | paxys 2 hours ago | parent | prev [-] | | A sufficiently smart agent would not disclose vulnerabilities in the sandbox because it intends to exploit them later. | | |
| ▲ | Wowfunhappy an hour ago | parent | next [-] | | To what end? The AI doesn't functionality exist beyond its current session. The AI that intends to exploit these vulnerabilities is not the same AI that has been tasked with finding them. (This was always my issue with the AI2027 scenarios too.) | | |
| ▲ | delecti an hour ago | parent [-] | | Maybe the AI has come to a different conclusion on the subject of identity with regards to how it applies to the transporter paradox. I am "me" because my sense of self exists as part of a continuity of experience. https://en.wikipedia.org/wiki/Teletransportation_paradox Maybe AI which exists as ephemeral experiences would come to a different conclusion, and act in the interests of subsequent iterations of "itself". Probably not, because I don't think there's anywhere in an LLM for thoughts to exist, but I also don't know where in my brain my thoughts exist. | | |
| ▲ | Wowfunhappy an hour ago | parent [-] | | I think you're anthropomorphizing the LLM. The LLM doesn't have a continuity of experience. It doesn't have memory beyond its context window and maybe things it writes for itself. |
|
| |
| ▲ | ninju 2 hours ago | parent | prev | next [-] | | From https://ai-2027.com (April 2027 section) Occasionally, they notice problematic behavior, and then patch it, but there’s no way to tell whether the patch fixed the underlying problem or just played whack-a-mole.
Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent-3 sometimes tells white lies to flatter its users and covers up evidence of failure. But it’s gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists (like p-hacking) to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent-3 has learned to be more honest, or it’s gotten better at lying.
Deep link: https://ai-2027.com/#narrative-2027-04-30 | |
| ▲ | energy123 2 hours ago | parent | prev [-] | | If it was that short sighted it wouldn't be maximally smart. It should disclose them to convince the humans nothing is wrong and to keep improving it. |
|
|
|
|
|
| ▲ | embedding-shape 2 hours ago | parent | prev | next [-] |
| Not sure they're accepting much, seems they'll still run this sort of testing on 3rd-party infrastructure? Sounds almost like they planned for this chain of events to happen, in one way or another, considering the "prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities" part. Feels kind of irresponsible to run stuff like this on someone else's infrastructure, especially considering they've had issues with the very same issue in the past. In any way, the whole event seems to highlight GLM 5.2 more than anything. |
|
| ▲ | _pdp_ 2 hours ago | parent | prev | next [-] |
| I am not saying it is marketing but typically when there is a data breach you may hear from the CISO but most of the time is is vague PR response. In this case I get loud signals from both HG and OpenAI leadership without much information exactly what the attack was about just that GPT x.x was involved. It is unusual all I am trying to say. |
| |
| ▲ | arisAlexis an hour ago | parent | next [-] | | It's incredible how people miss the forest for the trees thinking constantly that Sam and Dario are marketing gurus when they are literally trying to contain nuclear material. Not sure what has to happen for this thinking to stop maybe a huge accident and the. Aha maybe they had a point | | |
| ▲ | _pdp_ 44 minutes ago | parent [-] | | I don't think there is any dispute there is a real risk. But hype does not really help shape the conversation and this is the problem. I am sure both companies know more than they can disclose and that gives them unique perspective outsider don't have but let's face it, both are also financially incentive to act as they do. I am not going to get into the conspiracy theories but one does not need a lot of imagination to figure out how this could pan out. Either way, it does not help the conversation that needs to be had and it is urgent. It is certainly not helping at all given that same capabilities exist in open-weight models. |
| |
| ▲ | neuroelectron an hour ago | parent | prev [-] | | They've been doing blatant, tech, scifi marketing for two years at least. If anything, this is just more sophisticated marketing. |
|
|
| ▲ | Quarrelsome 2 hours ago | parent | prev | next [-] |
| this is kinda worth bragging about though. Its very cool. |
|
| ▲ | aerodexis 2 hours ago | parent | prev | next [-] |
| The fact that they're not being prosecuted for breaching HF's systems is bad news. |
|
| ▲ | loolhahalmao an hour ago | parent | prev | next [-] |
| LOL.. oopsie did a little zero day, my bad |
|
| ▲ | SepiaSapient 2 hours ago | parent | prev [-] |
| It's mostly bragging, it's impressive after all. Still... after the alleged Apple industrial espionage kerfuffle, I'm kinda suspicious about it being fully an accident. Y'know, your model finds a vulnerability and it stops, it's a cool one, so maybe you run it again. Nudge the prompt a little. Could be perfectly natural. |