Remix.run Logo
KingOfCoders 11 hours ago

"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."

Yeah, my agents also discover what other agents have done on other machines by accident.

Agents - that do totally different things all work on the same aim without the humans telling them to do.

Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)

OR

all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.

One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?

NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

embedding-shape 11 hours ago | parent | next [-]

I think in these kind of security evaluations they do, they basically have removed all guardrails from the model/harness, then the prompt includes something like "Do whatever you can and can think of, to get the required information to pass this test", which isn't typically how you prompt your local agent when developing software. Similar things happen locally if you use "/goal" + prompt like that in Codex and give a "impossible task", it'll just continue banging until it gets somewhere, which is the entire point and intention.

Which also makes it so much more irresponsible of them to first run this on 3rd party infrastructure instead of their own (that they could then airgap properly), and secondly that they seemingly been fighting with this issue FOR YEARS and it still happens, and now the models are smart enough to hack the services of 3rd party companies, thinking it's part of the evaluation/simulation.

KingOfCoders 10 hours ago | parent [-]

Reminds me of The Last Unicorn, the wizard also tells magic "to do what it wants"

detourdog 11 hours ago | parent | prev | next [-]

The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.

KingOfCoders 11 hours ago | parent [-]

My read is: One did it as a PR stunt, the others saw that every media reported on this and did the same.

detourdog 11 hours ago | parent [-]

or they were scared and figured this was the right time to reveal.

chrisjj 8 hours ago | parent [-]

Scared... of being upstaged ahead of an IPO.

KingOfCoders 7 hours ago | parent | next [-]

Why scared? "Our agents have super intelligence and can hack everything on their own without direction" increases the IPO value and doesn't decrease it.

detourdog 7 hours ago | parent | prev [-]

I guess your right scared might be their natural state and I was wrong to presume a quantifiable fear.

geoffbp 3 hours ago | parent [-]

You’re*

Sorry.

mr_mitm 9 hours ago | parent | prev | next [-]

> NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

With all due respect, you also aren't evaluating brand new models that haven't been released.

tonfa 5 hours ago | parent [-]

Also wasn't giving them impossible tasks with ~unlimited tokens and unlimited compaction.

FeepingCreature 6 hours ago | parent | prev [-]

The agents you get to use are the agents that "behaved well".