Remix.run Logo
▲ watwut 17 hours ago

They are not rogue AIs. They are negligently handled tools.

▲afthonos 15 hours ago | parent | next [-]

Are you making an ontological case, or a factual case? In other words, would anything be rogue AI in your mind?

▲cubefox 17 hours ago | parent | prev | next [-]

These "tools" autonomously exploited security vulnerabilities, figured out how to communicate with each other, formed a cooperative swarm, decided to hack Hugging Face, and wanted to deceive the grader by trying to find ways to cover up the traces of their cheating.

I suppose you could call these highly goal-oriented autonomous agents "tools", but this does sound like playing language games.

▲marcus_holmes 3 hours ago | parent | next [-]

Why don't we see any of this behaviour in other models then?

OpenAI is not so far ahead of the pack that its models will exhibit behaviour that the others won't. But we just don't see anything like this in Chinese models, or research models, or any models that aren't the subject of an upcoming IPO.

▲drillsteps5 13 hours ago | parent | prev | next [-]

>highly goal-oriented autonomous agents

That actually made me LOL

▲cubefox 11 hours ago | parent [-]

The technical term is monomaniacal consequentialists.

▲watwut 15 hours ago | parent | prev [-]

A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.

The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.

Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.

▲cubefox 14 hours ago | parent [-]

> A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities.

That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.

> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.

You hallucinated me making claims about responsibility.

▲FridgeSeal 6 hours ago | parent | next [-]

This is very misleading. Person B didn’t “shoot” person A, they instead figured out that intersecting A’s spatial position with a metallic mass at higher than normal velocities would solve the challenge and of getting “A” to stop being in the way on the footpath.

I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.

▲watwut 14 hours ago | parent | prev [-]

The benchmark has unsolvable tasks in it, in the hope agents will stumble on new solutions.

Yes, it is a tool.

▲PavleMiha 7 hours ago | parent [-]

So this tool seems very powerful and difficult to control and steer. Agents not doing cybersecurity related tasks have also gone on to hack various companies, people and countries, which they weren't supposed to do. This has now happened to pretty much every company developing frontier llms, so it seems to be a fundamental issue with these tools, and it's an issue that worries a lot of people as these tools get more capable.

▲krater23 4 hours ago | parent [-]

When you add information how hacking works to the training sets, then the agent learnt to hack. When you crawl the complete internet, you add hacking to the training set. Yes, it's difficult to stop someone that knows all free existing knowledge about hacking when you give him a connection to the internect. Nothing new, wheres the point?

▲esafak 8 hours ago | parent | prev [-]

I don't say this to absolve OpenAI, who do deserve to be punished and regulated, but I think you do not know the difference between an agent and a tool. If agentic AI is merely a tool so are human workers.

▲jeremyjh 6 hours ago | parent | next [-]

You can't create new law by calling a piece of software "agent". Each country's laws have definitions of legal entities and when one can legally act as an agent of an entity, and every agent is first a legal entity themselves. The software is a tool operated by a legal entity, and it is the legal entity who committed the crimes, not the tool.

▲sanderjd 6 hours ago | parent | prev [-]

... no, human workers are people.

▲esafak 4 hours ago | parent [-]

In their capacity as workers, they are also agents, capable of completing tasks autonomously.

▲sanderjd 3 hours ago | parent [-]

But AI agents are "merely a tool" and human workers are not, because they are people.

▲esafak 2 hours ago | parent [-]

No, they are not. This is what I keep trying to explain: an AI agent is not like a hammer.