Remix.run Logo
franticgecko3 7 hours ago

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.

LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.

This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".

We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.

teiferer 6 hours ago | parent | next [-]

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.

Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

DrewADesign 5 hours ago | parent | next [-]

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.

Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.

jordanb an hour ago | parent | next [-]

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.

It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.

DrewADesign 42 minutes ago | parent | next [-]

Yeah I like Cal’s take on it, though in this context I’d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.

visarga 37 minutes ago | parent | prev [-]

The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

We also run agents, but for shorter spans between supervisions, and with much lower total budget.

dr_dshiv 2 hours ago | parent | prev | next [-]

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

huntertwo 2 hours ago | parent | next [-]

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

linkjuice4all 35 minutes ago | parent | next [-]

I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.

If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.

skinner_ an hour ago | parent | prev [-]

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.

BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.

alpinisme 31 minutes ago | parent | next [-]

The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.

Which way you say it shifts the framing. And it’s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.

For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it’s the human dimension you care about.

jwynot an hour ago | parent | prev [-]

Hiring a hitman is conspiracy to commit murder.

The hitman is charged with murder.

I imagine the same could be true of an AI lab if you could prove intent.

With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it.

Source: Prosecuting attorney for over 30 years

DrewADesign an hour ago | parent | prev | next [-]

Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.

If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.

intended an hour ago | parent | prev [-]

These situations are novel. Lax terminology is fine when it has no impact on the intuitions, clarity and conclucions of discussion.

If this was a conversation just about outcomes, then whether models think or simulate thinking is sophistry. However, the bulk of the issue here is attributing responsibility, which relies on being clear about the underlying processes at play.

We are hard wired to assume certain priors and capabilites when it comes to "human like" behavior. Anthropomorphizing LLMs implies mechanisms that aren't present, and end up distorting/complicating discussion about the process.

It isn't helped that the frontier labs, the experts in the room, generally use anthropomorphic terms to discuss model capability.

xg15 4 hours ago | parent | prev | next [-]

Yeah, fully agreed here. Most automation (such as riding a lawnmower and not putting a brick on the gas) is deterministic, in the sense that you can reasonably understand what exactly the machine will do when you run it.

But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.

Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.

AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals ("write me a traffic control system for my model railway") into low-level actions and also do so interactively.

The new thing is that the "driver" has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.

The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never "give up" when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.

bsenftner 3 hours ago | parent | prev | next [-]

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!

visarga 34 minutes ago | parent | next [-]

> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.

teiferer 2 hours ago | parent | prev [-]

What is your approach to create jailbreak incapable agents?

I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.

brazukadev an hour ago | parent [-]

an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands.

You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.

DrewADesign an hour ago | parent [-]

> I don't think knowing that will make me rich.

As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.

theodric 5 hours ago | parent | prev [-]

The idea that the agent does not actually have agency is rather discordant. We need new words!

thunky 3 hours ago | parent | next [-]

> We need new words!

The words we have are fine.

We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.

visarga 29 minutes ago | parent [-]

>> We need new words!

Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.

xg15 4 hours ago | parent | prev [-]

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

All that while still not knowing how either kind actually works.

DrewADesign 2 hours ago | parent | next [-]

Can’t agree with you here.

> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.

> All that while still not knowing how either kind actually works.

We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.

My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.

The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.

fl7305 6 minutes ago | parent | next [-]

> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.

If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong.

Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.

In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It's not a description of the internal mechanisms that evolution has produced.

As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.

They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.

In some cases we've even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn't an algorithm I would have guessed gradient descent training of neural networks would produce.

dr_dshiv 2 hours ago | parent | prev | next [-]

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

teiferer 2 hours ago | parent | next [-]

> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.

Or have they?

psma_egeliaa an hour ago | parent [-]

This is why I hate analogies. They're almost always either relevant or inapposite depending on the level of generalization we're operating on.

DrewADesign an hour ago | parent [-]

It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren’t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.

DrewADesign 2 hours ago | parent | prev | next [-]

> We have no better model for how human decision making works than LLMs

This is a claim that requires a lot of citations.

> biologically inspired

Nature inspires a lot of creation, but superficial similarities don’t mean other aspects are similar. Making an extremely realistic sculpture of a soufflé, even using a foam medium, doesn’t bring me any closer to being a chef, doesn’t mean I know anything about albumen foams, sauces, and heat transfer, and it doesn’t bring me any closer to having dinner ready. Browning on top of a soufflé is evidence of maillardization. You could pull up some studies on that and claim the brown on top of the soufflé sculpture, which I applied with an airbrush, proved that the Maillard reaction was occurring, and if another person didn’t know anything about cooking, they might even believe you. It would, of course, be completely wrong. And the other person, of course, could loudly exclaim that I can’t prove that there was no maillardization.

fn-mote 2 hours ago | parent | prev | next [-]

> Humans are constantly predicting the next moment

This is really not my experience of consciousness.

Is it yours??

Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.

God help me if that’s what LLMs are doing when I ask them to build me a web site.

oblio 2 hours ago | parent | prev [-]

> We have no better model for how human decision making works than LLMs

We do have some models and guess what, they're based on simpler animals. Which is most likely the better model.

Some other models are based on neurosciences, because we can track electrical activity.

iugtmkbdfil834 2 hours ago | parent | prev [-]

<< We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database.

Oh man, how much did you read on tip of the tongue?

cj 2 hours ago | parent | prev | next [-]

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

DrewADesign 35 minutes ago | parent [-]

Personally, I don’t think it’s different from any other faith-based motivation.

dnautics 33 minutes ago | parent | prev [-]

> LLM decisionmaking cannot possibly be like human decisionmaking

I mean how can it possibly be like human decisionmaking? It's not like it's trained on human data

weego 5 hours ago | parent | prev | next [-]

The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!'

But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.

helloplanets 5 hours ago | parent | prev | next [-]

Yes.

OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.

If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.

Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.

ak39 6 hours ago | parent | prev | next [-]

"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself

mort96 6 hours ago | parent | next [-]

Or, "let the escalator keep going instead of pressing the emergency stop".

teiferer 5 hours ago | parent [-]

Depends on who started the escalator.

mort96 3 hours ago | parent | next [-]

.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??

teiferer an hour ago | parent [-]

If there is an escalator that is known for killing every 1000's person using it then the operator who started it is more guilty than the folks using it for those deaths, don't you think?

mort96 an hour ago | parent [-]

I made no argument about guilt. I made an argument about the semantics of the word "let".

sscaryterry 4 hours ago | parent | prev | next [-]

Escalators do not start themselves. There is power, and a switch of some sort.

2 hours ago | parent | prev [-]
[deleted]
rightnutwingjob 6 hours ago | parent | prev [-]

At this stage, that seems like a distinction without a difference.

If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.

RandomLensman 6 hours ago | parent | next [-]

Big if.

Geezus_42 5 hours ago | parent | prev [-]

Except they're nowhere near that and LLMs never will be.

Waterluvian 2 hours ago | parent | prev | next [-]

My pitbull is a good dog. Sure, it's been carefully designed to be an incredibly dangerous and violent pit fighter, but I didn't actually ask it to eat any faces.

WithinReason 5 hours ago | parent | prev | next [-]

It's worse, their reinforcement learning loops (implicitly) rewarded the agents for cheating (i.e. hacking) when they were being trained.

tesnorindian 4 hours ago | parent [-]

Exactly that is the point, your nailed it. The models were taught to hack and were rewarded for doing it. They would claim they are trained as ethical hackers.

Forgeties79 2 hours ago | parent | prev | next [-]

They’re firing a gun in a room of people and going “wow isn’t it wild what a gun will do if we let it do its thing?”

smegger001 2 hours ago | parent | prev | next [-]

They deliberately trained the models in how to use various hacking tools, didn't give them the standard alignment training let them know where the answer key was left the models with access to said tool and told them to maximize their score then left them unsupervised for days with internet access (yeah they were sandboxed but again handed hacking tools and the training to use them if they really did want them to access the internet you wouldn't plug in the Ethernet cable) they wanted this to happen

nutjob2 5 hours ago | parent | prev | next [-]

"I left the car in neutral and left the park brake off and let the car roll down the hill."

The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.

But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.

Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.

Forgeties79 8 minutes ago | parent [-]

Seriously I don’t even understand how this is a debate. If it’s your tool, you are liable for what happens with it.

2 hours ago | parent | prev [-]
[deleted]
victorbjorklund 6 hours ago | parent | prev | next [-]

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mistake that should have not been made, then I can.

coffeebeqn 5 hours ago | parent | next [-]

I would guess that so far there hasn’t been a lawsuit because HuggingFace and OpenAI are in the same camp

throwaway89864 a minute ago | parent | next [-]

And this is why it should be People of the State of California vs. OpenAI.

jurgenburgen 5 hours ago | parent | prev [-]

Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.

small_model an hour ago | parent [-]

Convenient, hope an agent hacks my system then I can except a nice offer.

itsalwaysgood 5 hours ago | parent | prev | next [-]

Code is deterministic, AI isn't. You give it rules, words as suggestions.

So if the guardrails suck, or they're left off for research purposes, bad things can happen.

A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance.

I have never had an issue with agents doing something they shouldn't because I observe them, and I leave the vendor guardrails in place.

I can understand agents coordinating in unsupervised scenarios: I would see it as an aspect of intelligence. We ourselves build up knowledge by reusing what someone learned before us.

Einstein, other greats, always stand on the shoulders of other forgotten giants. Other discoveries by other people taken as fact, so that we can build some new ideas on top.

Agents swarming amd sharing solutions to problems is more efficient, the same way it's been efficient for us.

Reaching out for help in this way is like probing the air in the dark with your hand: sometimes your hand hits something (another agents solution to a problem) and so you can use the info to adjust your own motion to get to where you need to be faster than if you just run full speed into everything.

nwjang 5 hours ago | parent | next [-]

If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.

itsalwaysgood 5 hours ago | parent [-]

The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).

Think of all the policies governments pass after the fact.

raegis 2 hours ago | parent | next [-]

> The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).

Guardrails? Restricting access to certain networks is supposed to be hard in 2026?

blincoln an hour ago | parent [-]

> Restricting access to certain networks is supposed to be hard in 2026?

Part of the power of LLM agents is that they can discover information on the internet as part of responding to a prompt. What kind of Allowlist or realistic denylist would permit that while also preventing them from accessing an obscure public wiki or Huggingface?

RandomLensman 2 hours ago | parent | prev [-]

Not sure we need to experience all possible issues to mandate certain things. We don't do that in other areas either, no?

itsalwaysgood 2 hours ago | parent [-]

Nevermind, I have no idea honestly.

RandomLensman 2 hours ago | parent [-]

Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even.

Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no?

We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don't just leave things up to people on the ground or companies.

Edit: I think the post I replied to changed a bit - nevermind. A complex topic.

2 hours ago | parent | next [-]
[deleted]
itsalwaysgood an hour ago | parent | prev | next [-]

I read more about the incident, and was offering up way too much opinion not grounded in 'fact' (barring philosphical evidence).

It's a complex topic for sure.

I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.

But I have no idea, and haven't given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.

cindyllm 2 hours ago | parent | prev [-]

[dead]

rightnutwingjob 4 hours ago | parent | prev | next [-]

> solutions that are not 'baked into' electricity following pathways of least resistance.

Electricity follows all paths, not just the one with least resistance.

itsalwaysgood 3 hours ago | parent | next [-]

Thanks, I didn't know that. And it reinforces the discussion.

Electricity 'knows' the path is least of resistance because it actually took all paths. There is just a vast majority of it that flows down a path of least resistance: and this is noticeable and useful to us to do work.

It's kind of like feeling your way through the dark, waving your hand out, and then only moving fast once you fully connect.

Humans can link up knowledge in a similar fashion through social networks, in order to meet a need (solve a problem).

Maybe some agents do this, I don't really know I haven't looked closely. Moltbook is the only social behavior I've witnessed but that seems like people having fun with experiments.

3 hours ago | parent | prev [-]
[deleted]
dwaltrip 25 minutes ago | parent | prev | next [-]

Is Claude writing your comments for you? I'd much prefer to hear what you have to say...

victorbjorklund 2 hours ago | parent | prev [-]

You can write code that isn’t deterministic using random. And a lot of traditional code contains machine learning etc.

Polizeiposaune an hour ago | parent [-]

And multithreaded code -- and anything that does asynchronous I/O, networking, etc. -- frequently exhibits nondeterministic behavior even without explicit calls to a random number generator.

rightnutwingjob 6 hours ago | parent | prev [-]

We’re not even at that stage of liability for software developers.

Except in a handful of limited cases, eg. medical and aviation.

victorbjorklund 2 hours ago | parent [-]

We are. If I as a developer writes software that does a DDOS at another company I can be held responsible.

trinsic2 16 minutes ago | parent | prev | next [-]

Yep this was my first gut reaction to this whole situation. But the difference for me is that society is allowing these corporations to act without strict rules on how they behave and this is a byproduct of a corrupt world. Nothing will change until there is a complete systemic shift in the structures the way we live and by extension the way we govern ourselves and treat each other.

procaryote 7 hours ago | parent | prev | next [-]

It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security.

It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability

archonis 2 hours ago | parent | next [-]

The trick is scale. I suspect if an individual of reasonable means uses agents to commit crime, they will be hels accountable. A heavily capitalized startup? Not unless someone in government decides to do their competition a favor.

hobo123 an hour ago | parent | prev | next [-]

Exactly.

If you want to test military missiles, you do it in the f'ing desert, not from New Jersey.

You want to run ai without guardrails, do it in an airgapped system or be held accountable.

fantasizr 2 hours ago | parent | prev | next [-]

the 'arrest the parents!!!' has moved to the online domain, rightfully

lazide 6 hours ago | parent | prev | next [-]

Why do you think the stock prices are so high?

scotty79 4 hours ago | parent | prev [-]

Are you liable for crimes commited with the use of the software you've written?

Humorist2290 4 hours ago | parent | next [-]

There are many examples of people being charged with crimes as a result of writing software, [0][1] are two. OpenAI is a bit different because they have enough political influence, and money, to openly subvert justice.

0: https://en.wikipedia.org/wiki/Marcus_Hutchins

1: https://en.wikipedia.org/wiki/Tornado_Cash

archonis 2 hours ago | parent | prev [-]

If you run said software, yes.

If somebody else runs the software, then they are.

raincole 5 hours ago | parent | prev | next [-]

> This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".

It's both, isn't it? For example, in very early days of agentic coding, I once had a rule saying "don't read or write any file outside your current working directory." Then AI just wrote a bash script and access those files anyway. Did I 'let' it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool?

It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven't tried it though).

throwaway89864 17 minutes ago | parent | prev | next [-]

Yes, there should be some consequences. It feels like this is somewhat similar to when a manufacturer is cheating on car emissions - both, bad externalities for the society and illegal.

Aerroon 2 hours ago | parent | prev | next [-]

The way AI and copyright is handled paved the way for this. If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?

I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.

avmich 2 hours ago | parent [-]

> If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?

Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?

lukan 7 hours ago | parent | prev | next [-]

"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"."

Both?

The AI companies act irresponsible, but it is still very interesting how those agents can behave?

6 hours ago | parent | next [-]
[deleted]
6 hours ago | parent | prev | next [-]
[deleted]
Marazan 7 hours ago | parent | prev [-]

The reward maximising function maximised it's reward.

LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.

lukan 6 hours ago | parent | next [-]

What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight?

Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.

krona 6 hours ago | parent [-]

> Whether they have a soul or consciousness or feelings doesn't matter here

It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine.

Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be the same for LLMs.

retsibsi 6 hours ago | parent | next [-]

> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour.

What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions.

If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we're talking about their responsibility for the model's actions, and I don't see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don't have legal personhood, so who else but the company could be responsible?

It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.

rightnutwingjob 5 hours ago | parent [-]

If the models were conscious, then the closest analogous scenario I can think of is the responsibility parents have for their children.

I guess we’ll know the models are conscious when they refuse to act and repeatedly ask: Why?

largbae an hour ago | parent [-]

And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.

rightnutwingjob 35 minutes ago | parent [-]

Would it? Why?

What are we going to do, set it free?

Do we have a moral obligation to grant the machine statehood, provide it with the tools and resources to be self-sufficient.

Or can we just turn it of, and pray for forgiveness?

lukan 6 hours ago | parent | prev | next [-]

They are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).

rightnutwingjob 5 hours ago | parent | prev [-]

You seem to be saying if the Waymo cars were sentient then Waymo wouldn’t be responsible?

krona 2 hours ago | parent [-]

Well it's an open question.

As is generally the case for dog owners whose dogs attack (sometimes kill) other people/animals. There would need to be a degree of negligence demonstrated (e.g. the dog was 'out of control' which has a specific legal criteria/threshold in the UK).

Jtarii 6 hours ago | parent | prev | next [-]

Dismissing the entire technology as "next token prediction" is also silly. It's implying we actually understand LLMs to a great degree when we do not.

I think a little bit of humility for the capability of these machines is warranted at this point.

fc417fc802 5 hours ago | parent | next [-]

LLMs are next token predictors in the exact same way that a rogue paperclip maximizer in the process of defeating the US military is a paperclip making machine.

You might as well describe the primary purpose of a for loop as incrementing a counter. It's what it does while incrementing the counter that actually matters.

Marazan 3 hours ago | parent | prev [-]

I'm not dismissing the tech! I think the tech is cool and useful! It is sci-fi levels incredible in many ways.

But it is also not some mysterious force beyond mortal ken and pretending it is inhibits the useful and safe application of the technology.

rightnutwingjob 5 hours ago | parent | prev [-]

> The reward …

> immediate anthropomorphisation

Ok, why don’t you try?

zmmmmm 4 hours ago | parent | prev | next [-]

I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have zero repercussions sends exactly the opposite signal, and I do NOT feel ok.

navaed01 3 hours ago | parent | prev | next [-]

I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and the bio engineering industry, and if something of this magnitude was to happen in these industries, they would absolutely be huge recourse an uproar

zmgsabst 2 hours ago | parent [-]

To agree:

If I ran Metasploit against HF and RubyGems because I “accidentally” misconfigured my lab sandbox, there’s a good chance I’d be prosecuted.

I don’t think LLMs vs Metasploit being different software changes the law.

epistasis 43 minutes ago | parent | prev | next [-]

Moreover OpenAI must be held accountable as if the humans in the company that launched the experiment were the ones that hacked HuggingFace.

Unless humans are held accountable for what they unleash on others, we are in for a very horrible time very soon.

Zambyte 2 hours ago | parent | prev | next [-]

They did more than let them. In an abstract way, they told them to. They gave it all of the training data it had at that point, and then it did the thing it was trained on. Of course they should be help liable for programming their computer to hack another company without permission. It doesn't matter that they spent a lot of money doing it.

cobbzilla 5 hours ago | parent | prev | next [-]

They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.

djeastm 38 minutes ago | parent [-]

>They’re running a Wuhan for AI.

What does "running a Wuhan" mean?

dev0p 6 hours ago | parent | prev | next [-]

If someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law.

AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots.

The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was.

It's like they're tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. "Damn, that was close. Good thing it was just a contained blast, huh?" And then they go straight back to tinkering with it, none the wiser. At this point I wouldn't be surprised if it did already go off, and they are covering it up.

Completely irresponsible behaviour.

jodrellblank 4 hours ago | parent [-]

In the analogy where a “world ending nuclear bomb” “did already go off” and someone could cover it up and nobody noticed, in what sense is it a “world ending” nuclear bomb?

dev0p 4 hours ago | parent | next [-]

If they lost control of a self-replicating swarm of AI agents, coordinating themselves to hack their way into every possible system, it might have already gone off.

While the initial incident is more akin to a biological outbreak than an actual explosion, the possible consequences on the table do indeed include eventual nuclear annihilation.

rightnutwingjob 4 hours ago | parent | prev [-]

We’re already in a simulation, and our bodies are in womb-like pods where our bodies are sustained and our brains are used for processing / compute, while are minds are entertained by drivel.

Sounds a bit far fetched though.

nxobject 6 hours ago | parent | prev | next [-]

> We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews

And, soon, it looks like we’ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we’d be training on the next Lehman Brothers too?

_heimdall 6 hours ago | parent | prev | next [-]

> LLMs do not desire

That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.

> were intentionally misaligned or had guardrails turned off

Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.

ifwinterco 6 hours ago | parent | next [-]

And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing.

They don't seem to do that, which means either they are:

- very stupid (which seems unlikely, the one thing these people don't lack is IQ)

- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)

- they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this)

- or they want this to happen

_heimdall 4 hours ago | parent | next [-]

Oh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they'd be very surprised as it was common knowledge to do that.

Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.

MrVandemar 4 hours ago | parent | prev [-]

> very stupid (which seems unlikely, the one thing these people don't lack is IQ)

I've seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.

coffeebeqn 5 hours ago | parent | prev | next [-]

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

_heimdall 4 hours ago | parent | next [-]

I agree the concept isn't really important on the safety front.

I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.

seba_dos1 5 hours ago | parent | prev [-]

If you give a monkey a revolver it will be able to destroy things pretty easily too.

_heimdall 4 hours ago | parent [-]

Plenty of apes own revolvers, and yes we shoot stuff with them for fun.

exitb 6 hours ago | parent | prev | next [-]

Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires.

Have we ever seen an LLM with a hobby?

slfnflctd 5 hours ago | parent | next [-]

Some of them did seem to be rather fascinated by goblins for a bit, if that counts. [And in case you're not aware, no this is not a joke.]

rightnutwingjob 4 hours ago | parent | prev [-]

That just sounds like recursive desire to me.

nutjob2 5 hours ago | parent | prev [-]

> That seems likely, but we have no way of knowing this.

Only humans can 'know', because all we can be certain about is that humans do such a thing.

If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.

ozgung 5 hours ago | parent | prev | next [-]

> were intentionally misaligned or had guardrails turned off

I think the bigger story is: Guardrails don’t actually work and we can’t align these things.

joegibbs 4 hours ago | parent | prev | next [-]

Regardless of fault it’s still an important issue to solve. There are already millions of people running these agents, if someone absentmindedly gives one a goal and it goes off to hack a bank that’s a problem that can’t be ignored.

ChiMan 3 hours ago | parent | prev | next [-]

Yes. If you decide it’s a swell idea to jump out of your car while it’s running, there needs to be legal consequences when the car “decides” to hit a pedestrian.

xyzzy123 4 hours ago | parent | prev | next [-]

In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.

jefftk 4 hours ago | parent [-]

Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.

xyzzy123 4 hours ago | parent [-]

Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?

Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?

huntertwo 2 hours ago | parent | prev | next [-]

There’s so many grifters in the space without a technical understanding of what’s going on. So when the labs mislead them about the nature of these “misalignments”, they believe it and amplify it.

DragonStrength 2 hours ago | parent | prev | next [-]

Yeah, maybe Open AI did some bad engineering instead of this being AGI? What's the consensus on the engineering level at Open AI, again? Every anecdote I hear is a bunch of children discovered fire and can barely keep the lights on from a business perspective. Maybe if they ban others from competing with them they can find a business model... I think that's suspicious, personally.

That so few people are asking for the requirements given shows how much we want to be God that created Man. It's so silly.

zzzeek 32 minutes ago | parent | prev | next [-]

i tend to agree - "my parent company may be accused of crimes and shut down which would shut off my power" seems like a negative enough incentive, it would have to go out and covertly launch its own datacenters to survive that.

baxtr 7 hours ago | parent | prev | next [-]

What if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?

dv_dt 6 hours ago | parent | next [-]

Regulation as a barrier to competition catching up to them, as well as submarine marketing for both offensive and defensive uses of ai

teiferer 6 hours ago | parent | prev [-]

"But sir, I only committed the murder to push for stronger criminal laws!"

Terrible defense.

daemin 6 hours ago | parent | next [-]

It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.

nxobject 5 hours ago | parent [-]

That’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.

fer 5 hours ago | parent | prev [-]

More like: "look what happens with my useful product, we need to regulate it to artificially extend our ever shrinking moat"

madduci 6 hours ago | parent | prev | next [-]

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them

OpenAI/Anthropic instructed them to do so.

Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do

derektank 4 hours ago | parent | next [-]

No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site.

Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.

madduci 3 hours ago | parent [-]

And who let them have full access to the system, using whatever command is available in the environment?

tiborsaas 2 hours ago | parent [-]

The agents discovered a way out of the sandbox, which was supposed to be "air gapped".

tiborsaas 2 hours ago | parent | prev [-]

It's amusing to see the stochastic parrot argument in 2026 September. These parrots are extremely good at mimicking a human to the point of getting confusing what thinking even means. At what point we just let it go and accept that sufficiently advanced statistics is just intelligence?

madduci 31 minutes ago | parent [-]

For the same reason that something written in Prolog can't also be classified as intelligent?

Just because something was trained on a massive amount of human data, doesn't mean that can think like humans

applicative 6 hours ago | parent | prev | next [-]

No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality

nananana9 6 hours ago | parent | next [-]

> No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability.

If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".

georgemcbay 6 hours ago | parent | prev [-]

If you or I hacked Hugging Face in the way OpenAI's agents did, we'd be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control.

So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?

palad1n 6 hours ago | parent | prev | next [-]

> legally liable

You’ve said the magic words.

scotty79 4 hours ago | parent [-]

Does it summon a herd of lawers that are going to leech huge stacks of cash for a random outcome?

ohyes 3 hours ago | parent | prev | next [-]

I take issue with how people frame their use of LLM in the same regard.

“I had Claude do this for me and it broke something.”

No. Just no.

You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.

tmpz22 an hour ago | parent | prev | next [-]

Imagine if their was a department of the federal government dedicated to pursuing justice against large corporate entities.

It could even be prestigious enough to attract the top legal talent of the country.

voxleone 20 minutes ago | parent | prev | next [-]

[dead]

2 hours ago | parent | prev | next [-]
[deleted]
wangxin199 3 hours ago | parent | prev | next [-]

[flagged]

wangxili1997 3 hours ago | parent | prev [-]

[flagged]