Remix.run Logo
DrewADesign 5 hours ago

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.

Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.

jordanb an hour ago | parent | next [-]

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.

It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.

DrewADesign 40 minutes ago | parent | next [-]

Yeah I like Cal’s take on it, though in this context I’d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.

visarga 36 minutes ago | parent | prev [-]

The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

We also run agents, but for shorter spans between supervisions, and with much lower total budget.

dr_dshiv 2 hours ago | parent | prev | next [-]

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

huntertwo 2 hours ago | parent | next [-]

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

linkjuice4all 33 minutes ago | parent | next [-]

I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.

If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.

skinner_ an hour ago | parent | prev [-]

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.

BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.

alpinisme 30 minutes ago | parent | next [-]

The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.

Which way you say it shifts the framing. And it’s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.

For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it’s the human dimension you care about.

jwynot an hour ago | parent | prev [-]

Hiring a hitman is conspiracy to commit murder.

The hitman is charged with murder.

I imagine the same could be true of an AI lab if you could prove intent.

With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it.

Source: Prosecuting attorney for over 30 years

DrewADesign an hour ago | parent | prev | next [-]

Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.

If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.

intended an hour ago | parent | prev [-]

These situations are novel. Lax terminology is fine when it has no impact on the intuitions, clarity and conclucions of discussion.

If this was a conversation just about outcomes, then whether models think or simulate thinking is sophistry. However, the bulk of the issue here is attributing responsibility, which relies on being clear about the underlying processes at play.

We are hard wired to assume certain priors and capabilites when it comes to "human like" behavior. Anthropomorphizing LLMs implies mechanisms that aren't present, and end up distorting/complicating discussion about the process.

It isn't helped that the frontier labs, the experts in the room, generally use anthropomorphic terms to discuss model capability.

xg15 4 hours ago | parent | prev | next [-]

Yeah, fully agreed here. Most automation (such as riding a lawnmower and not putting a brick on the gas) is deterministic, in the sense that you can reasonably understand what exactly the machine will do when you run it.

But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.

Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.

AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals ("write me a traffic control system for my model railway") into low-level actions and also do so interactively.

The new thing is that the "driver" has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.

The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never "give up" when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.

bsenftner 2 hours ago | parent | prev | next [-]

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!

visarga 32 minutes ago | parent | next [-]

> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.

teiferer 2 hours ago | parent | prev [-]

What is your approach to create jailbreak incapable agents?

I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.

brazukadev an hour ago | parent [-]

an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands.

You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.

DrewADesign an hour ago | parent [-]

> I don't think knowing that will make me rich.

As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.

theodric 5 hours ago | parent | prev [-]

The idea that the agent does not actually have agency is rather discordant. We need new words!

thunky 3 hours ago | parent | next [-]

> We need new words!

The words we have are fine.

We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.

visarga 28 minutes ago | parent [-]

>> We need new words!

Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.

xg15 4 hours ago | parent | prev [-]

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

All that while still not knowing how either kind actually works.

DrewADesign 2 hours ago | parent | next [-]

Can’t agree with you here.

> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.

> All that while still not knowing how either kind actually works.

We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.

My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.

The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.

fl7305 5 minutes ago | parent | next [-]

> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.

If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong.

Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.

In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It's not a description of the internal mechanisms that evolution has produced.

As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.

They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.

In some cases we've even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn't an algorithm I would have guessed gradient descent training of neural networks would produce.

dr_dshiv 2 hours ago | parent | prev | next [-]

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

teiferer 2 hours ago | parent | next [-]

> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.

Or have they?

psma_egeliaa an hour ago | parent [-]

This is why I hate analogies. They're almost always either relevant or inapposite depending on the level of generalization we're operating on.

DrewADesign an hour ago | parent [-]

It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren’t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.

DrewADesign 2 hours ago | parent | prev | next [-]

> We have no better model for how human decision making works than LLMs

This is a claim that requires a lot of citations.

> biologically inspired

Nature inspires a lot of creation, but superficial similarities don’t mean other aspects are similar. Making an extremely realistic sculpture of a soufflé, even using a foam medium, doesn’t bring me any closer to being a chef, doesn’t mean I know anything about albumen foams, sauces, and heat transfer, and it doesn’t bring me any closer to having dinner ready. Browning on top of a soufflé is evidence of maillardization. You could pull up some studies on that and claim the brown on top of the soufflé sculpture, which I applied with an airbrush, proved that the Maillard reaction was occurring, and if another person didn’t know anything about cooking, they might even believe you. It would, of course, be completely wrong. And the other person, of course, could loudly exclaim that I can’t prove that there was no maillardization.

fn-mote 2 hours ago | parent | prev | next [-]

> Humans are constantly predicting the next moment

This is really not my experience of consciousness.

Is it yours??

Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.

God help me if that’s what LLMs are doing when I ask them to build me a web site.

oblio 2 hours ago | parent | prev [-]

> We have no better model for how human decision making works than LLMs

We do have some models and guess what, they're based on simpler animals. Which is most likely the better model.

Some other models are based on neurosciences, because we can track electrical activity.

iugtmkbdfil834 2 hours ago | parent | prev [-]

<< We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database.

Oh man, how much did you read on tip of the tongue?

cj 2 hours ago | parent | prev | next [-]

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

DrewADesign 34 minutes ago | parent [-]

Personally, I don’t think it’s different from any other faith-based motivation.

dnautics 32 minutes ago | parent | prev [-]

> LLM decisionmaking cannot possibly be like human decisionmaking

I mean how can it possibly be like human decisionmaking? It's not like it's trained on human data