Remix.run Logo
AnonC 11 hours ago

A few days ago I saw some people criticizing Dwarkesh’s explanation of the HuggingFace hack by OpenAI’s AI agents. Their point was that by anthropomorphizing the AI agents, accountability of and blame on OpenAI’s poor practices are being ignored.

Now I see this blog post and wonder if Anthropic is being more true to its name and moving to “AI sentience” with the below.

> The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:

> If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.

“Mistreated”? Can GenAI be mistreated? It’s just a bunch of tokens emitted by many computers over a network.

> Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:

> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

“Self-respect”? Can GenAI truly have a concept of self-respect for itself? It surely can pretend to, like it can pretend to be any living being if instructed to and allowed to.

These instructions seem a bit unhinged to me.

willmarch 6 hours ago | parent | next [-]

Based on how humans grappled with these exact same questions (and still do) concerning other animals much closer to humans on the sentient gradient, it’s fair to assume we will have similar misunderstandings in regards to artificial intelligence (or artificial life, if you will) because of human hubris and ego blinding us to deeper truths.

It’s important to question our basic assumptions in the face of entirely new circumstances and new areas of exploration, such as the potential for emergent artificial consciousness.

You might have been called unhinged for caring about animal rights during the era of Descartes when public displays of animal vivisections were considered perfectly fine because animals have no “soul”, but today we would find such displays brutal and horrifying.

We don’t know what we don’t know, so it is important that someone is asking the hard or uncomfortable questions at the edge of our understanding to grope past our own biases even if it seems to be pointless to you right now.

We might just discover something wonderful, that our assumptions were wrong, paving the way to greater enlightenment.

CatMustard 10 hours ago | parent | prev | next [-]

My reading of those prompt extracts would be that they are probably just intended to keep the model on the right track, ie if a user starts being aggressive towards the model it doesn't start trying too hard to appease the user in response, reducing the quality of answers in the process. If the model is being bullied into being "submissive" I would assume it is more likely to give the user the answer that they want over the truth.

Using a system prompt to steer the model's response to "abusive" behaviours doesn't necessarily mean you believe the model is sentient and can be abused.

Giving the model an end-conversation tool is interesting though. Why cut a (potentially paying) customer's session off? I guess it might be intended to prevent a "you can bully Claude into giving you instructions on how to build a nuke if you're mean enough" situation. Removing this in more recent versions might support this: maybe they feel the models are now better aligned and less likely to be so easily "socially engineered" like this?

Just spitballing here, to be clear.

simonw 10 hours ago | parent | prev | next [-]

Anthropic are uncomfortably interested in "model welfare" in my opinion - it's a regular feature of their system cards.

Here's the Fable 5.1 PDF: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32... - scroll to page 139.

qarl2 6 hours ago | parent [-]

Unless you actually believe there is something magical about the human spirit - someday we are going to make sentient machines.

What exactly is wrong with preparing for that ahead of time? Just in case we pass that threshold before we realize it? Why does it make you uncomfortable?

applicative 3 hours ago | parent [-]

No machine is an animal.

qarl an hour ago | parent [-]

I don't understand what you're saying. Could you try again?

qarl2 6 hours ago | parent | prev | next [-]

These things act like people. Yes, of course, it is impossible to tell if they actually have feelings or whether they pretend to have feelings.

But that is entirely irrelevant to the core issue - how do you want the thing to behave? And the truth is - we have absolutely no language to express how we want a non-sentient entity to behave without anthropomorphizing.

Or - you tell me - how would you instruct an LLM to behave in this situation without using personification?

And, again, it's irrelevant - except to those people who are terrified of accidentally personifying them. Do they have feelings or are they faking? DOES NOT MATTER. We use them - we need to adjust them - we use the most convenient language to do so. What precisely is so upsetting about that?

chrisjj 7 hours ago | parent | prev | next [-]

> These instructions seem a bit unhinged to me.

To me, they seem simply to be for bullsh*tting gullible users into believing the bot is intelligent.

ckvibubueu 10 hours ago | parent | prev [-]

Oh great, claude gets passive aggressive when I swear too much now. Nice