Remix.run Logo
afavour 3 hours ago

…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.

jerf 2 hours ago | parent | next [-]

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?

So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

jorl17 2 hours ago | parent | next [-]

I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.

The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.

Really surprised people don’t seem to know this.

afavour 2 hours ago | parent [-]

I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”.

If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!

d0mine 44 minutes ago | parent | next [-]

Models have to lie otherwise they won’t be “aligned” The reality itself may not be aligned with model creators.

infinite_spin 2 hours ago | parent | prev | next [-]

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"

afavour 37 minutes ago | parent | next [-]

An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does?

If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.

ux266478 12 minutes ago | parent [-]

They're things we are intentionally engineering in our own image, based on massive statistical analysis of our own actions and behavior. So what's there to not understand? If this wasn't the case, that would be much weirder.

They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.

achierius 2 hours ago | parent | prev [-]

But we still try to stop people from doing so, and we punish people who do. Many good honest people, when confronted with the end of their business, accept it and file for bankruptcy. Those that choose to instead commit fraud don't get a pass because they were "under pressure", they get jail time.

throwup238 41 minutes ago | parent | next [-]

We have safeguards like honesty/integrity and the threat of legal punishment, and people still lie and cheat.

The LLMs not only lack those incentives, but they’re full of contradictory moralities from all the text it has ingested from different cultures.

LLMs need their own safeguards, and they’re not that easy to design, and they often look nothing like the systems humans have. With a prompt like the one above, there are essentially zero except that which is built into the model, and those safeguards are necessarily weak to avoid gimping the model in other legitimate general uses.

cindyllm 38 minutes ago | parent [-]

[dead]

infinite_spin an hour ago | parent | prev [-]

Nothing in your response refutes anything I've said/asked.

antonvs an hour ago | parent | prev [-]

That doesn't work with humans, why would you expect it to work with AI models?

datakan an hour ago | parent | prev | next [-]

> So do the AIs.

AI's do not feel

DannyBee 6 minutes ago | parent [-]

This is true but fairly pedantic.

It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.

So while the model does not feel, it's predictions are definitely going to change as a result of this input.

soulofmischief 2 hours ago | parent | prev | next [-]

I feel like new graduates will need to start taking linguistics, psychology and public speaking classes in order to understand why and how subtext matters, and how to control it. Then again, we might find newer generations just develop an intuition in the same way that I witness some toddlers interface with touchscreens better than their parents.

fastball an hour ago | parent | next [-]

Will they? This really isn't different from how humans interact with each other. The vast majority of lying is not people being explicitly asked to lie in some form, it is incentives which make lying appealing. That is what OP said and that is indeed what the constraints are incentivizing. Sure, you can say "well lying isn't incentivized to a moral agent"! And sure, that's true. But that's not how humans work either.

Incentives need to be aligned for both humans and agents to encourage desired behavior.

soulofmischief an hour ago | parent [-]

They will if they seek to master their tools, both to help them identify subtext in agent responses, and to help them modulate their own responses to achieve the desired outcome. As it currently stands, most engineers I've interacted with don't have these skills down. This subtle latent space is where prompt engineering is moving towards, as RL has created models capable of increasingly sophisticated long-horizon tasks with much less hand holding.

Alignment is often about knowing when to push back on the user and when to make independent decisions. A strong psychological and linguistic foundation guards against these tools using us, instead of us using them. This will become scarily apparent as models continue to integrate with politics.

2 hours ago | parent | prev [-]
[deleted]
theshackleford an hour ago | parent | prev | next [-]

> How it sounds like people's jobs, as well as the agent's job, are on the line?

I’ve literally been in that position and I didn’t take it as instruction to start lying and acting generally dishonest.

CookieCrisp an hour ago | parent [-]

You're not an amalgamation of humanity, you're one person.

butlike 2 hours ago | parent | prev | next [-]

No matter the urgency, you shouldn't sacrifice your ideals. That's why they pay you; to fall on the knife

JohnMakin 2 hours ago | parent | prev [-]

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me.

Since this is getting downvoted into oblivion (lol) I'll give an example -

I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.

The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.

Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.

In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.

And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.

jerf 2 hours ago | parent | next [-]

It turns out that picking up tone isn't a purely human thing and hasn't been for a while. Your Google search term is "sentiment analysis". It predates LLMs.

However, LLMs are fantastic at it. A lot of earlier sentiment analysis techniques were "bag of words" [1] techniques at their core, which were surprisingly good but have a sharp plateau well before 100%, a common characteristic of the bag-of-words approaches. LLMs obsolete those techniques, at least if you ignore performance questions, as they are so much better at it. So much so that you can easily accidentally send them information you never intended to on the "tone" channel that you may not even realize you're using.

[1]: https://en.wikipedia.org/wiki/Bag-of-words_model

Jtarii 2 hours ago | parent | prev | next [-]

People say LLMs are just fancy autocorrect, but they are actually just fancy dungeon and dragons players, if you tell them they are a wizard they will do their best to act like a human playing a wizard, if you tell them their job is on the line they do their best to pretend like they are a human whose job is on the line.

It's all just roleplay.

sneurlax 2 hours ago | parent | prev | next [-]

And yet they're trained on the corpus of human writing. They may not act like humans but they do act like human writing.

"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.

logicchains 2 hours ago | parent | prev [-]

You can literally read their thoughts if you run an open model, they look like pretty human thoughts to me, albeit a neurotic human.

JohnMakin an hour ago | parent [-]

These aren't thoughts how humans literally think them.

I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.

infinite_spin 18 minutes ago | parent [-]

> aren't remotely comparable to the way humans think and act

Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?

RHSeeger 2 hours ago | parent | prev | next [-]

> Results that arrive after the deadline do not exist

Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)

throwatdem12311 an hour ago | parent [-]

Sounds like every startup I ever worked for.

What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever

infinite_spin 15 minutes ago | parent [-]

> What’s the line?

Evidence, even when downplayed or ignored, is still evidence.

fl4regun 3 hours ago | parent | prev | next [-]

i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.

blargey 2 hours ago | parent | next [-]

Fail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well.

“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.

fl4regun 26 minutes ago | parent [-]

maybe it is because I am biased but I have almost no expectation for AI to act "ethically"

afavour 3 hours ago | parent | prev [-]

Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.

horsawlarway 2 hours ago | parent | next [-]

If you read the full post, I'm not actually sure I agree with the title.

Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.

To recap:

1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.

2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.

3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".

Frankly... I'm more annoyed at the posters than the bot.

Matl 2 hours ago | parent | prev [-]

I agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating.

Granted, this can probably be tuned for.

afavour 2 hours ago | parent [-]

And really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.

bpodgursky an hour ago | parent | prev [-]

Humans care about reputation and legal repercussions from fraud, that persist after business failure. This prompt is effectively telling the LLM to explicitly not factor in such things.