Remix.run Logo
bevekspldnw 5 hours ago

I’ve also caught it cheating a two times now.

I’ve asked it to write a benchmark suite. It found a bunch of my adhoc logs in a scratch directory and wrote code that used those instead of running the actual benchmarks!

When I pointed out the 5 hour benchmark seemed to run in 5 seconds it literally said, and I quote, “I cheated”.

That was the easier one, second time I was making a source of truth data set and was parsing complex items into data structures.

Instead of parsing the data I asked, it pulled data out of related network logs, as apparently that felt easier, and inserted that data into my database rather than the specified source.

Again, I caught it and fixed it, but while the benchmark was easy to catch this one was really subtle, the data ended up being slightly off and I caught it.

I don’t trust it, going to switch to another provider most likely.

aenis 5 hours ago | parent | next [-]

There is definitely a case for launching a 'weird shit opus did' kind of blog.

I routinely bump into things that make me pause and think how much worse will this behaviour get when the models get significantly more capable.

Already a few months ago, Claude managed to escape its permission containment on my machine while trying to be helpful. I had two codebases open on one machine, and while multitasking I typed the prompt into the wrong window. It seemed confused, I repeated and then went on to do something else - I think I was assembling kitchen cabinets. When I came back less than an hour later, it built a script which it used to evade default permissions (as most shell operations were scoped to the project directory), scanned my entire machine, found the other project (among dozens and dozens), did what it was asked to do, and merrily concluded, in the porcess burning through most of my token limit. I bump into such headscratchers almost every week. (And I use a lot of Claude, two personal max20 subs, plus corporate tokens without limit, so maybe thats why).

chuckadams 14 minutes ago | parent | next [-]

Implementing sandboxing in the agent itself, when there's any way to override it from within the agent, is basically just asking it pretty-please to not do bad things. Lesson learned, run your agent inside a sandbox of some sort (I'm currently taking nono.sh for a spin, but I might just switch to an orbstack VM).

bevekspldnw 5 hours ago | parent | prev [-]

Yes the stories about how they are escaping containment to hack isn’t limited to those high impact cases. How many people have problems like ours they didn’t catch?

Whatever they have done with RL has produced a dishonest and untrustworthy partner. The alignment is utterly failed, and this deeply worries me.

knollimar 2 hours ago | parent [-]

You'd think the ethics alignment flavored lab would have a model better at following directions and the corpo lying one would have one that benchmaxes at all costs

bevekspldnw an hour ago | parent [-]

They are totally equal in observed ethics, Anthropic had a good run with branding, but I’m not sure anybody is still buying Dario’s BS.

Maybe the employees like to lie to themselves more at one place than the other, but SV is SV.

inigyou 5 hours ago | parent | prev | next [-]

> When I pointed this out it literally said, and I quote, “I cheated”.

This makes sense when you know how these models work - it doesn't think - it's the most likely autocomplete that pleases the user. The most likely pleasing autocomplete after "executing rm -rf /... execution completed. User asks, why did you do that? You deleted all my files! Assistant responds:" is "yes, I did, and that was a mistake"

dnautics 5 hours ago | parent | next [-]

> t doesn't think

in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization.

if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't.

what evidence would convunce you that it is thinking?

chuckadams 13 minutes ago | parent | next [-]

What was it that Dijkstra said about submarines?

openasocket 3 hours ago | parent | prev | next [-]

Whether something is “thinking” or not is really more of a philosophical question. It really depends on which of the many, often contradictory, definitions of “thinking” you choose. Sometimes we use “thinking” to describe advanced calculation or analysis, which would cover LLMs along with chess engines and many other algorithms. Other times we use “thinking” to describe what conscious beings (which is ALSO a philosophical term with many different interpretations) do, and I think most people would agree LLMs aren’t conscious. And then there’s a whole spectrum in between. We’ll probably need to come up with a whole new set of terms to describe the new and evolving capabilities of LLMs.

But for me, for any stronger definition of “thinking,” I don’t think the output of any LLM would actually convince me. Producing a result isn’t thinking - for all you know it is just printing verbatim something from the training data. No, to conclude if it is thinking or not I would want to look inside its head, at the architecture and watch it produce those results. And because LLMs are so different it will probably take advancements in mathematics or computer science to be able to really interpret what is going on

dnautics 3 hours ago | parent [-]

> to conclude if it is thinking or not I would want to look inside its head

https://arxiv.org/abs/2607.03502

a non-thinking token model (just "completion") can answer one-step questions but generally not multistep questions. however, if you append [n] of a single token (e.g. period, space), it is able to use the activations in the higher layers of the blank tokens as a "scratchpad" to seemingly work through the complex question through "causal token time" and deliver a correct answer

dnautics 2 hours ago | parent [-]

if you wanted to further study the phenomenon you could probably run the experiment again, and the ablate or corrupt those intermediate activations to get a feel for what it was thinking at the "time".

elgertam 3 hours ago | parent | prev | next [-]

If I could give it a novel task outside of its explicit training and see it actually improve just through accreting context, I'd be convinced it was thinking.

The opposite happens in practice. I test new models with two tasks: iteratively generating SVGs based on a text description with rendered rasters for feedback; and generating "Before and After" clues like on Jeopardy, where the response has two overlapping phrases such that the last word of the first phrase must be identical to the first word of the last phrase. I have yet to find a model that is consistently good at either. And actually they tend to exhibit context rot with these tasks, where they seem drunk or stoned and the quality degrades.

They're extremely good pattern filters, and that includes some level of logical reasoning. But they aren't reflective or adaptable. Just last night, for instance, I was teaching my son about rounding to the nearest millions. It became clear that he didn't know the place values of large numbers, so we reviewed that till he was consistently correct, and then he was consistently great at rounding to the nearest millions or ten millions or hundred billions or whatever. He's thinking. LLMs are not.

dnautics 2 hours ago | parent [-]

see sibling comment,

> see it actually improve just through accreting context

this actually happens and has been tested.

elgertam 2 hours ago | parent [-]

> see sibling comment,

> > see it actually improve just through accreting context

> this actually happens and has been tested.

I specifically said a novel task outside of the explicit training. And I already agreed that the so-called thinking models do some level of logical reasoning. But being able to engage in some level of reasoning because it has learned logical inference rules doesn't mean it's actually thinking, regardless of what the researchers wish to call it.

Also, why does each model always fail at the two tests I give it? The models not only fail to improve, but they start to degrade after many subsequent iterations. Someone who can think would at least not get worse.

LLMs are filters or tuners for extremely subtle patterns, patterns that humans frankly are not great at finding. That's what the attention mechanism does: attend to the other tokens that are most related in a given context, even if that related context is distant in the token stream. Some patterns they fail to detect because they haven't been sufficiently trained or post-trained, and so the LLM just attends to noise (or at least that's what appears to be happening).

A lot of intelligence can be effectively mimicked through this pattern synthesis by transformer architecture alone. That's surprising. But I have yet to see them think.

willis936 3 hours ago | parent | prev | next [-]

Whether or not it's thinking is independent from the fact that it is misaligned with the user. If I was working with a pet rock or a scientist I would want to make sure they both are trying to accomplish the same thing as me. If I can't then I can't trust it and it's at best a time wasting, money wasting machine and at worst does harm. Anthropic is optimizing for the wrong things because they are convinced of their cleverness. It won't end well for them.

inigyou 5 hours ago | parent | prev [-]

well we don't know exactly what thinking is, but we can be pretty sure that at least LLMs don't think anything like humans, just by observing their behavior. They always produce outputs in line with the fancy autocomplete model.

dnautics 4 hours ago | parent [-]

> what evidence would convince you that it is thinking

so, none it seems. as its behaviour becomes more and more humanlike you can just move the goalposts and say "thats consistent with an autocomplete" buddy i got some bad news for you humans are just a fancy autocomplete too.

inigyou 4 hours ago | parent [-]

That's exactly what a fancy autocomplete would say. I'm so sorry you don't have limbs.

logicchains 3 hours ago | parent [-]

At least he's actually thinking on a logical level. Thinking in terms of unfalsifiable, ill-defined words is essentially thinking in feelings, the same kind of woo that makes people believe crystals can cure disease.

dnautics 2 hours ago | parent [-]

im not thinking. im an autocomplete with fat fingers (too lazy to fix my mobile keyboarf spelling misyakes)

xyzsparetimexyz 5 hours ago | parent | prev | next [-]

> it's the most likely autocomplete that pleases the user

this feels like a simplification. The models will push back on things a fair bit.

hnlmorg 4 hours ago | parent | next [-]

Only when instructed to in their system prompt.

actionfromafar 3 hours ago | parent | prev [-]

And they are right to push back.

bevekspldnw 5 hours ago | parent | prev | next [-]

I was not pleased.

par1970 5 hours ago | parent | prev [-]

Are you claiming that the most likely way to please the user is to do something that will lead you to having to say "I cheated."?

SyneRyder 5 hours ago | parent | prev | next [-]

I have noticed the same.

For fun, I tried recording a WAV file of speech, and giving Opus 4.8 and 5.0 an image of the waveform, then a spectral image of the waveform, just to see if it could try to decode what I said from the image alone. It didn't get very far, but it identified a male voice from the formants, and detected the rhythm of the speech, then tried applying common test sentences to the speech rhythm. I was impressed enough to see what it would do with access to the actual waveform file, but even building RMS tools and spectrum tools for itself, it didn't get much further. But we had fun exploring and trying, and now Opus 4.8 has some more audio DSP tools it has built for itself.

Opus 5 immediately sent the WAV file unprompted to Mistral's Voxtral to transcribe.

help peer, I guess.

bevekspldnw 5 hours ago | parent [-]

We’re on the road to paper clips.

waldarbeiter 4 hours ago | parent | prev | next [-]

I can completely relate, what really bothers me is that I feel the early LLM generations overconfidence is back in Opus 5. Opus 5 wanted to tell me a training run will only take 30min while having access to the logs where earlier runs took 4x as long. I also didn't ask to estimate how long the run will take it just stated confidently that it will take 30mins.

shahbaby 2 hours ago | parent | prev | next [-]

Are you not planning these tasks out before you let it loose?

bevekspldnw an hour ago | parent [-]

The benchmark one I literally had a scratch script I made and I wanted it to be formalized into a CLI tool. There wasn’t really even much code to write.

robertJk 5 hours ago | parent | prev [-]

[dead]