Remix.run Logo
disillusioned 12 hours ago

It's also routinely failing the car wash question across all models now, which wasn't the case a month ago. :-/

Seeing some things about how the effort selector isn't working as intended necessarily and the model is regressing in other ways: over-emphasizing how "difficult" a problem is to solve and choosing to avoid it because of the "time" it would take, but quoted in human effort, or suggesting the "easier" path forward even if it's a hack or kludge-filled solution.

tetraodonpuffer 3 hours ago | parent | next [-]

it does feel something in the hidden system prompt makes it try less hard, so many times in the past several weeks I have found divergences with what was in plan and looking back at the jsonl it's always some variant of "doing it this way would be too complicated, let me take this hardcoded way out". If asked to review the change, it will find it, and it will say also yeah I agree prompt said not to do this, but I did anyways, not sure why.

As others have said, anthropic is between a rock and a hard place, you can't scale compute as quickly, and the influx of new accounts has definitely made things tough for them: I think all the "how is claude this session 1/2/3/4" questions that keep coming up must be part of some a/b on just how far to quantize / lower thinking while still maintaining user satisfaction.

andai 9 hours ago | parent | prev | next [-]

> over-emphasizing how "difficult" a problem is to solve and choosing to avoid it because of the "time" it would take

I heard a while back Claude refused to attempt a task for days, saying it would take weeks of work. Eventually the user convinced it to try, and it one-shotted it in 30 seconds.

apetresc 8 hours ago | parent | next [-]

For days? Someone spent days trying to convince Claude to do something?

layer8 7 hours ago | parent [-]

If you asked yesterday, and asked again today, then you asked for days. OP might be trying to express that it wasn’t just a temporary fluke.

empath75 2 hours ago | parent | prev [-]

I have noticed refusals as context windows grow.

_blk 11 hours ago | parent | prev | next [-]

Awesome, I didn't know about the car wash question.

Totally true, also tokens seem to burn through much faster. More parallelism could explain some of it but where I could work on 3-5 projects at once on the max plan a month ago, I can't even get one to completion now on the same Opus model before the 5h session locks me up..

colechristensen 4 hours ago | parent | prev [-]

>“idgaf about risk you coward, waste some time just do it and stop bitching”

The above was a successful prompt to get Claude to stop whining about effort, difficulty, and time.

Unfortunately abusive language well placed is an effective LLM motivator.