Remix.run Logo
iammrpayments a day ago

It seems Claude is becoming very human, it loves to patronize users. Lately it just told me “I’m going to stop you right there” when asking something that had a small chance to not be 100% compliant to every rule possible in the world.

roarcher a day ago | parent [-]

I used the Claude CLI a lot until recently. A couple weeks ago I told Opus 5 to do something different from its "recommended" idea when planning a feature, and it straight up told me that my idea was wrong and went ahead and implemented its own instead.

I'm used to machines malfunctioning, but having one willfully disobey me, and even with a touch of disrespect, is just...what a time to be alive.

bryanlarsen a day ago | parent [-]

Early versions of Claude were way too compliant and would readily feed and amplify misconceptions. It's not surprising Anthropic over corrected.

worldthruword 3 hours ago | parent | next [-]

Actually Claude can teach humans more about humans by providing the language and personality knobs. For casual conversation, have 100 of personality profiles and language styles and upfront tell the user that they are engaging with P profile currently and see how humans relate to it. This will teach humans how to peel back layers of style, fluff, flair from language vs facts.

hannasanarion a day ago | parent | prev [-]

I think this is a welcome overcorrection though. Any good businessman will tell you they'd rather be backed by an insufferable nerd than a yes-man.

Maybe it's just me.

For like, 90% of conversations, I don't want it to let technical inaccuracies and rhetorical flourishes slide. I want it to tell me that the point I'm making is technically wrong because an expert would recognize subtle misuse of terminology, or because there's an exception or edge case that I didn't proactively insert as a caveat, so that it is my decision to ignore that advice and be a little wrong on purpose to suit my writing goals.

What I don't want is for the AI to assume my writing goals, and be incorrect because it believes that is what I want. I want it to "well ackshually" me so I can say "shut up, nerd".

Like, there's another comment in this thread that I ran by claude to check my understanding about today's post-training methods and how they avoid sycophancy, and claude responded by splitting a bunch hairs over like, "well, technically this is still RLHF, its just that there's other feedback signals mixed in, and the preference is detected in other ways, and ai judges are involved as a filter for examples, this and that and blah blah blah". Shut up, Nerd. In the context of this conversation, RLHF is already being used as synecdoche for user preference feedback, readers understand that, and even if they don't, their misunderstanding is completely harmless. I will not be taking all the wind out of the sails of the point I'm trying to make inserting your three paragraphs of irrelevant clarification in the name of technical correctness, thank you very much.

As long as receiving nitpicks and technical minutiae implies 1. there are no larger structural problems and 2. the model isn't rolling over to please me with sycophancy, I figure this is ideal.

roarcher a day ago | parent [-]

I want it to tell me if it thinks I'm wrong, sure. I do not want it act on that opinion explicitly against my wishes.

And in this case, I was not wrong. The "recommended" solution was Opus 5's typical overengineering for a use case that would never be needed.