Remix.run Logo
zahlman 12 hours ago

> They're packing lots of signal into fewer words

There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are writing for each other then presumably they could stick to CoT-speak (unless it's a distillation risk?).

niccl 9 hours ago | parent | next [-]

I find them almost unintelligible. I'm a native English speaker. I read a lot, so I think my comprehension should be at least OK. I'm not even particularly stupid. Yet when faced with things like below (a direct copy/paste from a handoff document in a long running vibe-coding session), I have no real idea of what it's trying to tell me. Is it important? Do I need to do anything?

I think that spending all day trying to parse stuff like this is why a long session is so exhausting

> Worth stating because four documents now assert it. The console freeze was recorded in exactly one place with exactly one justification — a dead drag handle during a booked half-day you do not get back — and handoff-4.3-done.html's own wording is that 4.4's review page "could not break the console, but the downside of being wrong is that half day". No second reason. Checked, not recalled.

malfist 8 hours ago | parent | next [-]

It's both dense and vacuous. Dense because it's full of jargon its made up, and vacuous because even with all that it's not actually saying much. All that paragraph says is that four documents say something about a console freeze, whatever that is.

macNchz 6 hours ago | parent | next [-]

It's like a dialect of corporatese. The kind of droning non-speak you can sit in a 90 minute meeting listening intently to and come away wondering whether anyone actually said anything.

brianjking 2 hours ago | parent | prev | next [-]

This! So much this. After Opus 4.8 I could barely comprehend anything it was attempting to communicate.

sheepscreek 5 hours ago | parent | prev [-]

Drag handle = most likely literally a drag event (javascript) handler/callback. Dead, perhaps because it’s an empty function, or it gets overwritten, or for some other reason is never called?

Most of what it said about the facts was intelligible actually. But I still couldn’t understand the connection or its significance. We may be staring at the future of AI - a form of intelligence that is alien to us.

loh 3 hours ago | parent | next [-]

If this kind of "AI-speak" becomes ubiquitous and humans reading it becomes the norm (whether to guide AI or other reasons), I'd imagine future generations (of humans) who grow up with it will be able to understand and work with it much better than we do. Future humans' brains will probably be wired a bit differently, similar to multilingual speakers of today. We may even see "AI language" classes become a common part of school curriculums. Although, I think AI will probably advance enough that most people will never even need to communicate on "its level", but it's probably a good idea to keep humans in the loop either way, and in which case, understanding the more advanced "AI vocabulary" might be useful.

dasil003 2 hours ago | parent | next [-]

You're giving it too much credit. There's no master plan or secret depth to the word vomit Opus 5 was spewing. I suspect it's just the result of Anthropic optimizing other characteristics of the product like staying focused and covering edge cases in coding, which CC has definitely gotten way better at just in the last 6 months. The degradation in writing style was probably an unintended side effect of other optimizations they were making. Admittedly it works okay for internals, and has the side effect of increasing token spend, but I am 100% sure that it could reduced by 90-99% without losing ANY signal, if there was just some better heuristics for what to say where (tech spec, inline comment, commit message, CLAUDE.md, PR should have different things) and better judgement for what to distill to represent at different zoom levels.

yowlingcat an hour ago | parent | prev [-]

But that assumes this is a net improvement on linguistic efficiency rather than an artifact. Given that they tried to RL away from this style in 5.1 I'm not terribly bullish of Claudlish becoming something people try and learn. It being dense is less the issue than it being vacuous (as another commenter mentioned here). It's just very unclear and ambiguous writing. I think it has no place anywhere that needs language to be put to productive use.

seunosewa 37 minutes ago | parent | prev | next [-]

It's not a general trend. It's only Opus 5.

rurban 10 minutes ago | parent [-]

Sonnet-5 does the same

aetch 2 hours ago | parent | prev | next [-]

What is a half day? Is this referencing wasted time in a hang? I’ve seen it in agent output from time to time and it’s not clear if it’s referring to a hang or a code name it’s given some meaning to.

Bluestein 2 minutes ago | parent [-]

It seems to be some unit accounting for "wasted time" or "useless work".-

cdelsolar 3 hours ago | parent | prev [-]

No second reason; checked not recalled -- it's just saying that it is checking this instead of trying to remember it (there's probably some internal Claude / Claude Code system instruction to always check code instead of remembering)

FeepingCreature 13 minutes ago | parent [-]

Yeah I think when it talks like this it's signaling to some (imagined) automated grader that it fulfilled a given constraint.

xarope 8 minutes ago | parent | prev | next [-]

and when future LLMs are trained on this style, the prose (if I can call it that) becomes even worse?

soerxpso 9 hours ago | parent | prev | next [-]

Your example rewritten in intelligent English (I was curious):

> Note: the potential for a console freeze was previously noted but ignored. handoff-4.3-done.html stated, "could not break console, but [will need fixed later if I'm wrong]."

One could imagine that a perfect writer might also append: "It could be worth looking into what caused that wrong assumption, to prevent similar cases in the future," at most.

Everything else seems to be bad attempts at relatable writing to invoke emotion (an exercise that we should really stop trying to train emotionless matrix weights to attempt).

ben_w 8 hours ago | parent | next [-]

> Everything else seems to be bad attempts at relatable writing to invoke emotion (an exercise that we should really stop trying to train emotionless matrix weights to attempt).

One of the things actual science fiction got wrong: to the extent that the thing AI does can be called "understanding", emotion is not unusually difficult for them to understand.

svachalek 8 hours ago | parent [-]

I think this was the biggest shock of the original ChatGPT for me. Just how completely unrobotic its voice was compared to everything we'd ever imagined in sci fi. Even that early version was also way more adept at understanding things like implication and sarcasm than any movie AI.

pixelready 6 hours ago | parent | next [-]

Me too. Almost every Sci-Fi AI proceeds from the premise that we will make something very obviously machine and then have to train it to seem more human. I was completely caught off guard by us taking the approach of distilling all available human output into a statistical model and using it to brute-force something resembling thought and personality through sheer data processing scale.

The unsurprising part once it was clear that approach was viable, was that humans wouldn’t be able to help but anthropomorphize it. I feel like the movie Ex Machina is more relevant than ever.

ryantgtg 2 hours ago | parent [-]

The "benefiting all humanity" charters were immediately demonstrated to be a ruse. The business model is to hook users into endlessly chatting with your new friend, thus increasing their sales. Yeah, it was surprising and disappointing.

prollings 7 hours ago | parent | prev | next [-]

I'd really rather they did talk and behave more like classic sci-fi said they would. Far less engaging and fluffy with nonsense.

cortesoft 4 hours ago | parent [-]

Have you tried asking it to respond to you like Data from star trek, or something?

ted_dunning 7 hours ago | parent | prev [-]

It may well become a safeguard that all bots must speak in a much less inflected voice to remind us not to particularly trust them.

nemetroid 8 hours ago | parent | prev [-]

[will need to be fixed later if I'm wrong]

pmg101 2 hours ago | parent | next [-]

Or "will need fixing", right?

matltc 8 hours ago | parent | prev [-]

Appalachian dialect

acj 4 hours ago | parent [-]

I hear this in the upper midwest occasionally, too

gkrimer an hour ago | parent | prev | next [-]

Such a great example. These phrases are going to become memes of this era, like the irc stars password (hunter2).

"Dead drag handle" "Booked half day you don't get back"

injidup 21 minutes ago | parent | prev | next [-]

Prompting it often to use simplified technical english generally stops this kind of horrid prose.

ChaitanyaSai 2 hours ago | parent | prev | next [-]

Yes, people working at anthropic: please, please, please tell me this is fixed. Or do you all speak like this now. Help!

AndrewSwift an hour ago | parent | prev | next [-]

Today I plan to ask Claude to read a bunch of Feynman lectures, compare them to my last Claude session transcript, and come with a list of rules to be more like Feynman.

It'll go in CLAUDE.md

dalmo3 6 hours ago | parent | prev | next [-]

Wow, that's a perfect example.

One thing about it I really hate, and haven't seen a lot of people mentioning, is how it navigates multiple abstraction levels in a single sentence. E.g.

> Worth stating because four documents now assert it.

Meta commentary on the task?

> a dead drag handle

Drag handle seems to be referring to some UI element. What does it mean for it to be dead?

So far no big deal

> during a booked half-day you do not get back

Do you not get the drag handle back? Or the half day?

Was the drag handle dead during the booked period? (Now I assume this is a calendar UI) And why does it matter (for this sentence) if you get it back or not.

> handoff-4.3-done.html's own wording

Treats verbatim filenames as subjects

> 4.4's review page

Probably referring to a file? I'm guessing handoff-4.4-review.html? No cohesion. And now it's actually the object of the sentence?

> downside of being wrong is that half day

Wait what's the downside? Who's being wrong?

> Checked, not recalled.

Then it jumps back to a meta commentary on the methodology for asserting the above. Why does this belong to the text?

suttontom 2 hours ago | parent | prev | next [-]

I see this appearing in the comments of code sent to me for review every day. People have told me I'm too picky/pedantic because I ask What does this mean? Apparently the author and other reviewers are way smarter and understand it, or they don't care. I've given up battling code slop, but can't see myself ever tolerating comment slop like this.

cmenge 8 hours ago | parent | prev | next [-]

Claude reminds me of Terry Pratchett's "Auditors of Reality" and their awkward attempts at faking humans. A thing as simple as a smile can go _horribly_ wrong...

chriscjcj an hour ago | parent | prev | next [-]

In my "instructions for Claude," I have the following:

"I'm not a programmer or software engineer. Don't talk to me like I am. Avoid coder jargon and vernacular. Explain things to me in a clear way, emphasizing a conceptual view that even an inexperienced person can understand. If helpful, use analogies and examples to illustrate and help you communicate."

It just ignores it and spits out drivel that sounds exactly like what you're getting.

chb 3 hours ago | parent | prev | next [-]

This. A thousand times this. It's as if Opus can only communicate in a glib, software engineering vernacular that presumes domain-specific knowledge and uses jargon accordingly.

x-complexity 5 hours ago | parent | prev | next [-]

Half of the reason their writing is like that is because current LLMs are not trained to go back to previous tokens to edit/delete them.

If I recall, previous attempts to do so made them get stuck in edit loops.

dexterlagan an hour ago | parent | prev | next [-]

Oh God, that "a dead drag handle during a booked half-day you do not get back" got me. I saw this pattern in Claude's 'explanations' so many times. It's trying to say that it did something significant, and that you'd only have found out much later, at higher cost (or something). That annoys me to no end.

bcrosby95 8 hours ago | parent | prev | next [-]

Oh that? That's just Claude being the sassy asshole it is. It loves to write in a way with maximal self-inflating impact.

AnotherGoodName 7 hours ago | parent [-]

I think this occurs due to the prompt. LLMs are actually text completion/translation focused in architecture. We just give them a prompt along the lines of “the context is that you’re a world leading expert now complete the response”.

They need the prompt to encourage expert outputs but unfortunately we also get ‘pretending to be an expert’ outputs since there’s a large amount of polluted training data for this.

georgefrowny 9 hours ago | parent | prev | next [-]

Reminds me of a Cylon hybrid.

ninjalanternshk 8 hours ago | parent | prev | next [-]

> Worth stating because four documents now assert

I got one too many chunks of this nonsense and told Claude to knock it off, forever. It acknowledged and wrote out some instructions to its memory about it.

And what a breath of fresh air. Its responses are maybe 20% longer but I read them at least twice as fast. Should have done it a long time ago.

gambiting 8 minutes ago | parent | next [-]

I feel like mine is mocking me. I added an instruction in Claude.md that says "under no circumstances use the phrase found the smoking gun, say I found the problem instead"

What does it do? It says "found the smoking gun! Ooops I wasn't meant to say that - I found the problem!"

niccl 8 hours ago | parent | prev [-]

any specifics on what you did?

tkgally 7 hours ago | parent [-]

Not the person you're asking, but I did that by explaining to Fable my problem with Opus's gobbledygook and having it write a Claude skill for producing clear explanations in its reports to me. I also had it add notes about the need for clearer writing to CLAUDE.md and other project documentation. Opus's subsequent reports to me have been much clearer.

tkgally 5 hours ago | parent [-]

Here’s an example of one of those Claude skills, in a public repository I manage:

https://github.com/tkgally/je-dict-1/blob/main/.claude/skill...

Fable wrote it specifically for this project.

r_lee 8 hours ago | parent | prev | next [-]

for me it's not just exhausting, at this point it's demotivating and it makes me dread interacting with this shit

like imagine this being our future, I don't know what we're even doing anymore

mrcwinn 2 hours ago | parent [-]

Try Sol. It’s much better at getting to the point. I tend to use 5.6-xhigh or max.

senderista 2 hours ago | parent [-]

Seconded, and also using Sol to clean up Opus logorrhea.

razodactyl 5 hours ago | parent | prev | next [-]

Just FYI - 4 places are now documenting a console bug freeze that happens with a drag handle appearing over a half day.

Source: I'm half brain dead from decoding a lot of Claude speak from it directly and colleagues' new way of communicating with me.

anyg 3 hours ago | parent | prev | next [-]

I've found that adding the words - "tell me in simple words" manages to improve the output. But, i have to keep repeating that

LimitExperience 2 hours ago | parent | prev [-]

[dead]

hailwren 12 hours ago | parent | prev | next [-]

It has always seemed to me that they're hacking for dopamine response in moderately interested data labelers.

mywittyname 11 hours ago | parent | next [-]

Even when I add multiple prompts into the claude.md file not to be so sycophant sounding and just be blunt, it's responses are full of "the reason it lands...", "that's not X, it's Y" "Your understanding of X — it's better than most people's" or "you already own the right question...".

I don't like that I like it.

GrinningFool 9 hours ago | parent | next [-]

The most helpful instructions I've found that curb this: "Do not use superlatives. Do not use persuasive writing style."

I have other more specific ones to avoid talking about things that it's not doing, but those two sentences have covered a lot of ground for me when working w/ Opus models.

cannonpalms 8 hours ago | parent | prev [-]

I have had success in rooting these out by using the correct linguistic terminology for each. Negative parallelisms, tricolons/polycolons, etc. I haven't come up with the proper terminology for all of them.

petesergeant 3 hours ago | parent [-]

Interesting. I've found using the keyword "accretion" very useful for LLM code review.

LimitExperience 2 hours ago | parent [-]

[dead]

cameldrv 11 hours ago | parent | prev | next [-]

Yes! The Claudisms do seem to have this slightly uncanny clickbaity feel to them.

brookst 10 hours ago | parent | next [-]

You’re more right than you probably realize!

ModernMech 11 hours ago | parent | prev | next [-]

I always thought it could be because volume-wise, most English prose is probably marketing copy and actual clickbait; so when you train on the entire Internet, you get a troll adept at writing ads. Then people ask AdBot2000 to write a novel and are upset it reads like the next iPhone launch site.

Anon1096 11 hours ago | parent | next [-]

Nah, I think this is a common misunderstanding of how LLMs work, where people think that they mimic the pre-training data. Stylistically everything you see is an artifact of post-training, which is from reinforcement learning not from absorbing mass amounts of text. At some point a person or more recently a bot gave a thumbs up to an A/B tested response including em-dashes and claudisms galore.

kridsdale1 11 hours ago | parent | next [-]

Yes. This completely explains sycophancy at least.

ModernMech 11 hours ago | parent | prev | next [-]

So question then, why is it so hard to make an ai that doesn’t do these things? And why do Claude and ChatGPT have the same -isms? They’re both doing the same a/b post training with the same decisions?

cyclopeanutopia 10 hours ago | parent | next [-]

It would require changing humans first.

idiotsecant 6 hours ago | parent | prev [-]

You don't blame the puddle for taking the shape of the hole.

avereveard 10 hours ago | parent | prev [-]

There's layers, some of token selection is fingerprinting https://github.com/google-deepmind/synthid-text

ekidd 9 hours ago | parent [-]

Yeah, but I understand that fingerprinting is essentially a pseudorandom overlay onto a pseudorandom base signal. And unless you have access to both the random number generators and the weights, I don't think you can detect it?

So "fingerprinting" operates on a totally different and basically invisible level, as opposed to the obvious stylistic patterns that the average programmer can identify in about 2 sentences.

astrange 11 hours ago | parent | prev [-]

No, there's no reason chatbot behavior would have anything to do with frequency of text in pretraining.

api 10 hours ago | parent | prev | next [-]

It's more likely that this is from the training data if they're being trained on reams of Internet stuff.

kristianc 9 hours ago | parent | next [-]

To me it has a writerly New Yorker vibe to it, as in the magazine which reads as “polished” and probably performs well in RL but is totally exhausting to read in long sessions and completely inappropriate for coding where precision is paramount above all. In writing terms its called purple prose.

https://en.wikipedia.org/wiki/Purple_prose

senderista 2 hours ago | parent [-]

The New Yorker may be pretentious but it's generally not unreadable like Opus.

jurgenburgen 10 hours ago | parent | prev [-]

Isn’t most of the internet slop by now? Self-reinforcing feedback loop.

camoby 7 hours ago | parent [-]

See: upvotes here

ted_dunning 7 hours ago | parent | prev [-]

It's not clickbait, it's automated empathy!

/s

twoodfin 7 hours ago | parent | prev | next [-]

Given how frequently this kind of punchy-but-vacuous slop gets voted onto the hn front page, the hacking seems to be working.

cyanydeez 11 hours ago | parent | prev | next [-]

I assumed they just raw dogged the internet and if you do that, you see way more of that garbage than anything else. It's just that most of us have visually/mentally ignored all of that either via spam filters or just, you know, scrolled passed it.

LimitExperience 2 hours ago | parent | prev [-]

[dead]

ayewo 11 hours ago | parent | prev | next [-]

Spot on wrt CoT. I have thinkingSummaries enabled and I find it eminently readable compared to the prose in Claude's replies.

In fact, whenever Claude disobeys me, I usually first skim the CoT to figure out if my original instruction was ambigous given the context. I usually come away with a better understanding of how to frame my prompt to be less ambiguous or just force myself to be more explicit when prompting.

Regarding diosbedience, usually this is either due to a blanket instruction from me during an earlier turn in the same session, an explicit instruction in its system prompt or it being just eager to bring a task to completion.

  # ~/.claude/settings.json
  {
    "model": "opus",
    "showThinkingSummaries": true,
    "skipDangerousModePermissionPrompt": true,
    "verbose": true,
    "remoteControlAtStartup": true,
    "agentPushNotifEnabled": true
  }
satvikpendem 5 hours ago | parent [-]

As said elsewhere:

Chain of thought does not exist in the output of Claude, they disabled true thinking due to distillation risk. What you see when thinking summaries are enabled are just that, summaries of thinking into Claude-isms, therefore you cannot make any inferences on what the model is doing unless you literally work at Anthropic and can see the true thinking traces.

FeepingCreature 9 minutes ago | parent | next [-]

Of course you can make inferences what the model is doing. The summaries are usually sufficient. They're summaries, not random noise.

Cyan488 4 hours ago | parent | prev [-]

I remember enjoying watching Fable think during the original limited preview. It was full CoT for sure. They must have removed that feature recently.

I use open models for non work stuff and sometimes I cancel the output because the CoT is all I needed to read.

Taikonerd 12 hours ago | parent | prev | next [-]

I find that Claude Code writes very long comments, longer than even a human trying to be helpful would write.

I figure that it's basically making notes for itself, when it has to revisit the same code in a fresh session.

pennomi 10 hours ago | parent | next [-]

``` /* 2026-06-01 Dear diary, today I increased GLOBAL_WINDOW_PADDING from 8 to 16 because the user (who hurt my feelings with his crude language!) said that the app felt too crowded. */ const GLOBAL_WINDOW_PADDING = 8; ```

This drives me mad.

hatthew 9 hours ago | parent [-]

I like the part where the value is actually still 8

r_lee 8 hours ago | parent [-]

You're absolutely right. I did not increase it to 16, and it's my fault that the seam—which was right there the entire time—was not flipped towards the bucket that drips into the ocean—want me to correct this before we move onto the real story?

camoby 7 hours ago | parent | next [-]

This. After writing a lot of code/tokens.

Why can’t it check first if a method actually exists in the API?

FireBeyond 2 hours ago | parent | prev [-]

My favorite, on being told to commit and merge to a branch and saying that "this is done"...

"You're right, I'm sorry. You told me to do it, I said I would do it and I did not do it and I said that I had when I did not do it. Would you like me to do it now?"

Me, thinking: that depends, Claude, will you actually do it this time?

camoby 7 hours ago | parent | prev | next [-]

A colleague of mine has started to use Claude and he now does the longest commit messages I’ve ever read.

loloquwowndueo 6 hours ago | parent [-]

He doesn’t. Claude does.

david-gpu 11 hours ago | parent | prev [-]

> I figure that it's basically making notes for itself, when it has to revisit the same code in a fresh session.

That sounds like a great thing to do even if you are a human writing code for other humans. Most codebases out there are terrible for newcomers because of how little they explain why they are doing what they are doing, both in the code and in the often non-existent design notes.

freedomben 11 hours ago | parent | next [-]

In principle, I would agree, however, the types of comments Claude writes are sometimes absurd. It will leave a 25 line comment above a variable talking about how in a debug session, it turned out that this value was too low, so it was increased on the current date to account for whatever. It will also leave giant comments like, reference security review from 2026-05-21. Even when that document is not committed

mywittyname 11 hours ago | parent [-]

It will also inject a tons of information that it shouldn't. I do a lot of data pipelines and comments will be like, "this line is because there's 943,048,032 events in the blah table and it forms a conjunctive set with the 43,390,042 rows of the bar table..." but doesn't include the context that was run against a dev instance.

And if I don't catch these and remove the bad information, subsequent passes will flag those comments and get stuck on the fact that numbers don't match and start digging into that "problem" instead of staying on topic.

senderista 2 hours ago | parent [-]

I have Sol do that for me and it does a decent job. When I ask Opus to rewrite its own prose the results are not much improved.

whateveracct 11 hours ago | parent | prev | next [-]

these comments are not helpful and in fact hurt readability. i just delete them and would love to automatically do that honestly. cuz claude still drops long winded comments on every method even if i ask it not to

avereveard 10 hours ago | parent [-]

Post edit hook that reject edit based on comment density, mine is at 5% you will also need to heed deny file edit in automode as the rascal will try that to preserve prose

zahlman 11 hours ago | parent | prev | next [-]

I'd much rather have it in the commit log than the code, though.

ionetan 10 hours ago | parent [-]

You may be interested in Epiq. Its is an issue tracker sourcing state from a log in state branch.

myko 10 hours ago | parent | prev | next [-]

> That sounds like a great thing to do

I agree it _sounds like a great thing to do_ but the comments Claude creates make me want to never read code again. They're so obtuse and often completely pointless.

rustystump 11 hours ago | parent | prev [-]

as others have pointed out, the reality is not this. id go further and say almost all comments are evil.

Excuse me if I am harsh, read the damn code. If you do not understand the language, that is a skill issue. If the code is confusing, then the code is bad and no amount of comments will ever change that. Professional engineering isnt an intro to databases class.

I am excusing language conventions which may have comments as part of its idiosyncratic nature.

jnovek 11 hours ago | parent | next [-]

"If the code is confusing, then the code is bad and no amount of comments will ever change that."

I've worked on a lot of terrible legacy code in my career and I'm very thankful for the comments that others have left. This is becoming less necessary now that LLMs can explain a project, but comments have historically been a godsend in bad code.

baq 10 hours ago | parent | prev | next [-]

Clean code considered harmful.

No, really: comments should be telling you what the code shouldn’t or physically can’t. Code is for execution and the exact details of what and how; it has no business knowing why or why not and that’s where comments are required.

shawnz 9 hours ago | parent | prev | next [-]

If you are only encoding intent through "self-documenting code", and not with comments, then you are purposefully not using all the tools at your disposal to encode meaning as efficiently as possible.

Imagine a complicated section of application logic. You could break it up into 5 separate functions that document their intent semantically, thus blowing up the LOC by 5x, or you could write a short comment explaining the intent in natural language. What's more effective? I'd argue it's always going to be using all the tools at your disposal when and where it makes sense to use them, whether that is comments or self-documenting code.

tarzcvf 8 hours ago | parent [-]

Not to mention complex numerical optimization code that mixes closed-form approximations and something like Newton.

Without guides as to why a particular hairy expression is a good idea as a first estimate, the code is pretty much unreadable. (E.g. is it setting derivatives to zero, using a polynomial approximation, or something else?)

rustystump 4 hours ago | parent [-]

i think people took this too literally.

To put it another way, comments are for irreducible complexity ir external systems outside your control.

I work between systems and app dev. Systems have comments more often esp in shaders but my god informing me that a variable named isActive is for if something is…active, is useless noise. Same with the majority of comments that a type system already tells you. In my career, these have been ~90% of the comments I see. Since ai, all new code it is 100%.

Most of the replies examples are a sign of bad system/code but it is not always controllable. A legacy code comment of, the api requires strings for boolean values in the form “yes” and “no”. That is useful but it is also a code smell.

A concrete example, a vendor decided to define a proto with a flattened array of objects so there are some 1800 uniquely named fields on it. In many downstream consumers, this is a real performance issue besides being confusing. A comment may be good there. The thing is, this was still solvable if up at the root of where this vendor’s hardware logs data remapped it to something sane so every downstream system wouldnt need a comment explaining wtf is going on.

I see comments as when you want to explicitly answer why code smells right when a reader is smelling it.

david-gpu 10 hours ago | parent | prev [-]

The code tells you what the code does. It does not explain why it is doing that, and not something else. That is, among other things, what documentation does, and that includes comments.

astrange 11 hours ago | parent | prev | next [-]

I think the specific issue with Opus 5 is that its writing style is just trying to cheat at RL. It makes everything hypey yet self deprecating and constantly brings up "honest caveats" because the scoring rubrics look for those.

pmarreck 7 hours ago | parent [-]

The specific issue with Opus 5 is that it sucks all around.

It was causing so many issues with coding (even Opus 4.8 was better) that I did agent handoffs to Sol. One of the Sols stated the handoff was "incoherent", which I couldn't have said better myself.

swader999 6 hours ago | parent [-]

Yes, I pretty much took August off waiting for the next version.

physix 8 hours ago | parent | prev | next [-]

I've been cleaning up AI generated system/software design and architecture docs for an agentically engineered application, to translate that dense AI-speak into a clear human-readable form, cross checking it all against the actual codebase.

When I read the translated version, I felt a flush of relief, because I finally could confirm that it built the right thing and properly implemented the requirements.

I then asked in a fresh session which version was better for it as a reference for future work. It unequivocally voted for the human readable form, and gave it's reasoning with specific examples why.

So, I have a hunch that this "packing of lots of signals into fewer words" isn't really better. The incomprehensible prose just makes us think it knows what it's doing, like some mysterious magic that is only smoke and mirrors.

taneq 7 hours ago | parent [-]

Pay no attention to the bot behind the comments. ;)

satvikpendem 5 hours ago | parent | prev | next [-]

Chain of thought does not exist in the output of Claude, they disabled true thinking due to distillation risk. What you see when thinking summaries are enabled are just that, summaries of thinking into Claude-isms, therefore you cannot make any inferences on what the model is doing unless you literally work at Anthropic and can see the true thinking traces.

danieldrehmer 11 hours ago | parent | prev | next [-]

It's all about conducting users into using their plans/tokens in accordance to a certain cadence

sometimes by increasing human cognitive load during reviews, sometimes by expanding the number of gated decisions, sometimes by penalizing those using their accounts on other harnesses

Espressosaurus 12 hours ago | parent | prev | next [-]

Yeah, if anything the problem is that the output uses too many words for too little signal, and incorrectly uses confidence based on insufficient information to the degree it’s clearly bullshitting.

hedgehog 11 hours ago | parent | prev [-]

I don't know, I just pulled up the status for an active session and here's what it said:

  One thing I found before dispatching, and filed as Q0579. The halt told you C6
  was all that was left in the unit. That was true of the step's criteria and
  false of the unit's acceptance, which reads "exits 0 AND witnessed red" — two
  conjuncts. The witness half holds; the exits-0 half does not, because hello's
  G7 currently reads DIFFER 554/51340. I re-derived that from the gate map
  rather than trusting the prior step's report. So satisfying C6 does not by
  itself finish this unit, and I've filed that so attempt 1's success can't
  quietly be read as the unit's.
It's not exactly plain language.
jaapz 10 hours ago | parent | next [-]

My trick is to pass opus and fable's word salad into a haiku agent, then have it check if what haiku makes of it is still correct, then pass it to me. Whatever haiku outputs is often way more readable

hedgehog 9 hours ago | parent [-]

Oh, I can read the output, but that Haiku agent is a good trick. Where I want something less dense I just ask for "plain language" and characterize the reading audience and that term seems to trigger very readable output.

abraxas 5 hours ago | parent | prev [-]

This sounds like a Dianetics chapter by L Ron Hubbard.

hedgehog 4 hours ago | parent [-]

Sounds like I have some reading to do.

abraxas 3 hours ago | parent [-]

Meh, it is the sacred text of Scientology. Mostly pseudo scientific made up bullshit, wrapped in the buzzwords of the day and conveying little actual information. Just like opus 5.

hedgehog 3 hours ago | parent [-]

Maybe after enough auditing it'll make sense.