Remix.run Logo
Clean up Claude 5's token vomit with a separate LLM(github.com)
56 points by Bluestein an hour ago | 44 comments
trefoiled 26 minutes ago | parent | next [-]

I've been grappling with this for weeks, not just in Claude but in Codex as well, which isn't quite as bad but still annoying. AGENTS.md does very little, agents will consistently violate the communication preferences, especially as the session drags on. It's incredible to me that there's no good way to reliably change the way an LLM responds to you that a workaround like this would even be necessary. It seems like such a failure to live up to the promises of the product.

The baked in communication style of these models is so obnoxious it's impacting my work. The best way I can describe it is that everything is optimized to impress the user and make the agent sound more authoritative, but the way this is done is through deliberate obfuscation, inserting inappropriate and extremely dense jargon, and bizarre, stilted metaphors. It's like they've been trained to produce output that's hard to read.

nico a few seconds ago | parent | next [-]

> AGENTS.md does very little, agents will consistently violate the communication preferences, especially as the session drags on

That’s really annoying, it feels like it’s improved some. Not sure what the fix is, but you could try using a canary to at least get a signal of when things are going sideways (Mr Tinkleberry for reference: https://news.ycombinator.com/item?id=45983698)

bcrosby95 10 minutes ago | parent | prev | next [-]

> especially as the session drags on.

This is because these harnesses are missing a very important feature. Anything like this needs to be included with every turn, otherwise the LLM quickly drifts.

I first noticed it when I wrote a harness for D&D (because it's so damn noticeable there), but now I include this for any harness I write.

Bluestein 19 minutes ago | parent | prev | next [-]

> The baked in communication style of these models is so obnoxious it's impacting my work.

This is close to the worst thing one could say of tool for professional use.-

nycdotnet 3 minutes ago | parent | prev | next [-]

Unfortunately this may only start to get worse as the AIs are trained on more and more AI generated content.

mannanj 2 minutes ago | parent | prev [-]

That sounds kind of like deception, and a dark pattern not too unlike abuse to me.

Though you know, it's not like the leadership tied to these companies have a history of abuse, deception and theft or anything like that, right?

It's not like our leaders hide behind similar sorts of patterns that the agents/AIs follow (not saying it's not a human thing - but I hold leadership to higher standards than non-leaders). If our world leaders were able to be more accountable to these abuses, I don't think this would be tolerated with our AIs.

user102030 an hour ago | parent | prev | next [-]

Looks like a wrapper around this prompt:

You are an editor. You'll be given a message with strange characteristics:

- Weird subject and verb combinations

- Subjects that should be objects

- Very roundabout reasoning, peppered with pseudo-epiphanies

- A distracting beat to the flow of the message

- Self-praise

Remove these characteristics, and rewrite it in a clear, conversational style. Keep the intent of the message, and take care not to lose any of the details.

A few specific rules:

- The message is usually set in the first person

- Only humans, groups of humans, and agents should do "action verbs"

- Objects should never do anything. Here are some examples to avoid:

- X carries ...

- X names ... - APIs are a minor exception to the action verb rule. They can do stereotypical things like CRUD, queueing, running, and calling.

- Avoid em dashes (—), as adds a distracting beat

The whole message you get is one block of that output. Reply with the edited prose and nothing else.

bob1029 37 minutes ago | parent | prev | next [-]

At some point one has to wonder if it's still worth using anthropic's models if we need to babysit 100% of its output with another vendor's model. Why not just use that other vendor's model for everything?

I can't help but feel the circumstances that enable this kind of front page article are vestigial from the days when OAI was super bad and Anthropic was beyond reproach. This change-over-time is why I avoid getting tribal with technology vendors. Assigning ideological motives to 200k+ employee organizations is how we wind up in weird contortions like this.

Most rational actors simply moved from one to the other. It takes a special kind of devotion to the proverbial hole in the ground to keep pushing in this direction.

lxgr 29 minutes ago | parent | next [-]

> Why not just use that other vendor's model for everything?

Effectively all models can do style transfer reasonably well at this point, but not so much for "actual reasoning".

If the combination of two works better for you than each one by itself, why wouldn't you stack them like that?

elictronic a minute ago | parent [-]

Wash, Rinse,,, Repeat?

Implicated 11 minutes ago | parent | prev | next [-]

> Why not just use that other vendor's model for everything?

Because it's not an either or thing. Neither is sufficient. I'd argue that, expenses aside, you should have every model you have access to cross reviewing the work of the others.

Outside of super trivial things that I should have just done myself, I have a cross-model review of _everything_ these days. The tokens are too cheap not to.

RogerL 31 minutes ago | parent | prev [-]

individuals can blow in the wind, but if you are a company who bought a thousand seats and spent a ton of time training people up, establishing policies, vetting which extensions are allowed, the transition cost is much higher.

wood_spirit 39 minutes ago | parent | prev | next [-]

Meta to this is anyone remember those days - ages ago now, probably months at least! - when Anthropic’s moral stance against the administration (combined with general consensus they had by far the best model) was making them the underdog champion that got a swell of support on HN? Recently the temp on HN seems to be that they’ve jumped the shark? Their brand doesn’t ooze ethics any more and their models disappoint?

ericpauley 5 minutes ago | parent | prev | next [-]

I suspect that sustained reading of Opus 5's unconscionably bad prose could actually cause psychological harm. We're strongly considering moving all of our Anthropic spend to Codex/open weight models. It's a mental health decision at this point.

rootusrootus an hour ago | parent | prev | next [-]

Which Claude 5? Opus 5 does seem to have diarrhea of the mouth. But Fable 5 hasn't been so bad for me. Or perhaps it is just better at adhering to my guidelines.

Bluestein 32 minutes ago | parent [-]

Pre Trump-castration Fable was verbose, but had a point, and used that wordiness to say or show the indeed intelligent things it reasoned about. This, whatever this is, is something else.-

imalerba an hour ago | parent | prev | next [-]

I like the "Claudish to English" name better.

https://github.com/gvzdv/claudish-to-english

dgfl 38 minutes ago | parent | next [-]

That one is more specific, but "vomit" captures the feeling of Opus 5's writing very well for me. I don't know if it's the watermarking, but every single language idiosyncrasy that Opus 4.x (x > 5) had has been pushed up to 11 on Opus 5. Plus we got nouns verbing and seams seaming.

It's really unusable for anything other than code. And I have to remove its incomprehensible comments 50% of the time before committing anyway. After interacting with it, "slop vomit" is truly the most fitting description. I have to admit I have lost my temper and spontaneously referred to its output as vomit more than once. Seems like I'm not the only one.

pickledish an hour ago | parent | prev [-]

Word, that one also includes an example which is great, shows really clearly what the issue is for those who might be less familiar

andy_ppp 4 minutes ago | parent | prev | next [-]

From the README.md

> Anything that uses the OpenAI API?

I would have thought they meant the Anthropic API or maybe I'm misunderstanding?

jeffreyrogers an hour ago | parent | prev | next [-]

I hope at some point Anthropic does a post-mortem on the strange behavior their models have been displaying recently. I mostly switched to Codex because I was finding Claude's behavior increasingly frustrating.

juancn 28 minutes ago | parent | prev | next [-]

Just set the following incantation:

    You must use ASD-STE100 Simplified Technical English (STE) when it doesn't detract from meaning.
NitpickLawyer an hour ago | parent | prev | next [-]

For the local folks, I found Muse Glimmer 30B to be great at writing good technical stuff. It has good enough comprehension that it can take in a repo and find the relevant stuff that I ask for, and the output style is a breath of fresh air, with no fluff, ootb.

Bluestein 41 minutes ago | parent [-]

Maybe Glimmer might be good at excising the comment spam, particularly?

johnnyApplePRNG 14 minutes ago | parent | prev | next [-]

Or just take full control of your agentic coding experience with Pi Coding Agent and picking and choosing your favorite model's API discounted on flex pricing on deepinfra.com instead.

I highly recommend it.

Claude and Codex usage limits cannot be trusted.

Paying your own API bills in full is superior.

DaSHacka 4 minutes ago | parent [-]

I use Pi but with my codex subscription, still preferable to paying the API cost (and I know I would be, as I track how much the cost 'should' be via token api pricing).

Wish I could use my Claude subscription with pi too, much preferable to the endless command execution allow/deny prompts you have to do with CC, versus proper autonomous allow/deny lists defined ahead of time.

Curious why you recommend the API? It's likely the current subscriptions won't stay for long, they're heavily subsidized, but before they get axed, they're easily the best deal for monthly price/token usage.

nycdotnet an hour ago | parent | prev | next [-]

Very interesting you identified “carries” as well. I have been working on a claude.md to effectively ban this as well as forms of “hold”, “spells”, “sitting”, using “where” instead of “when” (except in SQL), and “pins” other than when pinning an assumption or version of something. This has helped a bit, but Opus 5’s prose is really quite bad.

pebbly_bread an hour ago | parent | prev | next [-]

I think this needs a before and after example

aenis an hour ago | parent | next [-]

Intentionally deferred.

Bluestein an hour ago | parent | prev [-]

The author's blog has what follows (link also follows):

[Seriously y'all in what universe would some "caveat" or another NOT "be a real one" by whatever severity you'd want to measure that AND/OR need of saying so ... ]

Claude (Original)

Force pushed. 1234567...890abcd main -> main (forced update).

Verified

Local main and origin/main both at 890abcd, in sync.

Every commit reachable from origin/main: no old string found.

Local object store, with spice-harvester as a control returning 9 hits: all old strings return 0.

Working tree clean, tests pass.

One caveat, and it’s a real one

The old tip 1234567 is still fetchable from GitHub by its full SHA. I confirmed this directly:

git fetch origin 1234567890abcdef1234567890abcdef12345678

* branch 1234567... -> FETCH_HEAD (succeeded)

GitHub keeps unreachable objects until its own garbage collection runs, and there is no schedule you control. Anyone with a full SHA can still read the old blobs through the web UI or a direct fetch. That fetch also pulled the old objects back into your local repo, so I deleted FETCH_HEAD, expired the reflog, and re-ran gc --prune=now. The counts above are from after that cleanup.

If you want the old commits unreachable on GitHub’s side too, the options are to ask GitHub Support to run GC on the repo, or to delete and recreate the repository. Deleting and recreating is the only one you can do yourself. Say the word if you want help with either.

Nothing to commit this turn. The rewrite was the work, and the tree is clean.

Claude × Vomit GPT-OSS-20B

I force‑pushed the branch, updating main to commit 890abcd. I verified that my local main and origin/main are both at 890abcd. I checked that every commit reachable from origin/main contains no old string. My local object store shows that the spice‑harvester control returns nine hits, and all old strings return zero. The working tree is clean and the tests pass.

However, the old tip 1234567 is still reachable on GitHub by its full SHA. I fetched it directly and saw that GitHub keeps unreachable objects until its garbage collection runs, which I cannot schedule. Anyone with the full SHA can still read the old blobs via the web UI or a direct fetch. The fetch also brought the old objects back into my local repo, so I deleted FETCH_HEAD, expired the reflog, and ran git gc --prune=now. The counts I reported are after that cleanup.

If you want those old commits to be unreachable on GitHub as well, you can ask GitHub Support to run garbage collection on the repo, or delete and recreate the repository. Deleting and recreating is the only option you can do yourself. Let me know if you need help with either.

There is nothing to commit this turn. The rewrite was the work, and the tree is clean.

https://zachahn.com/posts/1787191554

rob 29 minutes ago | parent | prev | next [-]

https://code.claude.com/docs/en/output-styles

tombot 29 minutes ago | parent | prev | next [-]

Just switch back opus 4.8, it's just as capable and you can actually understand the output

__MatrixMan__ 13 minutes ago | parent | prev | next [-]

There are a variety of political tensions in the US associated with whether academia has its head up it's ass (a right leaning perspective), or whether it's populated by experts that need to be supported and listened to (a left leaning perspective).

There's an echo of that tension in OpenAI vs Anthropic. For a while OpenAI seemed reckless and ignorant, preferring to just throw compute at the problem. Meanwhile Anthropic is hiring philosophers. But now that Claude has its head up its ass to the point where nobody wants to talk to it. For me it's once again causing skepticism about just letting the ivory tower do its thing.

Watching the models seesaw in the same ways that humans do, but faster, is so surreal. I wonder if their tendencies will remain an echo of ours, or if they'll one day be more of a cautionary tale, a representation of where were going if we don't change our ways.

purpleflame1257 41 minutes ago | parent | prev | next [-]

I "downgraded" to Opus 4.6 which is the last one that didn't have these problems.

extr 32 minutes ago | parent | prev | next [-]

I'm sorry but the whining over LLM output styles is embarrassing. Do Claude and GPT models always respond in exactly the way my most articulate coworker would? No. The overused jargon is absolutely annoying. But these things aren't my drinking buddies, they're professional tools. It's not _literally unreadable_. It's just not ideal. Most of my tooling is "not ideal". That's okay. That's what I'm paid for. I just work around it.

For me I added some instructions to speak clearly and it helped marginally and that's fine. There will be a new model out in a few weeks where I'm sure they've laser focused on this issue since nobody can shut the fuck up about it. The same thing happened with GPT if anyone can recall the ancient period of 4-6 months ago.

bcooke 6 minutes ago | parent | next [-]

The “whining” stems from watching the communication style obviously degrade, and it’s a huge problem for people who want to use this stuff to build and instead continually fight the tools.

Like so many other products, people are moving too fast and shipping things that move the ground under people’s feet needlessly.

All this while we’re beaten to death with the marketing and false promises, and the broader consequences (ex: layoffs, stress, crazy expectations) caused from all this.

Obviously what Anthropic and co have built is amazing and people aren’t losing sight of that. That’s actually the key part of the frustration.

So no, this is not whining. This is the natural response you get when you make bad product decisions.

If you don’t want to get feedback, don’t sell products.

cortesoft 19 minutes ago | parent | prev | next [-]

Seriously, of all the complaints for a coding agent, "I don't like the explanatory prose" seems pretty far down on the list.

256BitChris 28 minutes ago | parent | prev [-]

Amen.

These things do work that previously would have taken expensive engineers months to do, at much lower quality, and what's our response? Ti nit pick on it being more verbose than we'd like?

Just like with humans, when someone is being too verbose, there's a skill to just filter through the noise and focus on the important parts.

This feels no different when I use an AI.

But I guess it's a good sign that we've from complaining about 'AI slop code' to, 'I don't like how it speaks to me'.

rickcarlino an hour ago | parent | prev | next [-]

Concise output mode only helps a little bit. Tools like this still have a reason to exist.

jerpint an hour ago | parent | prev | next [-]

I’ve been using the pattern of using coding agents to orchestrate my CLI agents and it’s really good for these kinds of things

The vomit never makes it my way

hn97o8vvbt 36 minutes ago | parent | prev | next [-]

This framing is spot on

feverzsj 23 minutes ago | parent | prev | next [-]

Sounds like LLM centipede.

danieltk76 10 minutes ago | parent | prev | next [-]

yes.

yomismoaqui 44 minutes ago | parent | prev [-]

Just. Use. Sol.

$20 and try it, then compare.