Remix.run Logo
mmastrac 7 hours ago

I've started giving these instructions and I think I've been much more successful in generating clear output:

Comment blocks are <= 7 words, function names <= 4 words. User-facing message strings should be <= 10 words. Use an active voice, no stage performances, and pick the most common word when choosing among alternatives.

Limiting the number of words is the strongest factor in cleaning up the output, IMO.

For older code I've instructed it to delete all the comments, and then I re-comment it using a new session and these guidelines, asking it to rejustify the need for every comment to itself.

graemep 4 hours ago | parent | next [-]

Claude not only writes verbose comments, it also writes comments about how things used to work when refactoring. That might have a place in version control comments, but not in the code.

bhelx 12 minutes ago | parent | next [-]

This speaks to the general problem with using LLMs for writing. The audience they are writing for us you, but you're trying to write for a totally different audience. In code, this manifests as comments in the code that are hyperspecific to the conversation you are having, and not the long term benefit of having those comments in the code.

I see this in docs a lot. I've been reading a lot of docs these days where it feels like the LLM is trying to hype up the person writing the docs. It's like it has no conception that the writing is meant for a 3rd party audience.

jorl17 an hour ago | parent | prev | next [-]

This!

Claude writes comments about how things used to work, which can be useful sometimes, especially if it's a big change that requires one to genuinely consider legacy behavior, but most of the time it shouldn't be there.

Two other somewhat related things it does:

- It writes as if someone reading the code and comments is aware of everything it is aware of (the current conversation, the code it has just looked at). It's really hard to make it understand that things need to stand on their own. A trick is to get a subagent to look at it with a fresh context, but it doesn't tremendously help

- It does all of this with user-facing strings too. Claude loves to write up tooltips and other labels that leak everything to the end user. Every single concern we have, every edge case we've meticulously made our code handle, it passes on to the user, so they don't "need to worry". But no sane user would think of these things. For them, a feature is a feature. The "dynamic scheduling" button should state what dynamic scheduling does plainly, and every edge case is handled by us. The "add" button does not need a label letting the user know that they will later be able to click the "delete" button, because the user will just realize it due to our adherence to proper design. Claude fails to understand good UX for the user cannot be replaced with endless labels and explanations.

It's an uphill battle and all attempts at solving this (or the brain-dead way new Anthropic models write) usually fail to work with me.

spooneybarger an hour ago | parent [-]

I've gone back to using Opus 4.6. It's quite nice along all these fronts.

pluralmonad 4 hours ago | parent | prev | next [-]

And will reference transient working docs in code comments.

// No retry was added here per AC 37b in FEATURE.MD.

jorl17 an hour ago | parent [-]

// The lesson from the Parse-dont-fail-era campaign

// Judged on merit from computed properties during the cursor saga

// Chop 6ms due to lenience and lax-constraints vs 18ms baseline April perf measurements

datsci_est_2015 an hour ago | parent [-]

Thanks, this sequence of 33 words alone was enough to give me a searing migraine.

jorl17 18 minutes ago | parent [-]

[flagged]

ErroneousBosh 44 minutes ago | parent | prev | next [-]

Presumably you're not just blindly copying down what Claude copies out for you, but actually reading, interpreting, and understanding it for yourself?

nrmitchi 4 hours ago | parent | prev [-]

I struggled with this for a long time, but actually seem to have gotten to a place where this is largely resolved. Copy/paste from my current claude.md:

The CC-5 rule specifically seems to be (just from reading through, nothing repeatable-eval based) the part that actually catches and prevents me from having to clean it up afterwards.

```

### Code comments

The failure this prevents: writing a comment that narrates the change I am making right now. That context is real, but it expires the instant the change merges — the defect it describes no longer exists, so the comment becomes a story about a problem no future reader can observe. It is a changelog entry in the wrong file, and a third copy of text already required in the commit body (3.b) and the PR description.

- *CC-1 (MUST NOT)* Write a comment describing a change, a fix, a defect, its cause, or what the code used to do. No "was/now/previously/instead of", no "this fixes", no "needed because otherwise", no "note that we no longer".

- *CC-2 (MUST)* Apply the survival test to every comment before writing it: would this still be true and useful to someone reading this file a year from now, who never saw the diff? If it only makes sense beside the diff, it is changelog — delete it and put it in the commit body.

- *CC-3 (MUST)* Default to zero comments. Declarative config — Terraform, DNS records, k8s manifests, CI YAML, Helm values — is self-describing and takes none. A resource named `dmarc-example-com` does not need a comment saying it is the DMARC record.

- *CC-4 (MAY)* Comment only when a future editor would actively break something without it: a non-obvious external constraint, a required out-of-band manual step, an invariant the surrounding code cannot show. One line. If it needs a paragraph it belongs in `plans/`, not inline.

- *CC-5 (MUST)* Before every commit, re-read the comment lines I added: `git diff --cached | grep '^+' | grep -E '#|//|/*'`. Each hit must pass CC-2 on its own. Deleting is always an acceptable outcome. "I already wrote it", "it is only one line", and "this one is genuinely useful" are not exemptions — the last one is the exact thought that precedes every violation.

- *CC-6 (MUST)* Applies to comments I edit as well as ones I add. When a change invalidates an existing comment, the default action is DELETE, not rewrite it into a new narrative.

```

Yes, I am aware that claude mostly generated this, and it can probably be better and/or more succinct.

vrosas 7 hours ago | parent | prev | next [-]

The problem is, when the context window grows, Claude tends to forget these kinds of rules. It will then do whatever it wants. I had to outright ban comments in the global claude.md, the local claude.md AND write a hook to catch any that still slipped through.

nater5000 7 hours ago | parent | next [-]

I think people really need to focus more on working with limited contexts rather than trying to work around it. I really try to keep my sessions as short as possible and it helps a ton with keeping Claude (et al) focused.

Specifically, I like the "canary" trick that people have discussed where you add a small, innocuous rule to your CLAUDE.md like "When responding to me, start every sentence with my name." so that when Claude stops doing this, you know you've used way too much context and need to start a new session.

faizshah 6 hours ago | parent [-]

This or you just repeat the initial prompt every 200k tokens

ianjbutler 3 hours ago | parent [-]

Which gets you to the point where the whole thing is.. still unreliable. Generative text engines are going to generate. This calls for real enforcement in deterministic pre-edit hooks.

And here is where naive people will say something like "Why do I care if robots shit all over the codebase? Code is for machines, I don't expect to deal with it much now". But really externalized CoT like this confuses machines too, wastes tokens, and eventually wastes exponentially many tokens. Agents tend to think it's more real grounding than prompts are, even for comments-in-code. One bad comment poisons everything, then gets copied around as a ground-truth assumption everywhere. Hooks are more real to them than prompts or comments, and even then if you add enforced limits and tell them to externalize CoT ONLY in scratch task-tracking docs.. they will violate comment-enforcement hooks about 25% of the time. That tells you everything you need to know: even with constant reinforcement, they just really want to break this kind of rule.

hectdev 26 minutes ago | parent | prev | next [-]

Yea, I've started making it write linters to check the code that goes out. Anything that can be deterministically measured, gets added to it once we lock it down.

mandeepj 7 hours ago | parent | prev | next [-]

> The problem is, when the context window grows,

You know the problem; then why not address it? Does Compacting the context not help?

adastra22 6 hours ago | parent | next [-]

Compacting the conversation almost never helps. It is uniformly worse than starting over with fresh context, or rewinding to a last-known-good state. It only exists because it increases engagement.

enraged_camel 2 hours ago | parent [-]

This does not match my experience. I use long-running orchestrator sessions. Each orchestrator is in charge of planning, writing kick-off prompts for implementers, answering questions from those implementers, doing code reviews and providing feedback, and answering side questions from me when I have them.

Depending on the initiative I might compact a session a dozen times, sometimes more. It is lossy, and the session certainly tends to forget earlier bits as more compactions happen, but overall it's a much better experience than starting fresh and having to re-explain everything.

The only time I compact is if the session goes wildly off-course and the context gets polluted with off-topic conversations.

Also worth noting: with Claude Code you can provide custom instructions when compacting, and instruct the LLM that is in charge of compacting the session to prioritize the retention of specific bits. It can help a lot.

cautiouscat 6 hours ago | parent | prev | next [-]

Compaction is a main cause of this problem.

troupo 7 hours ago | parent | prev [-]

Compacting context compacts context. So Claude forgets a lot during compaction.

Maxatar 6 hours ago | parent [-]

Compacting mostly gets rid of reasoning tokens, and honestly it would be nice of reasoning tokens did not constantly follow every follow up query. Asking even a simple/trivial question can have Claude use thousands of tokens. Compacting is good for getting rid of those.

troupo 6 hours ago | parent [-]

I've had Claude immediately fall back to its usual verbose style immediately after compaction.

To be fair, I've had it do that immediately after re-reading the output style instructions, too.

My chat history is filled with "Yes, I broke the language rule. Let me rephrase that and update my memory. — You already have that in memory — Yes, true, I ignored that" (because "Memory" is a yet another .md file)

strbean 3 hours ago | parent [-]

Claude Code supposedly supports a "post-compaction" hook, so you could have it automatically run the prompt "We just compacted the context, quickly refresh yourself on the rules in CLAUDE.md etc..".

Depending on what you've got in those files, maybe that will just use up all the context again though.

troupo 2 hours ago | parent [-]

> supposedly supports a "post-compaction" hook, so you could have it automatically run the prompt

Keyword "supposedly" :)

I've had it in my settings forever, and still...

Asking it to analyse and fix the issue it produced a plausible "my training supercedes/overrides settings especially if triggered by certain words in the phrase" (paraphrasing the long text)

strbean 2 hours ago | parent [-]

> Keyword "supposedly" :)

> I've had it in my settings forever, and still...

Checks out! I've never used it my self, so it I figured it likely didn't work at all.

aleksiy123 3 hours ago | parent | prev [-]

Hooks is the way.

Intermittent nudges

kybernetikos an hour ago | parent | prev | next [-]

I've been trying to understand this. It's been the universal experience of our team that claude code overcomments, and doesn't have very good adherence to prompts telling it to comment less. The team moved recently from cursor where we mainly used claude models and it was much faster and commented more appropriately. I figured that since the models were the same, maybe it was just the system prompt, so I went looking in tweakcc etc. All of the mentions of code comments in the system prompts seem to be also telling it to be terse and only use them when appropriate, which also seem to be ignored. I'm not sure where this overcommenting is coming from unless it's getting confused by other parts of the system prompt talking about other kinds of comments.

nottorp an hour ago | parent | prev | next [-]

I just gave up and edit the comments manually. However, I've had a surprise today.

I had it fix something then went and reduced one of the 3 line comments to 4 words. Then for some reason I told the bot to reload the source, it offered to make the other comments terse and did a passable job of it. Shocking!

Now how to get it to do that all the time...

transdev12 an hour ago | parent | prev | next [-]

The corollary here is to have Claude write tests to enforce this. The only thing it is consistently responsive to is test failures.

kanzure 6 hours ago | parent | prev | next [-]

Yep. Same here. I frequently tell agents things like "answer using only a single sentence" and "write no more than 10 words". They are excellent at writing code, so have them write code (and not English prose). Besides, most of the time we want them to make reusable software that doesn't require users (or future agents) to read too much text. Software should generally just work and do the obvious thing, without needing verbose explanation.

pbreit 3 hours ago | parent | prev [-]

Shouldn't this all be easily doable via (auto-)prompting? Surely I don't need to "install" anything?