| ▲ | hysan 6 hours ago | |||||||||||||||||||||||||||||||
Relevant anecdata because I've burned many a Claude sessions on this. If you're using Claude Code, then it's in the harness. At the close of many sessions, I would start a meta conversation over why the LLM would consistently break certain rules. What it found when debugging itself is that some of the "contradicting" rules that I had were in fact, not from my rules. Instead, the instructions from its own harness had phrases telling it to do things like that. When something contradicts, its own instructions would outweigh any custom ones you write. Every rule variant I had tested (including the one that says it overrides the harness instructions - and yes, I've actually tested all the ideas in your comment too) has ultimately been unsuccessful due to this according to the LLM. | ||||||||||||||||||||||||||||||||
| ▲ | vzmax 5 hours ago | parent [-] | |||||||||||||||||||||||||||||||
You can't trust it's account on why it did something, it does not "remember". It will just make up something plausible sounding. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||