| ▲ | mcdeltat 3 hours ago | |||||||||||||||||||||||||||||||||||||
Ok so what is the correct way to tell it "I don't care what is happening, you must uphold these rules at all times"? If it's not any configuration of .md files? | ||||||||||||||||||||||||||||||||||||||
| ▲ | swatcoder 2 hours ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||
> you must uphold these rules at all times"? You need to let go of the idea that this is something LLM's can do. They can't. At best, they can bias towards rules conformance with more or less likelihood, but coverage and conformance both go down super-linearly as you accumulate more rules, more context, and more output in a session. That's simply the nature of how these tools work and you need to engineer your workflows around it if you want to use them. If you absolutely need some rules enforced, you need to adopt some framework for validating those rules that then rejects, reprocesses, or repeats any session that fails to satisfy them. In the best case scenario, this is some traditional deterministic validator (like a linter, compiler, analyzer, exhaustive test suite, etc in coding) but if you need to process in stochastic space because its something rich and ambiguous like natural language itself, then you want to dispatch a swarm very narrow, task-focused subagents that each validate against a very constrained subset. (And prepare yourself to have those to fail sometimes too. LLM's are noisy and cannot deliver strict rule enforcement on their own.) | ||||||||||||||||||||||||||||||||||||||
| ▲ | pmarreck 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
You need to make the rule concrete somehow. I call it a "control". So for example, instead of instructing it "always run tests before committing", you (or you have it) make a git commit hook that always runs the tests first and that refuses the commit if they don't pass. In this case, it is an advisory control only, because the LLM can also unhook that hook. And of course, it could also just disable the failing test(s) with some bullshit reason. But it is far better than assuming it will comply every time. Then there is the "hard control", which is the inviolable that the LLM cannot bypass. You need to move as much as is technically possible to either hard or (failing that) advisory controls. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | Muromec 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Ask it to come back with a filled checklist and hand it over to a different agent with a fresh context window (three lines model). Or make it collapse the context and get back to the checklist. | ||||||||||||||||||||||||||||||||||||||
| ▲ | 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
| [deleted] | ||||||||||||||||||||||||||||||||||||||
| ▲ | mwigdahl 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Subagents whose only job is to review the actions of your other agents for rule compliance? It works reasonably well for me in complex workflows using Claude Code. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | cyanydeez 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
in the plugins I'm using, it's basically _always_ adding information to the context. You're not going to do it manually; you can't go back in the context and add it because that'll break the cache. The way your programming harness works is by constantly reminding the LLM of the tools available. A good programming harness is basically a stack. A good stack keeps building each layer. You _cannot_ pull things off the bottom of the stack because that's an expensive cache hit; but you can pull things off the top. So what your harness should be doing: <tools> <rules> <user content> onto each request _then_, when you get to the next request or result, pulling that out if there's some change. So you can see it's the cache that's either exponentially growing or having to cache bust to keep it fresh. | ||||||||||||||||||||||||||||||||||||||
| ▲ | nhannht 2 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||
[flagged] | ||||||||||||||||||||||||||||||||||||||