| ▲ | KronisLV 2 hours ago | |||||||
That's interesting! You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy? Also what about the cache read costs for longer context sizes? Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/ | ||||||||
| ▲ | surgical_fire an hour ago | parent | next [-] | |||||||
> You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy? For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system. | ||||||||
| ||||||||
| ▲ | cyanydeez 43 minutes ago | parent | prev [-] | |||||||
running local models, now with Qwen3.8-Flash-Next, they have 256k, but when they get up there their speed is just too slow. So i've taken https://github.com/Tarquinen/opencode-dynamic-context-prunin... and started improving it. It already worked well to get a lot of mileage out of just taking tool calls, code modifications, etc, and dumping them in favor of a summary. But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage. I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again. It's amazing how resilient it is if you give it a good plan. The work flow has basically been: 1. Write up an implementation document for some new set of features. 2. Rewrite the implementation as a TDD document 3. Set it to work. The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done. | ||||||||
| ||||||||