| ▲ | tekacs a day ago | |||||||
I've implemented a similar approach – although I'm surprised not to see mention of cache prefix busting in there! | ||||||||
| ▲ | planckscnst a day ago | parent | next [-] | |||||||
I did calculations on the prefix cache effect on costs of sessions where I used it and found that the removal of tokens from context had a much bigger effect on reducing costs than cache busting had on increasing them. I should re-do that and publish it. | ||||||||
| ▲ | jeremyjh a day ago | parent | prev [-] | |||||||
Yes the idea is cool but this could really hammer usage, especially just leaving it up to the agent to decide when to do it. I'm not surprised though, considering the github account is named "Vibecodelicious" and became active in December. Poking around in the repo the whole implementation is an unsupervised LLM fever-dream. | ||||||||
| ||||||||