| ▲ | brandall10 3 hours ago | |||||||
Seems like it might be more advantageous to just adjust reasoning effort to retain cache. Maybe in some exceptional cases where there will be a ton more inference to solve the problem, but going significantly dumber in that case seems counterintuitive. I can really only see the utility of things like spawning subagents to a lower tier model from another provider, and that's something harnesses can already handle (ie. give model specs for certain delegation roles). | ||||||||
| ▲ | rohaga 2 hours ago | parent | next [-] | |||||||
I agree that adjusting the reasoning effort to retain cache is a huge thing! But even doing that automatically is currently a challenge for people to figure out and do well, and costs mental energy when perhaps it doesn't need to. For example, there is GPT-5.6-Sol low, med, high, xhigh, max, and lots of "rules of thumb" that people develop on which one to use when. | ||||||||
| ▲ | verdverm 3 hours ago | parent | prev [-] | |||||||
caching is per model, it does not transfer between them | ||||||||
| ||||||||