| ▲ | andrewingram an hour ago | |
The value is that rather than an agent chaining together tool calls itself (which means each step sends the result back to the agent for it to analyse and work out what to do next), it writes a script for the harness to execute that chains together all the calls. The major benefits are: * speed - much fewer hops back to the LLM * fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does. * repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time. If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs. | ||
| ▲ | dools 26 minutes ago | parent [-] | |
Agents just do this anyway, how is it a “mode”? I always see the agent writing scripts in a tmp dir to execute or even just inlining bash and python scripts. | ||