| ▲ | ylxdzsw 4 hours ago |
| I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage). [1] https://github.com/ylxdzsw/mu |
|
| ▲ | Bonteq 3 hours ago | parent | next [-] |
| > However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be. The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Does your prototype overcome the limitations mentioned in the article? |
| |
| ▲ | alfiedotwtf 23 minutes ago | parent | next [-] | | Surely LLMs understand how to decide base64, but if that’s not possible you could convert the image to ascii and then send that in | |
| ▲ | andreypopp 3 hours ago | parent | prev [-] | | There's no limitation, just have bash commands `read` or `view_image` which communicate back to agent. | | |
| ▲ | codybontecou 3 hours ago | parent [-] | | But doesn't Armin mention the limitation? > If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support? (This is Bonteq, I was just logged into the wrong account.) | | |
| ▲ | andreypopp 2 hours ago | parent [-] | | it cannot use cat but you can make a command which communicates back to agent. Armin mentions that: > While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode. JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode. | | |
| ▲ | the_mitsuhiko 2 hours ago | parent [-] | | > But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode. It becomes much crummier when hands and brain are on different machines. | | |
| ▲ | andreypopp 2 hours ago | parent [-] | | Agree, but this orthogonal to JS or bash question. (I mean can always run some bash locally as well, even wasm compiled one). | | |
| ▲ | the_mitsuhiko 31 minutes ago | parent [-] | | > Agree, but this orthogonal to JS or bash question. It is not from my perspective because Codemode runs in the brain, and bash necessarily runs where the hands are. So if the hands need to reach into the brain, I need to set up a communication layer from the hands to the brain. |
|
|
|
|
|
|
|
| ▲ | the_mitsuhiko 4 hours ago | parent | prev | next [-] |
| > I'm not sure why almost all codemode implementations choose Javascript Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global. But a big reason is that code mode runs on the harness side so bash is a tricky target in particular. |
| |
| ▲ | ylxdzsw 2 hours ago | parent [-] | | sandboxing is indeed an advantage (can be an important one!), but 1. bash can also express concurrency easily, with sync and async (using standard & syntax) support for each command, and standard cancellation (kill, though crude).
2. "way fewer instructions": It requires no instruction for agents to use bash either, except for merely listing the custom commands (view_image, apply_patch, etc.). Also, bash has standard progressive disclosure mechanism (--help) that models will automatically use with no instruction. | | |
| ▲ | the_mitsuhiko 21 minutes ago | parent [-] | | We might be talking past each other here. The point of codemode is to orchestrate the LLM side tool calls, not to orchestrate scripts that it might execute within Bash. In a world where brain and hand are on different machines, getting the bash hands to reach back into the harness brain is something that requires a) putting tools in its hands that it does not know about b) are tricky to set up, usually involving some sort of socket based back channel. I tried this quite a bit, by having pi be always there on the hands side, but it causes a lot of complexity and the LLMs really do not understand it well at all. |
|
|
|
| ▲ | anilgulecha 3 hours ago | parent | prev | next [-] |
| Almost every model is fully trained on js. That does not need teaching either. Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling. |
|
| ▲ | searealist 4 hours ago | parent | prev [-] |
| Can do you make a tool or mcp call from bash? |
| |
| ▲ | clintonb 4 hours ago | parent [-] | | Yes. Invoking an MCP tool is just an HTTP call. You can do it with curl. | | |
| ▲ | searealist 4 hours ago | parent [-] | | That's one kind of MCP. Another is a local stdio server. Also there are things like subagents, etc (which may be considered tools). |
|
|