| ▲ | codybontecou 3 hours ago | |||||||||||||||||||||||||
But doesn't Armin mention the limitation? > If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support? (This is Bonteq, I was just logged into the wrong account.) | ||||||||||||||||||||||||||
| ▲ | andreypopp 2 hours ago | parent [-] | |||||||||||||||||||||||||
it cannot use cat but you can make a command which communicates back to agent. Armin mentions that: > While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode. JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||