| ▲ | ricardobeat an hour ago | |
To do this you need video/motion understanding, the intent cannot be judged from still images or state descriptions. We’ve had the tool to do this since mid 2025, V-JEPA2 [1], Yann Lecun’s last work at Meta. It runs at several FPS on a macbook and can even be trained locally. Chaining it with Jev for decision-making would probably work great! | ||