| ▲ | shken 5 hours ago | ||||||||||||||||
Every step here has an oracle: wall-clock, the profile, pass or fail from the verifier. I had an agent-built app audited task by task, 10 came back done and 7 worked, and the three misses were the ones needing a credential or a setting on someone else's dashboard. Nothing in the loop could tell the agent it had failed, so it said done and moved on. | |||||||||||||||||
| ▲ | phoghed 4 hours ago | parent [-] | ||||||||||||||||
I’ve tried it on simple UI tasks. Give it screenshot to work towards (or figma MCP), let it get screenshots from chrome to check its work. The existing models are surprisingly bad at it. | |||||||||||||||||
| |||||||||||||||||