| ▲ | drob518 2 hours ago | |||||||||||||||||||||||||
The article makes the point that on some benchmarks the AIs worked around the UI and then says that sometimes APIs won’t be available. The unwritten implication is that then we’ll be in trouble. But will we? If an API doesn’t exist, is the AI still able to perform the task? That’s not covered in any detail, nor is there any direct comparison between the frontier AI’s performance on these tasks and the model being sold by the company. At some level, I don’t fault the AI for taking the API path if it’s available. In fact, I’m impressed that it found the API and used it correctly. That doesn’t seem like an argument that the sky is falling. | ||||||||||||||||||||||||||
| ▲ | euphetar 2 hours ago | parent [-] | |||||||||||||||||||||||||
I get the point, but you can't always use the API. 1. Presumably you want to trust your agent to do no shady sheningans behind your back when you give it a simple task 2. Sometimes it just doesn't work. https://osworld-v2-monitor.xlang.ai/task/tasks/068 Does this look like an efficient way to solve the task to you? 500 steps of fiddling with a JS injection, followed by hacking the task. The funniest part is that it needs to achieve a score of 100, but puts 150 "just in case". I don't think I want it to take the same approach when e.g. fixing a customer's balance. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||