| ▲ | aesthesia 3 hours ago | |||||||
Where do you see the claim that "long-horizon goals in real world settings are now effectively settled"? The argument you put in their mouth would be a bad one, but I don't see anyone making it. | ||||||||
| ▲ | vector_spaces an hour ago | parent [-] | |||||||
https://openai.com/index/hugging-face-model-evaluation-secur... > UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings. I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess. We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting. | ||||||||
| ||||||||