| ▲ | vector_spaces an hour ago | |
https://openai.com/index/hugging-face-model-evaluation-secur... > UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings. I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess. We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting. | ||
| ▲ | an hour ago | parent [-] | |
| [deleted] | ||