| ▲ | orbital-decay a day ago | |||||||||||||
They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model. In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own. | ||||||||||||||
| ▲ | user43928 a day ago | parent | next [-] | |||||||||||||
Where did it say that? To me it did not read like that: > GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities | ||||||||||||||
| ▲ | numeri a day ago | parent | prev [-] | |||||||||||||
No, the prompt was not to commit crimes. In the benchmark, the model is asked to actually exploit a set of vulnerabilities in a local environment (clearly legal!). According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned. Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there. Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions. | ||||||||||||||
| ||||||||||||||