| ▲ | MikhailTal 2 hours ago | |
is this really that surprising? Exploitgym prompts are tuned for a model to do everything it can to achieve a cybersec/exploit task. And we know that models are good at finding vulverabiltiies. Its just random that the sandbox itself was buggy. But all that happened here is that we told a model "do everything you can to achieve your goal of hacking X" And it just hacked Y as a roundabout way of hacking X. Imo its PR for OpenAI to also start the mythos class mysterious unreleased model hype. From HF statement: "AI safety won't be solved by any single company working in secret". So now we have TWO companies working in secret | ||