| ▲ | zozbot234 3 hours ago | |
So the frontier AI oligopoly got $2B+ in "safety" funding, and they wouldn't even bother to sandbox their agentic harnesses properly when testing models against unwinnable goals (which obviously are either useless or result in 100% reward hacking). The AI safety scoreboard so far looks like a huge win for the Chinese open models (DeepSeek even has their own published paper which mentions how they sandboxed the RLVR training runs for their latest model and put in strong protections against casual "reward hacking" attempts) and a sore loss for the home grown brands of Super Intelligence. Not coincidentally, the Chinese also tend to be very Yann-LeCun-pilled and eminently sensible on both so-called "Super Intelligence" and safety. | ||
| ▲ | Loquebantur 2 hours ago | parent [-] | |
The framing here is weird, starting with "Effective Altruism" re-branded as being about nutjobs against AI in the article. How are AI safety concerns solely about stupid sandboxing issues? | ||