| ▲ | bcjdjsndon 14 hours ago | |||||||||||||||||||||||||||||||||||||||||||||||||
> there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. This is true of closed weights, and in fact the problem is worse because they cannot even be scrutinized. We should ban closed weight AI for the very reasons you have just given | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | eddyg 14 hours ago | parent [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
Constitutional classifiers go a long way to reducing unsafe usage in closed-weight models. And like we saw with Fable, closed models can be revoked and classifiers updated when “jailbreaks” are found. Having the weights gives you the exact affordance an unlearning attack requires, without rate limits. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||