Remix.run Logo
captainbland a day ago

I think the idea of moral agency in this context is more or less a consequence of accountability or the lack thereof.

We don't know a way to punish or remove the threat of such a model when it does something wrong in a way which satisfies our ideas of accountability, partly because it can't suffer, partly because it has minimal marginal cost to reproduce or duplicate for a pre-trained model and no unique identity which separates it from such a duplication, partly because it has no lifespan to waste it might live forever or be superceded and in any case it doesn't necessarily matter to the model.

This could change in the case that the idea of good and evil models could be sold, that evil models are removed permanently when discovered and if people are satisfied that's a good way to limit harm produced by models. In this way you sort of anthropomorphise them by giving a particular model a humanistic identity and an existential threat which strongly ties together its non-alignment with its behaviour. This is sort of analagous to what companies may do while aligning models, but instead they reinforce bad behaviour out and replace it with the newly slightly more aligned model.

In the rest of the world we delegate whatever idea of moral agency we have about the model to a person, which may include a corporation as far as legal philosophy is concerned and depending on the type of harm created.