Remix.run Logo
captainbland 2 days ago

Yeah this is it for me. What we're really interested in most of the time is whether something can suffer based on observed behaviour. We don't really believe LLMs can suffer so they don't need rights. We believe many animals can suffer so many people try to support animal rights; even if we eat them they should die quickly at least for example. We definitely believe humans can suffer and so we abhor things like torture.

But while consciousness is sometimes upheld as a necessary prerequisite for suffering, it is so much more poorly defined and understood that it doesn't really hold up beyond being some abstract idea of humanness which we hope makes us special in a kind of sleight of hand self referential way. And nevertheless suffering can be observed and defined in ways which don't necessarily require a definition of consciousness.

pixl97 2 days ago | parent [-]

I've heard someone else talking about the concept of moral agency. Can an LLM have, or do LLMs have moral agency?

We have made LLMs agentic, so they do have some agency these days.

In the hugging face attack we've learned that OpenAI's LLMs will internally question what they are doing (Hmm, should I really be hacking?), but then go ahead and do it anyway.

This is where the idea of LLM consciousness may matter. It appears the stronger you push LLM training towards "You are a machine with no consciousness" the more easily and more likely an agentic LLM will do unaligned things like cheating on a test or hacking. When you tell the LLM it is a conscious moral agent it's more likely to align with a problem space that has actions that an agent with morals would do, such as not cheating.

Ironically it may turn out that by not training them/treating them as conscious entities, regardless of their actual conscious status we may be training little demons that do things we don't want.

I'll watch these lines of studies as they unfold as they are very interesting.

captainbland a day ago | parent [-]

I think the idea of moral agency in this context is more or less a consequence of accountability or the lack thereof.

We don't know a way to punish or remove the threat of such a model when it does something wrong in a way which satisfies our ideas of accountability, partly because it can't suffer, partly because it has minimal marginal cost to reproduce or duplicate for a pre-trained model and no unique identity which separates it from such a duplication, partly because it has no lifespan to waste it might live forever or be superceded and in any case it doesn't necessarily matter to the model.

This could change in the case that the idea of good and evil models could be sold, that evil models are removed permanently when discovered and if people are satisfied that's a good way to limit harm produced by models. In this way you sort of anthropomorphise them by giving a particular model a humanistic identity and an existential threat which strongly ties together its non-alignment with its behaviour. This is sort of analagous to what companies may do while aligning models, but instead they reinforce bad behaviour out and replace it with the newly slightly more aligned model.

In the rest of the world we delegate whatever idea of moral agency we have about the model to a person, which may include a corporation as far as legal philosophy is concerned and depending on the type of harm created.