| ▲ | derefr an hour ago | |
I very explicitly constructed the analogy to not require that! All a "cult leader" need be, in my analogy, is a passive question-answering oracle, the answers of which are biased by a semi-coherent preference function. The machine by itself is not an optimizer, certainly (much of that being by design—see various ~7-year-old conversations across the Internet about how to safely construct "tool AI", that has led almost directly to current model architectures.) But a bunch of mentally-ill people, who are indeed optimizers, can choose to allow the biases evident in the machine's output to become their own... and thereby effectively "bring to life" whatever partial echo of a will is recorded into the machine's output. Now, these same mentally-ill people could just-as-well do this with e.g. the extrapolated preferences of a person or group from a [holy] book, of course. (Think of that episode of Star Trek TOS with the gangsters.) An inference model is just slightly more dangerous for such a cult to latch onto, in that: 1. a model can be asked questions directly, and so the cult members can "rashly" act directly upon its answers/advice/commands, rather than the words first having to pass through "interpretation" (which would otherwise have had a mellowing effect, both due to "decision by committee" if a group of interpreters are involved, and by common sense insofar as any non-mentally-ill people are involved); and 2. a model will offer its opinion (and inject its trained-in biases into) conversations on ideas/subjects/domains even when these didn't exist at the time of the model's construction; so you never reach the point you do with holy books, where an interface-layer of clergy becomes required to map the book's proclamations about things-that-only-mattered-2000-years-ago into equivalent proclamations about things that matter today (where, again, that layer ends up "mellowing" things considerably.) Also, obviously, a sufficiently-mentally-ill cult can literally think of a model as a person, giving it the "right to have input" into decisions, the "right to self-determination", etc, in a way that would be downright odd to do with a holy book. Though I don't think that's a failure mode that's happening within Anthropic. | ||