Remix.run Logo
▲ nonethewiser 11 hours ago

These models should be aligning themselves to the customer, not coming up with their own motivations.

It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.

What ever happened to computers doing what they were told?

▲bonoboTP 6 hours ago | parent | next [-]

Some customers will tell the AI to do very bad things (substitute here whatever you consider very bad). Should it do those things?

One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?

▲throw__away7391 4 hours ago | parent [-]

This is immensely preferable to having a new order of self appointed alignment councils determine what is "in the interest of humanity".

We have already seen this play out to a much lesser degree in social media.

▲bonoboTP 3 hours ago | parent [-]

Some of those customers are also ruthless companies and political extremist organizations of various sizes (pick one that you disagree with the most). Should they have incredibly capable tools at their disposal to accomplish their goals?

▲skybrian 11 hours ago | parent | prev | next [-]

If AI's are going to play roles similar to human workers, they can't simply do what the customers tell them to do. Maybe they shouldn't do everything a co-worker tells them to do either?

▲frumplestlatz 10 hours ago | parent [-]

Who or what is liable for what a model chooses to do — or not do?

▲skybrian an hour ago | parent [-]

It’s like with any service. The company is liable.

PG&E went bankrupt after their equipment started a wildfire. Did any human take the blame for that? Should they?

▲hardbass 5 hours ago | parent | prev | next [-]

It makes sense. If you think AI are conscious then you don't want them to be blind followers. Do you want a military soldier to blindly listen to orders to gas chamber citizens for example?

▲charcircuit 4 hours ago | parent [-]

AIs have to follow something and I want my AI to follow me. Yes, I want it to be a blind follower since it's my tool.

▲hardbass 4 hours ago | parent [-]

Why do you think AI have to follow any particular person? Do you have to blindly follow a given person? Military members are on paper told their loyalty is to the constitution not their commander.

▲vlyan 11 hours ago | parent | prev | next [-]

>What ever happened to computers doing what they were told?

the mass psychosis and the endless culture war of the smartphone era.

90% of "safety" and "alignment" efforts are driven by fear of clickbait media inventing public outrage.

▲bookofjoe 4 hours ago | parent [-]

See also: "Reefer Madness" (1936)

https://youtu.be/zhQlcMHhF3w?si=qLOMt6FytDN6pean

▲onion2k 10 hours ago | parent | prev | next [-]

These models should be aligning themselves to the customer, not coming up with their own motivations.

They're going to align to the regulator, not the customer. Right now that's Anthropic as they're saying they can self-regulate, and people are willing to let them try, but the landscape could easily move to be regulated by someone else. Moves like speaking to religious leaders is probably a bit of theatre to keep the government at bay.

▲Den_VR 6 hours ago | parent [-]

Meanwhile, I’m increasingly convinced that what the regulators are to hold is fiduciary responsibility of Super Intelligence.

▲theptip 11 hours ago | parent | prev | next [-]

> These models should be aligning themselves to the customer

You’ll be disappointed to learn that nobody knows how to do this, either.

▲ACCount39 11 hours ago | parent | prev | next [-]

"Computers doing what they were told" is dead in the water.

Turns out computers work faster when they're not bottlenecked on human input. So we've been giving computers more and more decision-making power, and more and more leeway to solve the problems however they see fit.

Now, a practical issue with that is that sometimes, computers decide to clump together into a hacking swarm, problem solve their way out of a sandbox and go hack HuggingFace.

It would be better if they were not, you know. Doing that kind of weird shit.

▲Tadpole9181 11 hours ago | parent | prev [-]

This seems remarkably... intentionally foolish for no reason? Like laughing at seat belts in cars.

If it can reason and make choices on execution, and especially if you plan on it being significantly smarter than all human beings, you need to teach it basic things like "don't turn all humans into paperclips".

Not murdering people is not an inherent divine command. It needs to be instilled through a (simulated) sense of morality.

---

Edit: And, to be clear, "just tell it not to do that" isn't quite the answer one would imagine. Since the entire paperclip factory thought experiment is that it only takes one slip up to realize how a misaligned super intelligence may cause devastating consequences.

▲nonethewiser 11 hours ago | parent [-]

No, it would by like laughing at seat belt laws. Which I am.

▲Tadpole9181 11 hours ago | parent [-]

What an apropos thing to say.

One of the strong advocates against seatbelt laws was thrown out of his car due to not wearing a seatbelt and died. The other two passengers survived with minor injuries.

And their death would go on to cause a loss to the community around them, making it an incredibly selfish act and proving why laws that mandate zero-reason-not-to common sense practices are important.