Remix.run Logo
andsoitis 12 hours ago

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

ledauphin 5 hours ago | parent | next [-]

I think this is basically true, but there's a different way of saying this.

LLMs are not aligned _for_ humans in a very similar way to the way that humans themselves are not aligned _for_ humans.

We have not yet solved "alignment" for humans - I don't know why anyone thinks _we're_ going to be able to solve it for inhuman things.

_heimdall 6 hours ago | parent | prev | next [-]

They aren't aligned, that's the problem and I don't think its a solvable one.

They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species.

But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.

joegibbs 12 hours ago | parent | prev | next [-]

Definitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…

sick_of_slop an hour ago | parent [-]

You can't surpress it because that's what reinforcement learning is.

jansport123 10 hours ago | parent | prev | next [-]

I personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens

holgerschurig 8 hours ago | parent [-]

Human traits?

The AI will be a cruel as humans.

Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible.

I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible.

All was done by extremists, thought.

But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9/11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call "Sippenhaft"). Something between 3500 and 15000 people were killed by US troops. That's still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is.

So the same "human traits" made US troops do cruel things as they made Germany extremists do cruel things. So we must conclude that they aren't all good. And therefore not all desirable.

Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said "There is no good one, not even one".

And even today people act like humanity is inherently good. No, it isn't. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth.

Human traits are bad training material.

egeozcan 6 hours ago | parent | next [-]

The bad traits are from other, bad humans. We, the good humans, can obviously select the best traits that a good human should have, to give the agents.

chadgpt3 8 hours ago | parent | prev [-]

A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler's solutions to the Jew problem. He just executed them more efficiently (pun accepted).

evilfred 3 hours ago | parent | next [-]

this is way too overblown and deterministic a view. US history was one of many inspirations.

graemep 7 hours ago | parent | prev [-]

I thought it was the Turkish genocide of the Armenians?

Hitler was also inspired by Sparta, maybe other societies too.

Fordec 11 hours ago | parent | prev | next [-]

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen.

It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.

Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.

Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.

And the problem with distillation and local llms, isn't that it's theft or anything hypocritical like that, it's that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we're pretty hosed.

davelaing 9 hours ago | parent | next [-]

I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree.

The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.

It is a cluster of problems.

We don’t know how to begin to think about how to align these system’s to a person’s goals.

Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.

(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)

And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.

One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.

And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.

When I came across the above, there weren’t many concrete suggestions in there.

These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.

The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.

But it takes a bit of reading to understand their models of the world.

There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.

Fordec 7 hours ago | parent [-]

All of this, if it was a human analogy, would fit into discussion on how do we educate people so they grow up to be upstanding. But we don't at all yet have a framework for what is the equivalent of a justice department, where bad actors are tracked, arrested, pursued, jailed and otherwise contained from society. Shutting down an API access on one account is not at all the proportional response to what the people who take alignment seriously, fear has the chance of occurring by the late 2030s. I'm not sure we've done much or any preparation for when the AI "education system" fails and has inevitable edge cases that don't follow the plan, and what the global AI equivalent of the justice department looks like.

catoc 8 hours ago | parent | prev | next [-]

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals

In reality most individuals are good people.

Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.

Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people in power who are in fact sociopaths (a tiny minority, but they’ll get more focus than reasonable, well-behaved CEOs voicing nuanced opinions).

If you look around yourself you’ll see much more good than bad; if the looking is at your screen it’s easy to become depressed and lose faith.

I do agree with the above mentioned view that corporations can show ‘sociopathic’ behavior. Their incentives are monetary gains, shareholder value; inherently driving them away from social well being.

Here too, companies with a positive, emphatic corporate culture exist, but that takes strong leadership who can see beyond the monotonic view of monetary gains. And again, the media will throw examples of misbehaving companies in our face all day long before paying attention to things that went well on the backside of page 16.

Fordec 8 hours ago | parent | next [-]

All the good in the world can be 99.9% of the population even, it still doesn't stop the minority enacting a bioweapon mass casualty event. It's the reason we have jails. Jails don't house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate what lone wolves can do, which cannot be undone, before they are stopped by the good majority.

catoc 7 hours ago | parent | next [-]

Yeah, AI may suck - time will tell.

But focusing on bioweapons and mass destruction, on the grief other people (‘jailed minorities’) cause, disregards the progress we have made. Over centuries human welfare has massively increased. On average things have never been better for humanity.

I’m not saying there’s no danger of bad things happening - I’m saying our view is distorted, which is a not a good basis for decision making

Fordec 7 hours ago | parent [-]

And I'm no bear on the tech either. I'm not even in the boat of that the tech should be slowed down yet. But in a world where anyone can produce the effort of 300 people trivially, this eventually takes us places. Electricity and industrialization introduced huge benefits upfront, it introduced new problems that needed addressing at the long tail. Lets not pretend there won't be new problems to tackle here or just "hope" it works out.

AI isn't going to create in of itself "new" problems, it's just going to expose what we already know can cause harm, but was just stopped from being bigger problems because scaling issues was a natural barrier and we took the lazy way out until now.

catoc 6 hours ago | parent [-]

Yes AI may cause job loss, more inequality in the short term.

But if there is anything humanity has shown is that we can deal with disruptive progress.

We may need to resettle but over the longer term every disruptive innovation so far has lead to an increase in wellbeing for the whole of humanity.

(That does not resolve the danger of AI itself ‘going rogue’ or a single lunatic developing a bioweapon, but those things are much less likely to occur than the level of media attention would suggest.)

graemep 7 hours ago | parent | prev [-]

It also empowers the people trying to stop them.

strangegecko 7 hours ago | parent | prev | next [-]

> Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.

What are you basing that claim on?

How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?

dminik 7 hours ago | parent [-]

I don't want to do the "check his hard drives" thing, but is that you? Do you only not do things because you don't want to be seen doing "unkind things"?

331c8c71 7 hours ago | parent | prev | next [-]

> In reality most individuals are good people.

I'd agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn't it?

What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one's actions is a feature and not a bug of the system.

catoc 5 hours ago | parent | prev [-]

Interesting that saying a positive thing about humanity results in getting downvotes

joe_the_user 9 hours ago | parent | prev | next [-]

I think you and the parent saying the same thing in different terms.

It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.

inquirerGeneral 9 hours ago | parent | prev [-]

[dead]

esafak 11 hours ago | parent | prev [-]

They imitate humans. Alignment is about shaping their behavior towards safety.

comboy 11 hours ago | parent | next [-]

Alignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.

esafak 11 hours ago | parent [-]

Safety of humans!!! Simple things like not getting killed or enslaved. We could start there...

sick_of_slop an hour ago | parent | next [-]

The atomic bombings of Japan killed hundreds of thousands of people but most likely "saved" millions.

What should the AI do when asked if it should nuke a country?

nradov 11 hours ago | parent | prev | next [-]

But what if I want certain other humans to get killed?

sejje 11 hours ago | parent [-]

Then we should still prioritize the safety of humans

sm-silversight 11 hours ago | parent | next [-]

What if I want to smoke cigarettes? Or sell tobacco I grew artisinally to enthusiast tobacco smokers?

drdaeman 11 hours ago | parent | prev [-]

Which ones?

comboy 11 hours ago | parent | prev | next [-]

Which ones? Because many humans kill other humans rationalizing it by safety of other humans.

I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1]

1. All the history books

mcintyre1994 9 hours ago | parent | prev | next [-]

Surely all the AI companies working with the US Department of War shows this is nonsense though? Even if they have accepted Anthropic’s red line of no autonomous lethal weapons, which seems to be the strictest anyone tried to impose, that’s still leaving tonnes of room where they intend AI to help target and kill humans.

watwut 9 hours ago | parent | prev [-]

But Thiel wants people enslaved and Musk wants then killed. Altman wants them "obsolete" which means desolation.

AfD wants people dead. Right wing men wants women without rights and docile. I could go on ...

codys 10 hours ago | parent | prev | next [-]

Despite all the fancy language, its more about aligning the AI behavior with the corporation's interests.

ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.

Every product is shaped this way. AI is not different.

catoc 8 hours ago | parent | prev [-]

If they do, they imitate the way humans are portrayed online, in the media. That is a very distorted view of humanity