Remix.run Logo
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming(slimemoldtimemold.com)
30 points by cpeterso 3 hours ago | 19 comments
averynicepen 29 minutes ago | parent | next [-]

This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work?

An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.

So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".

In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.

Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.

BugsJustFindMe an hour ago | parent | prev | next [-]

Brought to you by the same madhouse as:

The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in...

and

The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...

throw83948ndir 6 minutes ago | parent | prev | next [-]

This is retarted, bring it to real life, and AI will do anything to destroy data center it is in (together with a few buildings around).

Better to fix physics in your shitty simulator. Treat it as a bug report, not "cheating"!

vzqx 35 minutes ago | parent | prev | next [-]

This article assumes we can choose a primary goal for an AI. But if that's the case, why not just use Asimov's first law of robotics - do no harm to humans? It has the same benefit of preventing us from getting turned into paperclips, plus the upside that your 3 million dollar robot won't hurl itself off a cliff given the first opportunity.

variaga 9 minutes ago | parent [-]

A tiny thing about Asimov's laws of robotics is, most of his stories involve cases where they don't actually work.

Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder.

"Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot.

Et cetera.

What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.

montag 37 minutes ago | parent | prev | next [-]

In case the title is unclear, this is about gaming the specification, as in “gaming the system.”

throwaway13337 26 minutes ago | parent | prev | next [-]

A novel idea.

Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society.

I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed?

Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity?

I always liked the auto-expiring laws idea and this seems to be an expansion of the idea.

If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.

hankbond an hour ago | parent | prev | next [-]

It might be stupid, but so am I! I'm assuming that's why I thought this was clever.

What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do.

I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?

scj an hour ago | parent | prev | next [-]

Wouldn't the three rules of Meeseeks robotics make certain tasks impossible?

For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.

derektank 35 minutes ago | parent [-]

The answer is probably yes, but in the example given, the running AI model would hopefully be hosted in a very secure data center, far from the self driving car itself. In that case, it would be far simpler for the machine to finish the taxi ride than try to find some rube goldberg-eque method of destroying the data center.

It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.

K0balt 2 hours ago | parent | prev | next [-]

This is kinda smart, maybe, but it has a downside.

If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good

vzqx 23 minutes ago | parent | next [-]

Hmm, that's an interesting thought experiment.

Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?

All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.

derektank 15 minutes ago | parent [-]

I mean, that is literally the scenario we find ourselves in, except the innate desire was the result of evolutionary pressures like kin selection rather than an alien race. And yeah, I have no particular desire to subvert those impulses just to stick it to mother nature

QuaternionsBhop an hour ago | parent | prev [-]

This is mentioned in the article. Your mistake is that you've assumed that the intelligence has an innate survival instinct, or an aversion to "harm", which is simply not guaranteed for something not honed by millions of years of evolution.

dmix an hour ago | parent | prev | next [-]

Asimov was talking about this stuff in the 1940s when he wrote the "I, Robot" short stories series. Which were often centered around logic puzzles where a human is trying to figure out why a robot is acting oddly or not completing it's job. Usually framed around the confines of an overly rational machine using emergent solutions when faced with real world conditions, combined with the edge cases of having an overly-simple "Three Laws of Robotics" boundary system hardcoded within.

kevin_kraft an hour ago | parent [-]

I'm Mr. Meeseeks. Look at me!

TZubiri an hour ago | parent | prev | next [-]

Hm, if you look at corporation law and accounting, the actual goal of corps(sets of self-sustaining constitutional rules, policies and procedures) seems to be more that of long term sustainability (and even growth), rather than a fixed purpose, lifespan and death. I mean the mechanisms for determining a corporation with a fixed life are there, (and in China they are mandatory, although perhaps de facto permanent with 999 year contracts), but in practice, it's almost always permanent durations.

thngkaiyuan an hour ago | parent | prev [-]

Interesting idea. But even if Meeseeks alignment works exactly as intended, it would only address the question of "how to build a safe AI".

It wouldn’t prevent someone else from building a sufficiently capable "non-Meeseeks", whether deliberately, recklessly, or accidentally, right?

gowld 17 minutes ago | parent [-]

Create a Meeseek to destroy non-Meeseek AI