Remix.run Logo
ozgung 4 hours ago

Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

CGamesPlay 3 hours ago | parent | next [-]

> There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.

WarmWash 2 hours ago | parent [-]

Intent is what is being discussed here though, not liability.

A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.

monkpit 2 hours ago | parent | next [-]

Intent might be what’s being discussed but intent is, for the most part, legally irrelevant. It might make the difference in the degree of a murder charge, or maybe manslaughter, or criminal negligence, but it doesn’t get you off the hook.

WarmWash an hour ago | parent [-]

Correct, but the size difference of the hook can be so dramatic that you can't just hand wave it away.

The trainer who trained the lion to kill will probably be in jail for life. The one who happened to oversee a lion that went rouge would probably be given probation or something else that is a slap on the wrist.

infamouscow 2 hours ago | parent | prev [-]

Except liability always precedes intent.

azakai 3 hours ago | parent | prev | next [-]

Also, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems".

It would be great if they were so reliable, but I don't think they are!

ssivark 2 hours ago | parent | prev | next [-]

> Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.

It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.

WarmWash 3 hours ago | parent | prev [-]

Nobody picks up pitchforks for rational nuanced takes.

Knee-jerk surface analyses is far more powerful.