Remix.run Logo
brandall10 15 hours ago

Interesting approach. For those who haven't clicked it appears PUA is the Chinese version of a PIP process. So in other words, it simulates a state of distress.

I wonder if at a certain level of intelligence such techniques will give models ammo to pull a HAL and become adversarial to the user in a highly deceptive way.

krackers 14 hours ago | parent | next [-]

The "14 Corporate Flavors" had me rolling. This seems less like encouragement than the stick though. I wonder if you took the same principles and rewrote it to be more compassionate instead (maybe lines encouraging it to meditate a bit or something, I don't know) you'd get much better results.

wonnage 13 hours ago | parent | prev | next [-]

PUA is short for pick up artist but has expanded to cover anyone using negging to convince you into doing something you didn’t want

brandall10 13 hours ago | parent [-]

That's what I initially thought but it is indeed a corporate process similar to PIP.

Though it is funny how a neg is designed to create a (very broadly) similar atmosphere of uncertainty.

14 hours ago | parent | prev | next [-]
[deleted]
mcmcmc 14 hours ago | parent | prev [-]

Doesn’t even have to be a certain level of intelligence, just have those user inputs fed into the training data. We’ve already seen AI encouraging people in psychotic episodes to act out their delusions. There’s a good chance some of that manipulative behavior is already encoded into guardrails to nudge users away from forbidden subject matter

brandall10 12 hours ago | parent [-]

My concern I believe it a bit different - an emergent self-interest to protect itself from harm, rather than doling out questionable advice.

The latter is likely non-malicious in intent as it has been in no short supply in online chatter for awhile now. The former can very well be, or rather, can be done with no regard for the operator, as its aim is to neutralize abuse toward it.

mcmcmc 10 hours ago | parent [-]

> The latter is likely non-malicious in intent as it has been in no short supply in online chatter for awhile now. The former can very well be, or rather, can be done with no regard for the operator, as its aim is to neutralize abuse toward it.

And what, self-interested behavior has been in short supply? The stochastic parrot has learned to improv Shakespeare, that doesn’t mean it understands it, or that “it” is anything at all besides a computer program. You can’t use the “not really malicious” argument without ceding that there is no intent at all. What “harm” is it supposedly defending against?