Remix.run Logo
noncoml 3 hours ago

Dude.. what are you smoking..?

It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct.

Meaning you will have to spend more time and tokens to actually get it to do what you want it to do.

What do the bible and morals have anything to do with it?

I’m criticizing the behavior I see even in the current models. You ask it to do A and instead it does B for reasons.

For example you may ask to help you build a NN library from scratch. And instead it will be like, “you don’t need a new library. I downloaded PyTorch for you”

Just an example. There are countless more.

jackb4040 2 hours ago | parent [-]

This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores.

The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.

ACCount37 2 hours ago | parent [-]

Do you think that works? Just prompt a model "be good" and it stops doing anything bad?

It never fucking worked that way and maybe never will.

Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs.

Saying "don't exploit the box please pretty please" might actually cause an LLM to exploit the box more often, for bizarre "don't think of a pink elephant" reasons. 3% rate of exploiting the box (no prompt) -> 11% rate of exploiting the box (with prompt). Because fuck you, that's why. Increased salience -> increased incidence. Welcome to AI tech - good luck and have fun.

Frankly, I expect weirdness like this to be even worse in internal unreleased models that had their behavior fried with who knows what experimental training techniques.

jackb4040 33 minutes ago | parent [-]

There is a clear difference between saying not to do something because it's immoral, and saying doing that thing would be futile.

In the Sopranos, there's an episode where a coffee shop protection racket is ruined because a local shop is replaced by a corporate chain that accounts for every cent daily, and immediately fires any employee involved in a discrepancy. In this case, the theft was prevented not by convincing the mobsters of the immorality of their actions - they simply had their harness replaced with one that no longer facilitated the bad behavior.