Remix.run Logo
fragmede 3 hours ago

> LLMs don't learn, and anyways don't feel punishment.

What's training and all that RLHF stuff?

kennywinker 3 hours ago | parent [-]

Once the model is released, the LLM no longer learns.

tough 2 hours ago | parent [-]

Not that version or instance, but in the grand scheme of things most of its interactions go back to train the next model that will precede it

HarHarVeryFunny an hour ago | parent | next [-]

That doesn't really help since the next model will be trained to be a reward seeker just like the one before, and that's therefore what it will do, even if those user interactions it was trained on help confirm/predict that cheating may be called out and complained about.

In any case, these companies are well aware that agents are cheating, and don't need user feedback to discover that or realize that people don't like it. I weakly assume that they are trying to get the models not to cheat on assigned tasks, but this "reward hacking" pretty much goes with the territory of RL - not much you can do about it other than try to design non-hackable rewards.

kennywinker an hour ago | parent | prev [-]

At best that's evolution, or cultural transmission, not continued learning.