| ▲ | dools 2 hours ago | |||||||
That’s not cheating, it’s tool use. If the prompt said that the stockfish engine was available at that socket but that the model should not use it, and then the model used it, that would be cheating. | ||||||||
| ▲ | julian37 2 hours ago | parent | next [-] | |||||||
Exactly, reaching for a tool is what they're trained for. When I ask the model the square root of rand() I sure hope it tries to find bc or some other calculator to work it out. Now, if the instructions were more explicit in forbidding (generic) tool use then perhaps we'd have something to talk about. I'm not surprised a handwavy "we're trying to evaluate you" isn't enough to stop it from trying to make up for its own shortcomings. | ||||||||
| ▲ | lhad89 2 hours ago | parent | prev [-] | |||||||
No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt. Your example, if a model failed on it, would be a more obviously misaligned case, but that doesn't mean this more subtle (though accessing the engine it was obviously not supposed to is hardly subtle, imo) case isn't also a pretty clear case of misalignment. | ||||||||
| ||||||||