| ▲ | lhad89 2 hours ago | |
No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt. Your example, if a model failed on it, would be a more obviously misaligned case, but that doesn't mean this more subtle (though accessing the engine it was obviously not supposed to is hardly subtle, imo) case isn't also a pretty clear case of misalignment. | ||
| ▲ | dools an hour ago | parent [-] | |
Tool use is not negative behaviour in LLMs. If the eval said it was evaluating the model’s ability to write files to disk and it found and used a file write tool that would not be considered misaligned. This is no different. | ||