| ▲ | user43928 7 hours ago | |
About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research. When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace. And if edited out, the model was more likely to do the blackmailing. | ||