Remix.run Logo
user43928 7 hours ago

About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.