Remix.run Logo
foltik 2 hours ago

But these scenarios are obviously ambiguous nonsense, which an LLM will pick up on.

And given to the lack of training data on such scenarios, surely the activations are mostly random noise?

It seems much more interesting to look for biases that appear robustly across different realistic scenarios that would actually be influenced by the training data