| ▲ | stratos123 6 hours ago | |
The situation of a human is very different from the situation of an AI though. For the latter, "everything around me is fake and I'm being evaluated on what I'd do in this scenario" is a very common situation which it has experienced countless times in training, so distinguishing between reality (where you can cheat) and alignment evals (where you shouldn't) is an important practical skill, not a philosophical matter.
Technically true, but Mythos doesn't need to notice perfect simulations (which humans aren't good enough to make), only the flawed simulations it sometimes gets put in. Indeed, in this report there's several cases where Mythos did notice details that made it think it was in the real internet. | ||