Closest I know of are SycEval and the sycophancy evals in Anthropic's 2023 paper, both built on a user pushing back at a correct answer.