This is a pretty obscure and in-the-weeds benchmark, but to me the models’ interpretation feels quite reasonable.