It seems reasoning skills are declining rapidly here.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.