"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.