LLMs find flaws in human-published math papers _all the time_, usually in the process of formalizing them in lean.