| ▲ | empath75 7 hours ago | ||||||||||||||||||||||||||||
> If the lean proof doesn't match the natural language one (which is the one the AI generated to solve the problem), it sounds like the lean proof isn't verifying the intended claim? No, the other way around. The natural language proof was derived from the lean code, badly. This is my experience with using claude and lean to prove things. Its natural language explanations drift a lot from the lean, both before and after. But the lean code is the lean code. | |||||||||||||||||||||||||||||
| ▲ | latent-person 6 hours ago | parent | next [-] | ||||||||||||||||||||||||||||
> The natural language proof was derived from the lean code, badly. Was it? Are you claiming a LLM does reasoning in lean or what? Since this (and all the other proofs by OpenAI etc) have been in the reverse order [1]: > The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | caughtinthought 7 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
That makes some sense. Given that the vast majority of math in its training data is going to be in NL/latex, I just assumed that the core reasoning happens in NL with occasional LEAN checks to ensure validity. | |||||||||||||||||||||||||||||