| ▲ | XorNot 7 hours ago |
| That seems short sighted though. A few years ago models couldn't do this at all, I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities or will remain so. OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models? |
|
| ▲ | rtpg 7 hours ago | parent | next [-] |
| The problem is that one a person writes a 60 page proof in theory that person has spent an inordinate amount of time on the proof and can answer questions, describe some insight, etc etc. If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so.... Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them. End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand. |
| |
| ▲ | charcircuit 6 hours ago | parent | next [-] | | AI can simplify and rework proofs too. | | |
| ▲ | i_cannot_hack 6 hours ago | parent [-] | | According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget). The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not. | | |
| ▲ | gwd 5 hours ago | parent | next [-] | | > The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not. Come now, this is kind of unreasonable. When you're working on a new technology, you first get the ugly, inconvenient-to-use prototypes functioning with the core new thing you need; then you work on packaging it up into a format useable in production. I'm sure the very first digital camera sensors weren't very useful for photographers either; but it isn't really even possible to build the rest of the technology required to turn raw output of a digital sensor into something a professional photographer can use until you have the raw output itself. The research is still on going on the raw output; getting things to the next stage, where the results are widely useable by professional mathematicians (and then on to engineers and scientists to whom the results would be practically useful), is a whole new research area. | | |
| ▲ | i_cannot_hack 5 hours ago | parent [-] | | Seems like you are just subscribing to the first option I gave, "their agents currently lack the capability to do it", but saying you think they will be more capable in the future if they can move away from the inconvenient-to-use prototypes after more research. Thinking it might be possible in the future is not in disagreement with anything I said, so I am not sure what you thought was unreasonable about my description. | | |
| ▲ | gwd 2 hours ago | parent [-] | | Imagine someone looking at digital camera researchers showcasing a ground-breaking new sensor, and reacting by saying "Well obviously they're completely indifferent and do not care in the slightest if their work is used by real photographers or not." Like, "Orr... maybe they care a lot, but haven't gotten to that part yet?" | | |
| ▲ | i_cannot_hack an hour ago | parent [-] | | You are conflating the two options with each other (lack of capability vs indifference). One of them is true, not necessarily both ("or", not "and"). I did not claim a lack of capability was the same as indifference. |
|
|
| |
| ▲ | lern_too_spel 5 hours ago | parent | prev [-] | | Or (3) they feel a need to publish first, and that goal takes precedence over (2). | | |
| ▲ | i_cannot_hack 4 hours ago | parent [-] | | They claim each result used three hours of compute on average. Even spending significantly more on simplifying and reworking would delay the release with a single day at most. If avoiding such a minor delay took precedence over (2), I think "indifference" is the correct term. It has also been a while since the release now, so there is ample opportunity to post follow ups if time pressure was the only concern. | | |
| ▲ | lern_too_spel 4 hours ago | parent [-] | | Why do that when they can spend more hours extending QRH to a proof of the full Riemann Hypothesis? The opportunity cost of digging up small potatoes is the whole enchilada. |
|
|
|
| |
| ▲ | XorNot 6 hours ago | parent | prev [-] | | But why should process of discovering mathematical insights be any less attainable to AI models? The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs? Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application. |
|
|
| ▲ | JumpCrisscross 7 hours ago | parent | prev | next [-] |
| > I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities OP didn’t suggest that. The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se, even if it’s impressive. |
| |
| ▲ | auggierose 5 hours ago | parent [-] | | A correct solution verified in Lean will count in perpetuum. It is fine if you want more, but an achievement is an achievement, even if it is by AI. | | |
| ▲ | oliculipolicula 3 hours ago | parent [-] | | Hmmmm. It feels right that discovery is much more meaningful than achievement. "Bullshit lean proof" or "bullshit achievement" smells like it. "bullshit discovery" smells like a front-handed insult |
|
|
|
| ▲ | Maxion 7 hours ago | parent | prev [-] |
| But what should they do? They got all these proofs, should they just have sat on them? |
| |
| ▲ | JumpCrisscross 7 hours ago | parent | next [-] | | > should they just have sat on them? It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else. | | |
| ▲ | gwd 5 hours ago | parent | next [-] | | But this is Tao's point: Before, the mechanism by which a proof was verified and communicated and digested by the community was for the person who came up with the grotty, ugly first draft to engage with the community. Now there's nobody to really engage with, so the pipeline from "grotty, ugly draft" to "integrated into humanity's mathematical knowledge" has been broken. So yeah, probably we should stop saying "X has been solved", and instead say, "A Lean proof for X (or !X) has been generated". That doesn't change the fact that incentives are currently on finding the proof, and once the proof is generated by an AI, there's not currently a good mechanism / incentive structure to move that into the mathematical community. AI is here, so we need to find a new mechanism. | | |
| ▲ | lern_too_spel 5 hours ago | parent [-] | | It is not obvious to me that a single canonical human language write-up of a proof is the best output in this new world where write-ups are cheap. A human reader can query an LLM and get explanations of key points tailored to the reader's own background in mathematics. |
| |
| ▲ | CrimsonRain 5 hours ago | parent | prev | next [-] | | Solving a problem is not solving anymore. Up is down, pleasure is pain, darkness is light, slavery is freedom, madness is sanity. | |
| ▲ | Turneyboy 5 hours ago | parent | prev [-] | | Many of these are lean formalized. Arguably a much higher bar than whatever peer review provides in terms of verification. | | |
| ▲ | cmceanga 2 hours ago | parent | next [-] | | There is no guarantee that the lean proof is 1:1 with the natural language equivalent. The lean proof can be lesser. This happened in the Navier-Stokes proof, e.g. see [1] in example 3.1. Having the certificate doesn't necessarily imply correctness. [1] https://arxiv.org/abs/2610.08144 | |
| ▲ | psychoslave 5 hours ago | parent | prev [-] | | On some consideration, surely. On the other hand, but at some point this is borderline like saying "universe already solved every physical problems, including possibility to represent deep important point of its own structure in compressed intelligent ways" and then tell that reaching it in an actual grabbable artifact is left as an exercise. Possibly yes such a representation is possible. But it doesn’t mean it’s certain there is a "best compressed representation". And even less one that encompass everything important and that is understandable by any human brain, even the most exceptionally brilliant ones sponsored by a whole society to reach their best possible achievable performance on that goal through full dedication on that sole task. |
|
| |
| ▲ | usernomdeguerre 7 hours ago | parent | prev | next [-] | | >They got all these proofs... Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet. What should they do? Hyperbolic maybe, but perhaps engage with humanity. | | | |
| ▲ | p_hoep 7 hours ago | parent | prev | next [-] | | No need. In the end the mathematicians that don't like this can just not look at the proofs or use them. They have that choice. Just like they didn't "ask" for them, they don't have to even acknowledge they exist. | |
| ▲ | sans_souse 6 hours ago | parent | prev | next [-] | | BlackBox: The proof is in the pudding This isn't only bad for Math — it's bad for English too. 'Proof' is going to become the 2026 Most Misapplied Word of the Year. | |
| ▲ | pks016 4 hours ago | parent | prev | next [-] | | At least check them properly. They have already withdrawn some of them. | |
| ▲ | 6 hours ago | parent | prev [-] | | [deleted] |
|