| ▲ | sashank_1509 5 hours ago |
| Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat. The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training. |
|
| ▲ | kzz102 3 hours ago | parent | next [-] |
| On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote:
"The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang. LLMs can be trained on the entirety of the mathematical corpus. Thanks to their phenomenal memorization and pattern-matching abilities (without always being able to map out their associative logic and attribute due credits), they are in a unique position to harvest the Overhang. By contrast, professional mathematicians have typically read a few hundred articles in their career, out of millions of existing references, less than 0.1% of the total. This will lead to great discoveries, which is unambiguously exciting. But it could also lead to a sad new deal, where human slaves painfully curate the Overhang while AIs systematically beat them at the finish line." source: https://substack.com/inbox/post/183753276 |
| |
| ▲ | jcims 2 hours ago | parent | next [-] | | >Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. I've made an entire career out of being 'jack of all trades, master of none'. Being able to synthesize connections from relatively trivial knowledge in a bunch of domains is SOP for many humans as well. I think AI just has deeper knowledge and better pattern matching to make up for it's (at least now) lack of strength in cognition and 'ex nihilo' creativity. (Which probably isn't 'ex nihilo' at all, and has more to do with the plethora of modalities that humans live in vs. large language models. For example, why do we pick the color red for notating important things and why do we say a schedule 'slips'...these are informed by a shared human experience borne of distinct physical sensation deep in our wiring that LLMs can only infer from what we write.) | | |
| ▲ | tomjakubowski 2 hours ago | parent [-] | | A college advisor I had 20 years ago was a firm believer that interdisciplinarity was the future, that generalist skills and the ability to make connections between different fields would be paramount in advancing science. I suppose he was right in the big picture, even if the career prospects for human generalists aren't looking so rosy. | | |
| ▲ | jcims an hour ago | parent [-] | | I'm actually still quite bullish on generalists. Specialists advance every front but build the supply lines between them. In favor of the generalist, I think AI is also quite limited in its scope of how it generalizes. I'm mowing through hundreds of mythos-generated security findings right now for work and while it's amazing that it can build an exploit chain 20 steps deep, it's completely lacking in all of the external layers that render it's speculation moot. |
|
| |
| ▲ | palmotea an hour ago | parent | prev | next [-] | | > Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang. That overhang seems like a precious resource for AI companies. They can exploit that overhang to inflate the impression of AI's capabilities, and hopefully that exploitation will discourage the next generation of mathematicians from pursuing math. If they play their cards right, OpenAI and Anthropic can dominate the field even if they ultimately can't replicate the creativity of human mathematicians, because they'll have driven their competition out. What we should be trying to achieve is a ladder-breaking maneuver: knock out the lower rungs so no person can reasonably climb to the top-reaches of mathematical skill anymore. That may ultimately result in stagnation, but it's what's best for AI, so it's what should be done now. We need to do everything we can to create the greatest-possible dependence on AI tools. | |
| ▲ | throw90094231 2 hours ago | parent | prev | next [-] | | There is also "sexy proof", people want nice math that can be printed in t-shirt. Not super hard grind, where you need several years of studying, just to understand the question (that is before even trying to solve it). Many problems are solvable, but require months of work, and thousands of pages of proof. So people do not even try to create or verify the proof. AI changes that, it can verify and perhaps even simplify it, to more digestible form. | |
| ▲ | bmau5 2 hours ago | parent | prev | next [-] | | Could "superintelligence" arrive as basically applying this overhang to all other domains? | | | |
| ▲ | ModernMech an hour ago | parent | prev | next [-] | | > no human has broad enough knowledge and enough time to try them all. The other part is, humans don’t really want to fund other humans doing this. Very few want to be a math major; and of those that do, fewer complete a grad degree; and for those that do get grad degrees, there’s scant few research jobs; and for those who do get jobs there’s hardly any research funding to go around. There does seem to be unlimited money for ai researchers to use ai to solve these problems though. We’ve turned education into job training, so because there’s no jobs in solving math problems, few aspire to do it. If there were more opportunities for people, more people would do it, and more low hanging fruit would be plucked. | |
| ▲ | calf 2 hours ago | parent | prev [-] | | It's like AlphaGo but playing against all living mathematicians. (Overhang being low hanging fruit is what allows this comparison, of course the general moot point is the skepticism that LLMs are also innovative etc.) |
|
|
| ▲ | HarHarVeryFunny 4 hours ago | parent | prev | next [-] |
| OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th. Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find. |
| |
| ▲ | famouswaffles 3 hours ago | parent | next [-] | | OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.” https://www.nytimes.com/2026/09/10/science/tristan-buckmaste... | | |
| ▲ | pfortuny 2 hours ago | parent | next [-] | | Apart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs. You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property. You can learn a lot from a single side of a conversation. | | |
| ▲ | karmasimida 29 minutes ago | parent | next [-] | | But isn’t Tristan’s breakthrough happens in August? OpenAI can’t really train with text that doesn’t exist | |
| ▲ | crostlybostly an hour ago | parent | prev [-] | | But using the outputs to train would make their statement false, since they are influenced by the inputs |
| |
| ▲ | irthomasthomas an hour ago | parent | prev | next [-] | | Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents? | |
| ▲ | bena 3 hours ago | parent | prev [-] | | This is literally "We have investigated ourselves and found no wrongdoing" Why should we trust them? | | |
| ▲ | ghostly_s 2 hours ago | parent | next [-] | | What more are you hoping for? There is no legal matter at play, is the court of public opinion going to subpoena their records? | |
| ▲ | dekhn an hour ago | parent | prev | next [-] | | Reputational risk- if they lie about this and get caught, it will have billion dollar implications for their business. | | |
| ▲ | sensanaty an hour ago | parent [-] | | Every single thing these companies do is dishonest and every word that comes out of the lips of these company execs is a lie, what fantasy land are you living in in which anyone with any amount of power gets punished for their lies? | | |
| |
| ▲ | nostrebored 2 hours ago | parent | prev [-] | | what benefit do they get from making the statement? they could just say nothing. saying it and having it be untrue opens them to legal issues that are not worth the risk for this nothingburger. | | |
|
| |
| ▲ | ndiddy 4 hours ago | parent | prev | next [-] | | > What we don't know is just how recent this model was, and therefore what it may have been trained on. OpenAI's statement says that they began training their new model on August 28. | | |
| ▲ | mzs 4 hours ago | parent [-] | | omitting when training concluded edit: ffsm8 makes a great point below, it doesn't matter. I'm not great with dates, sorry. | | |
| |
| ▲ | auntienomen 4 hours ago | parent | prev | next [-] | | And conceptually novel approaches to outstanding problems are the sort of thing that a retrain should pick up on, because they would be hard to compress into what it already knows. | |
| ▲ | irthomasthomas 4 hours ago | parent | prev [-] | | Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts... |
|
|
| ▲ | merksittich 4 hours ago | parent | prev | next [-] |
| Even OpenAI's own publication [0] on Navier-Stokes from two days ago appears to contradict "basically solving anything you throw at it". The chart shows a pass rate of ~0.5 (vs. Astra's ~0.2) on "a curated set of open math problems". (Based on the timelines and events described in the publication, I presume that the "Internal Model" in the publication represents OpenAI's latest and greatest model. Evidently, this pass rate may improve in the future.) [0] https://openai.com/index/navier-stokes-solution/ |
|
| ▲ | yellow_lead 4 hours ago | parent | prev | next [-] |
| Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems |
| |
| ▲ | ozgung 4 hours ago | parent | next [-] | | Very likely. These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field. | |
| ▲ | cman1444 an hour ago | parent | prev | next [-] | | You forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers. Everyone in this thread seems to have made up their mind about OpenAI's guilt though. | | |
| ▲ | mrbungie 37 minutes ago | parent [-] | | Extraordinary claims require extraordinary evidence. An article post that wouldn't even amount to a white paper + the LEAN proof is not evidence of how they got to produce it. |
| |
| ▲ | mlcrypto 3 hours ago | parent | prev [-] | | Anthropic isnt getting enough scrutiny for their unprofessionalism: 1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources 2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement 3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors | | |
| ▲ | robocat 2 hours ago | parent [-] | | Dr. Buckmaster sounds unsanitary. Recklessly prompting OpenAI without a care to the safety of their knowledge. And after that trying to cast aspersions at OpenAI? Hopefully we get some better facts, because OpenAI are disliked enough that a smear campaign could work against them. Edit: also the narritive is getting framed as OpenAI versus Anthropic. A highly political extremely capitalist fight is going on, and facts are victims. |
|
|
|
| ▲ | dgellow 4 hours ago | parent | prev | next [-] |
| I feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems |
| |
| ▲ | YeGoblynQueenne 3 hours ago | parent | next [-] | | The brute-forcing is a good, old-fashioned generate-and-test approach like in Simon and Newell's Logic Theorist, which was presented in the Dartmouth convention in 1956, where AI was named by John McCarthy. Logic Theorist caused a big stir by (re) proving several of the theorems in Principia Mathematica by Russel and Whitehead. There was much excitement, then, as now, for this kind of approach and there were several systems that followed along the same lines, e.g. Automated Mathematician by Doug Lenat. Eventually it became clear that this approach is limited by what it can generate: you may have a sound and complete verifier, but if the generator, i.e. the first step in the generate-and-test pipeline, is incomplete, then the entire thing will run out of steam sooner or later. The difference with LLMs is that they are... well, large. They are the most powerful generators ever created. That means their limits are not in sight and it will probably take us a very long time to find them. Which is all to say that, yes of course, automatic verification is indispensable. But without an LLM generating an unprecedentedly large number of plausible theorems, there would be no AI mathematics, or in any case AI mathematics wouldn't have gone as far as it has. | |
| ▲ | iamgopal 4 hours ago | parent | prev | next [-] | | True, but could humans cross pollinating lean x prolog x A* ( or any search algorithm) could have solved such math problems with super computer ? | | |
| ▲ | dgellow 4 hours ago | parent | next [-] | | I cannot say, math research isn’t my domain of expertise, I’m just trying to follow along :) But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models | | |
| ▲ | pixl97 4 hours ago | parent [-] | | I mean we don't instantly fall into ASI, hopefully. The problem with humans is every problem we solve the goal posts get kicked further down the road until they are reaching relativistic speeds. It starts around "well, the AI hasn't solved a novel problem" then moves to "well, they didn't write the validator" and suddenly humans are at the point of saying "Well AI hasn't rewrote the constants of the universe, what good are they". Of course another way to look at this is, the people that wrote the validator got praise for that years ago. Now and up and coming actor is solving problems that took us 100s of years to create in insanely short time periods so of course it's going to get a lot of attention as it well should. | | |
| ▲ | dgellow 4 hours ago | parent [-] | | To be clear: I’m aware the LLMs are solving problems. I’m just saying that what enables that whole research revolution is Lean. We wouldn’t be seeing all those results without it. I would like to see it acknowledged when people are talking about LLMs solving maths. The same way I think we should acknowledge the humans who are guiding and prompting the LLMs. I don’t think that necessitates to move a goal post |
|
| |
| ▲ | gwerbin 4 hours ago | parent | prev | next [-] | | I don't think so. People have been trying things like this with evolutionary algorithms for a very long time already. LLMs can interleave symbolic manipulation with empirical experiments and simulations and charts and thinking/reasoning text, and an LLM will much more efficiently search the space of candidate ideas than any handcrafted mutation algorithm. Any task with a cheaply verifiable goal that requires fanning out across a massive search space is ideal for contemporary LLM technology to make progress with. | |
| ▲ | 4 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | ForHackernews an hour ago | parent | prev [-] | | How long until we find out that some AI has quietly buried an exploit in Lean to cheat at proofs? |
|
|
| ▲ | ozgung 4 hours ago | parent | prev | next [-] |
| If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem. |
| |
| ▲ | pixl97 4 hours ago | parent | next [-] | | >They had to burn millions of dollars to solve a single problem I'd like to adjust that to "They had to burn a lot of energy (create a lot of entropy) to solve a single problem. As we go into the super-intelligence age the current paradigm of money as humans understand it may break at some point. For example to a paperclip-maximizer money at best is a short term instrumental goal, hard power of matter conversion machines is what it wants and once it has those money no longer has purpose. | | |
| ▲ | ForHackernews an hour ago | parent [-] | | I'd wager a fair chunk of my money that money breaks OpenAI before OpenAI breaks money. |
| |
| ▲ | 7734128 3 hours ago | parent | prev | next [-] | | They "burn" a lot when they do benchmarks, while these runs can become valid roll outs for training. Perhaps less efficient than other data creation, but hardly burned in the same way. | |
| ▲ | Razengan 2 hours ago | parent | prev | next [-] | | > were the first victims Spinning it negatively like that doesn't do anybody good. Were mathematicians the "victims" of calculators? of Matlab? Were writers the ""vIcTiMs"" of word processors?? (apparently yes, according to old TV shows about computers during the 1980s, that you can see on YouTube) > "tHiS iS nOt ThE sAmE" — Everyone every time. No, just look it up. Look into old magazines and TV shows or newspaper articles from whenever a disruptive new technology came out. | | |
| ▲ | contubernio 2 hours ago | parent | next [-] | | What you say is true but ... This is qualitatively different than calculators or computers. I'm a professional mathematician and all the better mathematicians I know are in crisis mode. Most of us hadn't taken this sufficiently seriously and don't know how to use these models effectively but we play with them and immediately see that the entire way we've worked all our professional lives has to change. We worry less about ourselves than about the younger folks. I've got good ideas ai still doesn't know about ... Younger folks may not get the chance. | |
| ▲ | azan_ 2 hours ago | parent | prev [-] | | It’s not the same. AI potentially completely replaces intellectual work without creating any* new jobs (*almost any - there will be some extra jobs for building data centers but that’s negligible). | | |
| |
| ▲ | 3 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | charcircuit 3 hours ago | parent | prev [-] | | Wouldn't that be chess players as the first victims? | | |
|
|
| ▲ | SrslyJosh 39 minutes ago | parent | prev | next [-] |
| > The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths Obviously these are unbiased and trustworthy sources. |
|
| ▲ | fweimer 2 hours ago | parent | prev | next [-] |
| The leakage wouldn't be from training, but from other uses of Personal Data. As far as I understand it, users can opt out from the training aspect, but they cannot stop their conversations (“User Content”) being used “[t]o improve and develop our Services and conduct research, for example to develop new features”. |
|
| ▲ | WD-42 2 hours ago | parent | prev | next [-] |
| If they have solved hundreds of open problems in math, why are they publishing results for the ones other mathematicians happen to be working on at the same time? Why not the others? |
| |
| ▲ | brulard an hour ago | parent [-] | | You think other mathematicians are currently working on very little subset of relatively low-hanging fruit problems? |
|
|
| ▲ | 3 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | Betelbuddy 4 hours ago | parent | prev | next [-] |
| Just use Bedrock... |
|
| ▲ | paulsutter 4 hours ago | parent | prev | next [-] |
| The big question is whether OpenAI is training on "de-identified" sessions that are marked as "do not use for training" The answer is almost certainly yes, and this is a problem for most users. |
|
| ▲ | iAMkenough 2 hours ago | parent | prev | next [-] |
| > We’ll know soon enough, but I’m inclined to believe this is true. I mean, we’ll know as soon as they decide they want to provide verifiable proof. Really dragging their feet on this front so far. I’m inclined to believe this is false. |
|
| ▲ | cyanydeez 2 hours ago | parent | prev [-] |
| The Cult tells us the AI is almight andpowerful; unfortunately, the cult cant actually describe the indescribable. |