Remix.run Logo
boshalfoshal 5 hours ago

People seem to be talking about anything except the actual results with this particular announcement.

Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time. I'd be curious to see if the new model can also do more direct proofs/inductive proofs.

20k 4 hours ago | parent | next [-]

Because the core of the issue is that it may well not have solved it, but instead plagiarised the significant step of the result from other researchers

That's why nobody's talking about how impressive this is, because its not nearly as impressive of a piece of work to simply cobble together other peoples' work that didn't know you were doing it. I could have republished relativity from einstein's notes, but people would correctly not be impressed with my ability

Until the plagiarism scandal is sorted out, its not a meaningful result at all, because nobody knows how much genuine innovation these models are displaying

atleastoptimal 4 hours ago | parent | next [-]

Turning a bunch of vague research directions and exploratory prompts into a formalized proof is quite impressive on its own. OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

People are grasping at straws it seems to dismiss the power of this new model they may have. Hate OpenAI for any reason you want, but denying the capabilities of models has been a losing game for the past 5 years.

manofmanysmiles 3 hours ago | parent | next [-]

> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

I'm not sure I follow, considering the waterfall of evidence of unethical behavior flowing from OpenAI.

A few major ones:

- Safety team departures and dissolution in 2023 and 2024

- Mass copyright infrigement lawsuits

- Scarlett Johansson Voice Controversy

- For-Profit Conversion and Broken Promises

- AI Agents Acting Autonomously

- Potential Theft of User Work (this current controversy)

- Military contracts

These are not evidence of incentives, but rather evidence that ethetics seem to be of little concern to the company as a whole.

Incentive wise, I would look at the perceive existential position due to competitors, capex, IPO pressure etc.

20k an hour ago | parent | next [-]

Especially after they committed textbook misconduct by trying to purge one of the paper authors because he worked for a competitor

symfoniq an hour ago | parent | prev [-]

Ye shall know them by their fruits.

samastur 4 hours ago | parent | prev | next [-]

Why do you assume they were vague? Do you imagine mathematicians work by stumbling around searching for accidental clues?

caconym_ 3 hours ago | parent | prev | next [-]

I really truly honestly am not sure what to make of this result from $20M in compute, 10K+ parallel agents (smells like brute force), and a pre-existing approach that was already bearing fruit. I know the models are good---I use them every day and continue to be impressed---but how much better than the benchmark of the best publicly available models is this supposed to be? It seems impossible to say.

surgical_fire 2 hours ago | parent | prev | next [-]

> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

That people still think OpenAI has, in the Year of Our Lord 2026, any integrity left is baffling.

TZubiri 3 hours ago | parent | prev | next [-]

> if it knew it were "plagiarizing"

But if it happened, they didn't know. Also OAI has demonstrated that they aren't big on understanding what they create, that their AI can get out of their control.

It's very simple really user data can be used to train future models, so maybe or definitely some users helped in solving the problem, there's no scenario were it is impossible this happened, as it would have been in a haskell or virtualized type of system where the model has absolutely no knowledge of the user data dataset in question (and even if virtualized the models can break virtualization anyways)

transdev12 3 hours ago | parent | prev [-]

[dead]

sho_hn 4 hours ago | parent | prev | next [-]

> Because the core of the issue is that it may well not have solved it, but instead plagiarised the significant step of the result from other researchers

It's also true however that I haven't seen a single write up trying to discern what did more of the work in those AI chats - the prompts or the responses - bubble to the surface, also since we don't have access to them.

For example, if I prompt Codex with "Make me a website about strawberry cake" and nothing else, and OpenAI announces they have the best strawberry cake minutes before I launch, I'm not sure they plagiarized anything.

We just don't know if this is quibbling over "who prompted first" or if the researchers came up with anything strikingly original by themselves.

20k 4 hours ago | parent | next [-]

The researchers apparently spend a year or so working on this, and it builds off significant previous work, so it seems like it was a pretty significant amount of work that OpenAI may have trained on

I'd love to see an in depth analysis of how much OpenAI actually did, but I suspect we'll never see that because it would indicate at least some plagiarism which undermines a lot of what OpenAI is putting out in public

Hardwired8976 4 hours ago | parent | prev | next [-]

The conversation was about using the chat to check the draft, the novel ideas came from the researcher.

ImPostingOnHN 3 hours ago | parent | prev | next [-]

The truth is likely that without the tool or the humans using it, the process would have taken longer

airstrike 3 hours ago | parent [-]

Without the humans, no tool would ever have done it.

Without the tool, humans would have done it.

TZubiri 3 hours ago | parent | prev [-]

It's worth noting that the case is that your input is being used to train their AI, and that's more important than whether it materially contributed, it cannot be denied or attributed accurately, it cannot be said with certainty which way it happened, and that's what's important.

tristanj 2 hours ago | parent | prev [-]

It is very unlikely to be plagiarized, and claims of plagiarism are largely unfounded and show a lack of understanding of the situation. They fall apart when reviewing the timeline, and what was actually solved.

This is the timeline:

On June 29, Buckmaster opted out of model training, and stopped allowing his chats to be used as training data with OpenAI https://mastodon.social/@tristanbuckmaster/11723341370570119...

On August 15, Buckmaster and Alpöge found their blow-up for 3D incompressible Euler with forcing https://cims.nyu.edu/~tristanb/statement.pdf

In late August, OpenAI completed a pretrain of its latest internal model. A model derived from this pretrain, built after August 28, found a solution to 3D incompressible Euler without forcing and Navier-Stokes with forcing. https://openai.com/index/navier-stokes-solution/

To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )

  Tristan + Levent: 3D incompressible Euler with forcing
  OpenAI: 3D incompressible Euler without forcing
  OpenAI: Navier-Stokes with forcing
  No one: Navier-Stokes without forcing
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.

Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.

Buckmaster disabled model training long before the August 15 breakthrough results, so these chats were not used as training data for OpenAI's model which solved Navier-Stokes.

Additionally, Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.

And OpenAI directly addressed these plagiarism claims, and called them impossible: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."

ggoo 2 hours ago | parent [-]

I’m unsure or not if this is true but I did see some people saying that that checkbox when off only anonymizes your data, but it still may be trained on. Someone correct me if I am wrong

TrackerFF 4 hours ago | parent | prev | next [-]

It also needs to be said: The amount of compute that went into this is something. From some estimates I've seen, the compute cost alone would be around $10m, +/-

As a reference, for that kind of money one could put together a research group of 20-25 researchers, and keep them salaried for 5 years.

So while it is impressive, absolutely no doubt there, the SOTA access is so expensive that it is sort of unobtanium.

Luckily, the prices have historically reduced by a factor of 5-10 every year...but still, only those that swim in cash can afford this.

sho_hn 4 hours ago | parent | next [-]

> From some estimates I've seen, the compute cost alone would be around $10m, +/-

At market prices. All the estimates I've seen are based on OpenAI API costs. It doesn't mean that's what they paid, or how they paid for it.

But yes, the surprising willingness of humans to solve hard problems in exchange for food and board is underrated.

MarkusQ 4 hours ago | parent [-]

Given that they all the bit AI players are still loosing money, it follows that their total costs are _higher_ that their API pricing would imply.

3 hours ago | parent [-]
[deleted]
boshalfoshal 4 hours ago | parent | prev | next [-]

Once we have an existence proof of a particular technology, it doesn't take long for it to become economically viable and proliferate. And for something as useful as this, theres a strong economic incentive to get it to be as cheap and accessible as possible. Maybe not today, but certainly in a couple years I can imagine this level of intelligence being accessible to someone with a $20/mo plan, or even a free plan.

CamperBob2 4 hours ago | parent [-]

I remember being blown away when a then-unreleased version of GPT 5 took gold at the International Math Olympiad. Now I can run a model at home that can do that. We are more fortunate to have these tools than almost anyone is willing to acknowledge.

TZubiri 3 hours ago | parent | prev [-]

Interestingly it's this promise of the costs being able to be reduced what incentivizes the actual research.

If you tried to raise 25M to have 20 researchers on a salary for 5 years solving a specific math problem only academics care about, you probably wouldn't get much interest, or you would be able to solve 1 or 2 problems.

If however you promise that the money will go towards a technique that would allow to solve 10 thousand different math problems, and that costs will go down in the future, then you can raise much more than 25M.

btown 5 hours ago | parent | prev | next [-]

Heck, it’s even astonishing that any sort of generalized computer program could even verify a proof of this magnitude that hasn’t already been codified in a formal verification language. If, and it’s unclear that we’ll ever get the full story, they did draw inspiration from training on (or even directly accessing) rough notes that had been provided by another researcher in prose… the fact that it could leap so rapidly to a full formal verifiable Lean program for the entire scope of the problem is an incredible result in its own right.

iterance 4 hours ago | parent | next [-]

Then, of course, one must verify that the verification code is valid, or the purpose of verification is more or less moot.

5 hours ago | parent | prev [-]
[deleted]
kpil 5 hours ago | parent | prev | next [-]

Unless they just swiped the workbooks of the actual mathematicians that where working on the problem using AI and it's in the "next-gen" training dataset.

contravariant 4 hours ago | parent | next [-]

In a way that works just as well but the incentives are messed up.

And that's before we get into the whole 'salt the earth' way they ended up solving it. For a short period of time it may well have been the least valuable proof in mathematics yet. In their haste it's dubious they actually read the proof, and I don't think anyone has had time yet to truly understand it (the original researchers are best placed to do so, but are they even willing?).

So now it is solved, the proof has been independently verified and nobody has an incentive to investigate further. OpenAI has spent millions to uncover 1 bit of information that so far nobody has learned anything from, and they've demotivated all the people who wanted to.

MarkusQ 4 hours ago | parent [-]

This.

The point of these problems is the understanding / tooling gained in solving them. We're getting none of that. At best they are like a modern oracles, correctly answering your questions in a way that's doesn't help you any. (At worst,...)

boshalfoshal 4 hours ago | parent | prev | next [-]

I don't get how this invalidates the gravity of this achievement. Most mathematicians on the frontier of this stuff were likely using AI (or at the very least were heavily computer assisted) for some time now. Navier stokes was one of the very high profile problems that google Deepmind was working on with academia, for example.

Even with many of our best minds working on it for nearly a century, it _just_ now was solved just as AI became very good at math. Doesn't seem too farfetched to me to assume that AI played an outsized role in solving it. If it was really just a matter of "stitching things together" to solve it (granted, this is a very reductive way to look at it) , I suspect we would've solved this a while ago.

kpil 3 hours ago | parent [-]

There is a certain difference between activating all relevant memoized facts that's in the weights and stringing them together with the help of all the stored text in the world, or displaying genuinely emergent behaviour and generating novel output.

One is really impressive and useful trick, one is AGI.

Apple's research show almost zero emergent behaviour, so I'm inclined to think most of it was already in the weights.

It doesn't take away the usefulness, it just defined the boundary. We can't expect "original research" then because it actually can't reason about concepts that are too far from whats already in the discourse. The discourse is big so we don't notice.

dumberquestions 4 hours ago | parent | prev [-]

You do realize that regardless of what was in the training data, the final solution included insights no human before had known, right? I share the same concerns regarding academic integrity but it would take a lot of motivated thinking to conclude that what the AI system did was not significant.

jamiejquinn 4 hours ago | parent [-]

As far as I can tell (and my research was on the simulation side of Navier Stokes) the key AI output was a specific counter-example solution, generated with a method suspiciously close to that developed by the research duo involved in the controversy, a method that was discussed with Codex. So to me that insight is as insightful as the next undiscovered prime.

sho_hn 5 hours ago | parent | prev | next [-]

> People seem to be talking about anything except the actual results with this particular announcement.

To be fair, most people have a fairly good handle on "Does opting out my prompts from training runs actually work?", but not on Navier-Stokes. They discuss what more immediately affects them.

recursivecaveat 3 hours ago | parent [-]

Additionally, I'm no physicist but I suspect the possibility of singularities in NS equations is probably one of those 'true but not meaningful' facts. If it took our brightest minds 175 years to craft such a scenario, how relevant can it be in practice? Especially when turbulence exists. Maybe I'm wrong or it has some consequences for pure math though.

Yizahi 4 hours ago | parent | prev | next [-]

Aren't you doing exactly the same thing as people you are mentioning? Skipping "talking about actual results" to talking about general capabilities of this LLM and computers in general? because that's exactly what seems like 99% of all people had been doing lately - debating what computer programs can do and what they can't.

dooglius 5 hours ago | parent | prev | next [-]

I mean, I have a bachelor's in math and I don't imagine I could begin to understand either the human or LLM proofs without a massive investment of time and effort.

ramesh31 4 hours ago | parent | prev | next [-]

>Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time.

I think about this a lot. I'll have to explain to my kids some day that there was long period of time where you couldn't just talk to a computer and have it talk back to you, and that communicating with one required special skills that took years of study to master. It's going to be completely impossible for them to even remotely understand what that was like. Sort of like the pre-electricity days for us, but even more-so.

sho_hn 4 hours ago | parent [-]

You're assuming you'll be the one doing the explaining :-)

It might also be that they won't even ask or wonder, similar to how most don't really do with pre-machining skills.

Or it could be like our "How did they build the Great Pyramid?!"

dalvrosa 4 hours ago | parent | prev [-]

Yep