Remix.run Logo
arctic-true 8 hours ago

Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.

chilmers 7 hours ago | parent | next [-]

The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.

[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/

noir_lord 2 hours ago | parent | next [-]

Recursive self improvement of their upcoming IPO value maybe.

They are fluffy PR pieces otherwise.

piloto_ciego an hour ago | parent [-]

How can you possible say this sort of thing in context of what looks like a millenium prize being solved.

I swear there's nobody blinder than those who won't see.

samrus 28 minutes ago | parent | next [-]

You have to look at the incentives

42 minutes ago | parent | prev [-]
[deleted]
10xDev 7 hours ago | parent | prev [-]

Compute will always be the bottleneck even if this were true.

hgoel 5 hours ago | parent | next [-]

As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.

In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).

Fordec 6 hours ago | parent | prev | next [-]

If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.

mrbungie 6 hours ago | parent | next [-]

We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.

Fordec 5 hours ago | parent [-]

For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.

HenrikPontoppid 6 hours ago | parent | prev | next [-]

Yes. In other words: the singularity. I'll only believe it when I see it though.

Fordec 6 hours ago | parent [-]

I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.

piloto_ciego 43 minutes ago | parent | next [-]

I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.

But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.

From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"

The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.

supern0va 5 hours ago | parent | prev [-]

It doesn't imply that. The singularity is just the inflection point.

jsLavaGoat 4 hours ago | parent | next [-]

Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.

And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.

Fordec 4 hours ago | parent | prev [-]

Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.

dakolli 6 hours ago | parent | prev [-]

How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.

Fordec 6 hours ago | parent | next [-]

Hello good sir from the 1700s pre-industrial revolution who doesn't think that mines and factories can be automated.

dakolli 2 hours ago | parent [-]

The factories that supply the equipment, maintain the equipment, the energy inputs, the financials of those mines are not automated.

People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.

7373737373 4 hours ago | parent | prev | next [-]

Some mines are already heavily automated: https://youtube.com/watch?v=_Z9w-mUoUsY

https://youtube.com/watch?v=SRuht0QIprs

a2ff6eeb0 4 hours ago | parent | prev [-]

https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.

monster_truck 5 hours ago | parent | prev | next [-]

Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.

Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged

Fordec 5 hours ago | parent [-]

I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.

monster_truck an hour ago | parent | next [-]

It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.

I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.

xtracto 5 hours ago | parent | prev [-]

pi-fs will solve all our data compression problems.

Miner49er 7 hours ago | parent | prev | next [-]

Eventually recursive self-improvement includes reducing bottlenecks.

glenstein 6 hours ago | parent | next [-]

Which is to say, scalable and open-ended capability of ramping up physical infrastructure.

I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.

10xDev 7 hours ago | parent | prev [-]

Eventually the bottleneck might be people themselves.

ccozan 4 hours ago | parent [-]

Improbably, the real bottleneck is energy.

lijok 4 hours ago | parent | prev [-]

And the goalposts move again

danielmarkbruce 5 hours ago | parent | prev | next [-]

With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.

preommr 2 hours ago | parent [-]

It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).

So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.

Second biggest fumble after Google.

danielmarkbruce 2 hours ago | parent | next [-]

Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".

sebzim4500 30 minutes ago | parent | prev [-]

Don't they own a large portion of OpenAI? Things could be worse

magicalist 7 hours ago | parent | prev | next [-]

> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.

Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?

ameliaquining 7 hours ago | parent [-]

I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.

20k 7 hours ago | parent | next [-]

It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model

What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question

If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?

Edit:

OpenAI have admitted they were training on prompts at the time they made their breakthrough

https://mastodon.social/@tristanbuckmaster/11723647135247030...

ameliaquining 7 hours ago | parent | next [-]

If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.

If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.

20k 7 hours ago | parent | next [-]

The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools

The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem

orangecat 6 hours ago | parent | next [-]

In that it isn't able to genuinely solve problems

Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.

This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.

20k 6 hours ago | parent [-]

I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach

letmevoteplease 5 hours ago | parent [-]

You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.

ivory54321 6 hours ago | parent | prev | next [-]

I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?

Timwi 4 hours ago | parent [-]

If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.

lotsofpulp 6 hours ago | parent | prev [-]

Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?

20k 6 hours ago | parent [-]

OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct

lotsofpulp 4 hours ago | parent [-]

Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.

fwip 2 hours ago | parent [-]

Terms and conditions are almost entirely about the company doing things that would otherwise be illegal.

sdenton4 4 hours ago | parent | prev [-]

Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?

Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.

doctoboggan 6 hours ago | parent | prev | next [-]

I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?

user43928 5 hours ago | parent | prev | next [-]

All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.

That's it. The rest appears to be wild speculation.

jsw97 4 hours ago | parent [-]

Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.

Never ever touch those requests. If you get a side by side comparison just resend the prompt.

zem 4 hours ago | parent | prev | next [-]

even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?

flir 4 hours ago | parent [-]

Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.

Publication, though? Slimy is right.

But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.

TZubiri 4 hours ago | parent | prev | next [-]

>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.

TZubiri 4 hours ago | parent | prev [-]

>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.

an0malous 5 hours ago | parent | prev [-]

The things you don’t care about are highly relevant to that claim

ameliaquining 4 hours ago | parent [-]

Elaborate?

pama 7 hours ago | parent | prev | next [-]

Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.

danielmarkbruce 4 hours ago | parent [-]

If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.

mzhaase 7 hours ago | parent | prev | next [-]

The singularity happening under trump? We could have had star trek, instead we're getting the combine.

monster_truck 5 hours ago | parent | next [-]

pick up that can

Bluestein 6 hours ago | parent | prev | next [-]

"I love Singularities. I am the best at Singularities. Everybody knows it ..."

ccozan 4 hours ago | parent [-]

Beautiful Singularities.

dboreham 6 hours ago | parent | prev [-]

That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.

karmakurtisaani 6 hours ago | parent | next [-]

And then it will be shut down, proper guard rails put in place, and the new version will accelerate the cleptocracy.

dakolli 6 hours ago | parent | prev [-]

You think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi

E-Reverance 6 hours ago | parent | next [-]

The scifi pov has a good track record as this point, you people gotta be more open minded

sznio 4 hours ago | parent | prev | next [-]

it proved navier-stokes taking over the us government is easier imo, any idiot gets to be president

fc417fc802 4 hours ago | parent | prev [-]

Many present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?

dakolli 2 hours ago | parent [-]

[flagged]

_fizz_buzz_ 5 hours ago | parent | prev | next [-]

Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?

tristanj 5 hours ago | parent | next [-]

OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.

gcr 5 hours ago | parent | prev | next [-]

it's possible to do a RLHF or RLVR pass pretty quickly. I'm almost certain a full pretraining run isn't possible within that time frame.

lossolo 4 hours ago | parent | prev [-]

Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.

piloto_ciego an hour ago | parent | prev | next [-]

And... they found this "solution" in 88 hours or so.

It's all gas no brakes now boys and girls. Hold on to your hats!

naveen99 8 hours ago | parent | prev | next [-]

Astra was trained more than two weeks ago.

sashank_1509 7 hours ago | parent | next [-]

Astra was in use by OpenAI employees for more than 3 months internally from rumors I heard

credit_guy 7 hours ago | parent | prev [-]

The internal model they mention is different from Astra.

curt15 7 hours ago | parent | prev | next [-]

They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.

Aboutplants 7 hours ago | parent | prev | next [-]

I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works

stingrae 7 hours ago | parent [-]

the model is a set of weights, you can take a snapshot and test it. Reinforcement learning itself is largely testing and tuning.

blake__dev 7 hours ago | parent | prev | next [-]

Yeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch

cool_dude85 7 hours ago | parent [-]

The chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.

sebzim4500 34 minutes ago | parent | next [-]

Isn't the y axis just what portion of the open problems it could solve? The axis is unlabelled though, I'll give you that

merksittich 5 hours ago | parent | prev | next [-]

The x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?

blake__dev 7 hours ago | parent | prev [-]

That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.

bananaflag 7 hours ago | parent | prev | next [-]

Yeah it's Bel

refulgentis 6 hours ago | parent | prev | next [-]

Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.

vatsachak 7 hours ago | parent | prev | next [-]

Brain has loops and parallel connections.

Loops and parallel connections make transformer go brrr

irthomasthomas 4 hours ago | parent | prev | next [-]

Or they trained a LoRA on the victims chats in order to launder their plagiarism.

fer 4 hours ago | parent [-]

The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.

chinathrow 8 hours ago | parent | prev [-]

Pre-IPO marketing?

Aboutplants 7 hours ago | parent | next [-]

Even if it is, Anthropic better have a few things up their sleeve

jrflo 8 hours ago | parent | prev | next [-]

I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?

mrbungie 7 hours ago | parent | next [-]

They are highly capable, no doubt about that, but:

1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.

2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.

dsdf3 7 hours ago | parent | next [-]

"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."

Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.

And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.

scurnus 7 hours ago | parent | prev [-]

1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars. 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.

mrbungie 7 hours ago | parent [-]

> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.

Did I say otherwise?

> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

I know, but I don't know how that relates to my point, which is about the way they are doing it.

scurnus 6 hours ago | parent [-]

Sorry, I misinterpreted point 1), on X they said they didn't have people specialized in that specific field for prompting and steering the agents, just a group of mathematicians and physicists.

The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.

danielmarkbruce 4 hours ago | parent | prev | next [-]

Highly capable of writing math proofs, no doubt.

It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).

QuesnayJr 7 hours ago | parent | prev | next [-]

Of the seven Millenium problems, Navier-Stokes was the one most thought to be in reach.

I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)

anthonypasq 7 hours ago | parent | next [-]

the goalposts are on Pluto at this point.

dsdf3 7 hours ago | parent | next [-]

I'd put good money on the fact that we will have a lot of distilled intelligence and yet the world won't look much different.

anthonypasq 7 hours ago | parent [-]

i mean that is already true

QuesnayJr 7 hours ago | parent | prev [-]

I'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample).

I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.

ameliaquining 7 hours ago | parent | prev [-]

There were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathemati...

QuesnayJr 4 hours ago | parent [-]

I brought this up here at HN, and in the ensuing discussion Buzzard himself replied saying he was somewhat joking (https://news.ycombinator.com/item?id=49011950).

ameliaquining 4 hours ago | parent [-]

Certainly, but the key word there is "somewhat". Progress is now happening so incredibly fast that I no longer know what to consider implausible.

andrepd 6 hours ago | parent | prev [-]

Lmao my friend, the whole "drama" is that there are allegations of plagiarism.

eutropia 7 hours ago | parent | prev [-]

If pre-ipo marketing pushes them to train a model capable of resolving a millennium problem in mathematics in a weekend, then, to quote XKCD:

  "Mission. Fucking. Acccomplished."

https://xkcd.com/810/