Remix.run Logo
Ten advances in mathematics and theoretical computer science(openai.com)
173 points by milkshakes 5 hours ago | 129 comments
ultimatefan1 an hour ago | parent | next [-]

one of the early premises of how ai takeoff would go was that a system that could solve open problems in advanced mathematics would also discover novel advances in math and computer science that directly unlock drastically better software performance. we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B). we are also seeing incredible advances in software performance. open ai announced like 15% improvement by fixing gpu kernel issues. these are clearly linked in the sense of scaling laws and generalization of intelligence: a huge model gets capabilities in both math and software engineering that isn't possible at smaller scales.

but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)

aabhay 5 hours ago | parent | prev | next [-]

My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I want to know:

1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

azan_ an hour ago | parent | next [-]

> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.

einpoklum 5 hours ago | parent | prev | next [-]

Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

energy123 2 hours ago | parent | next [-]

Many less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.

traes 4 hours ago | parent | prev [-]

> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.

> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.

[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...

irthomasthomas 2 hours ago | parent [-]

Why you think that?

azan_ an hour ago | parent [-]

I guess that's because there are serious problems on which many professional mathematicians worked on years. If it was just a matter of hiring an expert, they would've been solved long time ago.

irthomasthomas 43 minutes ago | parent [-]

I guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?

dist-epoch 2 hours ago | parent | prev | next [-]

I don't think you want to bring cost into this argument.

Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.

Do you really think that if you paid that to humans, they will deliver the same results?

uh_uh an hour ago | parent | next [-]

It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.

mungaihaha an hour ago | parent | prev [-]

Grad students on zero pay solve problems like this everyday. What exactly is your point here?

mirzap an hour ago | parent [-]

Even if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.

simianwords 4 hours ago | parent | prev [-]

There are people who can’t grasp the universe without mandatory randomised controlled trial. Would tomorrow be a Sunday? Need an RCT for that boys!

My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.

nxpnsv 3 hours ago | parent | next [-]

No, this is valid criticism. Oai gives the impression anybody could get similar results at a similar price, but that’s very likely not true. This is marketing first, then mathematics.

traes 4 hours ago | parent | prev [-]

It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.

simianwords 4 hours ago | parent [-]

Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.

esperent 4 hours ago | parent [-]

> deliberate misleading “lack of transparency” etc

It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.

dist-epoch 2 hours ago | parent [-]

The results speak for themselves.

Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"

esperent an hour ago | parent [-]

Nobody is claiming the results are false.

We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.

This goes double since it's an internal secret model (Astra) so nobody else can verify the results.

simianwords 9 minutes ago | parent [-]

Would this be your reaction if OpenAI also solved millennium problems? The point we are trying to make is that the significance of this news is much larger than the skepticism you are providing.

robinhouston 3 hours ago | parent | prev | next [-]

In a way the most remarkable thing about this is that it isn't even at the top of the HN homepage. Even if this is a step up from what we've seen before, we're no longer astonished by the idea that AI can make significant advances in mathematics and computer science.

antirez 2 hours ago | parent | next [-]

This is not at the top as it is actively flagged by people that can't psychologically cope with the advances of AI. Hacker News is no longer a web site of an elite.

fg137 2 hours ago | parent | next [-]

Didn't know I was part of an elite.

pistoriusp 2 hours ago | parent | prev | next [-]

Interesting. I had no idea that a person could see what is flagged?

defrost 2 hours ago | parent [-]

If you page through the /newest listings you can see [flagged] and [flagged][dead] submissions.

eg. this: [flagged] A migrant surge tests Spain's open policies (economist.com) - https://news.ycombinator.com/item?id=49131860

is clearly marked as flagged.

Unlike the current submission: Ten advances in mathematics and theoretical computer science (openai.com) which isn't [flagged].

* https://news.ycombinator.com/newest

robinhouston an hour ago | parent [-]

That's true, but submissions are only killed in that way if they receive a ‘fatal’ number of flags. However, flags lower the rank of a story even at non-fatal levels. What antirez is suggesting here is that the rank of this story has been lowered by flags – and that seems plausible, if you compare its rank to that of other stories with a similar age and number of points.

defrost an hour ago | parent [-]

[flagged] submissions aren't [dead] (killed), they are still active and can be upvoted and commented upon.

> if you compare its rank to that of other stories with a similar age and number of points.

Ranking is complicated enough here even before weighting, speed of initial upvotes can play against ranking, number of comments and the shape of the comment tree also affect ranking. And yes, various subjects and submission sources do get weightings that impact ranking.

What's funny, to myself at least, is that any attention at all is paid to "HN front page ranking" - I've been on again off again active here since 2008 .. and can't recall ever really looking at a default HN "front page" ever.

( There's /newest /newcomments /active etc to browse and sites such as https://hckrnews.com/ )

revetkn an hour ago | parent | prev [-]

Couldn't have said it better myself.

curt15 an hour ago | parent | prev | next [-]

What about AI research itself? Is OpenAI close to automating its human staff out of a job?

yewenjie an hour ago | parent [-]

Yes, but they wouldn't publish that bit lest other companies steal the ideas.

schleck8 2 hours ago | parent | prev [-]

This is one of the most impactful mathematical publications in history by all accounts

I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient

No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.

Chance-Device 2 hours ago | parent | prev | next [-]

Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely.

The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.

christofosho 36 minutes ago | parent | prev | next [-]

I would love more time and money put into real-world problems by these companies. Climate, food insecurity, pollution, technology for convenience and/or to help people have a higher quality of life.

I'm sure they must do some of this type of work, right?

braneloop 27 minutes ago | parent [-]

Yes, but all of those are orders of magnitude harder than math.

amazingamazing 19 minutes ago | parent [-]

They are political problems, a computer could never solve them.

piker 5 hours ago | parent | prev | next [-]

I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.

Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.

[edit: deleted a distracting comparison to Chess]

energy123 4 hours ago | parent | next [-]

The old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.

traes 4 hours ago | parent [-]

If accomplishments can't be distinguished between talented people and untalented people with compute, is there really a point in trying? I suppose one can hope that talented people given compute will be more effective than untalented people with compute, but I despair that that may not be true for much longer.

aabhay 5 hours ago | parent | prev | next [-]

Given that we were nowhere near this state even two years ago, I think it’s a question of velocity more so than just distance.

traes 5 hours ago | parent | prev | next [-]

Every time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.

anematode 4 hours ago | parent | next [-]

Fully agreed. As someone who both loves chess and works on chess engines... these comparisons to chess needs to stop.

energy123 3 hours ago | parent | prev | next [-]

The distinction is mathematician vs mathematics. Mathematics is going to reach new heights beyond the wildest dreams of contemporary mathematicians. But perhaps without the participation of many paid mathematicians.

ratmice 4 hours ago | parent | prev [-]

Another noteworthy difference is that Stockfish is also gpl.

traes 4 hours ago | parent [-]

If there was any real money in it Stockfish would not be the best chess engine.

ratmice 3 hours ago | parent [-]

Thats not the point, if there were a better proprietary engine stockfish would still be there as a baseline. Anyone can access an engine as good as stockfish to practice against. Are any open models touting mathematical breakthroughs?

traes 3 hours ago | parent [-]

There is money in this, so of course the closed models are far ahead. The open models will likely catch up a bit at some point, just as Stockfish caught up to AlphaZero. That being said, there are already a couple. It seems Deepseek has a claimed proof to the "Ziegler's Cross-Polytope Conjecture" [0], but I can't speak to the significance of the result.

[0] https://arxiv.org/abs/2606.31640

baq 4 hours ago | parent | prev | next [-]

As in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels.

The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.

traes 4 hours ago | parent [-]

A fundamental difference being that no one was actually paid to find good moves in chess and go like they are to solve math problems and write code. You're comparing the digital camera and the automobile.

jibal 4 hours ago | parent | prev [-]

The chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating.

If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).

P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.

piker 4 hours ago | parent [-]

I’ve deleted it but no it’s not awful anymore than saying “we survived WWII, we can survive this.” The point was that change happens but humans find a way forward.

lifeisstillgood 4 hours ago | parent | prev | next [-]

On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.

Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.

Davidzheng an hour ago | parent | next [-]

RL training can use all of them - idk what needed means.

lwansbrough 3 hours ago | parent | prev | next [-]

For OpenAI, research is marketing. I’m sure they’ve got plenty of budget for that.

traes 4 hours ago | parent | prev [-]

Presumably it's a rounding error compared to their full output, and they're making sure they have enough compute set aside for research by limiting public models. The more datacenters they build the less they have to limit them.

pyentropy 44 minutes ago | parent | prev | next [-]

If they are targeting arithmetic circuit complexity, it doesn't take a genius to predict what their long term aim is... A permanent is much more than a determinant with plus signs...

zkmon 5 hours ago | parent | prev | next [-]

> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.

AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.

Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

raincole 4 hours ago | parent | next [-]

Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... [0]

A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

[0]: Not one of the proofs in the linked article, but from OpenAI too.

ben_w 4 hours ago | parent | next [-]

> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

I think you're over-estimating what a smarter highschooler could write.

A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:

  Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor:

  repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.

* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)

yaqubroli 4 hours ago | parent | prev | next [-]

The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.

Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.

esikich 4 hours ago | parent [-]

What gives the intention and ability to the human?

zkmon 4 hours ago | parent | prev | next [-]

When you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?

raincole 4 hours ago | parent [-]

Read the prompts in the PDF I link and see if your analogy makes sense in this context :)

zkmon 4 hours ago | parent [-]

Prompt quality should not matter. If a high-schooler operates the crane to lift a ton of weight 10 floors high, should the credit entirely go to the crane?

Anon1096 2 hours ago | parent [-]

When I type 56789*23456 into my calculator and get the result I don't claim to have solved the problem, the calculator did it.

ipnon 3 hours ago | parent | prev | next [-]

But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.

raincole 3 hours ago | parent [-]

If there aren't thousands of TPUs doing that [0] right now I'd be quite surprised.

[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".

mathisfun123 4 hours ago | parent | prev [-]

I don't disagree with you but there's no need for exaggeration; ain't no high school student writing this:

> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.

which is infact a very important part of the prompt.

don_esteban 2 hours ago | parent [-]

the fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)

esikich 4 hours ago | parent | prev | next [-]

Your brain also is physical. Electrochemical gradients flow between physical molecular constructs. Isn't it just chemistry? Do you attribute it to physics or some whole-is-greater-than-the-parts idea?

NitpickLawyer 5 hours ago | parent | prev | next [-]

A better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.

cure_42 4 hours ago | parent [-]

I'd say the creator is the one who created the 3d model, not the one who pushed the print button.

dgellow 4 hours ago | parent | next [-]

I would say „I made this gadget with my 3d printer, but the designer is someone else (I found the model online)“. The intent, the drive, the action comes from the human

traes 4 hours ago | parent [-]

"I made this proof myself, but the designer is someone else" is an extremely unconvincing claim to ownership.

dgellow 3 hours ago | parent [-]

Almost as if a proof isn’t the same as a 3d print. It’s just not a good analogy

NitpickLawyer 4 hours ago | parent | prev [-]

(let's assume that)My 3dprinter is special. It has a bunch of values + an algorithm (i.e. a neural network) that takes input as tokens and outputs a printed object.

ben_w 4 hours ago | parent | prev | next [-]

> Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.

I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".

naasking 4 hours ago | parent | prev [-]

> AI has no self-awareness

What is your mechanistic model of self awareness that yields this conclusion?

> It's a tool

Does your model suggest that tools can't have self awareness?

Delk 3 hours ago | parent | next [-]

I honestly don't think a language model is enough for self-awareness, regardless of the exact model of awareness.

A language model (or an image model or whatever) cannot even be sentient, and I think sentience is a prerequisite for awareness.

Even if we express a lot of our subjective experience with words, the language is just a symbolic representation of those experiences. The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.

You can't have an understanding of what hunger or physical pain feel like if you have no need for food or a sensory capacity for feeling pain. You can't understand what loneliness or pride at an achievement mean if you don't have a neural network wired to value social connection or status. We value connection because we're social animals that have needed each other for survival.

Even the more abstract of our subjective experiences are in some way rooted in our physical evolution.

I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.

AI awareness might actually be more believable if that awareness manifested itself in an entirely different way than in humans. But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience, I think we're seeing something that isn't actually there.

perching_aix 4 hours ago | parent | prev [-]

Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to, and influences future choices. This implies statefulness, which models are intentionally not at inference time (*).

(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").

You'll sometimes catch models mixing up who's who and how many who-s there even are for example.

(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.

I do wonder how reasonable it is to expect e.g. a single maintained identity though. Maybe it isn't?

(*) Another way to hack around this of course is to just precompute some internal "self-awareness states" and hop around between them. Probably the closest to what the models are actually "doing".

ben_w 3 hours ago | parent [-]

Before reading, know that I am uncertain in either direction.

> a hidden representation of self that is continually tended to

This sounds like a personality? They act like they have one of those. It may be an illusion, and even if it isn't an illusion it is unlikely to be anything like the source (us), but they act like it.

> I further fail to identify how it could be hidden or maintained, considering I control like half of it.

Indeed you control everything about a local model, and much of the context of even a remote model. But the state of activations and circuits in SotA AI is hidden in similar ways to those of synapses in your head: difficult to decipher even with probes monitoring the signals directly, and often not emitted at the normal output.

> The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").

While we can be confident that LLMs make up personas etc., it is insufficient to go from "that doesn't necessarily represent the model's internal state" to "therefore it doesn't have one".

> You'll sometimes catch models mixing up who's who and how many who-s there even are for example.

I've, unfortunately, also experienced this with humans. Perhaps they were losing their self-awareness at the time? I do wonder if old-age dementia does that by the end, though the person in question didn't ever get diagnosed with that.

> If you know of anything like this, your turn now, would be happy to learn.

Do you mean like these, or something else?

https://researchportal.hkust.edu.hk/en/publications/decoding...

https://aclanthology.org/2026.eacl-long.165/

https://transformer-circuits.pub/2026/emotions/index.html

perching_aix an hour ago | parent [-]

> This sounds like a personality?

Not quite what I meant, but it's also not entirely unrelated I guess? Personality to me is like a natural bias. It does also shift over time, and is also an internal bit of state. I guess in some respects it can also be self-referential, like personal convictions.

> Perhaps they were losing their self-awareness at the time?

I do think it is entirely possible for people's self-awareness to shift, yes. Or more precisely, I do model things that way.

> Do you mean like these, or something else?

They're adjacent, but I more meant something like these:

https://arxiv.org/abs/2410.03768

https://arxiv.org/abs/2310.18512

https://arxiv.org/abs/2605.26537

So basically, steganography. The difference is that these papers investigate from the perspective of separate LLM instances covertly exchanging information between each other. This is in contrast with the scenario I'm laying out, where an LLM's past state is exchanging information with its future state, continuously representing and modulating a concealed internal state of some sort. And then that state just so happening to be some sort of self-referential meta state.

And the best inkling I have towards this is basically: https://www.youtube.com/shorts/WP5_XJY_P0Q

But then I don't think there's enough covert channel bandwidth in the agent replies for anything interesting like this.

artninja1988 an hour ago | parent | prev | next [-]

Now that we've seen AI produce a fair number of proofs (and disproofs), I'm curious when we'll start seeing it build genuinely novel theory. Does anyone have predictions on when and how we'll get there and will it take new architectures/ training paradigms, or is the current approach enough?

laichzeit0 12 minutes ago | parent | next [-]

I’m personally hoping for the next big AI gangbanger to be theoretical physics. Boy does that field need a good reshuffle. I think when any novel mathematical theory can be done by AI you’ll see simultaneously theoretical physics getting wrecked as hard as pure math is. At that point we might see new physics or paradigm shifting technology emerging.

Davidzheng an hour ago | parent | prev [-]

There's no clean line between a collection of theorems and a theory.

artninja1988 an hour ago | parent [-]

I mean doing something like Grothendieck when he redeemed algebraic geometry or Galois when he invented group theory. We haven't seen that at all from LLMs.

avaer 4 hours ago | parent | prev | next [-]

What happens when OpenAI et al stop being open about these things, and just pack it into the training?

traes 4 hours ago | parent | next [-]

Not much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it.

Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.

asdewqqwer 3 hours ago | parent [-]

At this stage. No doubt calculus had plenty industrial benefit.

simianwords 4 hours ago | parent | prev [-]

What does this even mean lol. These are not solved questions. The solution never existed.

emil-lp 5 hours ago | parent | prev | next [-]

I wonder what the total cost of this research was, including the salary for their mathematicians and engineers.

kingstnap an hour ago | parent | next [-]

Why would you factor in salary unless they had to baby it through. You would only count the hours for setting up the harness and prompt and checking the result.

Training the model is going to be amortized over other uses.

emil-lp 22 minutes ago | parent [-]

> Why would you factor in salary

Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million.

Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it?

If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank.

What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent.

I'm just curious what the cost is.

traes 4 hours ago | parent | prev | next [-]

Given that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.

z7 4 hours ago | parent | prev [-]

> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices.

https://x.com/polynoamial/status/2083470822258467194

traes 4 hours ago | parent [-]

That's clearly just for the tokens, this doesn't really answer OP's question.

danielrmay 5 hours ago | parent | prev | next [-]

I'm enjoying learning about these hard problems, but this line about credit made me chuckle:

> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness

Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?

DroneBetter 4 hours ago | parent | next [-]

well, a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle, see https://infosec.exchange/@0xabad1dea/117002106099986943 and https://lipn.info/@mevenlennonbertrand/116997917683191056

traes 4 hours ago | parent | next [-]

That seems to have been more of a sensationalized joke. Even your link has a disclaimer in it now. Read this chat from the researcher who did this:

https://leanprover.zulipchat.com/#narrow/channel/270676-lean...

jibal 3 hours ago | parent [-]

It's not at all a joke ... that's a severe misunderstanding of the context.

traes 3 hours ago | parent [-]

There is no evidence that I can find for the claim "a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle."

As I currently understand it, all we know is that:

- a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kernel

- he claims that LLMs were involved somehow but pointedly refuses to specify how

- he admits that he knew about the bug before publishing the counterexample to his repository.

Perhaps not a joke (although it sure seems to me like they discovered a bug and thought falsely disproving the Collatz conjecture would be a flashy way to announce it), but at best extremely sensationalized by the above description. If you have additional context I would be happy to hear it!

danielrmay 4 hours ago | parent | prev [-]

Fascinating, and arguably an illustration of why the bifurcation of responsibility is interesting in the first place.

traes 5 hours ago | parent | prev | next [-]

I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)

emil-lp 5 hours ago | parent | prev [-]

No, the correctness isn't for the "inside the Lean proofs", but for the translation of "human language math" and its formal Lean variant.

danielrmay 5 hours ago | parent [-]

I see. It still feels like a bit of an oddly solemn way of saying "this is the part we admit responsibility for"

baq 4 hours ago | parent | next [-]

It’s more than you get from free software - you get no proofs, no warranties and any responsibility of its authors are their pure good will. Reminder lean proofs are software!

emil-lp 5 hours ago | parent | prev [-]

Well, to be fair, with Lean proofs, that's the only thing there is (unless I'm missing something).

bifftastic an hour ago | parent | prev | next [-]

Any advances in theoretical physics yet? Are there any fundamental obstacles? I would have thought not, but I haven't seen anything reported.

QuesnayJr 29 minutes ago | parent [-]

The Maxwell conjecture was a conjecture in theoretical physics (though not a particularly important one)

amazingamazing 24 minutes ago | parent | prev | next [-]

Can’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.

0x5FC3 5 hours ago | parent | prev | next [-]

How much do you all think it would cost to "buy" these advances from PhDs, practicing scientists?

traes 4 hours ago | parent | next [-]

This isn't really a productive way to think about these things, IMO. It's quite possible it would take hundreds of years for any specific group of PhDs to solve them. Or one individual PhD could have the correct flash of insight and solve it in a month. There's absolutely no way to predict this, besides trying to gauge the apparent simplicity of the proof or counterexample (which is likely to be misleading). Until someone actually runs an experiment like this it's not a viable metric.

0x5FC3 4 hours ago | parent [-]

I understand and I am not trying to deny the impressiveness or the velocity of AI in general. But at some point we have to ask how much do we trust the labs at face value without much transparency of how they got to the results when there is trillions of dollars on the line.

simianwords 4 hours ago | parent [-]

The level of conspiracy theory is nuts

0x5FC3 4 hours ago | parent [-]

I would say the lack of skepticism is nuts, honestly.

frozenseven 4 hours ago | parent [-]

Capabilities of this sort have already been demonstrated by independent parties, and models have consistently gotten better at this. Yes, insinuating that mathematicians and scientists are secretly solving decades-old problems on OpenAI's behalf is an insane conspiracy theory.

jgeralnik 3 hours ago | parent | prev [-]

A friend’s PhD advisor has been chasing non-sofic groups for 25 years (and was shown a preprint of the results by openai to verify them). He believed a solution would be Fields-worthy

This was not a problem that was for sale

DrBazza 2 hours ago | parent | prev | next [-]

Replace philosophers for mathematicians and Douglas Adams was spot on again.

Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.

--

"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"

"What's the problem?" said Lunkwill.

"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"

"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"

"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"

kingstnap 2 hours ago | parent | prev | next [-]

It's remarkable how you can manage to get these models to produce remarkable breakthroughs like an explicit construction of a non-sofic group.

And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.

Truly jagged beyond belief.

s_Hogg 4 hours ago | parent | prev | next [-]

I don't know why, but when I saw the source of this particular headline it reminded me of the album title 26 Mixes for Cash

defrost 4 hours ago | parent [-]

Ambient 0: Math for Airports

melagonster 2 hours ago | parent | prev | next [-]

Wow, so this is the end of science :(

xyzsparetimexyz 2 hours ago | parent [-]

It's just another tool that can help solve problems. It doesn't know _what_ problems to solve. It turns out that a lot of old problems are now low hanging fruit for these new models. In terms of 'expanding the frontier', we've just discovered dynamite and can now blast our way through mountains. The bottom of the ocean or space are still as hard to reach as ever.

readthenotes1 3 hours ago | parent | prev | next [-]

I wonder if Erdos would be saying " It's fine that y'all are answering my questions, but who is asking better questions??"

luciana1u 4 hours ago | parent | prev | next [-]

the real milestone isn't that AI solved ten math problems, it's that we now need a press release to tell us which ten problems count as important

baq 4 hours ago | parent [-]

I asked ChatGPT and it told me these aren’t not important /s

xyzsparetimexyz 2 hours ago | parent | prev [-]

Any implication of any of these findings? They seem like unimportant nerd snipes to me. If you want to do something actually relevant, get chatgpt to write a simulation of graphene nanotube construction and figure out how to do it at scale.

utopiah an hour ago | parent [-]

Very marketable nerd snipes indeed.

foobar10000 8 minutes ago | parent [-]

One - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to explore the adjacent fields, etc.

The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.