Remix.run Logo
▲ antman 6 hours ago

This argument implicitly makes a few assumptions which will probably not hold in the very near future.

One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.

The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse

▲Vetch 4 hours ago | parent | next [-]

Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.

This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.

Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.

Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.

Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.

The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.

▲p-e-w 3 hours ago | parent [-]

> Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally.

This is a feature, and a huge step forward.

If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.

▲dofm 5 hours ago | parent | prev | next [-]

> One is that AI will continue hallucinating in a manner that is not easy to verify

It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.

Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

▲user43928 4 hours ago | parent | next [-]

And why not?

For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.

Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.

Why would this not also be the case for mathematics?

▲catlifeonmars an hour ago | parent [-]

I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.

▲bonoboTP 31 minutes ago | parent [-]

That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.

▲user43928 24 minutes ago | parent [-]

To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.

"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.

But is running a malformed command that does not achieve the expected effect itself a hallucination?

▲antonvs 5 hours ago | parent | prev [-]

> It is an old saw at this point

An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.

▲Terr_ 6 hours ago | parent | prev | next [-]

> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify

Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.

At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.

I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.

▲user43928 4 hours ago | parent | next [-]

Hallucinations are no longer much of a practical problem in software engineering.

Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.

We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.

Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.

Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

▲ashkankiani 42 minutes ago | parent | next [-]

The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.

▲user43928 37 minutes ago | parent [-]

You're free to substantiate your comment by telling us about your apparently different experience.

▲ashkankiani 16 minutes ago | parent | next [-]

The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything.

And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.

Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...

▲jazzypants 24 minutes ago | parent | prev [-]

I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days.

This is just a fact. I'm sorry if it messes with your narrative.

https://arxiv.org/abs/2401.11817

▲user43928 20 minutes ago | parent [-]

Thanks for the link. What kind of hallucination are you seeing, and does it affect the end result?

▲catlifeonmars an hour ago | parent | prev [-]

It’s still a common occurrence.

It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.

This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).

I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.

> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.

▲hodgehog11 5 hours ago | parent | prev [-]

"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.

▲catlifeonmars an hour ago | parent | prev | next [-]

> but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds

You’re making the following assumptions:

1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)

2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

▲bonoboTP 35 minutes ago | parent [-]

> 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?

▲Retr0id 6 hours ago | parent | prev | next [-]

Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.

▲colordrops 5 hours ago | parent [-]

Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?

▲hodgehog11 5 hours ago | parent | next [-]

Yes. Most probably do not understand the notation involved in, and the statement of, the Lindeberg-Levy Central Limit Theorem. But every scientist uses this theorem in one way or another. These ideas have a way of trickling down because to people who work thanklessly to do so.

▲dofm 5 hours ago | parent | prev [-]

Something about this sentence makes me think about that Rob Auton bit, that before there were mobile phones, nobody had any reason to tell someone else that they were on a bus.

▲zer00eyz 2 hours ago | parent [-]

Before mobile phones you never called someone and asked "where are you" because a phone was tied to a location...

▲mathisfun123 6 hours ago | parent | prev | next [-]

> One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.

There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

▲bonoboTP 26 minutes ago | parent | next [-]

Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?

▲antman 5 hours ago | parent | prev | next [-]

The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing.

What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.

▲red75prime 5 hours ago | parent | prev [-]

> The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.

▲gottheUIblues 4 hours ago | parent | next [-]

I think people on here tend to somewhat fixate on the determinism issue. Even with a deterministic LLM - stabilising the floating point arithmetic, and choosing from the distribution by a fixed method, or just save the random seeds - there is still a kind of a chaotic unpredictability that can exist between its inputs and outputs. However maybe that is a price that needs to be paid to get creativity.

▲Vetch 4 hours ago | parent | prev [-]

Deterministic yes, robust deterministic no. The most likely conjunction is not always the best nor representative of what the model is considering unless its certainty is high.

▲red75prime 3 hours ago | parent [-]

I think determinism has nothing to do with it. If you mean sensitivity to word ordering and such, it's a generalization failure.

▲thereitgoes456 6 hours ago | parent | prev | next [-]

AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?

▲ApolloFortyNine a few seconds ago | parent | next [-]

Interesting claim. That's true for old models that simply have no way to explain, LLMs however can. [1]

[1] https://dev.to/natcher/researchers-develop-method-to-train-l...

▲jstummbillig 6 hours ago | parent | prev [-]

It's kind of astonishing that after all we have seen in the last years people still find the position that AI will not be able to do an obviously valuable thing likely and it requiring an explanation (instead of the other way around).

▲hodgehog11 6 hours ago | parent | next [-]

No, this is different, and this is coming from someone who has been studying deep learning for the last decade. We are talking about the difference between RLHF and RLVR strategies. The former benefits clarity and explanation, while the latter concerns only correctness. AI was moving in a particularly damaging direction by pushing on the first path, so it was natural to move to the second. But the second will come at the cost of clarity of explanation. It will likely get better at its explanations, but not fast enough to render its most advanced accomplishments readily understandable to the user. The chess example is a pretty good one (that is an RLVR approach).

▲__s 7 minutes ago | parent | next [-]

tbf GM explaining their 2700 elo moves are only understandable when vague, as elo goes up explanation becomes closer to "in this specific position there's these dpecific lines", why should 3500 elo moves have simple reasoning?

Maybe if we start with giving simple AI generated analysis of those clumsy humans with their measly 2700 elo moves

▲red75prime 5 hours ago | parent | prev [-]

The problem is that people strongly believe that this is an insurmountable problem that will persist indefinitely (or for a long time) and plan accordingly, while this, most likely, will be fixed soon by adding RLCAF (RL on conversational agent feedback) or something like that.

▲n6242 5 hours ago | parent | prev [-]

Some of us still remember 2016, when we had a couple of cars sorta half-driving themselves, and Tesla, Uber and others promised we were only one year or two away from three million people in the US working as drivers being out of a job. And here we are, a decade later. AI is pretty amazing, but companies have a tendency to severely and comically overestimate and oversell it's capabilities, and underestimate the challenges.

▲jstummbillig 4 hours ago | parent | next [-]

> And here we are, a decade later.

With Waymo and Tesla increasingly doing what they said they would do, and a small number of early adopters happily paying money for their services, that do work.

So what's the critique? That the timelines are not correct? Sure. And how about the timeline of the people who said "research level math, never in my lifetime" and the people inside the ai companies who are apparently increasingly spooked by how quick the progress is? How about the various levels of code/programming jobs that AI was supposedly never going to be able to do, but, in reality, now just does?

We are engaging in some very one-sided discrediting, and I am not sure, why.

▲jazzypants 20 minutes ago | parent [-]

And, those companies are notably avoiding wet climates because they still struggle with self-driving in inclement weather. It's probably going to be another decade before we get to the point where these things can handle every situation. Just like all other engineering, the first 90% is the easy part.

https://www.wsj.com/articles/self-driving-cars-dont-do-snow-...

▲p-e-w 3 hours ago | parent | prev [-]

Driving a car is unimaginably more difficult than proving the Riemann hypothesis.

You just don’t notice that because evolution has given you 99% of what is needed to drive a car before you were even born.

▲nullsanity 6 hours ago | parent | prev | next [-]

[dead]

▲wolvesechoes 5 hours ago | parent | prev | next [-]

Expression of tech-faith is not intellectually honest argument.

Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.

▲antman 5 hours ago | parent [-]

Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced

▲pegasus 4 hours ago | parent | prev [-]

Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.

Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...