Remix.run Logo
▲ againstapples 4 hours ago

As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

▲mapmeld 5 minutes ago | parent | next [-]

I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.

'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.

Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.

▲arctic-true 3 hours ago | parent | prev | next [-]

Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).

Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.

With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.

Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)

▲istjohn 3 hours ago | parent [-]

> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

See:

> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

▲edot 2 hours ago | parent | next [-]

Sure and if I make a half court shot after an hour of trying, the result only took 1 second.

▲somenameforme an hour ago | parent [-]

Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.

They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.

[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...

▲arctic-true 3 hours ago | parent | prev [-]

That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.

▲tehjoker 41 minutes ago | parent [-]

It’s very typical in human math that explaining the final result after years of searching looks very simple too.

▲doginasuit 3 hours ago | parent | prev | next [-]

I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.

Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.

When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.

▲mikestylz 2 hours ago | parent [-]

> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.

Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.

And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.

▲Kotlopou 2 hours ago | parent | prev | next [-]

I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

▲computably 2 hours ago | parent | prev | next [-]

Depends on your definition of doom.

If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.

If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.

I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?

▲pj_mukh 3 hours ago | parent | prev | next [-]

Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?

Or is it simply that you feel bad for Mathematicians.

▲againstapples 2 hours ago | parent | next [-]

I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.

I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.

▲lf88 an hour ago | parent | next [-]

A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.

▲frumplestlatz an hour ago | parent | prev [-]

In your imagined future, how do you imagine the AI would build, grow, improve, and operate its physical substrate independently of human intervention?

▲ndriscoll an hour ago | parent | next [-]

If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.

▲FuckButtons 41 minutes ago | parent | prev [-]

One step at a time - how reliant do you think the ai labs are likely to be on their own tools right now, today, let alone 1-5 years down the road?

▲Veedrac 3 hours ago | parent | prev | next [-]

Humans have one ecological niche. Soon we will have zero. That is worth worry.

▲postalrat 41 minutes ago | parent | next [-]

AI changes nothing for someone who believes aliens exist and may already be here on earth.

▲pj_mukh 2 hours ago | parent | prev [-]

>>ecological niche

As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?

▲whimsicalism 3 hours ago | parent | prev | next [-]

Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo

▲voiceeh 3 hours ago | parent | next [-]

So, you're worried about them breaking containment and deciding to do bad things?

▲orlp 3 hours ago | parent | next [-]

I'm more worried about them doing bad things at the behest of people who want them to do bad things.

That is 1. immediately technically possible, and 2. realistic.

If you need a source for 2 I'd suggest you open any history book.

▲whimsicalism 3 hours ago | parent | prev | next [-]

that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied

i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight

▲bamboozled 3 hours ago | parent | prev [-]

The rapid development of extremely dangerous bio-weapons?

▲pj_mukh 3 hours ago | parent | prev [-]

Misuse how exactly?

▲voganmother42 an hour ago | parent | next [-]

At a minimum its another force multiplier that enables a small(er) number of people to exert more control over more people.

▲whimsicalism 3 hours ago | parent | prev [-]

any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified

▲pj_mukh 2 hours ago | parent [-]

I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?

Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?

It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.

▲ewild 3 hours ago | parent | prev [-]

i feel bad for math guys yeah seems they are more cooked than CS

▲never_giveup 4 hours ago | parent | prev | next [-]

Try using AI for your work, whatever you do. You will quickly understand the limitations.

▲ggreer 3 hours ago | parent | next [-]

Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

▲psvv an hour ago | parent [-]

Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

▲ggreer an hour ago | parent [-]

Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.

▲psvv an hour ago | parent [-]

My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.

▲againstapples 2 hours ago | parent | prev [-]

It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?

▲jaykru 3 hours ago | parent | prev | next [-]

I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.

[0] https://dank.systems/posts/2026-09-15-ai-bear.html

[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...

▲red75prime 2 hours ago | parent | next [-]

> we can clearly specify what AGI or ASI is

We'll have plenty of time for this, while living off UBI.

▲p-e-w 2 hours ago | parent | prev [-]

> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains

But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.

▲somenameforme an hour ago | parent | prev | next [-]

Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

▲Rudybega 44 minutes ago | parent [-]

I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.

▲zeroonetwothree 3 hours ago | parent | prev | next [-]

I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.

▲pixl97 3 hours ago | parent [-]

You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"

The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

▲schleck8 4 hours ago | parent | prev | next [-]

From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.

So in other words, since deep learning is algorithmic research, we are now in the RSI era.

▲thereitgoes456 3 hours ago | parent [-]

> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches

"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)

How did you determine this in 1 hour? Are you a researcher in multiple of these areas?

Can you give an example, or explain more how you came to this conclusion?

▲scarmig 2 hours ago | parent [-]

One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.

▲besterman23 3 hours ago | parent | prev | next [-]

I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.

▲gizajob 3 hours ago | parent | prev | next [-]

Did AI beating humans at chess:

a) destroy chess and make it a pointless endeavour,

or

b) make humans much better at chess.

▲lf88 2 hours ago | parent | next [-]

Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.

▲Light_Hope 3 hours ago | parent | prev | next [-]

Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.

▲vouaobrasil 3 hours ago | parent | prev [-]

It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.

I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.

So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.

▲gizajob 2 hours ago | parent [-]

At the same time though, Magnus is Magnus because he’ll crush you in any endgame.

I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.

I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.

▲rcpt an hour ago | parent | prev | next [-]

Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.

▲skybrian 3 hours ago | parent | prev | next [-]

For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.

▲icepush 3 hours ago | parent | prev | next [-]

They can replace anyone but they can't replace everyone.

▲ForHackernews 3 hours ago | parent | prev | next [-]

AI performance has always been extremely spikey. It's great at some things and terrible at others.

Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?

▲againstapples 2 hours ago | parent [-]

I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.

I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.

▲psvv an hour ago | parent [-]

Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.

Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.

We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.

Does solving more theorems than before suddenly mean computers are capable of anything? No.

▲yk 2 hours ago | parent | prev | next [-]

I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.

So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.

▲outworlder 2 hours ago | parent | next [-]

Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.

▲lf88 an hour ago | parent | prev [-]

It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.

▲ijidak 3 hours ago | parent | prev | next [-]

For me it's a mixed bag.

There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.

At the same time we have to put what AI can do in perspective.

Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.

AI has incredible knowledge and in many areas approximates experience and wisdom.

But wisdom is harder to formalize than knowledge and skill.

For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.

To some extent advanced degrees try to certify maybe wisdom and experience.

In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.

Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.

Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.

Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.

But the world has been an especially volatile place over the last 10 years.

So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.

But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.

I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.

In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.

▲runarberg 3 hours ago | parent | prev [-]

AI hater here:

I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.

That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.

▲baoooooooooooo 35 minutes ago | parent | next [-]

A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here

▲runarberg 26 minutes ago | parent [-]

Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.

▲vouaobrasil 2 hours ago | parent | prev [-]

I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.

Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.

Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....

Personally, I think AI is a grand mistake.