Remix.run Logo
▲ bamboozled 8 hours ago

So OpenAI should be able to flood the world with AI pollution and ask scientists and mathematicians to wade through it all and tell us if there is any sense in it, then sit back and wait for them to report in?

Nice idea.

▲sanxiyn 7 hours ago | parent | next [-]

OpenAI in fact didn't know what to do with results and didn't want to flood the world, so they asked mathematicians. Mathematicians recommended OpenAI to release them. My guess is it would have been better for OpenAI if they didn't release them. OpenAI is basically doing this as a goodwill.

(See https://agmai.org/general-sep29/ for the recommendation in question.)

People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).

Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.

▲demibabs 7 hours ago | parent | next [-]

How is solving unsolved problems a good eval? Once a problem is solved and released, you can’t evaluate a future model’s ability to solve it.

▲sanxiyn 7 hours ago | parent | next [-]

As I said, a proper eval is better, but it measures something real that gives a good training signal, and there is lack of good alternatives for math eval. Since the result is 372/4000, it is also unsaturated.

▲jsrozner 2 hours ago | parent | prev [-]

I think this is right - it can be seen as trying to get free feedback from the community. In that way it’s reasonably described as exploitative, since it’s not a good faith effort. The problem of ai slop being submitted to conferences to get publication counts is similar.

▲dataflow 2 hours ago | parent | prev | next [-]

Any idea what sort of percentage the sentence

> supported by a clear plurality of respondents

was referring to?

▲Personbeing12 an hour ago | parent | prev [-]

"Mathematicians recommended OpenAI to release them."

This is a misleading characterisation of the mathematicians' position.

The very first paragraph of the AGMAI recommendations explicitly states:

"we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." You appear to have acknowledged this by saying “Mathematicians would prefer if OpenAI didn't use open math problems as an eval…”.

The mathematicians did not ask OpenAI to produce these results. They explicitly asked AI labs to stop producing them in this manner. Their subsequent recommendations concern what labs should do if they have already produced significant results, not an endorsement of the practice.

Furthermore, the recommendation was not simply to release the results, but to responsibly release already existing results. Section 2.B, Step I, explicitly recommends "...labs that have AI mathematical output that is not understood by the people who prompted the AI systems", to search the literature for relevant prior work, provide appropriate attribution, and improve the exposition of AI-generated proofs before releasing them, rather than leaving this work to mathematicians afterwards.

OpenAI published the results on GitHub while still exploring repositories that meet the committee's guidelines. So they followed some of the recommendations, but not all of them and hence, did not release the results as requested by the mathematicians.

I do not think it is a settled matter whether this was done out of goodwill. This is because releasing these results as they were can benefit OpenAI more than releasing them according to the AGMAI recommendations. AGMAI recommended in section 2.B, Step 1.5 that "Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen." If the results are released, it is easy to expect that the media will discuss the capabilities of the AI used in the work, as indeed happened. If this AGMAI recommendation was followed, the media would plausibly have also discussed the number of failed attempts and then the overall attitude would not be as favourable to OpenAI as it is now when it comes to the capabilities of the AI that was used. OpenAI did release on GitHub that approximately 4,000 problems were attempted and resulted in 719 manuscripts (after 3 containing suspected errors were removed by OpenAI) across 372 families of problems, but this does not give a calculable number of problems it failed to solve. I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.

AGMAI's October 6 statement explicitly clarified that its advisory role should not be interpreted as an endorsement of OpenAI's process, and that it was up to the mathematical community to assess how successfully its recommendations had been followed.

Recommending how to responsibly handle the outcomes of something you oppose is not the same as asking for it to happen.

▲spidersouris 7 minutes ago | parent | next [-]

It is also worth adding that the Association for Human Mathematics published a statement (which Tao reposted on his blog) in which they explicitly say the following:

> Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research — norms that guarantee that mathematics remains trustworthy, ethically researched, and in the public interest.

https://www.ahmath.org/

▲ 11 minutes ago | parent | prev [-]
[deleted]
▲ComplexSystems 8 hours ago | parent | prev | next [-]

Yes, they should. They have invented a magic button that can tell you the long-awaited answers to the burning mathematical questions that you've spent your life researching. The caveat is that the technology is still new, so the explanations "are not fun to read" like set theory papers usually are (lol). If you don't think that's a worthwhile tradeoff, that's your call, but it sure as hell isn't everyone's.

▲advael 8 hours ago | parent | next [-]

I don't think the claim is "they definitely have an oracle that solves the problem, and I reject it because it's hard to read". The claim is "OpenAI claims to have used an oracle to solve the problem. The proof is very difficult to read, and to even know if it does or not, we have to go through it with a fine-toothed comb, but they're going around claiming they definitely solved the problem (or at least getting press that claims that which they aren't pushing back against) and this might convince the people who sign grants even if it isn't true"

People who are invested in the idea that we've invented a general intelligence, now, which includes all these companies that are literally financially invested in this claim they are making, will tend to believe that its results can already be trusted in domains like this. Some mathematicians seem to believe some of the proofs written by their models, and some, like this one, don't. I do think it's valid for an expert to push back against the claim that the best use of their time right now is to verify the poorly written work of everyone who's claimed to solve the problem

▲famouswaffles 8 hours ago | parent [-]

Over the year, nothing they've released with a lean proof attached has turned out false (That's sort of the entire point. It's not impossible but it's really difficult). There's a reason most mathematicians, including the ones vehemently against OpenAI's dumping are not arguing the results are secretly false or have a high potential to be. And indeed, if that were the case, it would quickly become apparent and all this worry about grant signers would vanish into the wind. It's very easy to ignore nonsense. The problem is that it isn't nonsense.

▲advael 7 hours ago | parent [-]

I really don't know enough about it to know whether you're right or not, nor do I know whether or not you know enough to make the claim you're making, so I won't make an argument one way or another because it's non-sequitur to what I said anyway. The fact that you or I or Sam Altman or Terrence Tao believe the claim is irrelevant to whether this obligates the person who wrote the blog post to believe the claim, and it sounds like he's willing to consider the possibility that it is right, and would read the paper if it reached a threshold of comprehensibility expected of people making that kind of claim.

▲famouswaffles 7 hours ago | parent [-]

I's not a non sequitor because it cuts right to the point. He's under no obligation to read it sure, but that doesn't mean Open AI isn't justified in claiming to have proved it. The justification isn't Sam Altman's belief or Tao's or anyone else's authority. Results accompanied with lean-verified proofs whose formal statements match the problem at hand have arguably stronger justification than the vast majority of human math publications.

▲buriram 8 hours ago | parent | prev | next [-]

I don't understand. Why is this the onus of scientists and PhDs to review whatever results OpenAI had dumped out? If OpenAI had produced incomprehensible papers, surely any journals would just reject it, or demand the author to do a complete rewrite? Unless we are talking about a race to solve problems, which PhDs are afraid that they had been scooped up on?

▲bamboozled 8 hours ago | parent | prev [-]

The problem isn't just that the papers aren't fun to read. The problem is that a lot of the research that goes into solving these issues leads to other discovers, new fields to explore and people have to develop new approaches to solve them. The other part of it is, the quality and the enjoyment of working on these problems leads people to find new and other interest problems to work on.

If you just strip mine the answers and Sam Altmans magic button solves 100/100 problems, what's next? Who is left to come up with a new interesting question for the magic button to solve?

Lastly, life and the present moment is all there is, if there is no enjoyment in anything we do, then what's the point of all the "living for ever" Altman et al want to achieve.

We will live forever to read boring papers generated by LLMs? Literally sounds like an eternal hell.

▲coderenegade 7 hours ago | parent | next [-]

People are already finding stuff in the release to get excited about, and as the models get better at distilling proofs to make them more coherent, this will only amplify. Of all the things to worry about, human curiosity and the ability to run with new ideas probably aren't at stake.

In fact, I can't remember a time when I was more excited about the future of science. This could herald an end to the replication crisis, and kill off bullshit science completely. The danger of course is that we end up with two companies effectively dominating cutting edge research in every field, but it remains to be seen if that's even possible given the pace of improvement in open weight models.

▲vikramkr 8 hours ago | parent | prev | next [-]

Nuclear fusion is already proven by the universe to be a viable energy source by the fact that the sun exists but people still work on understanding and taking it and developing new approaches to accomplish it. People didn't stop experimenting with and developing programming languages because technically they're all turning complete and the first one was "enough." Y'all will be fine - every JavaScript framework that exists is someone looking at a theoretically correct and complete solution and deciding actually it sucks and they could do better. "I want to understand xyz but the proof is trash and I think it's ugly" will be plenty motivation for a lot of people to work on it.

▲inference-god 8 hours ago | parent [-]

How is the journey to understand fusion related to not wanting to spent your limited time on earth wading through AI slop?

▲vikramkr 3 hours ago | parent [-]

You don't have to wade through the slop. The point is a giant star in the sky figured it out and that doesn't demotivate you from figuring it out yourself

▲ComplexSystems 8 hours ago | parent | prev [-]

Here's a guy who's made a "beyond n log n" tracker: https://x.com/aurel_pr/status/2108214135179944096

He's tracking the community progress on sub-n log n multiplication. OpenAI started with 1 - 1.63e-55. The result has been now improved on 115 times, and the current record is "rohanarun"'s 1 - 9.87e-5. I'm sure by tomorrow it'll have improved again.

Does this look like people aren't having fun? Does it look like they aren't discovering stuff? It looks like it's spurred a cascade of interesting community activity. It doesn't really seem much different from what happened with the twin primes conjecture. Isn't that supposed to be the point of all this?

▲bamboozled 37 minutes ago | parent [-]

Some people are having fun doesn't negate the rest of the issues with this sort of thing.

▲LPisGood 8 hours ago | parent | prev | next [-]

They should certainly be allowed to share their findings. No one is forcing scientists and mathematicians to review the findings in general. It’s just the case that the findings are of such such a quality that it would not make sense to ignore them wholesale

▲Arainach 8 hours ago | parent [-]

> No one is forcing scientists and mathematicians to review the findings in general.

This is like "no one is forcing software engineers to use AI tooling" or "no one is forcing you to show your ID in the airport" or "no one is forcing you to own a car in your small midwestern city" - there can be no law requiring something and the practical consequences of not doing so can be so painful that you're effectively forced anyway.

▲LPisGood 5 hours ago | parent | next [-]

That is exactly what I’m trying to say. The findings are of such such a quality that it would not make sense to ignore them wholesale.

That’s why it doesn’t make sense to present AI companies as dumping or burdening the scientific community into doing labor for them; the scientific community is self motivated to do so.

▲Arainach an hour ago | parent [-]

It's not self motivated. The motivation is not "this is doing amazing things for us", it's "if we don't review this, the bullshit headline complex and the bullshit-spewing (sorry, marketing) departments of tech giants are going to misinterpret/misrepresent everything and our grant money will be taken away".

▲senderista 7 hours ago | parent | prev [-]

add "no one is forcing you to own a smartphone"

▲unknownian 7 hours ago | parent | prev | next [-]

The best part is that when the academics fix OAI’s issues, the model gets better and OAI shareholders get richer and more powerful!

As someone who uses LLM tech occasionally, this is why I prefer using open local models. If I’m making myself obsolete, at least I’m not making some asshole richer and their closed model better.

▲skeledrew 8 hours ago | parent | prev | next [-]

No, they don't have to ask. Those interested enough will jump at the opportunity, even if it's just to be "one of the first to get it".

▲latentsea 8 hours ago | parent | prev | next [-]

You know, I kinda relate to the feeling of not wanting look at those outputs if I think of it from a layman's perspective.

I just imagined that instead of math papers, they released 700+ feature length films, and the only way to tell if one of them is any good is to watch it in its entirety.

That feels pretty unappealing to me.

I know it's the same for human made films, so what's the difference right? But those are good enough most of the time that it's a decent bet, and the people that made them had real skin in the game.

Contrast that with something made by a nondeterministic slop machine with no skin in the game where small details can be off in a way that's jarring. Right out the gate I have an aversion to committing that much time to something that very well may waste it.

▲solyin 7 hours ago | parent [-]

That's actually a really interesting thought. Given Sora, and the amount of funding they have, they could have created started their own film festival and dropped 700+ feature length films, had they wanted to go in that direction. But they didn't. Hmm.

▲fbrncci 8 hours ago | parent | prev [-]

What else should they do? See these models get smarter and smarter, somewhat-solve things but only to the tune of 90% what mathematicians (or experts in any other field) would deem acceptable, and then gate keep the findings for the next few years going through peer review and paywalled journals? I for one welcome the flood, bring on more in every possible industry and see where all that progress lands up. Sure it will upset a lot. A lot of things also upset the luddites.

▲XenophileJKO 8 hours ago | parent | next [-]

Exactly.. this is an opportunity for people to pick up where the model left off and run with it. Like you don't have to, nobody is forcing you to.

However, there are always smarter, hungrier people out there and this is a buffet.

Some output is going to be wrong or incomplete. I am willing to bet even those have nuggets that can be used elsewhere.

▲fn-mote 8 hours ago | parent [-]

> this is an opportunity for people to pick up where the model left off and run with it

Like people enjoy racing in front of a stopped train? As soon as they turn on the engine again, they will run you over. The questions that remain will be only the low value ones, not worth the effort to vacuum up.

So no, the smarter, hungrier people are not the ones that are going to swoop in. It will be the most desperate.

> Some output is going to be wrong or incomplete

This is a very human take on the situation. No, the Lean proof is not going to be wrong, and it will be incomplete only in the sense that OpenAI didn’t try to push the results further.

▲ 7 hours ago | parent [-]
[deleted]
▲bamboozled 8 hours ago | parent | prev [-]

Are you going to be doing the work to verify the results ? Will you just expecting other people to wade through the flood and reap the benefits later on?

▲fbrncci 8 hours ago | parent [-]

Nobody is forced to verify the results. I am honestly not expecting anything other than AI to get smarter and smarter and people who are motivated and interested enough to pick up after it; and potentially reap all the long term benefits ahead of those who aren’t (seeing this happening with software development in my own field). But in the end; if you don’t like it, you’re not forced to do anything.

▲bamboozled 4 hours ago | parent [-]

Wondering if you read the article?