Remix.run Logo
aabhay 5 hours ago

My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I want to know:

1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

einpoklum 5 hours ago | parent | next [-]

Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

energy123 3 hours ago | parent | next [-]

Many less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.

traes 5 hours ago | parent | prev [-]

> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.

> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.

[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...

irthomasthomas 2 hours ago | parent [-]

Why you think that?

azan_ 2 hours ago | parent [-]

I guess that's because there are serious problems on which many professional mathematicians worked on years. If it was just a matter of hiring an expert, they would've been solved long time ago.

irthomasthomas an hour ago | parent [-]

I guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?

dist-epoch 3 hours ago | parent | prev | next [-]

I don't think you want to bring cost into this argument.

Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.

Do you really think that if you paid that to humans, they will deliver the same results?

uh_uh 2 hours ago | parent | next [-]

It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.

robotpepi 24 minutes ago | parent | prev | next [-]

it's still important. not everyone has access to 1 million USD. saying it "only" coat 2000 USD is highly misleading for the discussion and future. the concentration of power is a huge problem with AI.

kevinwang 41 minutes ago | parent | prev | next [-]

It would still provide better context to see the numbers that the parent proposes, though.

mungaihaha 2 hours ago | parent | prev [-]

Grad students on zero pay solve problems like this everyday. What exactly is your point here?

mirzap 2 hours ago | parent [-]

Even if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.

irthomasthomas 2 hours ago | parent | prev | next [-]

[flagged]

simianwords 5 hours ago | parent | prev | next [-]

There are people who can’t grasp the universe without mandatory randomised controlled trial. Would tomorrow be a Sunday? Need an RCT for that boys!

My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.

nxpnsv 4 hours ago | parent | next [-]

No, this is valid criticism. Oai gives the impression anybody could get similar results at a similar price, but that’s very likely not true. This is marketing first, then mathematics.

traes 5 hours ago | parent | prev [-]

It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.

simianwords 5 hours ago | parent [-]

Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.

esperent 4 hours ago | parent [-]

> deliberate misleading “lack of transparency” etc

It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.

dist-epoch 3 hours ago | parent [-]

The results speak for themselves.

Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"

esperent 2 hours ago | parent [-]

Nobody is claiming the results are false.

We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.

This goes double since it's an internal secret model (Astra) so nobody else can verify the results.

simianwords an hour ago | parent [-]

Would this be your reaction if OpenAI also solved millennium problems? The point we are trying to make is that the significance of this news is much larger than the skepticism you are providing.

azan_ 2 hours ago | parent | prev [-]

> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.