Remix.run Logo
simianwords 5 hours ago

There are people who can’t grasp the universe without mandatory randomised controlled trial. Would tomorrow be a Sunday? Need an RCT for that boys!

My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.

nxpnsv 4 hours ago | parent | next [-]

No, this is valid criticism. Oai gives the impression anybody could get similar results at a similar price, but that’s very likely not true. This is marketing first, then mathematics.

traes 5 hours ago | parent | prev [-]

It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.

simianwords 5 hours ago | parent [-]

Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.

esperent 4 hours ago | parent [-]

> deliberate misleading “lack of transparency” etc

It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.

dist-epoch 3 hours ago | parent [-]

The results speak for themselves.

Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"

esperent 2 hours ago | parent [-]

Nobody is claiming the results are false.

We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.

This goes double since it's an internal secret model (Astra) so nobody else can verify the results.

simianwords an hour ago | parent [-]

Would this be your reaction if OpenAI also solved millennium problems? The point we are trying to make is that the significance of this news is much larger than the skepticism you are providing.