Remix.run Logo
▲ pavitheran 4 hours ago

From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

▲password54321 4 hours ago | parent | next [-]

Oh cool, we will all now have a math genius on our computer.

▲jrflo 4 hours ago | parent | next [-]

It was using their internal math model, so not yet for us

▲password54321 4 hours ago | parent [-]

I used future tense. It was implied this will be available.

▲an0malous 4 hours ago | parent | prev [-]

Well, on their computers. But you can rent them for a price.

▲binlog 5 minutes ago | parent [-]

An open source model will reproduce it 6 months later

▲orlp 4 hours ago | parent | prev | next [-]

I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

▲timjver 4 hours ago | parent | next [-]

>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

That doesn't sound right

▲orlp 4 hours ago | parent [-]

Oops, edited.

▲machomaster 4 hours ago | parent | prev | next [-]

They did say that. "3 hours of ChatGPT Pro thinking compute"

▲orlp 3 hours ago | parent [-]

Yes, what does that mean?

▲pixl97 3 hours ago | parent | prev [-]

Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.

▲orlp 3 hours ago | parent [-]

I'm not denying that, but I'd still like to know what that cost.

▲scrlk 4 hours ago | parent | prev | next [-]

Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?

▲inferencecoder 4 hours ago | parent [-]

It doesn't imply that, it's just measuring the amount of compute.

▲bigmadshoe 3 hours ago | parent [-]

But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?

▲Jtarii 4 hours ago | parent | prev [-]

That estimate is obviously going to conveniently ignore all the failed runs.