| ▲ | wrsh07 2 days ago | |
It seems like they threw it a decently large battery of open math problems and probably limited it to something like $200-500 per problem: https://x.com/polynoamial/status/2083478171975082334 As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget -- The linked tweet from Noam Brown at OpenAI reads: > And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet). > But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further. | ||