| ▲ | Eridrus 2 hours ago | |||||||||||||
I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability. The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result. If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this. This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env. | ||||||||||||||
| ▲ | curt15 an hour ago | parent | next [-] | |||||||||||||
> The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result. Not to mention a century of theoretical foundations and innovations by humans, written for human understanding. | ||||||||||||||
| ▲ | alienbaby 25 minutes ago | parent | prev | next [-] | |||||||||||||
I've wondered how; "The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result." can possibly not also include many dollars worth of duplicated effort? I find it hard to believe every agent was doing something novel. What this means beyond wasting dollars I'm not sure of, but just flinging around the numbers I don't think should be considered impressive or even required to get the results found. | ||||||||||||||
| ▲ | hn_throwaway_99 an hour ago | parent | prev [-] | |||||||||||||
Regarding the Navier Stokes theorem, one of the best posts I've seen about this (and the other OpenAI math proofs) recently is Nestor Guillen's post where he introduces "the convex hull of ideas". I love this because I've had some vague ideas about the limitations of LLMs that aligned with this, but Guillen really clarified the idea and explained it with a great analogy: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t... Basically Guillen is arguing that LLMs are great at finding results within the "convex hull" of existing literature (i.e. their training data). They can make connections across different parts of the literature where it would be impossible for a human to be an expert in all these areas. But when it comes to truly novel, original ideas, there is no proof yet that LLMs are able to go there. That's honestly a limitation I'm really rooting for because otherwise I think the future of humanity is generally fucked. | ||||||||||||||
| ||||||||||||||