Remix.run Logo
hansvm 4 hours ago

Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will.

Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.

an hour ago | parent | next [-]
[deleted]
hollowcelery 38 minutes ago | parent | prev [-]

But the researchers also did their research using essentially the same models, so that isn’t a counterpoint to AI models being at the far frontier…

hansvm 4 minutes ago | parent [-]

> not a counterpoint per se

>> isn't a counterpoint

I don't think we're disagreeing? The observed behavior (solving Navier Stokes) can be generated in tons of ways:

- It's still not solved; the Lean formalism is wrong. I wouldn't bet a ton of money against this till I hear more.

- The models kind of suck, but they get lucky from time to time (better-than-average random search) -- this is correct but ultimately not very helpful; if the error rate is low enough then the fact that you're just sampling is less important than the fact that you're sampling from a nice space. For a one-off though, it's absolutely not to be discounted though that a team of monkeys with an appropriately armed yes/no button and enough time could've produced a Lean proof (if the problem is probable).

- The models suck, but the real failure is people not manually trying enough options, maybe because they think the problem is unsolvable and therefore not worth their time, maybe because they only act in the spirit of the problem and search for deep insights, but regardless of mechanism people haven't done the simplest of brute-force searches across easy solutions.

And so on. The implication (if you enjoy extrapolating from one-offs) of my observation is that the model got a little lucky, did some brute force stuff, might still be wrong, might actually be amazing, and we currently don't know much about its capabilities.

Maybe they're at some sort of far frontier. What's the evidence? A one-off which isn't yet known to be true and which could've been generated by monkeys isn't enough evidence by itself no matter the import of the result. More than that, what do you mean by "at the far frontier"? If you're saying that they're capable of superhuman tasks with sufficient context and prompting, yes, of course. Is that even your claim though? For the people who want to engage further with my ... whatever this is ... what are they engaging with?