Remix.run Logo
▲ charlieyu1 3 hours ago

The biggest problem is LLM tends to produce over engineered, very complicated proofs that are an eyesore even for relatively simple problems. Give it a beautiful Olympiad geometry problem and LLM will tear it apart into ugly algebraic calculations, turns all lines and circles into equations and calculate their intersection points that spans multiple pages because it is a guaranteed way to solve it. Correct, but hardly any use to the user.

▲Gigachad 2 hours ago | parent | next [-]

I’m not deep in to math but the op tweets make sense to me. In that it’s not just the final proof that mattered, but the mind and understanding of the person who arrived at the answer. An LLM dumping the answer can’t elaborate on it, can’t tell the story of how they got there, etc. But it also deprives someone else of that achievement and learning.

▲Ravus 2 hours ago | parent [-]

You can see it in the way we structure college courses: engineering curricula often cover in one semester what mathematicians study over one or two years.

This is because have fundamentally different goals: being able to use results in calculation versus having a deeper understanding of the subject matter.

▲KoolKat23 2 hours ago | parent | prev [-]

You know it's valid. You're not working on incorrect assumptions. Surely there's value in that?

▲datsci_est_2015 2 hours ago | parent | next [-]

How something is proven is often more important than what is being proven. There are underlying systems and patterns that, when understood properly, improve our model of mathematical (or physical) reality.

With convoluted and inelegant proofs, AI may fail to uncover those systems and patterns. As a most concrete example, it may fail to recognize some problems as isomorphic to other problems. Brute force solutions are a depth-first search.

To improve human mathematical understanding, AI is probably best used as a “copilot” (lol) rather than a black box oracle, like these AI companies appear to be doing.

▲KoolKat23 2 hours ago | parent [-]

I'm sorry this framing is just moving goal posts.

If you're after new methods. Then new methods is the goal, the answer to the question is not the goal then. The animated response indicates the answer wasn't just a byproduct.

There is still something to glean from the answer. You have a further constraint. Otherwise whatever "new method" proposed may as well be hallucination, potentially taking you in the wrong direction away from the answer.

This line of thought is not unique, stonemasons made obsolete by uniform brickword suddenly were "worried about the art and preserving traditions".

▲charlieyu1 2 hours ago | parent | prev [-]

I don’t know if it is valid. It is unverifiable. I still found some basic algebraic mistakes in top models as late as 3-4 months ago, not sure about it now. But that’s not what I want anyway, so I often put “Do not brute force” in my prompts.

▲KoolKat23 2 hours ago | parent [-]

Sorry I mean in very public releases such as this trove, where many have lean certificates attached and publicly scrutiny.

▲shakna an hour ago | parent [-]

Three of them have already been withdrawn. So I would say we have proof, that they cannot be implicitly trusted.

▲KoolKat23 an hour ago | parent [-]

So it turns out there is a, perhaps informal, working system in place and we can deal with it.

Peer reviewed and published insights are proven invalid all the time. This is the nature of research and how we learn.

▲shakna an hour ago | parent [-]

So instead of us "knowing it is valid", we don't. We need to put in extra effort, because someone felt like doing only half the work and dumping it on the community to fix.

We don't know the current system can work well enough at this scale, because that's un-knowable. We know it can find some of the problems. We don't know it can find all of them.

We do know it takes more effort - that's knowable. Increased data takes increased processing.

Whether the community has the required effort available, seems unlikely, considering the expertise required to be able to assess these things hasn't changed. Only the ability to generate them has increased.

▲KoolKat23 an hour ago | parent [-]

You don't have to go through it. You can ignore it if you wish. There is no obligation on you to check it.

I feel your concern stems from the risk that there is additional noise everyone needs to cut through.

In reality this isn't any tom, dick or Harry giving you their vibe code output. They have spent millions of dollars on this output, so there is a filter. The biggest filter of them all, funding.

Furthermore, LLM's have given us another gift semantic search, we can easily check your work against theirs, this is valuable insight so instead of researchers wasting decades and fortunes pursuing an avenue that shows no value (this includes methods), they can purse new avenues they know what to avoid, in the same breath they know what to work towards.