| ▲ | Position: LLMs Can't Jump(openreview.net) | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 105 points by theanonymousone 3 hours ago | 58 comments | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | quantum_mcts an hour ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
The popular retelling of how Einstein created Special Relativity to "Resolve the contradictions of Michelson-Morly experiments" is very reductive to the history of the question. The epitome is the quote from the paper: > From the two postulates, Einstein derived the Lorentz trans- formation ... If Einstein derived them, who is "Lorentz"? The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | setnone a minute ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
LLMs don't have legs yet. If you're smart but can't touch things you only keep being smart | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | defgeneric an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently: > A few reflections on my "LLMs Can’t Jump" paper: > My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things. > First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case. > This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing. > Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking. > Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere. > Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages! | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | yomismoaqui 7 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Why everybody is obsessed with replacing humans with LLMs when it seems like the most profitable use cases (like coding agents) rely on enhancing human capabilities? Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process). | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jvanderbot 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | bob1029 an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all. An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | kfarr an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Best comment on this from 6 months ago: https://news.ycombinator.com/item?id=46870575 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | dtj1123 6 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Neither can I, if I'm being honest. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | nativeit 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | GodelNumbering an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loops) For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make. [1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible [2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. " Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | pama an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
The early physics background is messy and incorrect. I didnt read the full position paper, but from its start: The Lorentz transformations were by Lorentz, well before Einstein’s paper on special relativity; the principle of relativity also existed before the Einstein paper. The math was all there, with steps taken by Maxwell, Voigt, Larmor, Lorentz, and Poincare. Einstein supplied a clean physical interpretation, making all inertial frames equivalent, making simultaneity frame dependent, and explaining length and time deformations without the need of the concept of ether. Skimming the end of the paper with the arguments about lack of abduction or inability to make the analogy without sensory experience, I see that this paper is unfounded speculation rather than solid/hard philosophical logic. As a position paper it is OK to appear, but i think it misses the point of how LLMs or other autoregressive learners of future states can build analogies and intuition that can help them formulate new theories of the world. Soon it will be more obvious to everyone, so I am not very worried about these writings. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | rsfern an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | sobiolite 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying "In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality." But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps? | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | zie1ony an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Interestingly, halucinations might be the way to achieve that. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Yopolo 40 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
We just put different things together and then we evaluate it. In math its simple: does the verification say its okay. If its mechanical: is any property better than what we have already. etc. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | ACCount37 an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
It's a very shaky position, and the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt. Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation. The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion. The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains. LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that. Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | redwood 27 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
For those who may not be aware this is a clever title derived from a film title https://en.wikipedia.org/wiki/White_Men_Can%27t_Jump | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | reliablereason an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Clearly LLMs cant do leaps of intuition since their "intuition" is locked after training ends. The only way a LLM can come up with new ideas if the "idea" appeared as a generalisation durring training or if it was achieved using reason in chain of thought. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | conartist6 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
This tracks for me as someone making keeps of intuition in little-explored areas. I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | brainless an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I have a weird thought experiment: If you give a GPT-2/3 level LLM tools to search the internet - any document, can it build bigger, better LLMs? You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses. Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Mistletoe 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I feel like you could just add some noise or randomness to the LLM and start approximating the leaps that the human mind uses to solve and understand unrelated things. Maybe that’s naive, it’s just coming from my organic computer in my skull. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | juleiie an hour ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
LLMs can’t but humans supplied with data and reasoning from an LLM can make novel jumps without absolute prior knowledge, or at least with reduced need for years of knowledge. And then such jump can be verified by a machine so human kind of plugs the intelligence gap. That’s pretty exciting. I always liked to provocatively call LLM „the new calculator”. Calculator for language. We are so focused on creating a standalone intelligence that we didn’t notice how we massively augmented our own. That could be considered transhumanism holy grail if only interface brain-LLM was faster. People need to understand that these things are tools. And every tool needs an operator to function. Tool doesn’t have its own goals, needs, wants or motives. It won’t do anything out of its own, it always exists in context of someone telling it what to do. In light of that most of the panic and fear mongering is rather ridiculous. Calculator won’t replace you. It wont take over the world. It is just a tool. You write a book with book generator? Cool, it can be used for this. We will judge output, not the methods. Sometimes we will judge people who have no taste in literature. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||