| |
| ▲ | AdieuToLogic 4 hours ago | parent | next [-] | | > Heck, just for fun I asked a reasonably smart LLM to ... LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. > You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path ... Again, LLMs do not "hallucinate." They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. Nothing more. See also anthropomorphism[0]. > More precisely it's that [LLMs] can't do the math internally but they're quite capable of producing the tool that does the math. This still falls under the purvey of statistical token generation. To wit, given enough variations of: bc -e '1 + 2'
bc -e '41 + 1'
...
LLMs can identify the addition expression in "What is 4 + 1?" and then emit a `'bc "4 + 1"'` command to produce a response. This is not "doing" or "understanding" math.It is pattern recognition, a task in which ANNs[1] excel. 0 - https://en.wikipedia.org/wiki/Anthropomorphism 1 - https://en.wikipedia.org/wiki/Neural_network_(machine_learni... | | |
| ▲ | bryanrasmussen 3 hours ago | parent | next [-] | | >LLMs are neither smart nor stupid. by that reasoning then neither are there smart or stupid designs, questions, answers, or any of the millions of things that were described as smart or stupid, that did not possess any brain to actually be smart or stupid long before LLMs showed up. The analogical process implied in many common English usages means that describing an LLM as smart or stupid is perfectly reasonable. | | |
| ▲ | UpsideDownRide 2 hours ago | parent | next [-] | | Nah it's not reasonable to use words that send you down a wrong concept path. | | |
| ▲ | martin- 2 minutes ago | parent [-] | | We've always used metaphors like that when talking about computers, without anyone having any complaints. |
| |
| ▲ | bryanrasmussen 3 hours ago | parent | prev [-] | | I'll just note here that sure, there are people who go around thinking that LLMs are actually endowed with the capacity to reason, but generally I find the people who think this do not know what an LLM and will just use the name "ChatGPT" | | |
| ▲ | 9 minutes ago | parent | next [-] | | [deleted] | |
| ▲ | SR2Z 2 hours ago | parent | prev | next [-] | | What would it take for you to say that an LLM can reason? The completions they provide are generally internally consistent. We're at the point where they can produce proofs that eluded human mathematicians for centuries. VLMs and self driving cars can handle ambiguity and run safely in a variety of situations. If it looks like a duck, walks like a duck, and quacks like a duck maybe it just makes sense to call it a duck and put off the philosophy for when it might make a difference. | | |
| ▲ | skygazer 12 minutes ago | parent | next [-] | | I think humans have to reason because we don’t already have a statistical embedding of the solution pattern built in. We have vastly less rote knowledge crammed into our heads and so require creative synthesis to span the gaps. With LLMs the trick is revealing their existing relevant embedded knowledge more reliably. They’ve almost literally seen it all before, and the trick is dialing it in. The reasoning tokens help shape the autoregressive attention lens that focuses on and enables recall of the already-experienced answer. It is interesting that “reasoning” has a similar outward appearance, but since LLMs are built to
mimic outward appearance from trillions of examples, you can’t infer underlying mechanism from appearance. | |
| ▲ | nevertoolate 36 minutes ago | parent | prev [-] | | It looks like a next token predictor, walks like a next token… You get my point. It definitely doesn’t look like my elderly neighbour, nor like my daughter, etc. It is confusing but very simple at the same time. |
| |
| ▲ | KPGv2 2 hours ago | parent | prev [-] | | Yeah and there are people who worship feces, but that doesn't stop the rest of us from freely saying "holy shit" and not correcting each other saying "technically it's not holy, and you shouldn't say that, because you might enable one of those poop worshippers." We can't tailor our linguistic shorthand to the lowest common denominator. Also we're on HN, not talking to an octogenarian US senator. | | |
|
| |
| ▲ | mapontosevenths 2 hours ago | parent | prev | next [-] | | By this logic a human is only $130-$160 worth of Oxygen, Carbon, Nitrogen and some trace elements. Perhaps structure sometimes makes things that are more valuable than their inputs? That said, this is also inaccurate at a technical level.LLM's are very capable of doing math and they ARE calculating internally. Most of what they do is calculation, not storage. It's just not done in a way that it's trivial to explain here. It's described in some detail below, though it's a bit dense. https://www.lesswrong.com/posts/E7z89FKLsHk5DkmDL/language-m... | |
| ▲ | hodgehog11 3 hours ago | parent | prev | next [-] | | During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. It encompasses virtually everything. It is totally meaningless. So to say "nothing more" is effectively also a tautology. This argument was asinine in 2024. It is insane to be saying these things in 2026. Where have you been? What have you been looking at? How many articles explaining why the "statistical parrot" analogy fails have you missed? How much mental gymnastics do you have to do to explain how a modern LLM can solve novel math problems that fall really far outside of its training set? It absolutely understands how to do math, by whatever reasonable definition you want to provide to the word "understand". For example, the identification of the addition expression is understanding, and no, it does not do tool calling for basic arithmetic any more than humans might. Isolation of individual concepts in intermediate layers can already be demonstrated, or else transfer learning wouldn't possibly work. Nobody is saying that LLMs are humans. But we need labels for some of the things that we observe and dismissing them because "statistical" is laughable. Look at the proof of this: https://github.com/anthropics/formal-math/blob/795efb86f1917... . Forget the Lean, look at the underlying argument construction. At the very least, this is continuing from an argument that was hinted at in the literature in 2024, but these proceedings were difficult enough that humans were not able to do them within two years. Do you attribute this to the harness alone? If so, that's a pretty sophisticated bit of engineering, I would say! Probabilities are far too small to argue infinite monkey theorem. If there was even a shred of a reasonable argument that LLMs were incapable of concept extraction and manipulation, I and my colleagues would be all over it. We would relish in it. It would bring us comfort. It is unbelievable that people think they can spew whatever basic garbage they think of as a gotcha, and think that minds all over the world haven't already considered that. This is like climate denial at this point. | | |
| ▲ | AdieuToLogic 2 hours ago | parent [-] | | > During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don't know what to say. | | |
| |
| ▲ | walrus01 3 hours ago | parent | prev | next [-] | | I'm not anthropomorphizing anything, I literally said that the training data for the formulas and equations is baked into it. It only "knows" things because a crawler and scraper acquired the information from an existing written source. In just about the same way that information is baked into a printed encyclopedia. | | |
| ▲ | hodgehog11 3 hours ago | parent | next [-] | | This is not even remotely accurate. "Baking information" like into a "printed encyclopedia" is memorization. It has been shown, time and time again, that LLMs do not merely memorize. It is not even possible for it to do so at scale. It can memorize some things, yes, but it is forced during the training procedure to bake general concepts into intermediate layers (this is why transfer learning works), analogous to compression. One can make several arguments that compression and intrinisic feature sparsity is the closest mathematical explanation to understanding that we have. | | |
| ▲ | walrus01 an hour ago | parent [-] | | It is completely possible to ask an LLM a series of increasingly more esoteric and discrete questions until you find precisely what information did, or did not make it into the model. If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually. | | |
| ▲ | AdieuToLogic 3 minutes ago | parent [-] | | >>> I'm not anthropomorphizing anything ... Yes you are, regarding LLMs at least. Here's why: just for fun I asked a reasonably smart LLM to ...
[be] capable of understanding if it's gone off on
a hallucinatory path ...
"Smart" in this context is a subjective value judgement.
"Hallucinations" are only experienced by living organisms.You then went on to state: > If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually. Again, "hallucinating" is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information. |
|
| |
| ▲ | astrange 3 hours ago | parent | prev [-] | | No, most of a modern LLM's training time is spent in RLVR, which does not "acquire information from an existing source". You can RL behaviors into a randomly initialized neural network. | | |
| ▲ | hodgehog11 3 hours ago | parent [-] | | This is true, but you're not going to get anywhere. The pretraining phase is necessary to immensely reduce variance in the RLVR stage. Once there, RLVR has a surprising tendency to only restrict the generated space further. This is not true of RLHF, by the way, which I find to be particularly fascinating, but I digress. |
|
| |
| ▲ | jibal an hour ago | parent | prev | next [-] | | Reductionistic fallacy, among other errors. "understanding" is best defined operationally. (Note that teachers and educational institutions test understanding operationally.) | |
| ▲ | Eisenstein 3 hours ago | parent | prev [-] | | > They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. You haven't demonstrated why this matters. > Nothing more. Are you contending that complex systems cannot be more than the sum of their parts? A market is nothing more than offers and counter offers. A ant colony is nothing more than scent trails. All life on earth is nothing more than reproduction with variation. > This still falls under the purvey of statistical token generation. Stating the mechanism does nothing to provide insight into capability. For instance: a nuclear power plant boils water by using fuel rods for heat. What does that tell us about the capability of nuclear power? > This is not "doing" or "understanding" math. Asserting something purely by stating it does not prove anything but that you intuitively believe it to be true. | | |
| ▲ | hardbass an hour ago | parent [-] | | I asked if you believe in souls recently to a few such people and didn't get straight answers. I think it is telling. | | |
| ▲ | jibal an hour ago | parent [-] | | Not responding to a troll isn't telling of anything. | | |
| ▲ | hardbass 23 minutes ago | parent [-] | | How is it a troll? Its a very easy question to answer in the yes or no. Why are people cagey about answering it? |
|
|
|
| |
| ▲ | jacobolus 6 hours ago | parent | prev | next [-] | | You know what also works to get the Karney formula into a program? You can download Charles Karney's free software (MIT license) implementation in several [1] programming languages and then just make a library call – the API is straightforward. If you have comments or questions you can read his several clearly written papers describing the problem, its history, and his algorithm, or you can directly email him: he's a very nice guy, and pretty responsive. [1] https://geographiclib.sourceforge.io/doc/library.html#langua... | | |
| ▲ | walrus01 6 hours ago | parent [-] | | Right, it was really more as a test of how much was contained in the training data set. For my purposes Vincenty is quite accurate enough. This isn't for millimeter level precision land surveying or measurements, but for distance in meters between microwave or millimeter wave band radio sites, point to point links. Even a distance difference of 4 meters plus or minus on a 12 km, 11 GHz band link is going to have no appreciable difference on link budget/reliability calculations, it can be that crude. But not so crude that I just want to throw Haversine at it when Vincenty exists and is not computationally expensive. As this was for a test of "what happens if..." I also watched to see if it did any web searches or external data retrieval to build the test script, and it didn't. I intentionally didn't give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see "hey what if I ask it to do this...". It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math. | | |
| ▲ | jacobolus 5 hours ago | parent [-] | | As an aside: I'm quite convinced that an extremely precise version can be implemented that is significantly faster than Karney's, roughly comparable in speed to simpler naïve approximations. But for most purposes where the precision matters Karney's implementation is not any kind of bottleneck, so it's not clear it's worth spending significant effort on trying to do better. Maybe that's something one of the big LLM companies might want to throw their machines at optimizing if they need to do a lot of geographical calculations. | | |
| ▲ | walrus01 5 hours ago | parent [-] | | One of the places where Karney does become computationally expensive (though still not ridiculous) is a scenario like this, working from a local in-RAM mariadb database that is a copy of the entire FCC radio license database: Draw a 400x400 km size bounding box on a map Find all FDD band plan (high/low split) microwave radio sites in that bounding box Find those sites which have azimuth aim column data which indicates that they are aimed at each other (corresponding halves of a point to point link). Do Vincenty (or Karney) calculation for distance and azimuth between all of them , treating the existing FCC column data for azimuth as suspicious (because it's hand entered by humans) to verify that each independent database rows for each site are actually corresponding halves of a PTP link. Use various other logic to group the successfully matched halves of links together as points A and B of PTP links, and write them out to a geojson file with placemarks and line drawn between them. Multiplied by the number of links that exist in an area like a 400x400km box drawn with Dallas, TX as the center, it's a lot to run through Karney. Actually does result in a lot of CPU load from combined db query due to the size of the db, and Karney calculation. But as I said, Karney isn't necessary, so it's instead implemented as Vincenty. | | |
| ▲ | jonah 4 hours ago | parent [-] | | Interesting project. I'm curious what the purpose is. (Having visited a number of sites with microwave antennas. (But there for VHF and UHF projects.) | | |
| ▲ | walrus01 3 hours ago | parent [-] | | To plan and license a new fdd band plan licensed point to point microwave link you need to first be able to verify the frequencies you want to use are available on a given azimuth and elevation (from the aim direction of the antennas at both ends) and won't conflict with a pre existing licensee. Which means you need data on everything licensed in the area and where it is, how it's aimed, what kind of antenna and gain it has. There's also business and market analysis purposes like knowing what corporate entity has which equipment on top of which tall office towers in a major metro area, and where their links go. |
|
|
|
|
| |
| ▲ | ragall 4 hours ago | parent | prev | next [-] | | > saying LLMs can't do math isn't really a hundred percent accurate anymore It's still accurate. Just because the LLM gave you a corect result doesn't mean it made a calculation. | | | |
| ▲ | Brian_K_White 7 hours ago | parent | prev [-] | | This just exposes that they don't even do the thing you said. Not only is it still true that they can't do math directly, but not even indirectly. They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments. Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to. That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen. It's nothing more than an sql query. | | |
| ▲ | astrange 3 hours ago | parent | next [-] | | GPT-6 can do math directly just fine. Fable apparently can't because they broke its self-estimate of thinking effort. https://x.com/maksym_andr/status/2100364212207837560 | |
| ▲ | bombela 6 hours ago | parent | prev | next [-] | | I don't know for you, but it would take me more than 30s to find and translate the open source code implementing the formulae/algo into small usable program. The more hesoteric the optimisation in the original code, the more time I need. So maybe it is more of a smart completion engine than a SQL answer. | |
| ▲ | walrus01 6 hours ago | parent | prev [-] | | > they found bits of code that are associated with "math" and the supplied arguments How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code? I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result. | | |
| ▲ | AdieuToLogic 5 hours ago | parent | next [-] | | >> they found bits of code that are associated with "math" and the supplied arguments > How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code? Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to... Wait for it... Understanding. | | |
| ▲ | hodgehog11 2 hours ago | parent [-] | | This doesn't make any sense at all. Was this supposed to be a gotcha? An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. The pattern recognition is also particularly compressed into its most sparse and fundamental components, as this is key to generalization. This is not a sensible difference between human and LLM learning, we do the same thing. | | |
| ▲ | UpsideDownRide 2 hours ago | parent | next [-] | | I'll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It's trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues. And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output. So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them. | |
| ▲ | AdieuToLogic an hour ago | parent | prev [-] | | >>> How is this different from a human using an algorithm they have memorized ... >> Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans ... > This doesn't make any sense at all. Was this supposed to be a gotcha? No, it was meant to be an explanation as to the difference between "memorization" and "understanding." In this context, people pick the algorithm they determine applicable and then the question of memorization is relevant. > An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. Funny that you make this argument here, where when I wrote elsewhere in this thread: [LLMs] are statistical token generators whose results are
dependent upon their training data set and involve a
degree of randomness.
Nothing more.
...
It is pattern recognition, a task in which ANNs excel.
To which you replied to the above with: During conversation, we are statistical token generators
whose results are dependent upon our training set.
Seriously, write that definition out rigorously. It
encompasses virtually everything. It is totally
meaningless. So to say "nothing more" is effectively also a
tautology.
This argument was asinine in 2024. It is insane to be
saying these things in 2026. Where have you been?
...
It absolutely understands how to do math, by whatever
reasonable definition you want to provide to the word
"understand".
So which is it?Are LLMs ANNs? Which themselves are pattern recognition algorithms (hint: they are)? OR (setting aside the ad hominems you kindly provided) Do LLMs possess "understanding" of concepts such as abstract mathematics (defined and interpreted by humans) and we, as simple humans, nothing more than statistical token generators as you assert? Because it cannot be both. |
|
| |
| ▲ | noduerme 4 hours ago | parent | prev [-] | | If by "result" you mean the final code, then just asking someone else who understood the math to write it would also have achieved the same result. On the other hand, if by "result" you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it's not the same result at all. I find a lot of the arguments that having LLMs write your code is no different from copy/pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right. |
|
|
|
| |
| ▲ | mbgerring 8 hours ago | parent | next [-] | | They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems. | | |
| ▲ | ericmay 6 hours ago | parent | next [-] | | Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus? > The LLM is not suited to giving deterministic answers to math problems. Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before. | | | |
| ▲ | vanuatu 7 hours ago | parent | prev | next [-] | | reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance) | | |
| ▲ | krapp 6 hours ago | parent [-] | | There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1. | | |
| |
| ▲ | versteegen 8 hours ago | parent | prev | next [-] | | It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.) TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-... | |
| ▲ | GaggiX 8 hours ago | parent | prev [-] | | Reasoning models can do math on their own without external tools. | | |
| ▲ | dcrazy 7 hours ago | parent | next [-] | | Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables. | | |
| ▲ | hodgehog11 2 hours ago | parent | next [-] | | I would say purposeful misremembering. The LLM can be run with zero temperature after all. | |
| ▲ | einichi 6 hours ago | parent | prev [-] | | I don't think anybody is arguing that LLMs do math better than a traditional processor | | |
| ▲ | unshavedyak 5 hours ago | parent [-] | | Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc. The nice thing about math is it can easily plug into a tool, making it even less of a concern. |
|
| |
| ▲ | demibabs 7 hours ago | parent | prev [-] | | Even without reasoning. 5.6 on Instant mode can knock out 3 digit multiplication just fine. |
|
| |
| ▲ | Isamu 8 hours ago | parent | prev | next [-] | | That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component. I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models. | | |
| ▲ | hodgehog11 8 hours ago | parent [-] | | No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago. | | |
| |
| ▲ | tjwebbnorfolk 8 hours ago | parent | prev | next [-] | | They can do math but not arithmetic, which I assume is what the commenter meant | | |
| ▲ | dcrazy 7 hours ago | parent | next [-] | | LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: https://arxiv.org/abs/2410.21272 | |
| ▲ | fasterik 8 hours ago | parent | prev [-] | | I just asked ChatGPT to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right. I'm sure it wouldn't have a 100% success rate, but saying it can't do arithmetic is just false. | | |
| ▲ | Xirdus 7 hours ago | parent | next [-] | | I tried prompt "6379 times 3875" and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1. | | | |
| ▲ | amluto 6 hours ago | parent | prev | next [-] | | I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult. | |
| ▲ | tremon 8 hours ago | parent | prev | next [-] | | Are you sure it honoured your stipulation of "without external help"? For all we know, it hacked its way into Wolfram Alpha and got the result from there. | | | |
| ▲ | guelo 7 hours ago | parent | prev [-] | | [dead] |
|
| |
| ▲ | okanat 6 hours ago | parent | prev [-] | | LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don't need to bring a calculator to count the letters in a sentence. It is a different neural machinery. | | |
| ▲ | fenomas 6 hours ago | parent | next [-] | | You're talking about doing arithmetic; GP was obviously pointing out that "do math" can refer to other things. | |
| ▲ | amluto 6 hours ago | parent | prev [-] | | LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this. |
|
|