Remix.run Logo
▲ mbgerring 8 hours ago

They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.

▲ericmay 6 hours ago | parent | next [-]

Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?

> The LLM is not suited to giving deterministic answers to math problems.

Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.

▲queenkjuul 2 hours ago | parent [-]

Pretty sure I've had 5x6 memorized accurately since i was 7 years old

https://x.com/maksym_andr/status/2100364212207837560

▲Dylan16807 an hour ago | parent [-]

I can't tell if you're joking but they're talking about 5 digit and 6 digit numbers.

▲vanuatu 7 hours ago | parent | prev | next [-]

reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)

▲krapp 6 hours ago | parent [-]

There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.

▲vanuatu 4 hours ago | parent | next [-]

functionally, LLM behavior indeed shares many parallels with human cognition.

▲Forgeties79 6 hours ago | parent | prev [-]

Law of Krap

▲versteegen 8 hours ago | parent | prev | next [-]

It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)

  TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
  next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
  pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...
▲GaggiX 8 hours ago | parent | prev [-]

Reasoning models can do math on their own without external tools.

▲dcrazy 7 hours ago | parent | next [-]

Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.

▲hodgehog11 2 hours ago | parent | next [-]

I would say purposeful misremembering. The LLM can be run with zero temperature after all.

▲einichi 6 hours ago | parent | prev [-]

I don't think anybody is arguing that LLMs do math better than a traditional processor

▲unshavedyak 5 hours ago | parent [-]

Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc.

The nice thing about math is it can easily plug into a tool, making it even less of a concern.

▲demibabs 7 hours ago | parent | prev [-]

Even without reasoning.

5.6 on Instant mode can knock out 3 digit multiplication just fine.