Remix.run Logo
fpaf 16 hours ago

Even assuming that inference is really profitable today's models don't learn new things by themselves, except in the limited sense of temporarily storing everything they need for a conversation in their context (and maybe leaving themselves little notes in md files like the guy from "Memento").

The difference with tyres is that If today's LLMs had been invented and had become "good enough" 50 years ago, you would have a cutover in their knowledge that excludes 50 years of information. Every "write me a program in Rust" conversation would involve LLMs filling up their context trying to learn Rust programming from scratch every time and probably doing a very bad job.

An example of that was when Fable disproved that mathematical conjecture and HN was full of other people who fed that information to other models (or Fable itself) and received incredulous answers from their LLMs. If something is proven true or false in math, the world of science moves on and that new piece of information can be used to build, prove or disprove other things. But an LLM is excluded from learning even from the very thing it just helped demonstrate and needs to re-discover it over and over again. In order to have an LLM that "lives" in a world where the Jacobian conjecture is false, you need to train a new model and add that information.