| ▲ | hodgehog11 an hour ago | |
This is not even remotely accurate. "Baking information" like into a "printed encyclopedia" is memorization. It has been shown, time and time again, that LLMs do not merely memorize. It is not even possible for it to do so at scale. It can memorize some things, yes, but it is forced during the training procedure to bake general concepts into intermediate layers (this is why transfer learning works), analogous to compression. One can make several arguments that compression and intrinisic feature sparsity is the closest mathematical explanation to understanding that we have. | ||