Remix.run Logo
gravypod 15 hours ago

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?

xyz100 12 hours ago | parent | next [-]

Presumably there is distillation or similar being used to transfer from a larger model to a smaller one.

scotty79 2 hours ago | parent | prev [-]

They innovated a lot.