Remix.run Logo
fastball 8 hours ago

The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.

michaelmior 5 hours ago | parent | next [-]

Some changes certainly can be. If the model produces the exact same output for a fixed seed across a variety of inputs after a code change, I think it's reasonable to expect that the change is correct. There are also mathematical transformations that can be applied in some cases that are provably correct. (Not suggesting there's necessarily anything of this nature that will lead to 1,000x improvement though.)

nixon_why69 7 hours ago | parent | prev [-]

Eh, linear algebra changes are still easy to measure correctness, it's just that you're competing with 50 years of research for most of them, less low hanging fruit.