| ▲ | swiftcoder 11 hours ago |
| So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table? |
|
| ▲ | fastball 10 hours ago | parent | next [-] |
| The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer. |
| |
| ▲ | janwas an hour ago | parent | next [-] | | hm, maybe not so trivially correct here. Do I understand correctly that incorrect results can happen as a result of a 42-bit hash collision?
That could happen after less than one MB of input, given the simple one-mul hash. BTW throughput is measured for a 12 GiB file. Would be interesting to see the throughput for something more like 32 KiB, with cold start (token cache not yet populated). | |
| ▲ | michaelmior 7 hours ago | parent | prev | next [-] | | Some changes certainly can be. If the model produces the exact same output for a fixed seed across a variety of inputs after a code change, I think it's reasonable to expect that the change is correct. There are also mathematical transformations that can be applied in some cases that are provably correct. (Not suggesting there's necessarily anything of this nature that will lead to 1,000x improvement though.) | |
| ▲ | nixon_why69 9 hours ago | parent | prev [-] | | Eh, linear algebra changes are still easy to measure correctness, it's just that you're competing with 50 years of research for most of them, less low hanging fruit. |
|
|
| ▲ | ProofHouse 10 hours ago | parent | prev | next [-] |
| the answer is many! This would take hours to write. Full teams and research on nearly every part. So many 'unlocks' coming. |
|
| ▲ | parineum 10 hours ago | parent | prev [-] |
| I'm sure there's been a lot more effort put into the other, more consequential, portions of inference time. |