| ▲ | Gigatoken: Fastest Tokenizer(twitter.com) | |||||||
| 25 points by convexstrictly a day ago | 6 comments | ||||||||
| ▲ | convexstrictly a day ago | parent | next [-] | |||||||
| ||||||||
| ▲ | ramon156 17 hours ago | parent | prev | next [-] | |||||||
Doesn't really explain why its faster. "Optimizing for every kind of CPU" is not really enough info. Cool project nonetheless, I will go through the code later tomorrow | ||||||||
| ▲ | benj111 18 hours ago | parent | prev | next [-] | |||||||
I kind of assumed the model would process the text 'directly', from what I understand, wouldn't this be biasing the input based on how you tokenise as it's lossy? I assume this tradeoff is purely for speed/compression. Or am I missing what's going on here? | ||||||||
| ||||||||
| ▲ | doosdom 21 hours ago | parent | prev [-] | |||||||
[dead] | ||||||||