| ▲ | cschmidt 3 hours ago | ||||||||||||||||
Can I say this seems to be fantastic work. I cloned your repo earlier today after seeing it on the tokenization discord. I know everyone in the tokenization community wants to absorb the lessons of how you got such a speedup. The caching and replacing the regex for pretokenization seem like generally useful ideas. And screw all the 0.1% haters on here, this is great stuff. | |||||||||||||||||
| ▲ | marcelroed an hour ago | parent | next [-] | ||||||||||||||||
Thanks for the kind words, Craig! I'm planning to do a technical writeup+paper and a presentation video on the project in the near future. Will make sure to share it with the Discord! | |||||||||||||||||
| |||||||||||||||||
| ▲ | gghfez 17 minutes ago | parent | prev | next [-] | ||||||||||||||||
>the tokenization discord Could I join this? | |||||||||||||||||
| ▲ | cs702 2 hours ago | parent | prev | next [-] | ||||||||||||||||
That is my reaction too. It looks like great work! Valuable not only for inference, but for training too (think proprietary datasets). I would add, a single individual did this. One person can make a difference :-) | |||||||||||||||||
| ▲ | victor106 an hour ago | parent | prev | next [-] | ||||||||||||||||
> tokenization discord How can I join this? Sounds interesting | |||||||||||||||||
| |||||||||||||||||
| ▲ | asdf88990 2 hours ago | parent | prev [-] | ||||||||||||||||
[flagged] | |||||||||||||||||
| |||||||||||||||||