| ▲ | famouswaffles 6 hours ago | |
It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization. Byte Latent Transformer - https://arxiv.org/abs/2412.09871 1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains. | ||