| ▲ | When Compilers Disagree About UTF‑8(nemanjatrifunovic.substack.com) | ||||||||||||||||||||||||||||
| 32 points by rbanffy 5 days ago | 6 comments | |||||||||||||||||||||||||||||
| ▲ | MiroslavPokorny 27 minutes ago | parent | next [-] | ||||||||||||||||||||||||||||
Title is misleading, the compilers do not disagree, running the compiled code produces the same output. The difference is one binary(instructions) are slightly different... | |||||||||||||||||||||||||||||
| ▲ | kstenerud 5 days ago | parent | prev [-] | ||||||||||||||||||||||||||||
You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||