When Compilers Disagree About UTF‑8

(nemanjatrifunovic.substack.com)

4 points | by rbanffy 7 hours ago ago

1 comments

  • kstenerud 6 hours ago ago

    You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop.

    Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.