Remix.run Logo
Dwedit 2 hours ago

FF bytes are an easy way to identify an invalid UTF-8 file. This idea doesn't have that property.

sph 2 hours ago | parent | next [-]

True, but not all non-UTF8 bytestrings contain 0xFF bytes, so it’s not very useful in practice.

da_chicken 2 hours ago | parent [-]

Yes, I agree.

It's more common for programs that say they support UTF-8 to not really do so at all. It wasn't that long ago that "UTF-8" support was often just single byte, so it was little more than ASCII. Even now it's common for programs to choke on the optional BOM. Yes, it is redundant, congratulations. The spec still explicitly allows it. Three and four byte character support is still not the best, too.

flohofwoe an hour ago | parent [-]

> "UTF-8" support was often just single byte, so it was little more than ASCII

"Single byte UTF-8" is ASCII. That's one of its most important properties.

> Even now it's common for programs to choke on the optional BOM

And they should... BOMs (and especially the hilarious UTF-8 BOM) are strictly a legacy Microsoft/Windows thing and should be abolished along with "extended" 8-bit ASCII encodings and UCS-2/UTF-16 (only UTF-32 makes sense, but should only be used at runtime to allow random access on UNICODE code points, but not for data exchange.

flohofwoe 2 hours ago | parent | prev | next [-]

It's still a joy to see how frigging elegant and extensible the UTF-8 specification is. And even without the esoteric 0xFF lead byte, the regular UTF-8 encoding with a 0xFE lead byte (11111110) would still have plenty of headroom (36 bits) compared to the current 21 bits for UNICODE.

beeforpork an hour ago | parent | prev [-]

As are FE, FD, FC, FB, FA, F9, F8, F7, F6 and F5.