| ▲ | Perseids 2 hours ago | |
You seemed to be deeply confused about encodings and character sets. > The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text. That is only true for the mapping of Unicode character (code point to be exact) to UTF-8, which is an encoding of Unicode characters. > Unicode are multibyte characters with variable byte length and endianess at play. That is only true of the UTF-16 encoding. UTF-8 does not have endianess, UTF-32 does not have variable length per code point. None of that is true for Unicode, because it is abstracted away from any byte representation. Furthermore, getting back to the start: > ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset. ASCII also depends on how you try to decode your text. If you interpret two bytes as one character, ASCII will never be correct. There is nothing magical about one byte mapping to one character. I'd even argue that the only reason ASCII support is so universal internationally is because of UTF-8. Otherwise many countries would default to encodings where ASCII is not a subset (as they did before UTF-8 became common). So IMHO, UTF-8 and Unicode are the only widely understood encoding and character set. | ||