| ▲ | drhagen 3 hours ago | |||||||||||||||||||||||||
This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0]. [0] Windows newlines not withstanding. | ||||||||||||||||||||||||||
| ▲ | bruce511 2 hours ago | parent | next [-] | |||||||||||||||||||||||||
Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better. As it is programs have to guess the following; A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian? B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right? C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy? D) what human-language is it in? E) CSV? Don't get me started... Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent. Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues. But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all... | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | layer8 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
> Windows newlines not withstanding “Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard. I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports. Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | tyromaniac 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
And file endings.. | ||||||||||||||||||||||||||
| ▲ | Analemma_ 2 hours ago | parent | prev [-] | |||||||||||||||||||||||||
If you write files in ASCII you’re already writing in UTF-8. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||