Remix.run Logo
▲ drhagen 3 hours ago

This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].

[0] Windows newlines not withstanding.

▲bruce511 2 hours ago | parent | next [-]

Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better.

As it is programs have to guess the following;

A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?

B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?

C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?

D) what human-language is it in?

E) CSV? Don't get me started...

Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.

Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.

But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...

▲hnlmorg 2 hours ago | parent | next [-]

ASCII is just a raw binary format. It isn't just for text.

The line endings problem really isn't a problem. Pretty much every text editor out there can handle different line endings.

I don’t think number nor date formats are relevant here. For example you could have that same problem entering text into MS Word. That’s really more of an issue if you want to use text as a database rather than a document format, which isn’t really something that even plain text advocates would generally recommend.

As for human language detection, that’s a much easier problem to solve than decoding a proprietary binary blob.

CSV definitely has its warts. But it’s not like that’s the only plain text option for serialising data. (JSON, jsonlines, YAML, XML, etc). Or you could use the actual ASCII codes reserved for records, if you really wanted something that didn’t require quoting and escaping in plain text. It’s actually a pity nobody does this.

▲dwaltrip 15 minutes ago | parent | prev | next [-]

The entire point of plain text is that it doesn’t have number formats, date formats, and so on.

This is the eternal tradeoff of specificity vs. generality.

▲conartist6 2 hours ago | parent | prev | next [-]

I've made a language called CSTML to do some of this: https://docs.bablr.org/guides/cstml

Basically it's a text-based language for embedding semantic metadata into some kind of underlying text stream. Is this interesting to you?

▲jimbokun 2 hours ago | parent | prev [-]

[flagged]

▲layer8 2 hours ago | parent | prev | next [-]

> Windows newlines not withstanding

“Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard.

I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports.

Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign.

▲hnlmorg 2 hours ago | parent [-]

UNIX (and Linux) is even more annoying because pseudo TTYs will require CRLF when in raw mode but requires only LF when in line mode.

▲tyromaniac 2 hours ago | parent | prev | next [-]

And file endings..

▲Analemma_ 2 hours ago | parent | prev [-]

If you write files in ASCII you’re already writing in UTF-8.

▲sillysaurusx 2 hours ago | parent | next [-]

That’s true, though if you’re writing UTF-8 you have to handle the corner cases when reading UTF-8, of which there are many.

Fortunately they’re easy to test for and most languages have standard libraries that make this painless.

▲bell-cot 2 hours ago | parent | prev [-]

ASCII has, in principal, infinitely many superset.

And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII