| ▲ | boomlinde 2 hours ago | |
The article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains. For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS). There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly. Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes. | ||
| ▲ | adiabatichottub an hour ago | parent | next [-] | |
Text is wonderful, but note that the one keyword not found in this text is 'parse'. How much extra work is generated writing parers for text that would have been so much easier to deal with if it had just been generated in a structured binary format? You still get the pleasure of designing a new wheel with every format. | ||
| ▲ | gwbas1c an hour ago | parent | prev | next [-] | |
This morning I had co-pilot generate a LINQPad script that fixes a lot of that. Granted, I knew that all files are UTF-8, so I didn't need to have it "guess" what the encoding was without the BOM. | ||
| ▲ | Liquid_Fire an hour ago | parent | prev [-] | |
Sorry, can you clarify about Microsoft having adopted \n? This is the first time I hear of it, and I can't find anything online about it. | ||