Remix.run Logo
eviks 4 hours ago

> The original character (𡚴) was not added to JIS or Unicode until much later and doesn't display on most sites for me

Why didn't they simly replace the original bad one?

> nine hundred pages. Imagine tracking down a single character without a page reference

Not that hard to imagine, OCR existed back then?

gucci-on-fleek 3 hours ago | parent | next [-]

> Not that hard to imagine, OCR existed back then?

How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.

Izkata an hour ago | parent [-]

(Talking about Japanese here)

I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.

For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).

IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.

Kye 4 hours ago | parent | prev [-]

OCR was slow and unreliable and was for a very long time.

eviks 3 hours ago | parent [-]

"slow" - wasn't like they were pressed for time. It took them almost 20years to even start the investigation! Also not certain you needed that much reliability compared to what was available to match one pic to a quality scan to get a reasonable number of candidates