| â–˛ | stingraycharles 2 hours ago | |
Back in the day - maybe two decades ago - I implemented language detection like this. I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages. I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language. Absolutely totally not the best approach, but very fast and super simple to implement. | ||