| ▲ | numpad0 an hour ago | |
Latin alphabet transcription of East Asian languages tend to be incredibly lossy and there tends to be a lot of ambiguities. That Vietnamese surname of Nguyen is formally written as Nguyễn in the current script, or as 阮 in the old adapted Hanzi script, and practically pronounced somewhere around "gwen". Most other EA countries tend to do equivalents of "阮" -> "gwen", and East Asian cultures generally prefer Last-First ordering, which tend to put more information in the first name, and so there are going to be tons of last names in Latin that end up being barely unambiguous to be a form of identification. These can be a problem if email addresses are going to be formulated as e.g. "阮 必成" -> "Nguyễn Tất Thành" -> "Gwen Tat Than" -> "t.gwen@example.net", even if the surname is going to be that much long(although Nguyen in particular is just identical even in Vietnamese). (Japan may be less affected thanks to having default 2+ char random surnames farmers made up in modernization efforts, but there are still the 灘(nada) highscool, the 津(tsu) city, etc etc that can be NLP problems) | ||