Well, in theory, with proper UTF-8 support you don't need to filter out anything.
Unsupported unicode code points are just rendered using a place holder.