Remix.run Logo
▲ acgourley 3 hours ago

Very cool.

I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.

Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

▲yannis 3 hours ago | parent | next [-]

>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.

▲hypfer 2 hours ago | parent | prev [-]

Please just make sure to keep the ethical implications of any such work in mind.

I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.