Remix.run Logo
qingcharles 17 hours ago

Not everyone should scan stuff. If you spend any serious time looking through stuff that randos on the Internet have scanned the quality fits the Bell Curve perfectly.

Biggest problems:

  - scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
  - becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.

Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.

Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:

https://imgur.com/a/vgMkM7b

(preview size, they sent a 500MB TIFF)

Now I can reassemble the issue and upload it.

I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.

I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.