Remix.run Logo
zmmmmm 9 hours ago

> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image.

It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)

mkagenius 9 hours ago | parent [-]

Can split and feed?

throwaw12 9 hours ago | parent [-]

that's difficult as well, how do you k ow where to split?

johndough 8 hours ago | parent | next [-]

There are models specifically for splitting an image into text regions, e.g. PP-DocLayoutV3 https://huggingface.co/PaddlePaddle/PP-DocLayoutV3

I am using a stripped-down minimal version of it which I uploaded here, since I am not a fan of huge dependency trees: https://github.com/99991/simple-pp-doclayoutv3

Another recent model for this task is Unlimited-OCR: https://github.com/baidu/Unlimited-OCR

kgwgk 8 hours ago | parent | prev | next [-]

Text is often written as separate lines (and paragraphs) at least in some languages.

wongarsu 8 hours ago | parent | prev | next [-]

Let the model do the splitting. A 800x800px image should be enough to make those decisions

grog454 8 hours ago | parent | prev | next [-]

Overlap the splits?

vrganj 8 hours ago | parent | prev [-]

Presumably a small cheap model could do that part?