Kinda liking this idea :)
I’ve not played with vision models much. I did some experiments with OpenCV and Processing many years ago, but times have changed.
I think you just wrote me next weekend project ;)