Remix.run Logo
▲ minimaxir 11 hours ago

Finally. I was getting annoyed that there's been an inflection point in how LLMs/agents work but there hasn't been a good moderate-size embeddings model, and this one is multimodal too! 270M for text only is great compared to older embedding models, and a total 440M for text + vision is also fair.

I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.

▲alberto467 7 hours ago | parent | next [-]

Not just vision with video, but also audio, it really seems amazing.

I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.

▲treetalker 5 hours ago | parent [-]

> Building local multimodal search with this would be amazing.

(lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?

▲reacharavindh 4 hours ago | parent [-]

The embedding model stays loaded in memory. It is used for turning your search keywords into embeddings.

The index you’re thinking of is made once per file.. then you compare and search in embeddings. Add/modify files = asynchronous updating or adding corresponding embeddings using the model in memory onto wherever you persist those embeddings(say SQLite)..

▲alberto467 3 hours ago | parent [-]

Also how you turn a file into one or more items is a separate question and would likely need tuning to circumstances, usually big documents are split (chunked) sometimes at paragraph or even more granularly. Where to optimally chunk alone is not easy. This also allows you to then search for a specific part of the document, at the expense of not taking the wide context into account, but embeddings usually struggle with too many tokens anyway.

▲minimaxir 4 hours ago | parent | prev [-]

Numbers update on my hypothetical tool after Opus 5.5 added support for image and audio input (plus video support via ffmpeg, why not) using my M3 Pro:

- Text: 78 embeddings/second on small texts

- Image: 4 embeddings/second

- Audio: 6 embeddings/second for 30 second chunks (which is how the encoder works)

- Video: 0.2 embeddings/second per minute of video (not entirely surprising since the encoder does 1fps).

Not bad numbers for a local laptop running a model this large, although the M3 Pro is a few years old and I suspect a M5 Ultra will 5x them at minimum.