| ▲ | Better Vector Search for Long Documents: Chunking Inside Manticore Search(manticoresearch.com) | ||||||||||||||||||||||||||||
| 65 points by GloriaVinogrado 6 hours ago | 10 comments | |||||||||||||||||||||||||||||
| ▲ | entrope 4 hours ago | parent | next [-] | ||||||||||||||||||||||||||||
A lot of the article focuses on problems induced by a 512-token input limit. For example, one needs a lot more chunks with such a small input, especially with overlap. I realize that some embedding models do have input contexts that small, but 8K and 32K are fairly widely supported and reduce chunking-related problems. For languages like English, there's also usually a lot of redundancy within a text, so 512 tokens might not give a very clear indication of the context. Lots of documents have similar introductions (like "#include <foo.h>\n") that make short contexts and truncation particularly harmful. Also, "Nothing in the document past that point can ever be retrieved, and nothing anywhere told you." This is user-hostile behavior, even if they didn't want to admit to users that the auto-embedding support was poor. Finally, the paragraph later on about truncation being "what you already have" reads like Claude talking to the developer, not like a vendor talking to users. But sure, maybe this is a good default for a database searching page titles, chat logs and Xeets? | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | hn45e7pbij 3 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Bigger context windows help but they don't remove the need to chunk. Embedding 8K tokens into one vector smears everything, retrieval quality drops even though nothing got truncated. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | Chance-Device 3 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Pretty interesting, I’m sure it will be useful for anyone who is rolling their own RAG. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | mistrial9 3 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
BGE-M3 has an input window of 8,192 tokens | |||||||||||||||||||||||||||||