Remix.run Logo
▲ blakeashleyjr 5 hours ago

This sounds like the Postgres vs. InnoDB argument 10 years later. Postings pointed at physical location (the ANN slot), so every SPFresh rebalance rewrote every index touching that doc. InnoDB solved this by pointing secondary indexes at the PK and eating an extra lookup on read. Curious what that extra lookup costs you when it's an S3 GET instead of a B-tree hop.

"Updating one vector can move hundreds of attributes and their indexes" is basically Uber's 2016 Postgres write amplification post, but for search. Same fix too: stop pointing indexes at where the row lives.

So ANN becomes a secondary index that points at a doc ID, and vector search now needs a hop to complete. Do clusters keep their own copy of the vectors so the search itself stays local, and only result fetch pays the indirection? Otherwise cold p99 seems like it gets worse.

▲alfiedotwtf an hour ago | parent [-]

I haven’t read it, but “ stop pointing indexes at where the row lives” sounded interesting. So if not the row, what does the index point to instead?