| ▲ | robrorcroptrer 2 hours ago | |||||||||||||
What about splitting bigger content into chunks before embedding? | ||||||||||||||
| ▲ | freakynit 2 hours ago | parent | next [-] | |||||||||||||
How are you gonna handle the relations that span across individual chunks... if a later chunk refers something from 2 chunks before using `it`, rather than proper name, how will you handle that? Because at query time, that later chunk would not match. | ||||||||||||||
| ||||||||||||||
| ▲ | mdp2021 37 minutes ago | parent | prev [-] | |||||||||||||
What member freakynit said nearby about chunks and relations between chunks, plus the storage and information efficiency problem: make some calculations about storing vectors - for paragraphs and for collections of paragraphs -, then compare the needed space with the original data... Because you could have clever ideas about vectors related to more paragraphs related in the document structure - but that would multiply the vectors. The index can become much bigger than the corpus. | ||||||||||||||