| ▲ | imtringued a day ago | |
If the AI companies ever get the idea that it would make sense for the AI model to modify its KV cache, it would end up adding a write step to the KV cache per token (moderately expensive) and increase the cost of inference by the width of the update (arbitrarily expensive). This in itself could double or triple the demand for both memory bandwidth and compute. I honestly don't see a future where AI demand disappears... | ||