My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.