I am not sure the author of the comment you are replying to understands that LLM systems have prompt caches