So YES if you're using Cowork or Chat
No, because they are cached, the inference cost is paid once per model, does not scale linearly per user or use.