Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.