I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter.
And before that, businesses will be able to get decent results with dedicated inference hardware.