Remix.run Logo
fweimer 2 hours ago

Centralized inference can easily increase batch size, leading to huge efficiency gains in the usual scenario where most users have just one or very few session. Using local resources efficiently requires some way to increase the batch size. I'm not sure if we are there yet.

janalsncm an hour ago | parent [-]

I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter.

And before that, businesses will be able to get decent results with dedicated inference hardware.