Sidebar: single threaded inference isn’t good enough anymore
What about do you mean by single threaded? Each token is predicted by using parallel computation on the GPU.
[delayed]