What about do you mean by single threaded? Each token is predicted by using parallel computation on the GPU.
[delayed]