| ▲ | brokencode 4 hours ago | |||||||
Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM. It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings. | ||||||||
| ▲ | tito 3 hours ago | parent [-] | |||||||
Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation. | ||||||||
| ||||||||