| ▲ | noahbp 10 hours ago | |||||||||||||||||||||||||||||||
Time to first token, especially for smaller models, can be sharply reduced. Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras. | ||||||||||||||||||||||||||||||||
| ▲ | fastball 10 hours ago | parent [-] | |||||||||||||||||||||||||||||||
Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||