Remix.run Logo
marcelroed 9 hours ago

Time to first token refers to the time until the model outputs one token, which includes the time to process the entire prompt (doing prefill). The GPU time per token is much lower when doing prefill, so the significance of tokenization is higher.

dingdingdang 8 hours ago | parent [-]

Have you done preliminary numbers on replacing tokenizer on, say, llama-server?

marcelroed 7 hours ago | parent | next [-]

Added numbers here: https://news.ycombinator.com/item?id=49015014

marcelroed 7 hours ago | parent | prev [-]

Running the numbers now