| ▲ | reverius42 6 hours ago | |
It's been a while now that for "thinking" or "reasoning" models, most of the tokens generated are "thinking" tokens, and depending on what goes into that "thinking" token stream, it "decides" whether and how many output tokens to produce that the user actually receives as output. It's a bit more sophisticated than just "what's the next token" in a tight loop. Anthropomorphizing words in scare quotes for those who don't appreciate attributing thinking to machines. | ||