Remix.run Logo
navigate8310 3 hours ago

What stops the LLM not to iteratively think and expound upon before emitting the final tokens?

samatman 3 hours ago | parent [-]

Nothing at all. They're not designed to, so they don't. Change that, and they would.

The question is the wrong one. The right question: why aren't frontier models designed to work that way? The answer: it's slow and expensive.

The other answer: that's basically what you're selecting with "Medium", "High" and so on, how many tokens they'll blow on muttering to themselves before they get back to you with an answer. There's more to it, but not that much more.