Remix.run Logo
▲ theapadayo 4 hours ago

Not just structured output. Dropping down to logprobs, prompting the model to emit one word as the answer, and then ranking the output tokens to pick your answer works great on small Qwen & Gemma models.

The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.