| ▲ | petercooper 4 hours ago | |
Smaller models have been able to do these sorts of tasks, but a little slower, for a while now. Give a small Qwen 3.8 model a classification task and force a structured output, and it'll do a good job. I've used Qwen 0.8b for basic image classification in <500ms on my local machine for a while now. There are a few technical details that can reduce the latency significantly (covered in the post) but the real insight has been from watching the reaction to Jev and seeing that there's enough of a market interest to offer it as a distinct thing. The underlying concept/approach was already there. | ||
| ▲ | theapadayo 4 hours ago | parent [-] | |
Not just structured output. Dropping down to logprobs, prompting the model to emit one word as the answer, and then ranking the output tokens to pick your answer works great on small Qwen & Gemma models. The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer. | ||