Remix.run Logo
▲ porridgeraisin 4 hours ago

They are not too difficult to train if you already have infra to train regular LLMs. You can typically replace a few layers train them alone and you're off to the races.

Getting training data that works well for calibrated classification objectives is difficult.

I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.

But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.

So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.