Remix.run Logo
▲ hbrn an hour ago

And what do you think Jev confidence score is?

Here's a hint: confidence is not generated by a model.

▲adrian17 15 minutes ago | parent | next [-]

Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.

▲jgilias an hour ago | parent | prev [-]

Thanks, fixed my understanding!

Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?

▲hbrn 19 minutes ago | parent [-]

Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration.

Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).

Unfortunately calibration is hard to benchmark.