Remix.run Logo
▲ jgilias 2 hours ago

Isn’t the OpenAI decisions API basically just Luna cosplaying a decisions model and pretending the confidence score isn’t just a hallucination?

▲hbrn an hour ago | parent | next [-]

And what do you think Jev confidence score is?

Here's a hint: confidence is not generated by a model.

▲adrian17 14 minutes ago | parent | next [-]

Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.

▲jgilias an hour ago | parent | prev [-]

Thanks, fixed my understanding!

Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?

▲hbrn 19 minutes ago | parent [-]

Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration.

Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).

Unfortunately calibration is hard to benchmark.

▲phalangion 2 hours ago | parent | prev [-]

What’s the difference?