| ▲ | jgilias an hour ago | |
Thanks, fixed my understanding! Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs? | ||
| ▲ | hbrn 18 minutes ago | parent [-] | |
Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration. Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one). Unfortunately calibration is hard to benchmark. | ||