Remix.run Logo
▲ amluto an hour ago

Two reasons:

1. It’s ridiculous. Not only is the question prima facie absurd for a model of this type, it sort of doesn’t fit into the whole training model. An LLM (charitably) predicts token probabilities, which one might generalize to mean that the LLM operates on probability distributions over strings. So asking for the probability of “fraud” versus “not fraud” makes sense. But asking for the probability of “1.23” versus “2.7” is kind of out of distribution - those are numbers, no one is training on an entire continuum of two-decimal-place real numbers, and similar numbers can have wildly different representations (“25.4” vs “25.40” vs “2.54e1”).

2. It’s barely a classification problem as written. If I wanted it to be a classification problem, maybe I would try:

“There is a machine that receives little sealed containers of air. In each container one molecule is painted red. The machine measured the interior temperature of one particular container and determined that it was 300K.” Question: in what range was the velocity of the red molecule at the instant that the container entered the machine. Choices: 0-100m/s, 100-200m/s, etc.

I maintain that this question is a weird thing to train a Jev-like model on and that I don’t think that a classifier I use would need to answer it well.

I do find it disappointing that Jev conflates “the probabilities are all equal” with “I have no clue”, and I think it would be better if it were at least clearly documented how the model’s ability to figure something out relates to the API response probability.