| ▲ | robotresearcher an hour ago | |
Why? LLMs (along with other DNNs) model the distribution seen in their training data. Does the training data have dice roll examples being mainly 1? Maybe so! If that’s the case it’s an interesting example of LLM fragility since it’s failed to reason from the many (millions of?) times it’s seen stated in training data that each outcome has probability 1/6. | ||