| ▲ | applfanboysbgon 13 hours ago | |
Prediction, by definition, can extend beyond what has been literally seen in the training data. With the tokens "2 + 2 = ", the overwhelming prediction is going to be 4, but with enough samples, you can generalise the prediction to apply to more numbers. However, that is all that it is - a prediction. Humans are capable of engaging in prediction, using heuristics as a method of conserving mental energy, because always engaging in full logical reasoning would be a waste of the body's resources. However, humans can also follow a set of logical rules and arrive at their conclusion deterministically, something which is completely outside of an LLM's programming. I don't really care to publicly write about my tests because they will become training targets and not be usable for future internet arguments anyways, but there are a great number of trivial 2~3 sentence logical prompts that will completely fuck an LLM's prediction algorithm and result in incoherent replies that a human, or really anything with a theory of mind, would never generate. Not that a human would always answer correctly on the first try, but the failure methods happen to be completely different, eg. Sol will short-circuit and repeat the prompt verbatim (when the instructions don't remotely suggest doing anything of that nature), even on Max. Prediction can superficially resemble reasoning when there's sufficient training data, but it breaks down severely when confronting a task that is OoD. | ||
| ▲ | 9 hours ago | parent [-] | |
| [deleted] | ||