| ▲ | gdiamos 3 hours ago | |
I wish I could get a model to state its assumptions. | ||
| ▲ | reichstein an hour ago | parent [-] | |
Models do not have assumptions. They have probabilities for what the next token should be. With enough context, in the context window and built into the model, that next token isn't completely random, it's correlated with something someone might choose to write. But people write all kinds of crap manually. So far, the data people have been writing has tended to be denser around what people could agree on (there are many lies, but only one truth), so the model is more likely to go there. If we start putting AI generated text into the training data, it's not clear what that means for the resulting model. It's already clear that some actors are trying to influence models by putting large amounts of content out there that agree with them. Figuring out which content is safe to train from is the real problem for future model trainers. | ||