| ▲ | SiempreViernes 10 hours ago | |
Is it? I'd expect most of the training set to be synthetic data extrapolated from a small set of human authored texts. | ||
| ▲ | TeMPOraL 6 hours ago | parent [-] | |
Most of the training set is half of the Internet. LLMs are pre-trained on general set of human biases and patterns of thinking. | ||