Remix.run Logo
SiempreViernes 10 hours ago

Is it? I'd expect most of the training set to be synthetic data extrapolated from a small set of human authored texts.

TeMPOraL 6 hours ago | parent [-]

Most of the training set is half of the Internet. LLMs are pre-trained on general set of human biases and patterns of thinking.