Remix clone Hacker News

new | show | ask | jobs Github

	▲	sigbottle 10 hours ago
		What would unsupervised mean, would unsupervised be something like alphago playing against itself trillions of times? Whereas self-supervised, allows learning without explicit annotation of data ; but it doesn't matter if the models already trained on the entire Internet, and it's not like a game where it can come up with effectively new training data for itself?
	▲	jmalicki 5 hours ago \| parent [-]
		Unsupervised is basically clustering. Alphago is RL - winning or losing a game is a form of supervision. Unsupervised is something where there is no intrinsic reward signal. In pre training, predicting the next token and seeing that it matches is a reward signal, hence it is self supervised.