Remix clone Hacker News

new | show | ask | jobs Github

	▲	gpjt 3 days ago
		OP here: one thing that surprised me in this experiment was that the model trained on the more curated FineWeb-Edu dataset was worse than the one trained on FineWeb. That is very counterintuitive to me.