Remix clone Hacker News

new | show | ask | jobs Github

	▲	wavemode a day ago
		He tests several Claude versions as well
	▲	causal a day ago \| parent [-]
		Ah you're right, scrolled past that - the most salient contrast in the chart is still just GPT-5 vs GPT-4, and it feels easy to contrive such results by pinning one model's response as "ideal" and making that a benchmark for everything else.