Remix clone Hacker News

new | show | ask | jobs Github

	▲	zozbot234 an hour ago
		True, ARC is mostly an artificial "human-like AGI" benchmark that doesn't really reflect any plausible workload. Very different from things like Humanity's Last Exam that reflect real-world knowledge and are now getting closer and closer to saturation even with open models.