Remix clone Hacker News

new | show | ask | jobs Github

	▲	twotwotwo 21 hours ago
		Yeah--it's difficult to go from a benchmark involving the model attempting things alone to the effect assisting people on real tasks because, well, ideally you'd measure that with real people doing real tasks. Last time METR tried that (in early '25) they found a net slowdown rather than any speedup at all. Go figure!