| ▲ | seizethecheese an hour ago | |||||||
These results don’t just contradict more serious benchmarks, they are wrong on an entirely different axis. This is a saturated benchmark. Haiku gets 96%. The results here are “not even wrong” and this being #1 on HN right now is a massive smell of either bots or massive ignorance or both. | ||||||||
| ▲ | urams an hour ago | parent [-] | |||||||
People REALLY want the open models to be better than the frontier labs'. | ||||||||
| ||||||||