| ▲ | oersted 4 hours ago | |
That’s a good visualization, although I am a bit mistrustful of Arena’s scores. It does get around the fact that models are getting trained for the benchmarks, but the methodology of letting random people compare outputs side-by-side is a very shallow judgement method in my opinion. EDIT: Indeed looking at the overall rankings for text again, the list is rather strange, a lot more about writing style than intelligence. | ||