Remix.run Logo
▲ spottedmarley a day ago

I built my own benchmarking arena that tests local models on all of things the I need a model to do well. I don't look at any of the existing benchmark data that is out there. When a new model drops, I run it through my arena and see how it compares to previous models. If I talk about a model being good I am referencing my own accumulated knowledge on how a model performs for me on tasks that I care about. I generally will never be heard talking negatively about a model (except maybe a frontier/hosted model, they all suck in their own ways) because if a model sucks it just gets deleted and I move on to other things. I suppose I'd consider myself somewhat of an 'expert' when it comes to analyzing local model performance, but I don't really listen too much to what anyone else says about them, or which benchmarks tell them which things about a model. Just test them on the things that are important to you.