Remix.run Logo
▲ peterbell_nyc 4 hours ago

You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.