| ▲ | michaelbuckbee 2 hours ago | |
The eval world is split into: 1. Long form task based examinations like this that test the ability of the model+harness to remain on task, tool calling, overall effectiveness and taste. 2. More direct 1:1 and qualitative comparisons that you might get with a tool like https://evvl.ai/ - which also uses OpenRouter and does similar one off model comparisons (or lets you use it as a MCP from your dev env to be like: "take the prompt from this loop and try it against these other models") | ||