Remix.run Logo
busymichael 4 hours ago

I think you're underweighting the Pelican test.

Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.

Simon explains it well: https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l...

Simon, you should put up a summary table page that you update after every release.