Remix.run Logo
demibabs a day ago

No one’s talking about how good the final product is.

Edit: someone else commented that as I was typing this, lol.

blackhaz a day ago | parent [-]

I wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already.

ziofill 20 hours ago | parent | next [-]

That’s a fair question, but it seems that it’s not yet necessary. See here

https://dylancastillo.co/posts/pelicanmaxxing.html

https://simonwillison.net/2026/Jul/22/

stymaar 14 hours ago | parent [-]

I don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”.

Anon1096 12 hours ago | parent | next [-]

That is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability.

simonw 5 hours ago | parent | prev | next [-]

Gemini have done exactly that.

(I doubt it's because of my stupid benchmark, though!)

flexagoon 9 hours ago | parent | prev [-]

https://xkcd.com/810/

20 hours ago | parent | prev [-]
[deleted]