| ▲ | demibabs a day ago |
| No one’s talking about how good the final product is. Edit: someone else commented that as I was typing this, lol. |
|
| ▲ | blackhaz a day ago | parent [-] |
| I wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already. |
| |
| ▲ | ziofill 20 hours ago | parent | next [-] | | That’s a fair question, but it seems that it’s not yet necessary. See here https://dylancastillo.co/posts/pelicanmaxxing.html https://simonwillison.net/2026/Jul/22/ | | |
| ▲ | stymaar 14 hours ago | parent [-] | | I don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”. | | |
| ▲ | Anon1096 12 hours ago | parent | next [-] | | That is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability. | |
| ▲ | simonw 5 hours ago | parent | prev | next [-] | | Gemini have done exactly that. (I doubt it's because of my stupid benchmark, though!) | |
| ▲ | flexagoon 9 hours ago | parent | prev [-] | | https://xkcd.com/810/ |
|
| |
| ▲ | 20 hours ago | parent | prev [-] | | [deleted] |
|