| ▲ | stymaar 14 hours ago | |
I don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”. | ||
| ▲ | Anon1096 12 hours ago | parent | next [-] | |
That is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability. | ||
| ▲ | simonw 5 hours ago | parent | prev | next [-] | |
Gemini have done exactly that. (I doubt it's because of my stupid benchmark, though!) | ||
| ▲ | flexagoon 9 hours ago | parent | prev [-] | |