| ▲ | simonw 8 hours ago |
| Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.) |
|
| ▲ | onlyrealcuzzo 8 hours ago | parent | next [-] |
| This is in comparison to Fable: > https://tools.simonwillison.net/markdown-svg-renderer?url=ht... > Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30! So 50x cheaper - and how much faster? |
| |
| ▲ | simonw 8 hours ago | parent [-] | | The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757 | | |
| ▲ | scosman 7 hours ago | parent | next [-] | | Community effort happening here to build the ideal dataset: https://github.com/scosman/pelicans_riding_bicycles | | |
| ▲ | isoprophlex 5 hours ago | parent | next [-] | | Wow nice. If an llm could replicate these excellent examples, I'd consider the pelican benchmark fully saturated. | |
| ▲ | MadameMinty 6 hours ago | parent | prev | next [-] | | Nice. Really high quality SVG pelicans riding bicycles here. | |
| ▲ | phatfish 4 hours ago | parent | prev [-] | | Come on, don't provide the smoking gun that shows how to draw a pelican riding a bicycle. If it's on the public internet it will end up in training data and invalidate this important LLM capability benchmark. |
| |
| ▲ | dieortin 8 hours ago | parent | prev | next [-] | | I don’t know if you’re joking, but I don’t see anything in the linked tweet which suggests that is the case | | |
| ▲ | simonw 8 hours ago | parent [-] | | Watch the video. It's from then-Gemini-lead Jeff Dean and the video shows off an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine. | | |
| ▲ | JacobAsmuth 6 hours ago | parent [-] | | I'm confused. This doesn't mean they trained on it. | | |
| ▲ | simonw 5 hours ago | parent [-] | | I didn't say trained on, I said "trained for SVG output". Gemini team members have publicly stated that they have trained for SVG: https://twitter.com/sunjiao123sun_/status/202455551655137292... > I’ve been developing the SVG generation capabilities for Gemini 3.1, and the complexity of the SVGs is stunning. > This allows UX designers to transcend pixel constraints and directly output structural, production-ready code! |
|
|
| |
| ▲ | uif124 6 hours ago | parent | prev [-] | | Yesterday's transcript of the best version from Fable looked like Fable already knew exactly what it should do without "thinking". In other words, there were no passages like "on the one hand I could do this, on the other hand ...". It saw the fish in the basket from some other previous attempt but completely missed the gap between the tires and the rims where the background shines through (now it knows after scraping this comment and watch the next transcript). |
|
|
|
| ▲ | anigbrowl 2 hours ago | parent | prev | next [-] |
| These are becoming unreadable as the reasoning chains expand. I think you should consider reformatting them and either putting the image first or else folding the COT output. |
|
| ▲ | mrdependable 7 hours ago | parent | prev | next [-] |
| Why are the SVGs getting more detailed rather than just more correct than previous models? |
| |
| ▲ | trentor 7 hours ago | parent [-] | | Because people tend to like fidelity more than correctness. | | |
| ▲ | aesthesia 6 hours ago | parent | next [-] | | It bugs me a little that "fidelity" has connotations other than "faithfulness to an original"---fidelity should be basically the same as correctness here! | | | |
| ▲ | neuronic 2 hours ago | parent | prev [-] | | If correctness would matter anymore, people wouldn't be using LLMs in the first place. |
|
|
|
| ▲ | hughw 7 hours ago | parent | prev | next [-] |
| It's about to squash a tiny baby pelican |
|
| ▲ | jpadkins 7 hours ago | parent | prev | next [-] |
| The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills. Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak. |
| |
| ▲ | neuronic 2 hours ago | parent [-] | | LLMs are not intelligent and don't actually understand the concept of a bicycle. Parrots also don't understand human language but they're really good at pretending otherwise. |
|
|
| ▲ | 8 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | EugeneOZ 4 hours ago | parent | prev | next [-] |
| Impressive pelicans! |
|
| ▲ | lern_too_spel 8 hours ago | parent | prev | next [-] |
| The fenders are a nice touch, but putting the fenders through the tires seems like a design flaw. |
|
| ▲ | world2vec 8 hours ago | parent | prev [-] |
| I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore. |
| |
| ▲ | wongarsu 8 hours ago | parent | next [-] | | If everyone agreed with you, the comment would disappear near the bottom of the thread I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash) | | |
| ▲ | IshKebab 2 hours ago | parent [-] | | > If everyone agreed with you, the comment would disappear near the bottom of the thread If only it were true that things that are tiresome are unpopular. But witness "6 7", "first post", ... remember the "in soviet Russia" jokes on Slashdot"? It seems like there are a subset of people that simply don't get tired of tiresome things. |
| |
| ▲ | bitexploder 8 hours ago | parent | prev | next [-] | | It is more fun than serious at this point. Don't overthink it :) | |
| ▲ | simonw 8 hours ago | parent | prev | next [-] | | Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point. (Next up is the comment saying that the labs are clearly training for the benchmark.) | | |
| ▲ | world2vec 8 hours ago | parent [-] | | The labs are clearly training for the benchmark. | | |
| ▲ | WarmWash 8 hours ago | parent [-] | | This has been addressed endlessly, for a few years now, and is just as much of a trope as "this benchmark is useless". | | |
|
| |
| ▲ | anentropic 7 hours ago | parent | prev [-] | | it's a tradition |
|