| ▲ | simonw 7 hours ago |
| Surprisingly it only supports reasoning "none" or reasoning "high". That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none. The high bicycle frame is better then the none one though. Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... ) |
|
| ▲ | rahen an hour ago | parent | next [-] |
| The benchmarks against Opus 5.5 and GPT-6.1 Sol look pretty good for 3D generation:
https://x.com/atomic_chat_hq/status/2107516529608700383 |
| |
| ▲ | vippy a minute ago | parent | next [-] | | Got an example hosted on a platform that isn't owned by Twitter's loser owner? | |
| ▲ | morningsam 26 minutes ago | parent | prev [-] | | I wonder if the other models were worse than usual for that particular video because they "dislike" making an advertisement for another company's model. A test with a generic video might be more meaningful. |
|
|
| ▲ | defjm 4 hours ago | parent | prev | next [-] |
| This is such a pristine pelican. Let me say it here first folks, AGI is here. |
| |
| ▲ | search_facility 2 hours ago | parent | next [-] | | > AGI is here If AGI is "Attractions to Get Investments" then yes, it's happening | | |
| ▲ | paimapi 40 minutes ago | parent [-] | | this is fun, I love initialisms :) let's see what five minutes of end-of-day brain can crank out: Automated Grift Infrastructure Absurdly Glorified Interpolation Avoid Genuine Investigation Always Great In-theory |
| |
| ▲ | AlexCoventry an hour ago | parent | prev | next [-] | | Pelican benchmark is saturated, anyway. :-) | |
| ▲ | lofaszvanitt an hour ago | parent | prev | next [-] | | Nah, its beak is still too small to hold a capybara. | |
| ▲ | wellthisisgreat an hour ago | parent | prev | next [-] | | I wonder when will we see a photorealistic pelican on a bicycle in SVG format. | |
| ▲ | alsetmusic 2 hours ago | parent | prev [-] | | > AGI is here Far from it. This shows a strong ability to generate an image known to be frequently used as a model test. This isn't a measure of thought. | | |
| ▲ | xeyownt 2 hours ago | parent | next [-] | | don't know about AGI, but humor is gone. | |
| ▲ | roarcher 2 hours ago | parent | prev | next [-] | | Pretty sure the parent comment was sarcasm. | | |
| ▲ | EGreg 2 hours ago | parent | next [-] | | > was sarcasm Far from it. This is an example of Poe’s law, a very frequent occurrence on the internet. This isn’t a clear example of sarcasm any more than the pelican is a clear example of AGI! | | |
| ▲ | phlakaton 2 hours ago | parent | next [-] | | I agree, it wasn't clear at all to me... until I clicked the links. Now it's clear to me. Surely we live in the AGI times that were prophesied. | |
| ▲ | roarcher an hour ago | parent | prev [-] | | I guess I thought the sarcasm was obvious because I can't imagine a serious person looking at that pelican and considering it proof of AGI. But you're right, this is the internet, anything is possible. |
| |
| ▲ | unconscionable an hour ago | parent | prev [-] | | Was this a sarcastic comment? |
| |
| ▲ | throw310822 2 hours ago | parent | prev | next [-] | | You just failed the Turing test for sarcasm. I can't ask you how it feels because that would require subjectivity. | | | |
| ▲ | brumar 2 hours ago | parent | prev [-] | | I am all for the /s marker. Downvoting to hell first degree interpretation is a bit punishing for people who do not have a radar for sarcasm. |
|
|
|
| ▲ | danbrooks 6 minutes ago | parent | prev | next [-] |
| That's one heck of a pelican! |
|
| ▲ | pilaf 2 hours ago | parent | prev | next [-] |
| I think it's curious that it has so many shared elements with the latest Astra pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - Sun on top right - Cloud on top left - Three "speed lines" - Two feathers on top of the head - Eye rendered as a black circle with smaller white circle inside I wonder if the pelican benchmark is converging across models due to past results being used in training. |
| |
| ▲ | ricardobeat 2 hours ago | parent | next [-] | | These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration. | |
| ▲ | russellbeattie an hour ago | parent | prev [-] | | This comes up every thread. I think we've all noticed how similar they are becoming. I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark. Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something. |
|
|
| ▲ | RGS1811 44 minutes ago | parent | prev | next [-] |
| That beak is CHONKY. |
|
| ▲ | inknight 6 hours ago | parent | prev | next [-] |
| Why are pelicans almost identical across different models? |
| |
| ▲ | comboy 5 hours ago | parent | next [-] | | I recently was testing something, I asked some models to provide me a single random word: claude-opus-5: Lantern
claude-opus-5-5: Lantern
claude-fable-5-1: Lantern
claude-fable-5: Lantern
gemini-3.8-flash: Zephyr
gemini: Petrichor
qwen3.5-dashscope: Zephyr
glm-5.1: Lantern
gpt-6-astra: Lantern
grok-4: octopus
mimo-v2.5-pro: Breeze
minimax-m2.5: serendipity
kimi2.6-or: Gossamer
grok-4.20: luminescent
deepseek-v4-flash: serendipity
deepseek-v4-pro: Endurance
deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out. | | |
| ▲ | lossyalgo 4 hours ago | parent | next [-] | | Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt: GPT 6 Astra High: Flabbergasted
GPT 6.1 Sol High: Petrichor
GPT 6 Sol High: Kaleidoscope
GPT 6 Sol Med: Firefly
GPT 6 Sol Light: Persimmon
GPT 6 Luna High: Tumbleweed
GPT 5.6 Sol High: Kaleidoscope
GPT 5.6 Terra High: Liminal
GPT 5.6 Luna High: Mellifluous
GPT 5 mini Medium: Serendipity
GPT 5.3 Codex Med: Nebula
Junie: Flourishing
Claude Haiku 4.5 Med: Serendipity
Claude Sonnet 5 Med: Banana
Claude Sonnet 5 High: Banana
Claude Sonnet 5.5 Med: Serendipity
Gemini 3.7 Flash: Zephyr
Gemini 3.8 Flash: Kaleidoscope
Grok 4.5 Medium: nebula
Grok 4.6 Medium: Serendipity
Grok 4.7 Medium: Quasar
Kimi K3 Low: Lantern
Kimi K3 Max: Lantern
MAI Code 1.1 Flash Med:Peregrine
| | |
| ▲ | jsw97 2 hours ago | parent | next [-] | | I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries. | | |
| ▲ | nomel 38 minutes ago | parent [-] | | > You could expand on this by giving programming tasks and measuring code similarity. But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge. |
| |
| ▲ | varjag 2 hours ago | parent | prev | next [-] | | I got Peregrine out of GPT-6 too. Huh. | |
| ▲ | billnad 3 hours ago | parent | prev [-] | | Just tried M365 Copilot with a premium account. Petrichor | | |
| |
| ▲ | timschmidt 40 minutes ago | parent | prev | next [-] | | This feels uncannily like the ancestor of the Voight-Kampff test[0] 0: https://www.youtube.com/watch?v=Umc9ezAyJv0 | |
| ▲ | search_facility an hour ago | parent | prev | next [-] | | Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho. So not something internal to model thinking. | |
| ▲ | aktenlage 5 hours ago | parent | prev | next [-] | | That is a cool idea. That astra gave the same word as claude is highly unexpected. | |
| ▲ | smokel 2 hours ago | parent | prev | next [-] | | What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment. "Zephyr" and "breeze" might be related to forgetting everything, starting fresh. So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these? | | |
| ▲ | ricardobeat an hour ago | parent [-] | | just “a random word” gives you Zephyr in Gemini, and “Lantern” in Claude and ChatGPT. | | |
| |
| ▲ | Rebelgecko 2 hours ago | parent | prev | next [-] | | I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities | |
| ▲ | jacereda 5 hours ago | parent | prev | next [-] | | Just tried Mistral Large 4: Serendipity. | |
| ▲ | vunderba 5 hours ago | parent | prev | next [-] | | I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean. The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel. I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”) | | |
| ▲ | Lord-Jobo 4 hours ago | parent [-] | | Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both. |
| |
| ▲ | russellbeattie an hour ago | parent | prev | next [-] | | Muse Spark 1.3: lighthouse The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math. Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt? I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test. | |
| ▲ | Gracana 2 hours ago | parent | prev [-] | | The eqbench creative writing "slop profiles" do something similar. https://eqbench.com/creative_writing.html Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases. |
| |
| ▲ | wren6991 6 hours ago | parent | prev | next [-] | | The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. | | |
| ▲ | simonw 3 hours ago | parent | next [-] | | > The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. OK well I couldn't resist this one: llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... | | |
| ▲ | viraptor 2 hours ago | parent | next [-] | | I'm getting 403 inside the tool for this one. (The pelican bike on top works) This has been happening a lot recently. | | | |
| ▲ | flowardnut 3 hours ago | parent | prev | next [-] | | The rover honking is pretty silly, opus has a good sense of humor | |
| ▲ | largbae 2 hours ago | parent | prev [-] | | Gemini wins this one clearly. Honey please! |
| |
| ▲ | YawningAngel 3 hours ago | parent | prev [-] | | https://chatgpt.com/s/m_6ac53d4e5b0c8191949050dbf1f402d7 Not sure I'd call it jaywalking exactly but pretty good |
| |
| ▲ | simonw 5 hours ago | parent | prev | next [-] | | I think they're still visually pretty different. The most common shared details are: - Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is. - Bicycle is usually red. No idea! Red ones go faster? | |
| ▲ | whyenot 2 hours ago | parent | prev | next [-] | | Also, why are they almost always riding from let to right? | | |
| ▲ | Sharlin 11 minutes ago | parent | next [-] | | It's been discussed many times. The reason is bikes are almost without exception depicted that way in order to show the drivetrain. | |
| ▲ | ricardobeat an hour ago | parent | prev [-] | | Ever seen a movie chase scene where cars are going right to left? |
| |
| ▲ | advisedwang 6 hours ago | parent | prev | next [-] | | They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even. | |
| ▲ | Tade0 6 hours ago | parent | prev | next [-] | | Everyone is stealing from everyone else. | |
| ▲ | peder 4 hours ago | parent | prev [-] | | Because it's a terrible benchmark |
|
|
| ▲ | BeetleB 2 hours ago | parent | prev | next [-] |
| The difference between high and none is the bicycle. |
| |
| ▲ | mcv 2 hours ago | parent | next [-] | | The bicycle looks significantly better in high. And feet and hands are actually where they should be. The road looks worse, though. No flowers either. And in neither is the pelican sitting on the saddle, but I can understand it's hard for a pelican to ride a bicycle properly. Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that. | | |
| ▲ | stymaar 2 hours ago | parent [-] | | > hands are actually where they should be. If you don't mind the fact that a pelican shouldn't have hands, of course. |
| |
| ▲ | senderista 2 hours ago | parent | prev [-] | | Two pelicans, one shape. The difference is load-bearing, and that's the big unlock. |
|
|
| ▲ | deflator 6 hours ago | parent | prev | next [-] |
| Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican |
|
| ▲ | XCSme 5 hours ago | parent | prev | next [-] |
| I tried testing it, but reasoning effort indeed seems to be broken somehow. |
|
| ▲ | dizhn 4 hours ago | parent | prev | next [-] |
| High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc. Both are riding on the left side of the path for some reason. |
| |
| ▲ | kingstnap 4 hours ago | parent [-] | | This is entirely stochasticity. The entire reasoning trace was: > Create a cartoon pelican riding a bicycle. Need SVG only output. |
|
|
| ▲ | nicolamanzini 5 hours ago | parent | prev [-] |
| [dead] |