| ▲ | simonw 4 hours ago |
| Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right. The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds. The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model: llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/ |
|
| ▲ | rotis 28 minutes ago | parent | next [-] |
| thinking_effort: max
Reasoning trace: This is the classic pelican-on-bicycle SVG test. Opus 5.5 had similar response on max: This is a classic test request https://tools.simonwillison.net/markdown-svg-renderer?url=ht... I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;) |
|
| ▲ | ozgung 3 hours ago | parent | prev | next [-] |
| Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right. |
| |
| ▲ | kenhwang 4 minutes ago | parent | next [-] | | The grass too, everyone knows bikes ride on grass. | |
| ▲ | accrual 3 hours ago | parent | prev | next [-] | | I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack. > "What's up with the pelican?" Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ... | | | |
| ▲ | cortesoft an hour ago | parent | prev [-] | | The medium thinking effort one doesn't have a sun at all? |
|
|
| ▲ | Jcampuzano2 2 hours ago | parent | prev | next [-] |
| I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane. Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max. |
| |
| ▲ | mudkipdev an hour ago | parent | next [-] | | Benchmarks are the entire reason why max exists | |
| ▲ | LoganDark an hour ago | parent | prev [-] | | I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general) I'm definitely not the average person though. | | |
| ▲ | abustamam an hour ago | parent [-] | | I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround? I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use. | | |
| ▲ | LoganDark an hour ago | parent [-] | | I thought the thinking effort was specified out of band from that, though maybe it's not. Not sure if the model was trained to listen in other areas. The biggest issue is, it's difficult to tell if it works because you can no longer see the thinking! Though I guess if you can't tell a difference in the output, was there any point to max in the first place? |
|
|
|
|
| ▲ | zahlman an hour ago | parent | prev | next [-] |
| I'm still getting network errors. Seems to be CORS-related. |
|
| ▲ | ijidak 3 hours ago | parent | prev | next [-] |
| I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time. The pelicans all start to look the same after a while. But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference. |
| |
|
| ▲ | vinni2 4 hours ago | parent | prev | next [-] |
| I thought Anthropic models didn’t generate images. |
| |
| ▲ | simonw 3 hours ago | parent | next [-] | | This is SVG, but recent Anthropic models have got extremely good at other forms of visual data. Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%... And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%... | | |
| ▲ | noduerme an hour ago | parent | next [-] | | Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file. | |
| ▲ | khanan 4 minutes ago | parent | prev | next [-] | | the "monkey island"-musc makes me so sad. why would anyone make an AI do this when it took such craft. i hope you all die :'( | |
| ▲ | lastdong 2 hours ago | parent | prev | next [-] | | This is great! Love the pixel art and tunes. | |
| ▲ | jansan 3 hours ago | parent | prev | next [-] | | They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive. They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great. | | | |
| ▲ | hazelnut 3 hours ago | parent | prev [-] | | Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯ |
| |
| ▲ | vunderba 3 hours ago | parent | prev | next [-] | | I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images. This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style. https://kq-styles.specr.net | |
| ▲ | abirch 4 hours ago | parent | prev | next [-] | | They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success. | |
| ▲ | sixtyj 3 hours ago | parent | prev | next [-] | | They don’t do raster images. | |
| ▲ | 1ucky 4 hours ago | parent | prev | next [-] | | Those are SVGs not images. | |
| ▲ | jonshariat 4 hours ago | parent | prev | next [-] | | SVG is code | |
| ▲ | sixothree 3 hours ago | parent | prev | next [-] | | I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg. Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only. https://www.youtube.com/watch?v=2EqMplbt0gU | | |
| ▲ | zyberzero 2 hours ago | parent | next [-] | | To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all?
Then I think it is impressive! Are you able to share the prompts you used? | | |
| ▲ | sixothree 2 hours ago | parent [-] | | Source code is linked from the video! Scan the QR code. I will try to /resume tonight and give you some prompts. | | |
| ▲ | cruffle_duffle an hour ago | parent [-] | | I mean Claude code sessions are all jsonl files that it can interrogate on its own. Get a new agent to capture how it was made and what the prompts were. No need for tedious /resume’ing and prompting. |
|
| |
| ▲ | Bluestein 2 hours ago | parent | prev [-] | | The Purple Screen of Death at the end :) |
| |
| ▲ | fakedang 4 hours ago | parent | prev | next [-] | | They're SVGs | |
| ▲ | ghoshbishakh 3 hours ago | parent | prev [-] | | Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations. |
|
|
| ▲ | jsolson 2 hours ago | parent | prev [-] |
| [dead] |