Remix.run Logo
▲ simonw 2 hours ago

Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Here's how the thinking effort levels compare:

  low
  27 input, 1,623 output, thinking_tokens: 0
  1.6284
  Duration: 10138ms (10s)
  
  medium
  27 input, 1,796 output, thinking_tokens: 0
  1.7914 cents
  Duration: 11266ms (11s)

  high
  27 input, 2,334 output, thinking_tokens: 745
  2.3394 cents
  Duration: 17376ms (17s)

  xhigh
  27 input, 5,730 output, thinking_tokens: 2535
  5.7354 cents
  Duration: 41882ms (41s)

  max (failed to return response)
  27 input, 128,000 output, thinking_tokens: 128000
  $1.28
  Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.
▲thefourthchime an hour ago | parent | next [-]

It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...

2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...

Up until very recently, all models struggled with this.

All results: https://jonclegg.github.io/pacman-bakeoff/

▲sixtyj 6 minutes ago | parent | next [-]

I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.

Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…

▲judge2020 an hour ago | parent | prev | next [-]

Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.

▲thefourthchime an hour ago | parent [-]

Yes! The point being that up until yesterday, every model struggled with this, and now they don't.

▲copperx 6 minutes ago | parent | next [-]

"this" being recreating Pacman specifically, or games?

▲thefourthchime 21 minutes ago | parent | prev [-]

Your welcome!

▲ilamont 39 minutes ago | parent | prev | next [-]

Thank you for doing this. It is very helpful not just for capabilities but also for costs.

▲russellbeattie an hour ago | parent | prev [-]

Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.

▲thefourthchime 21 minutes ago | parent [-]

Thanks!

▲platinumrad an hour ago | parent | prev | next [-]

The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.

▲croemer 2 hours ago | parent | prev | next [-]

This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.

▲gumby271 2 hours ago | parent [-]

If it was trained on HN, there would be a 60% chance of it just saying "I'm so tired of this request, can we please move on"

▲parkersweb 40 minutes ago | parent | prev | next [-]

I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!

▲keeeba an hour ago | parent | prev | next [-]

Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?

▲TomGarden 2 hours ago | parent | prev | next [-]

Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?

▲petu an hour ago | parent | next [-]

That's max output tokens per response limit, separate from context length

▲simonw an hour ago | parent | prev | next [-]

It's the output token limit, which has been 128,000 for Claude models for quite a while note

▲croemer an hour ago | parent [-]

Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?

▲simonw an hour ago | parent [-]

I think this is a bug. I've not seen this problem from any of the other frontier models.

▲Insanity 42 minutes ago | parent [-]

Do other models put a hard cap on the output tokens it can generate?

▲ 2 hours ago | parent | prev [-]
[deleted]
▲heyjstn an hour ago | parent | prev | next [-]

I think the next models will be benchmaxxing on the Pelican benchmark tbh

▲pelicanmaxer 30 minutes ago | parent | prev | next [-]

that pelican one-pedaling

▲aimaxxed an hour ago | parent | prev [-]

“Pelicans are solved.”