Remix.run Logo
▲ TomGarden 3 hours ago

Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?

▲petu 3 hours ago | parent | next [-]

That's max output tokens per response limit, separate from context length

▲simonw 2 hours ago | parent | prev | next [-]

It's the output token limit, which has been 128,000 for Claude models for quite a while note

▲croemer 2 hours ago | parent [-]

Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?

▲simonw 2 hours ago | parent [-]

I think this is a bug. I've not seen this problem from any of the other frontier models.

▲NewJazz 19 minutes ago | parent | next [-]

I would also consider this a bug. I think ajy reasonable consumer would.

▲Insanity 2 hours ago | parent | prev [-]

Do other models put a hard cap on the output tokens it can generate?

▲simonw 37 minutes ago | parent [-]

Yes, the OpenAI GPT-6 Astra limit is 128,000 as well: https://developers.openai.com/api/docs/models/gpt-6-astra

Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...

▲ 3 hours ago | parent | prev [-]
[deleted]