Remix.run Logo
▲ croemer 2 hours ago

Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?

▲simonw 2 hours ago | parent [-]

I think this is a bug. I've not seen this problem from any of the other frontier models.

▲NewJazz 17 minutes ago | parent | next [-]

I would also consider this a bug. I think ajy reasonable consumer would.

▲Insanity 2 hours ago | parent | prev [-]

Do other models put a hard cap on the output tokens it can generate?

▲simonw 35 minutes ago | parent [-]

Yes, the OpenAI GPT-6 Astra limit is 128,000 as well: https://developers.openai.com/api/docs/models/gpt-6-astra

Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...