Remix.run Logo
cyanydeez 12 hours ago

llamacpp has the basics: reasoning-budget limits thinking token output and reasoning-message thats injected when budget is exhausted. client can set these per request so dynamics are possible.

models dont seem to care if you cut them off.