Remix.run Logo
msdz 5 hours ago

> At such output speed, I wouldn’t expect reasoning.

As the sibling comment to yours mentioned, if they had a reasoning model “hardware-ified” onto a custom chip (as is their plan for IIRC this or next year, a new ASIC), it’d output fast decode speeds for the regular output as well as reasoning sections. Both would be ≈equally fast.

senordevnyc 2 hours ago | parent [-]

Yeah, I thought reasoning was literally just chain of thought in the output token stream, with the model itself adding delimiters to indicate what part of the output is internal reasoning, and what part is an answer to the user. Is that wrong?

beering an hour ago | parent [-]

You are right, reasoning is unrelated to tokens per second.