Remix.run Logo
sixtyj 5 hours ago

At such output speed, I wouldn’t expect reasoning. (But I didn’t know it, thanks.)

700 TPS with reasoning is awesome and it speeds things up.

Cerebras as public traded company is worth keeping an eye what they produce.

msdz 5 hours ago | parent [-]

> At such output speed, I wouldn’t expect reasoning.

As the sibling comment to yours mentioned, if they had a reasoning model “hardware-ified” onto a custom chip (as is their plan for IIRC this or next year, a new ASIC), it’d output fast decode speeds for the regular output as well as reasoning sections. Both would be ≈equally fast.

senordevnyc 2 hours ago | parent [-]

Yeah, I thought reasoning was literally just chain of thought in the output token stream, with the model itself adding delimiters to indicate what part of the output is internal reasoning, and what part is an answer to the user. Is that wrong?

beering an hour ago | parent [-]

You are right, reasoning is unrelated to tokens per second.