| ▲ | Argonautlabs 4 hours ago | |||||||
Fair, and we didn't measure it. Decode was flat from 128 to 512 generated tokens (0.926 → 0.923 tok/s drafter-off), but that's a 6-token prompt plus the output — total context under a thousand. The current configuration admits about 4.4k tokens of context at all, and at anything like 200k the killer wouldn't be decode, it would be prefill: today it reads each layer's experts once per 64-row pass, so 200k tokens of prompt would be measured in days, not minutes, until the scheduling fix. | ||||||||
| ▲ | NooneAtAll3 3 hours ago | parent [-] | |||||||
what's the main limitation on context size? 4.4k seems... I just realized I have no sense of scale whatsoever | ||||||||
| ||||||||