| ▲ | andy_ppp 2 hours ago | |
Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that. | ||
| ▲ | cubefox 8 minutes ago | parent [-] | |
According to SemiAnalysis, both inference and post-training (RLVR) is mostly memory bandwidth bound. Only pre-training is compute bound, but it now only takes a small share of overall data center capacity. | ||