Remix.run Logo
villgax 4 hours ago

Lol, such a lazily written article by wafer.ai

GPUs. 8× MI355X (TP8) B300 (TP8+DCP8)

Decode tok/s per stream 118 tok/s 172 tok/s

Peak aggregate. 952 tok/s 1,568 tok/s

Peak aggregate per GPU 119 tok/s 196 tok/s

On every row the B300 beat the MI355X

The B200 is being forcefully compared against something which is not gonna fit within it's memory in a single node & not much details about multi-node interconnectivity, disagg or not. As expected of a shoddy slop.

The only point it won is of cost per hour is one aggregation website for rentals, the premium a B300 commands against the $3/hr AMD chip which no provider has in abundance. Never bothered to do TCO of owning the hardware either.

throwa356262 2 hours ago | parent [-]

Did you see this section?

    To the B200’s defence, its numbers are somewhat deflated by the fact that it pays a cross-node all-reduce on the decode critical path (RoCE v2 at ~195 Gb/s) — it’s the only config here that spans two nodes