| ▲ | Aurornis a day ago | |
Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving. RAM was probably the bottleneck for the amount of context they were offering. I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful | ||