| ▲ | BillStrong a day ago | |
I mean, why would we wait til then? AMD is buying a company that put LLama 3.1 7b(or something close) on silicon, and can give you 15k t/s, Qwen just released Qwen 3.8 Next that seperates the intelligence from the database, so the intelligence is 6b and the rest can be done on CPU. You can put that 6b on the AMD silicon, and then use another accelerator to speed up the database if you need fast PP, and you are now much faster than the B300 after optimizations such as caching the Prompt once, and keeping it around, and many others. Suddenly, you can put that on the CPU next to an NPU, or on its own card, on the motherboard or on a GPU. | ||