Remix.run Logo
▲ librasteve 3 hours ago

As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org

▲feffe an hour ago | parent [-]

I think tenstorrent architecture is more like this. A grid of RISC-V cores with local SRAM and vector units.