| ▲ | warkdarrior 3 hours ago | |
> Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses. Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack. | ||
| ▲ | Marha01 an hour ago | parent [-] | |
Burning the weights into static mask ROM is pretty trivial. Taalas is doing it. | ||