Remix.run Logo
▲ warkdarrior 3 hours ago

> Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.

▲Marha01 an hour ago | parent [-]

Burning the weights into static mask ROM is pretty trivial. Taalas is doing it.