| ▲ | Getting 50 GB/S Back from the Apple Neural Engine(eiln.github.io) | |
| 28 points by eiln 3 days ago | 3 comments | ||
| ▲ | VladVladikoff 19 minutes ago | parent | next [-] | |
This website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave. | ||
| ▲ | Neywiny 43 minutes ago | parent | prev | next [-] | |
Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL? | ||
| ▲ | eiln 3 days ago | parent | prev [-] | |
RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s. | ||