| ▲ | ifz 6 hours ago | |
It's really simple, basically just the magnitude of the value vector, weighted by QK dot product, summed across all attention heads and layers. When I started, I expected I'd have to experiment a lot to find something comprehensible. But this simple computation can already show some patterns. | ||
| ▲ | stared 6 hours ago | parent [-] | |
Nice! Sometimes the simplest approaches work the best. | ||