| ▲ | E-Reverance 3 days ago | |||||||||||||||||||||||||
I know it goes a tiny a bit against the spirit of what y'all are doing, but applying a few layers of pixel-wise local attention (so 1x1 "patch", with 3x3 or 5x5 attention window, basically treating it as a dynamic conv) has worked way better than both linear and conv unpatching in my recent experiments. Diagram for reference https://x.com/1rreverant/status/2107546198093730287 (In my most recent recent experiment I actually removed the MLP and just used a linear project on the pixel's hidden states) | ||||||||||||||||||||||||||
| ▲ | schopra909 3 days ago | parent [-] | |||||||||||||||||||||||||
That sounds like an interesting idea! Can you confirm I'm understanding correctly? 1) Linear unpatchify as usual to go from hidden states to pixel space 2) Attention within a local window (e.g. 3x3, 5x5) to "blend" pixel space data and come up with a better image (as an alternative to MLP or Convolution) And follow up questions: 1) How do you handle boundaries between your "attention windows"? Do you move the window just like a convolution does or are the "attention windows" all mutually exclusive from one another? 2) How much faster/slower is this operation vs. a linear layer + MLP? | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||