indeed, can't wait for it to be supported by llama.cpp (or other engines)?
Then we probably have to wait a little for them to optimize it.