| ▲ | mchinen 10 days ago | |||||||||||||||||||||||||
The audio side is even more interesting, as it seems they totally got rid of positional embedding are just doing a single linear transform to match the LLM input dimension and that's it. > Audio: We simplified audio processing even further. We removed the audio encoder entirely and projected the raw audio signal into the same dimensional space as text tokens. | ||||||||||||||||||||||||||
| ▲ | make3 10 days ago | parent [-] | |||||||||||||||||||||||||
I guarantee you there's positional information one way or another. they just don't mention it because positional embeddings are extremely cheap computationally, not worth mentioning | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||