Remix.run Logo
▲ sharktheone 8 hours ago

I am hoping for portable SIMD so much. But I still think that often a manually rolled SIMD will be faster.

Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.

▲dwattttt 7 hours ago | parent [-]

I guarantee you with my lack of skill, my attempt at using portable simd will exist while using manual intrinsics I doubt I'd get there.

The question for me is whether portable simd will result in faster code than plain auto-vectorisation; for the simplest loops auto has me beat (the few times I've tried it), but I imagine as the complexity grows I'll be more likely to try do something that breaks auto-vectorisation, and it'll be more obvious to me when I do that in portable simd.