| ▲ | forrestthewoods 12 hours ago | ||||||||||||||||||||||
auto-vectorization is not nearly as good as you would hope it to be. The best SIMD optimizations likely require changing your data format from AoS to SoA. | |||||||||||||||||||||||
| ▲ | nylonstrung 12 hours ago | parent | next [-] | ||||||||||||||||||||||
The one feature in Jonathan Blow's Jai language I really envy is a a single keyword to switch AoS to SoA and visa-versa at comptime | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | ethin 11 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Either this or you have to do special tricks like pairwise tree reductions and hand-unroll certain portions of loops. | |||||||||||||||||||||||
| ▲ | raegis 12 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
What are AoS and SoA? | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | formerly_proven 12 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
And -march=native or at least -march=x86-64-v3 or similar, alternatively identifying relevant functions and manually invoking FMV and uarch specialization via target_clones. Plus non-integer code can generally not be autovectorized in normal-math mode since FP is non-commutative. | |||||||||||||||||||||||
| ▲ | Joker_vD 12 hours ago | parent | prev [-] | ||||||||||||||||||||||
Well, then I just prompt Claude and get SIMD without having to learn it /s | |||||||||||||||||||||||