| ▲ | marginalia_nu a day ago | |
That still just gets you autovectorization, and generally locks you out of the performance you could have with direct SIMD intrinsics. Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck. | ||
| ▲ | louthy a day ago | parent [-] | |
> That still just gets you autovectorization Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something? If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly. [1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime.... | ||