| ▲ | modulovalue an hour ago | |
I’m not sure I agree on the impossible part, I feel like a sufficiently smart interprocedural analysis that also implements range analysis interprocedurally could prove a lot to where it becomes useful. I guess what I would like to see is SIMD libraries being able to confidently say nobody needs to use intrinsics (or differentiate between relaxed/normal SIMD on the user API level) because the language + high level SIMD APIs are smart enough to choose the right implementation. IIRC IEEE min/max with proper NaN handling needs 8 instructions on x86 vs 1 on arm64 I find it very sad that we apparently haven’t really solved that yet without forcing the user to use different APIs. | ||
| ▲ | athrowaway3z 13 minutes ago | parent | next [-] | |
Anything without range analysis is not worth it. Note that: NonZerof32 * NonZerof32 -> NonNanf32 NonZerof32::from_bits(1) multiplied with itself is zero. Doing range analysis needs the language to support it at compile time, and the dev to specify what range it is. The only 'stable' thing i can think of is a type for 'greater-eq-one' using only addition and multiplication. Practically every other operation breaks most of the type knowledge up to that point. | ||
| ▲ | orlp 19 minutes ago | parent | prev [-] | |
You can do it with just 3 instructions for IEEE 754-2019 minimumNumber (ignores NaN):
If you want proper IEEE 754-2019 minimum (propagate NaN, -0.0 < +0.0, NaN bitpattern picked in the usual way) you can do it in 6:
I personally find this a load of nonsense I don't care about.If you want propagating NaNs but don't care about signed zero or NaN payload/sign, you can use
What I do in Polars is a bit different, there for propagating NaNs I do
this isn't fully optimal on x86-64 but it's fairly simple and autovectorizes decently on various platforms, here's AVX2: | ||