Remix.run Logo
▲ glouwbug 40 minutes ago

Learn which instructions SIMD nicely (sqrt / fabs, etc). Use ternaries in loops for masking. Use trig identities and lookup tables (don't recompute sin(3t) when you can use two vector multiples using a table of sin(t) eg. sin(t) * sin(t) * sin(t)). Use divisible constexpr constants in loops to eliminate the SIMD tail. Be careful with type casts and floats. `float x; x += 0.5` will introduce *cvt instructions even if the compiler statically knew better otherwise (use 0.5f). Compile with --fast-math and friends so errno doesn't invalidate your SIMD pipeline.

▲MaxBarraclough 26 minutes ago | parent | next [-]

That has a similar problem to the article, it's trying to fit far too much into too small a format.

What you've written mostly makes sense to someone who already has a solid understanding of SIMD and of C++ (although I can't say I follow all of it), but the target audience is people who don't. For them, each point needs a much lengthier explanation.

▲glouwbug 4 minutes ago | parent [-]

Likely the best tip would to `objdump -d` and inspect the assembly. For something like floating point, a good starting point would be to write and fiddle with small functions until you get rid of (most) branch labels and cvt instructions. A useful metric could be grepping for and eliminating ss instructions. Prepending (__attribute__((used)) will allow you to inspect your functions.

A quick restrict example:

    #define fn __attribute__((used))

    fn void copy1(int* to, const int* from, const int size)
    {
        for(int i = 0; i < size; i++)
            to[i] = from[i];
    }

    fn void copy2(int* to, const int* from)
    {   
        constexpr int size = 1024;
        for(int i = 0; i < size; i++) 
            to[i] = from[i];
    }

    fn void copy3(int* restrict to, const int* restrict from)
    {
        constexpr int size = 1024;
        for(int i = 0; i < size; i++) 
            to[i] = from[i];
    }

    gcc test.c -c -O3 && objdump -d ./test.o
copy1 is 52 lines, copy2 is 28 lines, copy3 is 2 lines (just a call to memcpy).

This is a good starting point for self teaching. The impact of your TLB, L1, and overall instruction count (with IPC) can further be measured with `./perf stat -d -d -d ./a.out`

▲creata 29 minutes ago | parent | prev [-]

Most applications (including most applications that care about numerical performance) should not use -ffast-math.