| ▲ | Memory Ordering in CPUs(fgiesen.wordpress.com) | ||||||||||||||||||||||
| 66 points by ibobev 4 days ago | 7 comments | |||||||||||||||||||||||
| ▲ | BretonForearm 4 days ago | parent | next [-] | ||||||||||||||||||||||
Would be useful for the article to start with a word on what memory ordering is. Edit: it means, (re)ordering of memory access operations esp. when multiple execution units (cores) operate in parallel. Weakly ordered (ARM, RISC-V): if Core 1 writes Memory 1 then Memory 2, but for some reason writing into Memory 2 is quicker, then Core 2 may see Memory2 written before Memory 1 is written. While strong ordering (x86) retains the original order. | |||||||||||||||||||||||
| ▲ | gopalv 4 days ago | parent | prev | next [-] | ||||||||||||||||||||||
The good part is that at least there's some explicit C++ std::memory_order and std::sync::atomic::Ordering lets me pick where it really matters. For example, I built a skew handling model which needed low overhead cross-thread counters, where Ampere and Graviton was different from the M1 mac in benchmark - even down to the same assembly on different systems (cmov specifically). | |||||||||||||||||||||||
| ▲ | StillBored 4 days ago | parent | prev | next [-] | ||||||||||||||||||||||
Yes, but the contended case is often the critical part of an application’s performance. Weakly ordered CPUs therefore need fast barrier mechanisms, which in turn means re-creating much of the store-ordering, writeback buffering, and commit logic anyway. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | j_seigh 3 days ago | parent | prev [-] | ||||||||||||||||||||||
I'd be more concerned with instruction reordering by the compiler. So even if the hardware guarantees some ordering, you will still need those memory fences. | |||||||||||||||||||||||