Remix.run Logo
8organicbits 4 hours ago

How did I miss zstd?

Here are my benchmarks for 2.3 GB of jsonl, on a laptop. Compressed size, compress time, decompress time; using defaults.

    gzip  7.3%  21s  9s
    bzip2 4.6% 251s 50s
    bzip3 3.3%  82s 69s
    zstd  6.9%   2s  3s
    lzma  4.7%  51s  3s
vlovich123 4 hours ago | parent | next [-]

At what levels? There’s no guarantee that the default compression level is comparable. You have to normalize by time spent compressing.

kccqzy 3 hours ago | parent | prev | next [-]

But zstd is super tunable. Where gzip gives you compression levels from 1 to 9, zstd gives you up to 22 for ultra compression and negative compression levels for ultra fast. The ultra fast options so fast that they are great as a substitute for memcpy if your CPU is already waiting for other things, like DRAM.

handsome_jack_ 3 hours ago | parent [-]

Comparing it to memcpy is idiotic.

kccqzy 3 hours ago | parent [-]

No it’s not. The pace of improvement of CPU compute speed is far greater than that of DRAM throughput. And in fact compression algorithms geared towards speed aims to outperform memcpy (on suitable machines).

praseodym 3 hours ago | parent | prev | next [-]

zstd has a built-in benchmark mode to compare different compression levels, e.g. `zstd -b1 -e9 [FILE]` to test levels 1 to 9 (try up to 22 if you have enough spare time)

out_of_protocol 4 hours ago | parent | prev [-]

zstd with better compression level would be nice - these numbers are not really comparable since both time and compression level are too different