| ▲ | Farmadupe 4 hours ago | |
> 2G of random garbage is written directly onto one member device (behind the filesystem's back, offset 1G — python injector; uutils dd mis-seeks on dm devices), caches dropped, then a full scrub: btrfs scrub -B, zpool scrub + wait, bcachefs scrub, md/lvm sync-action 'check' (which can only COUNT mismatches — no checksums to know which copy is right). I'm not sure that nuking 2G of the underlying block device is a recoverable error on any filesystem that I'm aware of? Can you confirm if any ofthe filesystems really came out of the other side in a usable state after scrubbing? ----- > Trivial-op p99, idle (ms) # A trivial operation — one 4k write + fsync every 200ms (like a shell appending history or an editor updating its swap file) — run alone for 10s. p99 of the fsync completion In fact, if it's OK for me to ask, are any of the metrics tht you used standard industry metrics? It looks like several of the tests are bypassing the kernel's page cache? -- which I worry may fall into the trap of "I modified the system to be unrepresentative of reality and then tested it". ---- > kernel 7.0.0-1012-azure Can you confirm if you tested on a bare metal machine? were you the only tenant? | ||
| ▲ | matja 2 hours ago | parent | next [-] | |
> I'm not sure that nuking 2G of the underlying block device is a recoverable error on any filesystem that I'm aware of? ZFS and btrfs were designed from the start to handle this, by using checksums on every piece of (meta)data and redundancy to return the same data as was stored to the kernel, and rewrite the bad data. I've tested my own machines running ZFS by random writes out of band from the filesystem/kernel and it has always found and fixed them. | ||
| ▲ | vlovich123 4 hours ago | parent | prev | next [-] | |
1 device out of the replica set I’m assuming so all of them should recover. | ||
| ▲ | hlieberman 4 hours ago | parent | prev [-] | |
The integrity check is only on the tests which are either RAID or the filesystem equivalent. | ||