From 7fba2f39969f16261a1046596080f5d783a49ad9 Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 21:33:37 -0500 Subject: [PATCH] BENCHMARKS: streaming FileBuilder::write A/B (the 1 MiB buffer regression and the fix) Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 36 ++++++++++++++++++++++++++++++++++++ 1 file changed, 36 insertions(+) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index c81101a..f0ec068 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th same; files grow by about 36 bytes per chunk. The default therefore stays the 1.10 format; the 1.8 format is opt-in. +### Streaming `FileBuilder::write` (2026-09-29, tank) + +`FileBuilder::write` now streams the file through a buffered writer +instead of building it in memory first (`feat/huge-chunk-dims`, the change +that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on +tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the +start of each run): `main` at `4260af4` against the stacked branch at +`549e442`, alternating, three full runs each; median (min-max) of +criterion's estimate, ms. + +> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3` + +The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower +on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g. +`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's +128 KiB mmap threshold, and after the smaller cases it was mapped afresh on +each write (1 629 368 minor page faults over the `write_2d_chunked` group +against 12 888 on `main`). Run alone, the case showed no difference. With a +64 KiB buffer (`549e442`): + +| benchmark | main | candidate | ratio | +|---|---:|---:|---:| +| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x | +| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x | +| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x | +| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x | +| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x | +| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x | +| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x | +| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x | +| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x | + +The other 18 cases are within 2%. Minor page faults per full run: 59 602 to +64 486 on both sides. The write path is as fast as before; what changes is +peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29). + ### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank) Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load