Chunk dimensions of 2^32 or more; copy-free unfiltered 4 GiB writes #32

Open
osobh wants to merge 7 commits from feat/huge-chunk-dims into main
Showing only changes of commit 7fba2f3996 - Show all commits
+36
View File
@@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th
same; files grow by about 36 bytes per chunk. The default therefore stays same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in. the 1.10 format; the 1.8 format is opt-in.
### Streaming `FileBuilder::write` (2026-09-29, tank)
`FileBuilder::write` now streams the file through a buffered writer
instead of building it in memory first (`feat/huge-chunk-dims`, the change
that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on
tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the
start of each run): `main` at `4260af4` against the stacked branch at
`549e442`, alternating, three full runs each; median (min-max) of
criterion's estimate, ms.
> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3`
The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower
on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g.
`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's
128 KiB mmap threshold, and after the smaller cases it was mapped afresh on
each write (1 629 368 minor page faults over the `write_2d_chunked` group
against 12 888 on `main`). Run alone, the case showed no difference. With a
64 KiB buffer (`549e442`):
| benchmark | main | candidate | ratio |
|---|---:|---:|---:|
| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x |
| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x |
| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x |
| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x |
| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x |
| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x |
| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x |
| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x |
| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x |
The other 18 cases are within 2%. Minor page faults per full run: 59 602 to
64 486 on both sides. The write path is as fast as before; what changes is
peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29).
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank) ### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load