docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks; known-issues moves "the writer does not produce output HDF5 1.8 can read" to Fixed (history) as an opt-in fix and keeps what is still open (Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs libhdf5's chunk cache); README's capability table lists 1.8 output and v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the two formats on the read harness (tank, 2026-09-28, under load), which is why the default stays the 1.10 format. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -536,6 +536,68 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
|
||||
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
|
||||
explain the slower windows.
|
||||
|
||||
### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank)
|
||||
|
||||
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), branch `feat/libver-v18`
|
||||
at `3c61635`, to decide whether the writer's default should become the
|
||||
HDF5 1.8 format (`libver_bounds(LibVer::V18, LibVer::V18)`: version-1 B-tree
|
||||
chunk indexes) instead of the 1.10 format (Fixed Array indexes for these
|
||||
datasets). **Not an idle machine:** two other agents were building; the
|
||||
1-minute load average was 7.0 to 7.6 throughout (the rule is below 2), so
|
||||
treat differences under about 20% as noise. Default and `--v18` runs
|
||||
alternated, three of each per chunk size; medians of the three.
|
||||
|
||||
> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]`
|
||||
> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset)
|
||||
|
||||
Write (the whole 3-dataset file) and file size:
|
||||
|
||||
| chunks | 1.10 write ms | 1.8 write ms | 1.10 file bytes | 1.8 file bytes | size |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| 256 x 256 | 467-500 | 467-483 | 134 916 080 | 134 933 848 | +0.013% |
|
||||
| 32 x 32 | 490-495 | 489-506 | 138 879 336 | 139 465 112 | +0.42% |
|
||||
|
||||
Reads (chunked datasets; the contiguous one does not change), ms. The
|
||||
window labels are the harness's, which count 256 x 256 chunks: with 32 x 32
|
||||
chunks the 64 x 64 window covers 9 chunks and the 512 x 512 one 289.
|
||||
|
||||
| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 |
|
||||
|---|---|---|---:|---:|---:|
|
||||
| 256 | chunked + deflate | full (first) | 6.70 | 7.30 | 1.09x |
|
||||
| 256 | chunked + deflate | full (repeat) | 4.80 | 3.80 | 0.79x |
|
||||
| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 | 0.16 | 1.07x |
|
||||
| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 2.40 | 1.25 | 0.52x |
|
||||
| 256 | chunked + deflate | one row | 0.95 | 0.95 | 1.00x |
|
||||
| 256 | chunked + deflate | one column | 1.93 | 1.94 | 1.01x |
|
||||
| 256 | chunked | full (first) | 11.30 | 11.80 | 1.04x |
|
||||
| 256 | chunked | full (repeat) | 6.20 | 5.60 | 0.90x |
|
||||
| 256 | chunked | 64 x 64 window (1 chunk) | 0.04 | 0.04 | 1.00x |
|
||||
| 256 | chunked | 512 x 512 window (4-9 chunks) | 1.50 | 0.37 | 0.25x |
|
||||
| 256 | chunked | one row | 0.05 | 0.05 | 1.00x |
|
||||
| 256 | chunked | one column | 0.44 | 0.45 | 1.02x |
|
||||
| 32 | chunked + deflate | full (first) | 11.50 | 12.20 | 1.06x |
|
||||
| 32 | chunked + deflate | full (repeat) | 9.30 | 9.30 | 1.00x |
|
||||
| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.45 | 0.89 | 1.98x |
|
||||
| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.97 | 2.41 | 1.22x |
|
||||
| 32 | chunked + deflate | one row | 0.69 | 1.13 | 1.64x |
|
||||
| 32 | chunked + deflate | one column | 1.26 | 1.70 | 1.35x |
|
||||
| 32 | chunked | full (first) | 11.80 | 13.00 | 1.10x |
|
||||
| 32 | chunked | full (repeat) | 6.30 | 6.20 | 0.98x |
|
||||
| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 | 0.83 | 2.31x |
|
||||
| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 | 1.16 | 1.66x |
|
||||
| 32 | chunked | one row | 0.38 | 0.84 | 2.21x |
|
||||
| 32 | chunked | one column | 1.05 | 1.21 | 1.15x |
|
||||
|
||||
With 128 chunks per dataset the two formats read and write alike (the 512 x
|
||||
512 windows' 0.25x and 0.52x are not explained by the index and are likely
|
||||
the load). With 8192 chunks, full reads stay within 10%, but a selection on
|
||||
a freshly opened file costs about 0.4 to 0.5 ms more through the version-1
|
||||
B-tree (1.2x to 2.3x). Each timed selection opens the file anew, so the
|
||||
likely cause (not profiled) is walking the B-tree's nodes (2.6 KB each,
|
||||
about 150 per dataset here) against a Fixed Array's few blocks. Writing costs the
|
||||
same; files grow by about 36 bytes per chunk. The default therefore stays
|
||||
the 1.10 format; the 1.8 format is opt-in.
|
||||
|
||||
## Local file speed after range reads
|
||||
|
||||
### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)
|
||||
|
||||
Reference in New Issue
Block a user