docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks; known-issues moves "the writer does not produce output HDF5 1.8 can read" to Fixed (history) as an opt-in fix and keeps what is still open (Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs libhdf5's chunk cache); README's capability table lists 1.8 output and v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the two formats on the read harness (tank, 2026-09-28, under load), which is why the default stays the 1.10 format. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -55,6 +55,62 @@
|
||||
skips (opaque, references). Affected v2.1.0 to v2.7.0.
|
||||
`docs/known-issues.md`.
|
||||
|
||||
### Writing files HDF5 1.8 can read (2026-09-28)
|
||||
- New `LibVer` (`V18`, `V110`, `V112`, `V114`, `V200`, `Latest`; re-exported
|
||||
as `clawhdf5::LibVer`) and `FileBuilder::libver_bounds(low, high)`
|
||||
(`FileWriter::libver_bounds` in the format crate), as libhdf5's
|
||||
`H5Pset_libver_bounds` and h5py's `libver=(low, high)`. The default,
|
||||
`(V110, Latest)`, writes exactly what was written before (the HDF5 1.10
|
||||
format, which HDF5 1.8.23 refuses to open).
|
||||
- A low bound of `V18` writes what libhdf5 2.x writes for h5py's
|
||||
`libver=('v108', 'latest')`: a version-2 superblock (version 3 only for a
|
||||
paged file), version-3 layout messages for contiguous, compact and chunked
|
||||
datasets, and a version-1 B-tree chunk index for every chunked dataset,
|
||||
resizable and multi-unlimited ones included, instead of the single chunk,
|
||||
Fixed Array, Extensible Array and version-2 B-tree indexes. The B-tree
|
||||
writer (`btree_v1_write`) replays `H5B_insert` with the chunk callbacks of
|
||||
`H5Dbtree.c` for chunks arriving in row-major order: libhdf5's split
|
||||
ratios, right keys moved exactly when `H5D__btree_cmp3` moves them, the
|
||||
root kept at its address. Its trees are libhdf5's node for node (levels,
|
||||
child counts, keys) for 1-D, 2-D and 3-D datasets with two- and
|
||||
three-level trees, deflated or not, against libhdf5 2.0 writing without a
|
||||
chunk cache (`chunk_btrees_match_libhdf5`). An empty chunked dataset has
|
||||
no tree (undefined address), as in libhdf5.
|
||||
- A high bound refuses, with the new `FormatError::LibverBound` and before
|
||||
anything is written, what needs a newer format: virtual datasets and the
|
||||
paged file-space strategy (1.10), the 1.12 reference types (datatype
|
||||
version 4), native complex numbers (datatype version 5, HDF5 2.0), and a
|
||||
low bound above the high one. `(V18, V18)` therefore writes a file HDF5
|
||||
1.8 reads, or fails; `(V18, Latest)` writes such objects in their newer
|
||||
format, as libhdf5 does. `Datatype::max_encoded_version` reports the
|
||||
version a type needs.
|
||||
- Checked against a real HDF5 1.8: `scripts/build-hdf5-1.8.sh` builds
|
||||
1.8.23 (the last 1.8 release) with its tools. `clawhdf5-tools`'
|
||||
`tests/libver_v18.rs` writes every writer feature under `(V18, V18)`
|
||||
(contiguous, empty, scalar, compact, f16/f32/f64/i64, chunked with
|
||||
deflate + shuffle + Fletcher-32 and edge chunks, resizable with one and
|
||||
two unlimited dimensions, finite maxshape, 100 000 one-element chunks for
|
||||
a three-level tree, a 100 x 100 grid of chunks, fill values, fixed-length
|
||||
strings, compound, enum, array, compact/dense/creation-ordered groups,
|
||||
soft, hard and external links, dense attributes); HDF5 1.8.23's h5dump
|
||||
dumps the whole file exactly as h5dump 1.14 does and returns our bytes
|
||||
for every numeric dataset (`-b LE`), and h5py, clawhdf5 and
|
||||
`h5rs check --data` read it. Then `FileEditor` appends to the resizable
|
||||
datasets 300 times (splitting B-tree nodes), grows the three-level one,
|
||||
rewrites a deflated row, grows a 2-D dataset and sets compact and dense
|
||||
attributes; h5py appends too; and every reader checks again. The 1.8
|
||||
checks are skipped where no HDF5 1.8 is found (`CLAWHDF5_H5DUMP18` or
|
||||
`~/.cache/hdf5-1.8.23`), as in CI.
|
||||
- The default stays the 1.10 format: with 8192 chunks per dataset, small
|
||||
selections on a freshly opened file read 1.2x to 2.3x slower through a
|
||||
version-1 B-tree (read harness, tank, 2026-09-28, under load; full reads
|
||||
and writes within 10%, files about 36 bytes per chunk larger).
|
||||
`read_harness` gains `--v18` and `--chunk N`. `BENCHMARKS.md`, "HDF5 1.8
|
||||
format".
|
||||
- Not yet: the Python bindings' `'w'` mode has no `libver` argument, and
|
||||
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
|
||||
written. `docs/known-issues.md`.
|
||||
|
||||
### A dropped `FileEditor` releases its lock at once (2026-09-28)
|
||||
- `FileEditor`'s `flock` could outlive the editor for a moment when
|
||||
another thread forked to spawn a process: the child shared the locked
|
||||
|
||||
Reference in New Issue
Block a user