docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s

CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 23:44:48 -05:00
co-authored by Claude Opus 5.5
parent b5a5041655
commit 01d2a5dc5d
5 changed files with 176 additions and 6 deletions
+51 -2
View File
@@ -197,8 +197,16 @@ wrong data.
`filter_registry::register_filter`.
- **Writer:** in dense storage (more than 8 attributes on an object, or
more than 8 links in a group) one attribute or link message over 65 515
bytes is an error (no huge fractal-heap objects). The writer does not
produce output that HDF5 1.8 can read.
bytes is an error (no huge fractal-heap objects). Files HDF5 1.8 reads
are opt-in (`libver_bounds(LibVer::V18, LibVer::V18)`, since 2026-09-28,
[history](#hdf5-18-could-not-read-the-files-we-wrote)): by default the
writer uses the HDF5 1.10 format. The pre-1.8 format (version-0
superblock, symbol-table groups; h5py's `libver='earliest'`) cannot be
written, and the Python bindings' `'w'` mode has no `libver` argument.
Under the 1.8 bound, chunk B-trees equal libhdf5's node for node when
libhdf5 inserts the chunks in row-major order; libhdf5 with its chunk
cache inserts the small chunks of a multi-dimensional dataset in eviction
order, which fills the nodes differently (same chunks, same keys).
- **Checks we deliberately do not make:**
- a float sign bit position outside the type, and a size-0 string type:
clawhdf5 up to v2.7.0 wrote them;
@@ -515,6 +523,47 @@ Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files.
## HDF5 1.8 could not read the files we wrote
**Status:** fixed 2026-09-28 (branch `feat/libver-v18`), as an opt-in.
Affected every release (v2.1.0 to v2.7.0): the writer only ever wrote the
HDF5 1.10 format. Users who need HDF5 1.8 to read their files call
`FileBuilder::libver_bounds(LibVer::V18, LibVer::V18)` (format crate:
`FileWriter::libver_bounds`) and write them again; the default output is
unchanged. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
HDF5 1.8.23's h5dump refused a default clawhdf5 file outright ("unable to
open file": its superblock is version 3; tank, 2026-09-28). The files
also use version-4 layout messages and the 1.10 chunk indexes (single
chunk, Fixed Array, Extensible Array, version-2 B-tree). libhdf5 2.0 made
the 1.8 format its default low bound (`H5F_LIBVER_V18`).
With a low bound of 1.8 the writer now writes what libhdf5 2.x writes for
h5py's `libver=('v108', 'latest')`: a version-2 superblock, version-3
layout messages, and a version-1 B-tree for every chunked dataset,
resizable ones included, built the way `H5B_insert` builds it (checked node
for node against libhdf5's trees). Everything else the writer emits was
already 1.8's (version-2 object headers, link messages, dense storage in
fractal heaps and version-2 B-trees, filter pipeline version 2, fill value
version 3, datatypes up to version 3). With a high bound of 1.8, what 1.8
cannot read is `FormatError::LibverBound` before anything is written:
virtual datasets, the paged file-space strategy, the 1.12 reference types,
native complex numbers.
`crates/clawhdf5-tools/tests/libver_v18.rs` writes every writer feature
under the 1.8 bound; HDF5 1.8.23's h5dump (built by
`scripts/build-hdf5-1.8.sh`; skipped where it is missing, as in CI) dumps
the whole file exactly as h5dump 1.14 does and returns our bytes for each
numeric dataset, and h5py, clawhdf5 and `h5rs check --data` agree; again
after `FileEditor` grows and appends to it (splitting B-tree nodes) and
sets attributes, and after h5py appends.
The default stays the 1.10 format: on the read harness (tank, 2026-09-28,
under load from other builds) a freshly opened file with 8192 chunks
per dataset reads small selections 1.2x to 2.3x slower through a version-1
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release