Write files HDF5 1.8 can read: libver_bounds #28

Open
osobh wants to merge 3 commits from feat/libver-v18 into main
Owner

Merge order: 2 of 3. This PR is stacked on #27.

clawhdf5 can now write files that HDF5 1.8 reads. libhdf5 2.0 made 1.8 its default low bound.

  • New API: LibVer (V18 … V200, Latest) and libver_bounds(low, high) on FileWriter and FileBuilder, named after H5Pset_libver_bounds.
  • What a 1.8 low bound writes:
    • a version-2 superblock and version-3 layouts;
    • a version-1 B-tree for every chunked dataset, resizable ones included. It comes from a new B-tree writer, and the trees are node-for-node identical to libhdf5's (1-D, 2-D and 3-D; two and three levels).
  • What a 1.8 high bound refuses (FormatError::LibverBound, before anything is written): virtual datasets, a paged file, 1.12 references and native complex.
  • Default unchanged: (V110, Latest) writes the same bytes as before, and a unit test checks that.

Default kept, based on measurement. read_harness A/B on tank, idle re-run in #29's BENCHMARKS section: with 8192 chunks per dataset, a small selection on a freshly opened file is 1.2x to 2.3x slower through the v1 B-tree. Full reads and writes are within 7%, and files are about 0.4% larger. So the 1.8 format is opt-in.

Checks:

  • HDF5 1.8.23, built from source by scripts/build-hdf5-1.8.sh, prints every feature's file the way h5dump 1.14 does.
  • For every numeric dataset, 1.8's binary dump returns the exact bytes we wrote.
  • h5py, clawhdf5 and h5rs check --data agree, and FileEditor and h5py edits keep the file readable by 1.8.
  • CI has no HDF5 1.8, so those checks skip there.

Not done: the Python bindings' 'w' has no libver argument, and the pre-1.8 (earliest) format can't be written.

🤖 Generated with Claude Code

**Merge order: 2 of 3.** This PR is stacked on #27. clawhdf5 can now write files that HDF5 1.8 reads. libhdf5 2.0 made 1.8 its default low bound. - **New API:** `LibVer` (`V18` … `V200`, `Latest`) and `libver_bounds(low, high)` on `FileWriter` and `FileBuilder`, named after `H5Pset_libver_bounds`. - **What a 1.8 low bound writes:** - a version-2 superblock and version-3 layouts; - a version-1 B-tree for every chunked dataset, resizable ones included. It comes from a new B-tree writer, and the trees are node-for-node identical to libhdf5's (1-D, 2-D and 3-D; two and three levels). - **What a 1.8 high bound refuses** (`FormatError::LibverBound`, before anything is written): virtual datasets, a paged file, 1.12 references and native complex. - **Default unchanged:** `(V110, Latest)` writes the same bytes as before, and a unit test checks that. **Default kept, based on measurement.** `read_harness` A/B on tank, idle re-run in #29's BENCHMARKS section: with 8192 chunks per dataset, a small selection on a freshly opened file is **1.2x to 2.3x slower** through the v1 B-tree. Full reads and writes are within 7%, and files are about 0.4% larger. So the 1.8 format is opt-in. **Checks:** - HDF5 1.8.23, built from source by `scripts/build-hdf5-1.8.sh`, prints every feature's file the way h5dump 1.14 does. - For every numeric dataset, 1.8's binary dump returns the exact bytes we wrote. - h5py, clawhdf5 and `h5rs check --data` agree, and `FileEditor` and h5py edits keep the file readable by 1.8. - CI has no HDF5 1.8, so those checks skip there. **Not done:** the Python bindings' `'w'` has no `libver` argument, and the pre-1.8 (`earliest`) format can't be written. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
osobh added 3 commits 2026-09-29 05:00:59 +00:00
clawhdf5-netcdf4: variables' dimensions come from the file
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m24s
00b6f76ee0
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.

variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.

Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.

With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).

The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.

Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s
01d2a5dc5d
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
All checks were successful
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s
You are not authorized to merge this pull request.
This pull request can be merged automatically.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/libver-v18:feat/libver-v18
git checkout feat/libver-v18
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: quantumclaw/clawhdf5#28