Write files HDF5 1.8 can read: libver_bounds #28

Merged
osobh merged 2 commits from feat/libver-v18 into main 2026-09-29 11:57:36 +00:00
5 changed files with 176 additions and 6 deletions
Showing only changes of commit 01d2a5dc5d - Show all commits
+62
View File
@@ -536,6 +536,68 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
explain the slower windows.
### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), branch `feat/libver-v18`
at `3c61635`, to decide whether the writer's default should become the
HDF5 1.8 format (`libver_bounds(LibVer::V18, LibVer::V18)`: version-1 B-tree
chunk indexes) instead of the 1.10 format (Fixed Array indexes for these
datasets). **Not an idle machine:** two other agents were building; the
1-minute load average was 7.0 to 7.6 throughout (the rule is below 2), so
treat differences under about 20% as noise. Default and `--v18` runs
alternated, three of each per chunk size; medians of the three.
> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]`
> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset)
Write (the whole 3-dataset file) and file size:
| chunks | 1.10 write ms | 1.8 write ms | 1.10 file bytes | 1.8 file bytes | size |
|---|---:|---:|---:|---:|---:|
| 256 x 256 | 467-500 | 467-483 | 134 916 080 | 134 933 848 | +0.013% |
| 32 x 32 | 490-495 | 489-506 | 138 879 336 | 139 465 112 | +0.42% |
Reads (chunked datasets; the contiguous one does not change), ms. The
window labels are the harness's, which count 256 x 256 chunks: with 32 x 32
chunks the 64 x 64 window covers 9 chunks and the 512 x 512 one 289.
| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 |
|---|---|---|---:|---:|---:|
| 256 | chunked + deflate | full (first) | 6.70 | 7.30 | 1.09x |
| 256 | chunked + deflate | full (repeat) | 4.80 | 3.80 | 0.79x |
| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 | 0.16 | 1.07x |
| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 2.40 | 1.25 | 0.52x |
| 256 | chunked + deflate | one row | 0.95 | 0.95 | 1.00x |
| 256 | chunked + deflate | one column | 1.93 | 1.94 | 1.01x |
| 256 | chunked | full (first) | 11.30 | 11.80 | 1.04x |
| 256 | chunked | full (repeat) | 6.20 | 5.60 | 0.90x |
| 256 | chunked | 64 x 64 window (1 chunk) | 0.04 | 0.04 | 1.00x |
| 256 | chunked | 512 x 512 window (4-9 chunks) | 1.50 | 0.37 | 0.25x |
| 256 | chunked | one row | 0.05 | 0.05 | 1.00x |
| 256 | chunked | one column | 0.44 | 0.45 | 1.02x |
| 32 | chunked + deflate | full (first) | 11.50 | 12.20 | 1.06x |
| 32 | chunked + deflate | full (repeat) | 9.30 | 9.30 | 1.00x |
| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.45 | 0.89 | 1.98x |
| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.97 | 2.41 | 1.22x |
| 32 | chunked + deflate | one row | 0.69 | 1.13 | 1.64x |
| 32 | chunked + deflate | one column | 1.26 | 1.70 | 1.35x |
| 32 | chunked | full (first) | 11.80 | 13.00 | 1.10x |
| 32 | chunked | full (repeat) | 6.30 | 6.20 | 0.98x |
| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 | 0.83 | 2.31x |
| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 | 1.16 | 1.66x |
| 32 | chunked | one row | 0.38 | 0.84 | 2.21x |
| 32 | chunked | one column | 1.05 | 1.21 | 1.15x |
With 128 chunks per dataset the two formats read and write alike (the 512 x
512 windows' 0.25x and 0.52x are not explained by the index and are likely
the load). With 8192 chunks, full reads stay within 10%, but a selection on
a freshly opened file costs about 0.4 to 0.5 ms more through the version-1
B-tree (1.2x to 2.3x). Each timed selection opens the file anew, so the
likely cause (not profiled) is walking the B-tree's nodes (2.6 KB each,
about 150 per dataset here) against a Fixed Array's few blocks. Writing costs the
same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in.
## Local file speed after range reads
### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)
+56
View File
@@ -55,6 +55,62 @@
skips (opaque, references). Affected v2.1.0 to v2.7.0.
`docs/known-issues.md`.
### Writing files HDF5 1.8 can read (2026-09-28)
- New `LibVer` (`V18`, `V110`, `V112`, `V114`, `V200`, `Latest`; re-exported
as `clawhdf5::LibVer`) and `FileBuilder::libver_bounds(low, high)`
(`FileWriter::libver_bounds` in the format crate), as libhdf5's
`H5Pset_libver_bounds` and h5py's `libver=(low, high)`. The default,
`(V110, Latest)`, writes exactly what was written before (the HDF5 1.10
format, which HDF5 1.8.23 refuses to open).
- A low bound of `V18` writes what libhdf5 2.x writes for h5py's
`libver=('v108', 'latest')`: a version-2 superblock (version 3 only for a
paged file), version-3 layout messages for contiguous, compact and chunked
datasets, and a version-1 B-tree chunk index for every chunked dataset,
resizable and multi-unlimited ones included, instead of the single chunk,
Fixed Array, Extensible Array and version-2 B-tree indexes. The B-tree
writer (`btree_v1_write`) replays `H5B_insert` with the chunk callbacks of
`H5Dbtree.c` for chunks arriving in row-major order: libhdf5's split
ratios, right keys moved exactly when `H5D__btree_cmp3` moves them, the
root kept at its address. Its trees are libhdf5's node for node (levels,
child counts, keys) for 1-D, 2-D and 3-D datasets with two- and
three-level trees, deflated or not, against libhdf5 2.0 writing without a
chunk cache (`chunk_btrees_match_libhdf5`). An empty chunked dataset has
no tree (undefined address), as in libhdf5.
- A high bound refuses, with the new `FormatError::LibverBound` and before
anything is written, what needs a newer format: virtual datasets and the
paged file-space strategy (1.10), the 1.12 reference types (datatype
version 4), native complex numbers (datatype version 5, HDF5 2.0), and a
low bound above the high one. `(V18, V18)` therefore writes a file HDF5
1.8 reads, or fails; `(V18, Latest)` writes such objects in their newer
format, as libhdf5 does. `Datatype::max_encoded_version` reports the
version a type needs.
- Checked against a real HDF5 1.8: `scripts/build-hdf5-1.8.sh` builds
1.8.23 (the last 1.8 release) with its tools. `clawhdf5-tools`'
`tests/libver_v18.rs` writes every writer feature under `(V18, V18)`
(contiguous, empty, scalar, compact, f16/f32/f64/i64, chunked with
deflate + shuffle + Fletcher-32 and edge chunks, resizable with one and
two unlimited dimensions, finite maxshape, 100 000 one-element chunks for
a three-level tree, a 100 x 100 grid of chunks, fill values, fixed-length
strings, compound, enum, array, compact/dense/creation-ordered groups,
soft, hard and external links, dense attributes); HDF5 1.8.23's h5dump
dumps the whole file exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset (`-b LE`), and h5py, clawhdf5 and
`h5rs check --data` read it. Then `FileEditor` appends to the resizable
datasets 300 times (splitting B-tree nodes), grows the three-level one,
rewrites a deflated row, grows a 2-D dataset and sets compact and dense
attributes; h5py appends too; and every reader checks again. The 1.8
checks are skipped where no HDF5 1.8 is found (`CLAWHDF5_H5DUMP18` or
`~/.cache/hdf5-1.8.23`), as in CI.
- The default stays the 1.10 format: with 8192 chunks per dataset, small
selections on a freshly opened file read 1.2x to 2.3x slower through a
version-1 B-tree (read harness, tank, 2026-09-28, under load; full reads
and writes within 10%, files about 36 bytes per chunk larger).
`read_harness` gains `--v18` and `--chunk N`. `BENCHMARKS.md`, "HDF5 1.8
format".
- Not yet: the Python bindings' `'w'` mode has no `libver` argument, and
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
written. `docs/known-issues.md`.
### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked
+4 -3
View File
@@ -13,7 +13,8 @@ It reads superblocks v0–3, every group and chunk-index structure libhdf5
writes, the standard filters and the common plugin filters,
variable-length data and virtual datasets, and follows files a SWMR writer
is appending to. It reads files from libhdf5, h5py and netCDF-4 and writes
files they read. The same library opens files over HTTP and in object
files they read, in the HDF5 1.10 format or, on request, in one HDF5 1.8
reads. The same library opens files over HTTP and in object
stores by range requests, runs in the browser as WebAssembly, and has
Python bindings with an h5py-shaped API.
@@ -112,10 +113,10 @@ Limits and open issues, with dates, are in
| Area | Supported | Read only | Not supported |
|---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers | Metadata cache images | Writing files HDF5 1.8 can read |
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does); fill values; resizable datasets; virtual datasets (read limits in known-issues) | Chunk indexes v1 B-tree and implicit (the editor also changes them) | External raw data files (explicit error) |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
+3 -1
View File
@@ -36,7 +36,9 @@ clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
`type_builders` (datasets, groups, attributes, compound and enum types,
links, virtual datasets, creation-order tracking); chunk indexes and
dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`,
`ea_writer`). Output is read by h5py and h5dump.
`ea_writer`, and the version-1 chunk B-tree of `btree_v1_write`). Output
is read by h5py and h5dump; `FileWriter::libver_bounds` (`libver`) picks
the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other
IDs can be registered at run time with `register_filter`). Built in:
deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4,
+51 -2
View File
@@ -197,8 +197,16 @@ wrong data.
`filter_registry::register_filter`.
- **Writer:** in dense storage (more than 8 attributes on an object, or
more than 8 links in a group) one attribute or link message over 65 515
bytes is an error (no huge fractal-heap objects). The writer does not
produce output that HDF5 1.8 can read.
bytes is an error (no huge fractal-heap objects). Files HDF5 1.8 reads
are opt-in (`libver_bounds(LibVer::V18, LibVer::V18)`, since 2026-09-28,
[history](#hdf5-18-could-not-read-the-files-we-wrote)): by default the
writer uses the HDF5 1.10 format. The pre-1.8 format (version-0
superblock, symbol-table groups; h5py's `libver='earliest'`) cannot be
written, and the Python bindings' `'w'` mode has no `libver` argument.
Under the 1.8 bound, chunk B-trees equal libhdf5's node for node when
libhdf5 inserts the chunks in row-major order; libhdf5 with its chunk
cache inserts the small chunks of a multi-dimensional dataset in eviction
order, which fills the nodes differently (same chunks, same keys).
- **Checks we deliberately do not make:**
- a float sign bit position outside the type, and a size-0 string type:
clawhdf5 up to v2.7.0 wrote them;
@@ -515,6 +523,47 @@ Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files.
## HDF5 1.8 could not read the files we wrote
**Status:** fixed 2026-09-28 (branch `feat/libver-v18`), as an opt-in.
Affected every release (v2.1.0 to v2.7.0): the writer only ever wrote the
HDF5 1.10 format. Users who need HDF5 1.8 to read their files call
`FileBuilder::libver_bounds(LibVer::V18, LibVer::V18)` (format crate:
`FileWriter::libver_bounds`) and write them again; the default output is
unchanged. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
HDF5 1.8.23's h5dump refused a default clawhdf5 file outright ("unable to
open file": its superblock is version 3; tank, 2026-09-28). The files
also use version-4 layout messages and the 1.10 chunk indexes (single
chunk, Fixed Array, Extensible Array, version-2 B-tree). libhdf5 2.0 made
the 1.8 format its default low bound (`H5F_LIBVER_V18`).
With a low bound of 1.8 the writer now writes what libhdf5 2.x writes for
h5py's `libver=('v108', 'latest')`: a version-2 superblock, version-3
layout messages, and a version-1 B-tree for every chunked dataset,
resizable ones included, built the way `H5B_insert` builds it (checked node
for node against libhdf5's trees). Everything else the writer emits was
already 1.8's (version-2 object headers, link messages, dense storage in
fractal heaps and version-2 B-trees, filter pipeline version 2, fill value
version 3, datatypes up to version 3). With a high bound of 1.8, what 1.8
cannot read is `FormatError::LibverBound` before anything is written:
virtual datasets, the paged file-space strategy, the 1.12 reference types,
native complex numbers.
`crates/clawhdf5-tools/tests/libver_v18.rs` writes every writer feature
under the 1.8 bound; HDF5 1.8.23's h5dump (built by
`scripts/build-hdf5-1.8.sh`; skipped where it is missing, as in CI) dumps
the whole file exactly as h5dump 1.14 does and returns our bytes for each
numeric dataset, and h5py, clawhdf5 and `h5rs check --data` agree; again
after `FileEditor` grows and appends to it (splitting B-tree nodes) and
sets attributes, and after h5py appends.
The default stays the 1.10 format: on the read harness (tank, 2026-09-28,
under load from other builds) a freshly opened file with 8192 chunks
per dataset reads small selections 1.2x to 2.3x slower through a version-1
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release