diff --git a/docs/known-issues.md b/docs/known-issues.md index 0b25635..45e42de 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -1,189 +1,38 @@ # Known Issues -Bugs found during development or downstream use, tracked here because this -repository's issue tracker is disabled. One entry per bug; when an entry is -fixed, record the fix in `CHANGELOG.md` and update its status here rather than -deleting it. +Bugs and limits found during development or downstream use, tracked here +because this repository's issue tracker is disabled. One entry per bug; when +an entry is fixed, record the fix in `CHANGELOG.md` and move it to +[Fixed (history)](#fixed-history) with its date, the releases it affected and +what users must do, rather than deleting it. + +Releases referred to below: v2.1.0 (2026-06-03) to v2.7.0 (2026-09-20). +Everything fixed after v2.7.0 is on `main` and unreleased. + +## Open issues + +Checked against `main` at `9b5803f` on 2026-09-28. + +| Issue | Kind | Since | +|---|---|---| +| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | +| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | +| [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 | +| [HDF5 features still unsupported](#hdf5-features-still-unsupported) | errors, never wrong data | 2026-09-25 audit | +| [External links and external raw data are not followed](#external-links-and-external-raw-data-are-not-followed) | error, by design for now | 2026-09-19 | +| [Range reads (`File::open_storage`) limits](#range-reads-fileopen_storage-limits) | cost; zero-copy APIs need an in-memory file | 2026-09-26 | +| [Remote files (`clawhdf5-remote`) limits](#remote-files-clawhdf5-remote-limits) | untested backends, fixed block size, timeouts | 2026-09-26 | +| [`clawhdf5-wasm` (browser) limits](#clawhdf5-wasm-browser-limits) | memory, round trips, unsupported types | 2026-09-26 | +| [The Node.js package does not work](#the-nodejs-package-packagesclawhdf5-node-does-not-work) | broken, unpublished, not in CI | 2026-09-25 | --- -## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) - -**Status:** fixed 2026-09-27 (`96086ad`, branch -`perf/header-parse-and-last-files`): the remaining cost was the call to the -version-1 message loop, kept out of line by `#[inline(never)]`; inlined, -`object_header_parse_x401` is 1.0% and 2.6% below `8f59b2e` in two idle -A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's -speed"), with every chunk-queue check unchanged. - -Was: open (speed only; values are correct). Measured on an idle tank -(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) -against `main` `7a8fae0`, separate binaries alternating, 3 rounds -(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"): -`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges -23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read -continuation chunks from a bounded queue since `a69c5be` (bounding what a -crafted header can make it read); `4313917` removed its per-header -allocations but not all of the cost. Listing the 400-group file through the -facade, which parses the same headers, is 1.7% faster, so no user-visible -path is slower. - -An earlier run the same day under load also listed single-thread contiguous -hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with -overlapping ranges (noise), so that item is withdrawn. - -## Conformance: the last non-ok files (checked 2026-09-27) - -**Status:** classified with evidence — none is a clawhdf5 bug. The -conformance report had 3 our-errors and 2 mismatches left, each documented -as "not ours" but still counted against us. Re-checked on tank -(h5py 3.16 / HDF5 2.0.0, h5dump 1.14.6): - -- **2 mismatches, h5py's big-endian VL bug:** - `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64` and - `hdf5/tools/test/testfiles/tcomplex_be.h5` - `/VariableLengthDatasetFloatComplex`. h5py returns the elements with the - file's big-endian bytes under a little-endian dtype (a `vlen('>f4')` - holding `[1.0, 2.0]` reads back as `[4.6e-41, 9.0e-44]`). h5dump 1.14.6 - prints `(1, 2), (3, 4, 5), (42)` for `vlen_uint64`, which is what we read; - it cannot print `tcomplex_be.h5` (complex types are HDF5 2.0). The - reference side (`conformance/ref.py`) now checks that the installed h5py - has the bug and relabels such elements with the file's byte order, so the - values are compared rather than excused: both objects are identical to - ours, and both files are now *ok*. -- **3 our-errors, objects HDF5 2.0 reads only by over-reading memory:** - `cve-2025-2308.h5` `/Scale_offset_long_long_data_le` (first chunk records - `minbits` 11; its 12 values need 17 bytes of codes and the 26-byte chunk - holds 5 after its 21-byte header), `cve-2025-44904.h5` - `/Scale_offset_float_data_le` (unfiltered chunks stored as 38 and 37 bytes - for 48-byte chunks; 1.14.6's `H5D__chunk_lock` reads the stored bytes - into a buffer of that size and uses it as the whole chunk) and - `bad_nbit_parms_walk.h5` `/Nbit_int_data_le` (N-Bit parameters - `(7, 0, 40, 1, 4, 0, 20)`: an integer needs 8, and the decoder takes the - bit offset from `cd_values[7]`, past the list). **Could we match - libhdf5?** No: its output is not determined by the file. h5py's values for - the first two change from one run to the next (three plain runs gave three - different results), and all three change with `MALLOC_PERTURB_` or with - whether numpy is imported before h5py (`bad_nbit_parms_walk` then fails - with "filter returned failure during read" or reads all zeros); h5dump - 1.14.6 prints yet other values for the over-read parts (`1280, 0, 0` where - h5py gave e.g. `1283, 749, 1713`). libhdf5's develop branch refuses all - three. We keep - refusing them. `conformance/ref_bugs.py` repeats the check in every - conformance run (six reads per object in differently set-up processes) and - such a file is classified *ref-bug* only while its values keep changing. - -Result (`conformance/run.sh --no-fetch`, tank, 2026-09-27): 602 of 697 ok -(was 600), 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read. - -## Files a SWMR writer had open could not be read past a stale end of file - -**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any -release: reads have been bounded by the recorded end of file since -`7d7a7e7` (2026-09-26), which no release contains. - -A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the -superblock's SWMR-write flag and does not keep its end-of-file address up to -date: a copy h5py made of its own file mid-write records 715 in a -6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only -ignored the recorded end when it lay *past* the end of the file, so every -reader bounded such a file at 715 bytes: it listed, but every chunked read -failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check` -reported the chunk indexes past the end of the file. Never wrong data. - -**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the -end of the file, as libhdf5's SWMR reader reads it. **Test:** -`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy, -`tests/fixtures/swmr_mid_write.h5`, through every open path and against -h5py's SWMR reader). A file still being written is read with -`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file -at its length at open and is not meant for files that change while open. - -## Shrinking a chunked dataset with no recorded maximum scrambled it - -**Status:** fixed 2026-09-27, before any release (`FileEditor::resize` -shipped on main in PR #18, a4c2ace). Files the writer produced before the -fix still lack the maximum; see *Existing files*. - -clawhdf5's writer stored no maximum dimensions for a chunked dataset -created without a `maxshape` (libhdf5 always stores one, equal to the -dimensions when none is given). With none recorded, the maximum is the -current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index -places chunks by the maximum. `FileEditor::resize` changed only the -current dimensions, so a shrink moved every existing chunk and every -reader returned wrong values; after a shrink the dataset could not grow -back. Found by the review of the Python editing work. - -**Fix:** before the first resize of such a dataset the editor records the -maximum libhdf5 would have written (the dimensions the index was built -with); the writer now records it for every chunked dataset. **Test:** -`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:** -datasets written before the fix have no recorded maximum. The fixed -editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`) -does not** — it scrambles them the same way, and lets them grow past -their Fixed Array. Resize them once with a fixed `FileEditor` (a resize -to the same shape changes nothing; shrink and grow back) before letting -libhdf5 resize them. A dataset already shrunk by the unfixed -editor holds misplaced chunks; rewrite it from a good copy. - -## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 - -**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to -v2.7.0) is affected**, in both directions. - -Our Fletcher-32 reduced its two running sums with `% 65535`; libhdf5's -`H5_checksum_fletcher32` (H5checksum.c) folds them with -`(s & 0xffff) + (s >> 16)`. Both are arithmetic mod 65535, but where a sum -is a non-zero multiple of 65535 the fold leaves 0xffff and the modulo 0, so -the checksums differ — for random data about one chunk in 32768 (each of -the two sums hits it with probability about 1/65535). Found by the review -of the editor work: a random-edit fuzzer with gzip + Fletcher-32 hit it on -13 of about 100 seeds. - -- Chunks we wrote (`FileBuilder`/`FileWriter` `with_fletcher32`, and the - unreleased `FileEditor`) with such a sum are refused by h5py and libhdf5: - "filter returned failure during read". h5py writing `[1, 0xfffe]` as - big-endian `u2` stores checksum `0x0001ffff`; we computed `0x00010000`. -- Chunks libhdf5 wrote with such a sum were refused by every reader here - with `Fletcher32Mismatch`; the data itself was never wrong. - -**Fix:** `clawhdf5_format::checksum::fletcher32`, a port of -`H5_checksum_fletcher32`, used by the filter for writing and verifying. It -also accepts a checksum whose 16-bit halves are byte-swapped, as libhdf5 -does for files from 1.6.2 and earlier, and the `% 65535` form clawhdf5 -v2.7.0 and earlier wrote (the two differ only in a half that is 0xffff). -**Test:** -`crates/clawhdf5/tests/fletcher32_interop.rs` (libhdf5's own function -through ctypes on every 1- and 2-byte input plus 40 000 random and -fold-heavy inputs; h5py reads fold-case chunks from `FileBuilder` and -`FileEditor`; we read h5py's). **Existing data:** a Fletcher-32 dataset -written by v2.7.0 or earlier may hold chunks libhdf5 cannot read; a fixed -build reads them. Rewrite such datasets with a fixed build (read, then -write them again) before handing the file to libhdf5 or h5py. - -## LZF/Blosc chunks written with a stale filter mask - -**Status:** fixed 2026-09-26, before any release (the LZF and Blosc writers -were added the same day; v2.7.0 and earlier write neither). - -`FileBuilder` stored every chunk of an LZF or Blosc dataset through the -filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller -than the chunk as a failure of the (optional) filter and stores the chunk -raw with the filter's mask bit set. When a chunk's LZF stream was exactly -the chunk's size, the first libhdf5 rewrite of it stored raw data at the -same size and left our mask 0 in the index, so h5py could no longer read -the dataset. `FileEditor` had the same bug, fixed earlier the same day. -Both now use `clawhdf5_format::filters::compress_chunk_masked`, and every -chunk index the writer builds records the real mask (see `CHANGELOG.md`). -Files written before the fix read correctly; rewrite them before letting -libhdf5 modify them. - ## In-place modification (`FileEditor`) limits -**Status:** open (documented 2026-09-26, updated the same day when -version-2 B-tree chunk indexes, shrinking, dense attributes and space -reuse were added). `clawhdf5::FileEditor` refuses, with -`Error::Unsupported` and without writing anything: +**Status:** open (documented 2026-09-26, updated when version-2 B-tree +chunk indexes, shrinking, dense attributes and space reuse were added). +`clawhdf5::FileEditor` refuses, with `Error::Unsupported` and without +writing anything: - new chunks in an **implicit** index (it has all of its chunks from the start; they are written in place, and allocated/filled on growth under early allocation as libhdf5 does); @@ -201,11 +50,7 @@ reuse were added). `clawhdf5::FileEditor` refuses, with random-edit harness (120 runs of 150 random edits, `earliest`/`v110`/ `latest`, about 5600 `set_attr` calls of 8 bytes to 6 KiB): 2.2% of `set_attr` calls are refused, every one the last-attribute-in-a-block - replacement; before blocks could be skipped (an attribute needing a heap - block larger than the next one — any attribute of about 1 KiB or more - once a heap has started, or at the move to dense storage), 24% were, - since an object whose move to dense storage was refused kept refusing - every new attribute; + replacement (24% before heap blocks could be skipped); - version-1 object headers asked for an attribute larger than a header message (they have no dense storage); - partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external @@ -219,18 +64,19 @@ heap's replaced blocks) is reused by later edits of the same `FileEditor`; what is left when it is dropped is leaked, as libhdf5 leaks it without a persistent free-space manager (`h5repack` reclaims it). A chunk that is the last thing in the file grows in place, which covers the usual append. -Measured 2026-09-26 on tank with `cargo test -p clawhdf5-tools --test -edit_interop -- --ignored --nocapture measure_append_waste` (one editor for -the whole workload; file sizes are deterministic): 1000 appends of 100 `f8` -values to a 1-D dataset with 1024-element chunks give 810 504 bytes -unfiltered, as libhdf5's file (`h5repack` of either: 810 360), and 306 780 -bytes with gzip (307 210 before reuse; libhdf5's file: 306 058; `h5repack` -of the editor's file: 306 104, of libhdf5's: 305 954); 2000 appends of 10 -values with 4096-element gzip chunks give 79 829 bytes (119 684 before -reuse) against libhdf5's 50 292 (`h5repack` of the editor's file: 49 930, -of libhdf5's: 50 188): the chunk being appended to -is followed by new index blocks and moves each time it grows, and the -space it leaves is too small for its next, larger version. +Measured 2026-09-26 on tank (`cargo test -p clawhdf5-tools --test +edit_interop -- --ignored --nocapture measure_append_waste`, one editor for +the whole workload; sizes are deterministic): + +| Workload | `FileEditor` | libhdf5 | `h5repack` of ours | +|---|---|---|---| +| 1000 appends of 100 `f8`, 1024-element chunks, unfiltered | 810 504 B | 810 504 B | 810 360 B | +| same, gzip | 306 780 B | 306 058 B | 306 104 B | +| 2000 appends of 10 values, 4096-element gzip chunks | 79 829 B | 50 292 B | 49 930 B | + +In the last case the chunk being appended to is followed by new index +blocks and moves each time it grows, and the space it leaves is too small +for its next, larger version. **No journal.** A crash while an edit patches existing structures can leave the file inconsistent; see the `FileEditor` documentation. @@ -288,711 +134,94 @@ before anything is written. On top of them: ## Selection reads that decode more than the selection -**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so -the Python `ds[...]`) materialises only the selection's bounding box when -that box covers at most half the dataset (`partial_read`). It decodes the -whole dataset and extracts the selection instead when: +**Status:** open (documented 2026-09-26; re-checked 2026-09-28 in +`Dataset::read_selection`, `crates/clawhdf5/src/reader.rs`). A selection +read (and so the Python `ds[...]`) materialises only the selection's +bounding box when that box covers at most half the dataset +(`partial_read`). It decodes the whole dataset and extracts the selection +instead when: - the bounding box covers more than half the dataset — which a strided selection across a chunked dataset (`ds[::100]`) always does, although it may touch few chunks; - the dataset is compact or virtual, or has no storage; - it is chunked with a non-default fill value (the box path does not fill unallocated chunks, so the fill-aware full read is used). + Values are correct in every case; this is cost only. Selections other than -`Selection::All` also bypass the file's chunk cache. The bounding-box -heuristic's other cost is measured under "Concurrent and contiguous read -performance" below. +`Selection::All` also bypass the file's chunk cache. -## Concurrent and contiguous read performance (measured 2026-09-26) +## HDF5 features still unsupported -**Status:** fixed (2026-09-26). Re-measured on tank at `c5334b1`: full -reads of chunked deflate data at 16 threads run at 4944 MB/s against 3135 -for 16 h5py processes (1.58x; 0.69x-0.76x before), and contiguous reads -are 1.29x h5py on one thread (`BENCHMARKS.md`, "Results after in-place -chunk decoding"). The history below is kept for reference. Measured on -tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, -"Concurrent reads"): -- **Partly fixed 2026-09-26.** Full reads of chunked datasets from several threads - through one `File` stop scaling at about 4 threads (880 MB/s on deflate - data vs 4424 MB/s for 16 h5py processes). Hyperslab reads, which skip the - chunk cache, scale to 1244 MB/s, so the `File`'s shared chunk cache is the - suspect. *Cause:* not the cache. Those numbers were taken with - `--decode-threads 1`, a one-thread rayon pool, and every full read handed - its chunks to that pool, so all reader threads queued behind its single - worker (per-thread CPU time: one thread did all the decoding, the 16 - readers almost none). Hyperslab reads touch one chunk each and never used - the pool. Reads now decode on the calling thread when the pool has one - thread (`tests/single_thread_decode_pool.rs`); re-measured at - `408f69e`, 8 threads went from 887 to 2943 MB/s (h5py processes: 3042). - **Still open:** at 16 threads full chunked reads reach 2142-2341 MB/s, - 0.69x-0.76x 16 h5py processes (3083 MB/s in the same run), with the - default pool as with a one-thread one; with a small pool (2-4 threads) - readers outside it still wait on its workers. Datasets larger than the - cache's budget were already read without inserting into it, and skipping - its lookups entirely gained only a few percent at 16 threads. Remaining - per-read overhead: each full `read_f32` of a chunked dataset faults in - about three times its size in fresh pages (the output, the `f32` copy of - it, and a new buffer per decoded chunk). - **Fixed 2026-09-26** (both causes; see `CHANGELOG.md`, "Chunked full - reads"): chunks are decoded into buffers each thread reuses and copied - straight into the output, and the typed readers decode into the `Vec` - they return, so a full read no longer faults in a buffer per chunk or a - second copy of its output; and the reading - thread decodes its own chunks with pool workers helping when free, so no - reader waits on a small or busy pool - (`crates/clawhdf5/tests/busy_decode_pool.rs`). The 16-thread comparison - with h5py processes has not been re-measured yet (tank was busy with - other work); this item stays open until it is. -- Contiguous datasets read 4x slower than h5py on one thread (2.5 vs - 9.8 GB/s full, 0.12x for 256 x 256 hyperslabs). - **Fixed 2026-09-26** (re-measured on tank at `408f69e`: 13665 MB/s - full and 31991 MB/s for 256 x 256 hyperslabs on one thread, 1.44x and - 6.3x h5py; `BENCHMARKS.md`): full - reads were dominated by 4 KiB page faults on the fresh output buffer, - which is now backed by transparent huge pages as numpy's is; hyperslab - reads copied the selection three times, element by element, and now copy - each contiguous run once, straight from the file into the output (see - `CHANGELOG.md`). The chunked-read scaling item above is still open. -Values are correct; this is speed only. +**Status:** open. What remains of the gaps the +[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit) +found (the rest are [fixed](#gaps-found-by-the-2026-09-25-hdf5-audit-fixed-parts)), +plus later residuals. Each is an error or a documented difference, never +wrong data. -## Scale-offset data read back wrong values - -**Status:** fixed 2026-09-26, after v2.7.0. **Every release that decoded -the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on ordinary -files h5py writes, with no error. - -Found by the review of the 2026-09-26 conformance work: h5py wrote 1480 -scale-offset datasets (every integer type `i1` .. `u8`, `f4` and `f8`, -little- and big-endian, with no fill value and with the type's minimum, -maximum or another value as fill, random, constant, all-fill and extreme -data, `scaleoffset` 0, 1, 3, full width less one and full width for -integers, decimal scale factors 0, 1, 2, 4 and 7 for floats). v2.7.0's -decoder read 332 of them differently from h5py: **151 returned wrong -values with no error**, 181 failed to read. - -| Case | Datasets | v2.7.0 | -|---|---|---| -| integer, `scaleoffset=0` (libhdf5 picks the bits), data spanning most of the type's range | 82 | wrong values | -| integer, `scaleoffset` = full width: `i4`/`u4` (10), `i8`/`u8` (41) | 51 | wrong values | -| integer, `scaleoffset` = full width, the other datasets (every `i1`..`u2` one, most `i4`/`u4`, some `i8`/`u8`) | 181 | "truncated minval" / "implausible minbits" | -| `f4` D-scale, factor 4 or 7, values up to about 10^6 | 18 | wrong values | - -In every case libhdf5 stored a chunk at full width (`minbits` equal to the -type's width): the chunk then holds the elements as they are, and they -were decoded as offsets from `minval`. Two more differences were found on -crafted files and fixed with them: the packed codes start at byte 21 -whatever size the chunk records for `minval` (`cve-2025-44905` -`/Scale_offset_short_data_be`), and a chunk with `minbits` 0 and a fill -value is all fill values (it read as `minval`). - -**Fix:** `clawhdf5_format::filters` decodes a scale-offset chunk as -`H5Z__filter_scaleoffset` does. **Test:** the whole matrix is -`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at test -time and compared dataset by dataset; on v2.7.0's decoder it reports the -332. **Existing data:** the files were always right; only reads were -wrong, so re-reading with a fixed build gives the correct values. - -## Silent wrong data found by the 2026-09-25 HDF5 audit - -**Status:** fixed after v2.7.0 (2026-09-25). **Every release up -to and including v2.7.0 is affected.** - -An audit on tank checked clawhdf5 against libhdf5 in three ways: -- a sweep of 686 public files: the libhdf5 test files, the HDF Group's - `cve_hdf5` reproducers, and the pyfive, netcdf-c, netcdf4-python, h5wasm, - h5py and xarray corpora; -- 567 read cases generated with h5py 3.16 / HDF5 2.0; -- 96 write cases checked with h5py builds linking HDF5 1.10, 1.12, 1.14 and - 2.0, plus h5dump 1.14.6. - -It found these cases where a value came back wrong **without an error**: - -| Area | What happened | Who is affected | -|---|---|---| -| Chunk index (read) | Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place | any file with a max shape larger than its shape and `libver='latest'` (h5py `maxshape=(10, None)`, `(20, 10)`) | -| Chunk index (write) | Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled | files we wrote with one unlimited dimension and > 244 chunks, or e.g. `maxshape=(20, None)` | -| 4-byte offsets | unfiltered chunked datasets read as zeros | files created with `sizeof_addr = 4` | -| Filter mask | any skipped filter skipped the whole pipeline | files with partially filtered chunks (optional filters, direct chunk writes) | -| Numeric reads | float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half | `read_i32`/`read_i64`/`read_u64` callers on float or wider data; HDF5 2.0 bf16 data | -| SZIP | garbage or zeros | every libhdf5-written SZIP dataset | -| Scale-offset | float values 1 ULP off | libhdf5 D-scale float data | -| Shared fill value | read as zero fill | fill values stored as shared messages | -| VL sequences | `read_vl_bytes` truncated non-byte base types | VL int/float sequences | -| Chunk cache | two threads reading two chunked datasets through one `File` could get each other's chunks | multi-threaded readers, including Python with the GIL released | - -The audit also found files we wrote that libhdf5 **refuses**, now fixed: -- Fixed Array datasets with more than 1 024 chunks. -- Header messages over 64 KiB (large attributes). -- Reference, Opaque, BitField and Time datatypes. -- Files written with `with_page_size`. -- Several unlimited dimensions. -- A finite max shape larger than the shape. -- An empty-string attribute, which broke every attribute on its object. -- `FillTime` codes, which were rotated. - -Our LZ4 and Zstd output could not be read by libhdf5's registered plugins, and -our pcodec filter used Granular BitRound's ID. The details are in -`CHANGELOG.md` under Correctness and Interop. - -Before the fix, 419 of the 686 files read correctly and 43 differed from h5py. -After it, 448 read correctly and 23 differ. Of those 23: -- 17 are N-Bit float files. The probe compares raw file-type bytes; the typed - reader returns libhdf5's values (`nbit_custom_float_decodes_like_libhdf5`). -- 2 are an h5py bug: VL data with a big-endian base type comes back - byte-swapped in h5py, and h5dump agrees with us. -- The rest are object or attribute listing differences. - -**Update 2026-09-25:** the sweep is now in the repo (`conformance/run.sh`, -corpora pinned by commit) and its current numbers are in `CONFORMANCE.md`, -regenerated nightly by `.gitea/workflows/conformance.yml`. The probe now -compares N-Bit floats as the values libhdf5 converts them to, so the N-Bit -files above count as identical. Its file list is defined by -`conformance/list_files.py` (697 files: netCDF classic files are left out, -and 11 HDF5 files the ad-hoc sweep missed are in). On 42b81d9: 467 identical, -123 our-error, 15 mismatch (2 are the h5py bug above), 92 that libhdf5 cannot -read, and no panics, hangs or crashes. - -There were no panics, hangs or crashes before or after, including on all 147 -CVE and fuzzer files. On some of those files, h5dump 1.14.6 and h5py/HDF5 2.0 -segfault or abort. - -## Gaps found by the 2026-09-25 HDF5 audit (open) - -**Status:** open. These fail with an error; none returns wrong data (the VDS -fill-value item that did is fixed). - -- **Layout message versions 1 and 2** (HDF5 1.6-era files): 84 of the 686 - sweep files, `InvalidLayoutVersion`. This is the largest single gap. - **Fixed 2026-09-25:** versions 1 and 2 are parsed (compact, contiguous, - chunked via the v1 B-tree). -- **Compound datatype version 1 array members** (found with the layout - fix; pre-1.4 files such as `tarrold.h5`): **wrong data** — the legacy - per-member dimensions were skipped, so an array member read as one scalar. - **Fixed 2026-09-25.** -- **Virtual datasets:** - - ~~**Wrong data:** unmapped regions read as 0 instead of the fill value.~~ - Fixed 2026-09-25: unmapped elements and missing sources read as the - virtual dataset's fill value. - - ~~`%b` printf-style source names are not expanded.~~ Fixed 2026-09-25: - printf-style and unlimited mappings are read, and the extent is - recomputed from the sources as libhdf5 does. Still open: the - "first missing" view and a printf gap other than 0 (libhdf5 access - properties we always read at their defaults), source-to-virtual type - conversion other than a byte swap, nested virtual sources, and source - files outside the virtual file's directory (refused with an error). - - ~~Hyperslab selection versions 1 and 2 are refused.~~ Fixed 2026-09-25: - versions 1-3 and irregular hyperslabs are decoded. - - ~~The version-1 mapping list written with a 2.0 low bound (flags byte, - shared names) was misparsed.~~ Found and fixed 2026-09-25. -- **Files with a user block:** the base address is not applied. - **Fixed 2026-09-25:** every reader views the file from the superblock on - (`twithub.h5`, `twithub513.h5`, `h5clear_fsm_persist_user_*.h5`; the - `twithub` files still stop at the user-defined link type below). -- **Old-style shared messages (version 1)** read the wrong address. - **Fixed 2026-09-25:** the address follows the length-sized link-name - offset of the embedded symbol table entry (`tcompound.h5`, `tcompound2.h5`). -- **Groups and links:** - - Groups with a user-defined link type (e.g. 187) cannot be listed. - **Fixed 2026-09-25:** user-defined links are skipped; the rest of the - group lists. - - Dense groups with more than about 22 000 links cannot be listed. - **Fixed 2026-09-25:** two bugs — fractal-heap child indirect blocks had - the wrong row count, and v2 B-tree internal nodes at depth 3+ were read - with the wrong pointer widths. - - Soft links are left out of `datasets()`. **Fixed 2026-09-25:** soft - links are listed as their targets; dangling ones are left out. - - **Wrong data (found while fixing user blocks):** an old-style group whose - local-heap free list points outside the heap listed garbage names where - libhdf5 refuses the heap. **Fixed 2026-09-25** - (`InvalidLocalHeapFreeList`, checked when a name is first read, as - libhdf5 does). -- **Dense attributes:** a large attribute stored as a fractal-heap "huge" - object makes every attribute on the object fail. This affects real NetCDF - files (`issue671.nc`). **Fixed 2026-09-25:** huge and tiny heap objects, - and filtered heaps, are read; and an attribute that still cannot be read - is left out of `attrs()` (reported by `attrs_with_errors()`) instead of - failing the others. -- **Other readers:** - - VL-string datasets are not readable through `File`. **Fixed - 2026-09-26:** `read_string` reads them (also `read_string_bytes`, - `read_string_selection`, and on `MmapFile`/`LazyFile`), with h5py's - values: strings end at a NUL, null elements (heap address 0) are `""`, - and an element at the undefined heap address is an error as in libhdf5 - (it read as `""` until 2026-09-26); VL sequences of - numbers read with `read_vlen::()`, and VL values inside compounds or - `AttrValue::Raw` attributes decode with `File::decode_strings` / - `File::decode_vlen` (`crates/clawhdf5/tests/vl_data_interop.rs`). - - Variable-length values inside a compound (and VL-string attributes) in - a file with 4-byte offsets (`sizeof_addr = 4`) fail with - `GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume - the 16-byte element of an 8-byte-offset file. The datatype itself reads - (it was refused as "member overlaps with previous member" until - 2026-09-26). **Fixed 2026-09-26:** a VL type's element size is the one - its datatype message stores (12 with 4-byte offsets), and the global - heap is read with libhdf5's header padding - (`crates/clawhdf5/tests/vl_offset4_interop.rs`). - - Metadata cache images are not supported. **Fixed 2026-09-26:** the - image is applied at open, as libhdf5 loads it over the file's metadata - (`clawhdf5_format::superblock_ext`); `h5clear_mdc_image.h5` reads - (`crates/clawhdf5/tests/metadata_cache_image.rs`), without copying the - file (a private copy-on-write mapping takes the image's entries; - `tests/cache_image_memory.rs`). A file whose image libhdf5 cannot load - (`cve-2025-6269-*`, `cve-2025-6516`) opens, as in libhdf5, and every - object lookup fails with the image's error. Differences from libhdf5 - that remain: libhdf5 fails only the first metadata read and then reads - the file's own (possibly stale) metadata, where we keep failing; an - image entry that runs past the end of file is refused (libhdf5 checks - only its start); a flush-dependency parent flag is checked against the - child count as libhdf5's debug build checks it (HDF5 2.0 release - builds refuse every entry that has children, even in images they - wrote); the superblock extension's driver-info and shared-message table - messages are not decoded at open. - - x87 long double and binary128 are refused. - - N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail. - **Not our bug (checked 2026-09-26):** both are corrupt files HDF5 2.0 - reads only by reading past a buffer. `cve-2025-2308`'s - `/Scale_offset_long_long_data_le` has scale-offset codes that run past - the end of the chunk (libhdf5's develop branch refuses it, "Buffer too - short"); `bad_nbit_parms_walk.h5` has an N-Bit parameter list one value - short (libhdf5's own `test_filter_bad_params` in `test/dsets.c` now - requires that read to fail). We refuse both; `CONFORMANCE.md` lists them - under *Known not-our-bug* (since 2026-09-27: the *ref-bug* class, with - the evidence re-checked every run — see *Conformance: the last non-ok - files* above). Scale-offset did decode three cases - differently from libhdf5 (codes after a `minval` of recorded size other - than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed - 2026-09-26, and the full-width case was silent wrong data on ordinary - h5py files (see *Scale-offset data read back wrong values* above). -- **Filters:** blosc, blosc2, bitshuffle, bzip2, LZF and zfp are not - implemented. **Fixed 2026-09-26** for LZF (default-on `lzf` feature), - bitshuffle, bzip2 and Blosc 1 (`bitshuffle`, `bzip2`, `blosc`, or - `plugin-filters` for all), read and write, pure Rust; h5ex_d_lzf, - h5ex_d_bshuf, h5ex_d_bzip2 and h5ex_d_blosc now read (conformance 573 of - 697 ok). **Fixed 2026-09-26** for Blosc2 (32026, `blosc2` feature, also - in `plugin-filters`), read only: hdf5plugin's frames and B2ND arrays, every - codec and filter it offers; h5ex_d_blosc2 now reads (conformance 576 of - 697 ok, tank, `conformance/run.sh --no-fetch`). Blosc2 frames using - dictionaries, lazy chunks, variable-length blocks, user-defined codecs or - registered filters (e.g. bytedelta) are refused with an error. - **Fixed 2026-09-26** for ZFP (32013, `zfp` feature, also in - `plugin-filters`), read only: every H5Z-ZFP mode and type, bit-exact - against h5py + hdf5plugin 7.1 (`crates/clawhdf5/tests/zfp_interop.rs`); - h5ex_d_zfp now reads (conformance 600 of 697 ok, tank, - `conformance/run.sh --no-fetch`). **Still open:** clawhdf5 cannot write - Blosc2 or ZFP. Other filters (32023, Granular BitRound, too, since - 2026-09-26 even with the `pcodec` feature) can be plugged in with +- **Virtual datasets:** the "first missing" view and a printf gap other + than 0 (libhdf5 access properties we always read at their defaults), + source-to-virtual type conversion other than a byte swap, nested virtual + sources, and source files outside the virtual file's directory (refused + with an error). +- **Datatypes:** x87 long double and binary128 are refused. Revised + (version-4) region and attribute references are recognised but not + decoded, and external references are an error (object references + decode). A multi-dimensional numeric attribute is returned as a flat + array (its shape is not reported; `AttrValue::Raw` carries the shape). +- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in + that: libhdf5 fails only the first metadata read of an image it cannot + load and then reads the file's own (possibly stale) metadata, where we + keep failing; an image entry that runs past the end of file is refused + (libhdf5 checks only its start); a flush-dependency parent flag is + checked against the child count as libhdf5's debug build checks it (HDF5 + 2.0 release builds refuse every entry that has children, even in images + they wrote); the superblock extension's driver-info and shared-message + table messages are not decoded at open. +- **Filters:** Blosc2 and ZFP are read-only (clawhdf5 cannot write them). + Blosc2 frames using dictionaries, lazy chunks, variable-length blocks, + user-defined codecs or registered filters (e.g. bytedelta) are refused. + Other filters (32023 Granular BitRound among them, even with the + `pcodec` feature) can be plugged in with `filter_registry::register_filter`. -- **Wrong data: a chunk whose filters decode to fewer bytes than the chunk - read with zeros for the missing bytes** (any filter; found reviewing the plugin - filters). **Fixed 2026-09-26:** it is an error naming the chunk. A corrupt - chunk must never read as zeros. Unfiltered chunks are read at their stored - size and are not checked this way. **Fixed 2026-09-26** for unfiltered - chunks too: in a dataset without filters a chunk the index records at - other than the chunk's size is refused, as libhdf5's develop branch - refuses it (`cve-2025-44904`, where HDF5 2.0 fills the rest from its - buffer). -- **Crash:** a hostile Blosc chunk (frame size below its header) panicked in - builds with overflow checks. **Fixed 2026-09-26**; the new decoders are - fuzzed in the unit tests. -- ~~**Header checks:** on 12 CVE datasets libhdf5 rejects a corrupt header and - we read data anyway. We need stricter header checks.~~ **Fixed - 2026-09-26** (counted again: 18 objects on the CVE corpus that libhdf5 - refuses; some read as wrong data, e.g. a zero chunk dimension read as all - fill values): object headers, datatypes, chunk dimensions and chunk-index - offsets are checked as libhdf5 checks them, and truncated files are - refused. 17 of the 18 now - fail as in libhdf5 (conformance on tank, `conformance/run.sh --no-fetch`, - 2026-09-26: 571 of 697 ok). Still read where libhdf5 refuses: - - ~~`cve-2024-32624.h5` `/Dset_OBJREF`: a dataspace whose storage size - overflows 64 bits. `File::dataset` and `shape()` succeed (libhdf5 - refuses at open); reading the values fails.~~ **Fixed 2026-09-26:** - `File::dataset` (and `MmapFile`, `LazyFile`) refuse it at open - (`FormatError::InvalidDatasetStorage`), as they do contiguous storage - past the end of the file. - - ~~`cve-2020-10810.h5`, `cve-2020-10812.h5` (whole files libhdf5 cannot - open, not among the 18): libhdf5 decodes the superblock extension's File - Space Info and metadata-cache-image messages at open and refuses these - files; we do not decode those messages at open.~~ **Fixed 2026-09-26:** - the superblock extension is decoded at open with libhdf5's checks, and - both files are refused. - - Deliberately not refused, because clawhdf5 up to v2.7.0 wrote them: a - float sign bit position outside the type, and a size-0 string type. - - Not refused because current libhdf5 reads it though HDF5 2.0.0 - (h5py 3.16) refuses it: a v4 chunked layout whose dimensions are - encoded in more bytes than they need (HDFGroup/hdf5@e124c36, - 2026-06-05, relaxed that check; clawhdf5 wrote such layouts until - 2026-09-26). - - Not refused because HDF5 2.0 (h5py 3.16) reads them though newer - libhdf5 refuses them: bit-field offset/precision outside the type, an - unknown variable-length kind, an array type whose stored size is not - its element count times its base size. - - (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF` - attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.) - - `h5rs check` validates with the library's parsers, so it inherits what - they accept: of the 150 CVE and fuzzer files, `check --data` passes 15, - and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26; 28 and 21 - before these checks, 16 and 9 before a VL type's stored element size - was checked, which flags `cve-2024-32608`). -- **Writer:** - - ~~Nested groups beyond one level: path-like names are now refused, not - created.~~ **Fixed 2026-09-26:** groups nest to any depth (path names - create intermediate groups, as h5py does), with soft, extra hard and - external links at any depth and optional creation-order tracking; - h5py, h5dump and `h5rs check --data` read them - (`crates/clawhdf5/tests/writer_groups_interop.rs`, - `crates/clawhdf5-tools/tests/h5rs_interop.rs`). ~~Still missing: - attribute creation order is not tracked.~~ **Fixed 2026-09-26:** - `track_order` (file default, `GroupBuilder`, and the new - `DatasetBuilder::track_order`) tracks and indexes attribute creation - order as h5py's `track_order=True` does; h5py lists the attributes in - the order they were set, inline and dense (20 000 on one dataset), and - keeps numbering them in "r+" mode (tank, - `cargo test -p clawhdf5 --test writer_groups_interop - track_order_lists_attributes_in_creation_order`). libhdf5 numbers at - most 65 535 attributes on such an object, so more is an error. - - ~~A group with more than 65 535 links, or an object with more than - 65 535 dense attributes, is an error (the index is one B-tree leaf).~~ - **Fixed 2026-09-26:** the dense indexes are v2 B-trees of any depth - (libhdf5's 512-byte nodes once the records outgrow the one-leaf layout, - which smaller indexes keep byte for byte). Tested on tank with 100 000 - links in one group (short names, and 111-byte names with creation order - tracked) and 70 000 attributes on one object: h5py lists them in order - and reads the values, h5dump reads spot checks, `h5rs check` reads every - record, and h5py in "r+" mode adds and deletes thousands of links and - attributes in those trees (`cargo test -p clawhdf5 --test - deep_btree_interop`; `cargo test -p clawhdf5-tools --test h5rs_interop - check_files_with_deep_btrees`). The name indexes are ordered by hash and - then name, as libhdf5 needs for names whose hashes collide. - - ~~Dense link or attribute storage past 512 KiB of messages was written - unreadable (child indirect blocks of the fractal heap written as direct - blocks).~~ **Fixed 2026-09-26** (it affected 2.7.0 too): tested with - 20 000 and 65 535 links and with 8 MB of dense attributes, read by - h5py, h5dump, `h5rs check` and clawhdf5, and h5py can add links to - such groups. - - ~~libhdf5 could not add a link to a group we wrote (no Group Info - message).~~ **Fixed 2026-09-26.** - - Huge fractal heap objects: in dense storage (more than 8 attributes on - an object, or more than 8 links in a group) one attribute or link - message over 65 515 bytes is an error. - - Output that HDF5 1.8 can read. - - ~~A B-tree v2 chunk index larger than one leaf, so datasets with - several unlimited dimensions are limited to 65 535 chunks.~~ **Fixed - 2026-09-26:** more chunks get libhdf5's 2048-byte nodes with internal - nodes above the leaves. Tested on tank with 200 000 chunks (and 80 000 - deflated): h5py and clawhdf5 read every value, and h5py resizes the - dataset and writes 4 510 new chunks into the tree (same commands as - above). +- **Writer:** in dense storage (more than 8 attributes on an object, or + more than 8 links in a group) one attribute or link message over 65 515 + bytes is an error (no huge fractal-heap objects). The writer does not + produce output that HDF5 1.8 can read. +- **Checks we deliberately do not make:** + - a float sign bit position outside the type, and a size-0 string type: + clawhdf5 up to v2.7.0 wrote them; + - a v4 chunked layout whose dimensions are encoded in more bytes than + they need: current libhdf5 reads it (HDFGroup/hdf5@e124c36, + 2026-06-05) though HDF5 2.0.0 (h5py 3.16) refuses it, and clawhdf5 + wrote such layouts until 2026-09-26; + - bit-field offset/precision outside the type, an unknown + variable-length kind, an array type whose stored size is not its + element count times its base size: HDF5 2.0 reads them though newer + libhdf5 refuses them. ---- - -## Compound datatype message version 5 is not parsed (HDF5 2.0) - -**Status:** fixed on `main` in `a13ff51` (2026-06-03); **not in the v2.1.0 -tag**, which was cut five commits earlier. Ships in the next release. - -**Reported by:** M. Scot Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0. - -**Summary:** `clawhdf5-format` v2.1.0 rejects any dataset with a compound -(struct) datatype written by an HDF5 2.0 library in `libver='latest'` mode: -`InvalidDatatypeVersion { class: 6, version: 5 }`. - -**Reproduction** (h5py 3.16.0 / HDF5 2.0.0): - -```python -import h5py, numpy as np -dt = np.dtype([('x', 'f8'), ('y', 'f8'), ('id', 'i4')]) -data = np.array([(1.0, 2.0, 10), (3.0, 4.0, 20)], dtype=dt) -f = h5py.File('compound.h5', 'w', libver='latest') -f.create_dataset('particles', data=data) -f.close() -``` - -Committed as `crates/clawhdf5-format/tests/writer_h5py_tests.rs::read_h5py_generated_compound` -(`#[ignore]`d; needs `python3` with h5py on `PATH`). Run with -`cargo test -p clawhdf5-format --test writer_h5py_tests -- --include-ignored`: -v2.1.0 gives 25 passed / 1 failed; `main` passes everything. - -**Root cause:** the compound (class 6) branch of `Datatype::parse` -(`crates/clawhdf5-format/src/datatype.rs`) accepted only versions 1–4. Datatype -message versions 4 and 5 changed only the Reference and Complex classes, so a -v5-tagged compound uses the unchanged v3 member-list layout. - -**Fix:** versions 3–5 are accepted for compound (class 6) and array (class 10) -datatypes, and data layout message version 5 is accepted too (needed for every -chunked dataset written by HDF5 2.0). Byte-level regression tests: -`test_compound_v5_from_hdf5_2_0`, `test_array_v5_from_hdf5_2_0`. - -## Native complex datatype (class 11) is mis-parsed (HDF5 2.0) - -**Status:** fixed 2026-09-18. Found while validating the report above. - -**Summary:** HDF5 2.0 native complex types (`H5T_COMPLEX_IEEE_F64LE` etc.) -were parsed as if they carried a compound-style member list. The properties are -actually a single base floating-point datatype, so the parser produced a garbage -datatype, or `UnexpectedEof` when the complex type was a compound member. h5py's -default numpy-complex mapping is unaffected (it writes a `{r, i}` compound); -only files using the native type through the C API / h5py low-level API hit this. - -**Fix:** class 11 parses its base type and is surfaced as the equivalent -`{r, i}` compound. Tests: `test_complex_v5_from_hdf5_2_0`, -`test_compound_with_complex_member_from_hdf5_2_0`, -`writer_h5py_tests.rs::read_h5py_generated_native_complex`. - -## Revised reference datatype (class 7, version 4) is not parsed - -**Status:** fixed 2026-09-19 for object references; region and attribute -references are recognised but not decoded. - -**Summary:** HDF5 1.12+ `H5T_STD_REF` references use datatype message version 4 -with reference types 2-4 (object / region / attribute), which `Datatype::parse` -rejected with `InvalidReferenceType`. h5py still writes the legacy references, -so no file had been available to test against. - -**Fix:** a real file was produced by driving the libhdf5 bundled in the h5py -wheel through ctypes (`tests/fixtures/gen_std_ref.py` -> -`std_ref_hdf5_2_0.h5`). The three new types parse as -`ReferenceType::{Object2, DatasetRegion2, Attribute}`, and -`read_object_references` decodes `Object2` elements (type, flags, token size, -token = target object header address). External references (flag bit 0) and -the region/attribute payloads are errors rather than misreads. - -## `clawhdf5-gpu` `gpu_tests` can hang under the default parallel test runner - -**Status:** fixed 2026-09-19. - -**Summary:** during `cargo test --workspace` the `gpu_tests` binary sat idle for -25+ minutes. Every test created its own `wgpu::Instance` + device (requesting -adapter-maximum limits) concurrently, and readback used an unbounded -`device.poll(Wait)`. - -**Fix:** tests hold a process-wide lock while they own a device, and -`GpuAccelerator` readback waits time out after 30 s with `GpuError::BufferMap`. - -## Compound datatype versions 1 and 2 are mis-parsed (default libver files) - -**Status:** fixed 2026-09-19. Found by adding a default-libver axis to the h5py -interop tests. - -**Summary:** any compound dataset written with default libver bounds (plain -`h5py.File(path, 'w')`, datatype message version 1) failed to read, typically -with `Overflow("compound member 'x': byte_offset(0) + field_size(4136977) ...")`. -Only `libver='latest'` files (version 3+) and files written by clawhdf5 itself -worked, which is why the existing tests never caught it. - -**Root cause:** `Datatype::parse` skipped 24 bytes of legacy per-member array -fields for v1 where the format has 28 (dimensionality 1 + reserved 3 + -permutation 4 + reserved 4 + 4 dimension sizes 16), and treated v2 like v1 minus -name padding, whereas v2 keeps the 8-byte name padding and has no array fields. - -## Attributes with unsupported datatypes are silently dropped - -**Status:** fixed 2026-09-19. - -**Summary:** `Dataset::attrs()` / `Group::attrs()` returned only attributes -convertible to `AttrValue` and omitted the rest without any indication — every -Python `bool` (an HDF5 enum), complex, compound and reference attributes. Unsigned -64-bit arrays were also cast to `I64Array`, turning values above `i64::MAX` -negative. - -**Fix:** booleans decode as 0/1 integers, `AttrValue::U64Array` keeps unsigned -arrays unsigned, and `AttrValue::Raw { datatype, shape, data }` carries any other -attribute verbatim. Both new variants are writable. Still lossy: a -multi-dimensional numeric attribute is returned as a flat array (its shape is -not reported). - -## B-tree v2 chunk index (layout v4, index type 5) is not supported - -**Status:** fixed 2026-09-19. - -**Summary:** a chunked dataset with **two or more unlimited dimensions** written -with `libver='latest'` indexes its chunks with a version-2 B-tree, and reading it -failed with `unsupported chunked layout version=4, index_type=Some(5)`. - -**Fix:** record types 10 (unfiltered) and 11 (filtered) are decoded — address, -stored size, filter mask, scaled offsets — through the shared chunk-listing -function, so full reads, cached reads, partial reads and fill-value handling -all work. Covered by an h5py interop test (plain, gzip+shuffle, a 2500-chunk -tree with internal nodes, a sparse dataset with a fill value, a hyperslab). + `h5rs check` validates with the library's parsers, so it inherits what + they accept: of the 150 CVE and fuzzer files, `check --data` passes 15, + and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26). +- **VL sequences:** a file may point many elements at one large global + heap object, and a VL-sequence read then returns that object once per + element, as h5py would (memory is bounded otherwise; see + [Crafted global heaps](#crafted-global-heaps-exhaust-the-variable-length-readers-memory)). ## External links and external raw data are not followed -**Status:** open (by design for now); both are explicit errors. - -**Summary:** a path through an external link returns -`FormatError::ExternalLinkUnsupported { filename, object_path }`, and a dataset -created with `external=[...]` storage returns -`FormatError::ExternalDataFilesUnsupported`. Neither is resolved. If support is -added, file names must be confined to the opened file's directory, as the -virtual-dataset resolver now does. - ---- - -## Python interop suites skip silently when no interpreter has h5py - -**Status:** fixed on `main` in `a29c1b2` (2026-09-19). - -On a system where `python3` is a PEP 668 "externally managed" interpreter, -h5py cannot be installed into it at all, and every interop suite — the h5py -writer round-trips, the facade suite, netCDF4, and the reference files — -returned `false` from its availability probe and skipped without failing. CI -reported `SKIP` and a green run. This is the same class of gap that let the -compound-datatype v5 bug above reach a release. - -The probes now read `CLAWHDF5_PYTHON`, and `scripts/ci-test.sh` picks up -`.venv/bin/python` automatically. To restore the coverage on a fresh checkout: - -```bash -python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 -``` - -Set `CLAWHDF5_REQUIRE_INTEROP=1` in any automated runner so a missing -interpreter is a failure rather than a skip. - - ---- - -## Crafted B-tree v2 structures crash or exhaust the reader - -**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to -and including v2.6.0 is affected.** - -B-tree v2 traversal (`clawhdf5-format`, `btree_v2::collect_btree_v2_records`) -recursed one frame per level with the depth taken from the file, and followed -child addresses without checking whether they were shared. Two consequences -for anyone reading untrusted files: - -- A node that is its own child, under a header claiming 65 535 levels, overflows - the stack and aborts the process. The file is under 100 bytes. -- Levels whose children all point at one node below make the traversal visit it - fan-out^depth times: ~30 million records from ~5 KB, and memory exhaustion one - level deeper. - -B-tree v2 backs dense attribute storage, v2 groups, shared object header -messages and chunk indexes, so opening an object that uses any of them is -enough. Both are now errors: depth is capped at 64, and traversal stops once it -has produced more records than the file could physically hold. - ---- - -## Crafted global heaps exhaust the variable-length reader's memory - -**Status:** fixed on `feat/p2-vl-strings` (2026-09-26). Not a regression of -that branch: every earlier release is affected through `read_vl_strings`. - -Reading variable-length values kept an owned copy of every object of every -global heap collection visited, for the whole read. A file whose collections -nest inside one another's object data (32 bytes apart, each element pointing -at a different one) made retained memory O(elements × file size): a 744 KB -file reached 1.58 GB. Letting every collection's object chain jump to one -shared run of tiny objects made the parse time O(elements × objects) too. -libhdf5 refuses such files. - -Now `VlResolver` caches where each object lies instead of a copy, drops its -cache past a 32 MiB budget, and refuses a collection that overlaps one it -has already read (libhdf5 gives each collection its own block, so only a -crafted file has them). `GlobalHeapCollection::parse` (and the new -`parse_index`) also refuse a collection that runs past the end of the file, -or an object that runs past the end of its collection. Guarded by -`crates/clawhdf5-format/tests/vl_heap_bounds.rs`, which measures peak heap -use with a counting allocator. Still open: a file may point many elements -at one large heap object, and a VL-*sequence* read then returns that -object once per element, as h5py would. - ---- - -## Extensible Array chunk indexes read back wrong data past the inline elements - -**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to -and including v2.6.0 is affected.** - -A dataset created with exactly one unlimited dimension (`maxshape=(None, ...)`, -the usual append-only/resizable case) is indexed by an Extensible Array. Its -index block holds the first `idx_blk_elmts` chunk entries inline — 4 by -default — and everything after that lives in data blocks and super blocks whose -layout `clawhdf5-format` computed incorrectly. - -Consequences, by dataset size (1 chunk per element): - -| chunks | result before the fix | -|---|---| -| <= 36 | correct (inline, plus two data blocks that happened to line up) | -| 37 | 1 element wrong | -| 400 | 364 elements wrong | -| >= ~1000 | `invalid Extensible Array data block signature` | - -The dangerous case is the middle one: values were returned from the wrong -chunks rather than an error being raised. Any reader that accepted the data at -face value saw plausible but incorrect numbers. - -The root causes were the super block sizing formulas (`ndblks` and -`dblk_nelmts` each double every *other* level, a half-step apart), a missing -block-offset field in the super block, and a page-init bitmap read from the -wrong structure. All four are fixed and covered by interop tests against -HDF5 2.0 at sizes that cross each boundary, including paged data blocks. - -Files written by this crate were not affected by *this* read bug, but the -writer had its own: it indexed only the first 244 chunks, so later chunks -read back as 0 in libhdf5 and in clawhdf5. See "Silent wrong data found by -the 2026-09-25 HDF5 audit" below. - -## Every `f32` dataset we wrote was unreadable by h5py / libhdf5 - -**Status:** fixed 2026-09-23, after v2.7.0. **Every -release up to and including v2.7.0 is affected** — the encoder was already -wrong in v2.1.0. - -The floating-point datatype message carries the position of the sign bit -(bits 8–15 of its class bit field). `clawhdf5-format` wrote 63 for every -float, which is correct only for `f64`. libhdf5 validates the field, so opening -any `f32` dataset written by this crate failed: - -``` -KeyError: 'Unable to synchronously open object (sign bit position out of bounds)' -``` - -That covers every agent store (`/memory/embeddings`, `norms` and -`activation_weights` are `f32`). `clawhdf5` itself ignores the field on read, -and the interop suites only ever wrote `f64` from our side, so nothing here -noticed. - -**Fix:** the sign position is computed from the type (`bit_offset + -bit_precision - 1`: 15, 31, 63 for half, single, double). Regression tests: -`float_sign_location_is_the_top_bit_of_the_value` (byte level), -`clawhdf5_writes_f32_h5py_reads` and the agent's -`h5py_reads_every_dataset_of_an_agent_store`. - -**Existing files:** an agent store is rewritten in full at every checkpoint, so -it becomes readable by h5py at its next checkpoint with a fixed build. Other -files with `f32` datasets need to be rewritten. - -## Empty datasets we wrote were unreadable by h5py / libhdf5 - -**Status:** fixed 2026-09-23, after v2.7.0. Every -release up to and including v2.7.0 is affected. - -A dataset with no elements was written with a real file address and a storage -size of 0. libhdf5 guards contiguous storage with an overflow check -(`addr + size <= addr`) that is always true when the size is 0, so it rejected -the dataset: - -``` -KeyError: 'Unable to synchronously open object (invalid dataset size, likely file corruption)' -``` - -In practice: every agent store without sessions or a knowledge graph — the -`/sessions` and `/knowledge_graph` datasets are empty until something is added -— could not be read by h5py even once the `f32` bug above was fixed. Found by -the same agent-store interop test. - -**Fix:** an empty contiguous dataset gets the undefined address (all `0xff`), -which is what libhdf5 itself writes. +**Status:** open (by design for now; re-checked 2026-09-28); both are +explicit errors. A path through an external link returns +`FormatError::ExternalLinkUnsupported { filename, object_path }`, and a +dataset created with `external=[...]` storage returns +`FormatError::ExternalDataFilesUnsupported`. If support is added, file +names must be confined to the opened file's directory, as the +virtual-dataset resolver does. ## Range reads (`File::open_storage`) limits -**Status:** open (added 2026-09-26, milestone M2 of -`docs/design/range-reads.md`; remote backends added by M3). `File::open_storage` +**Status:** open (added 2026-09-26 with milestone M2 of +`docs/design/range-reads.md`; updated for M3-M5). `File::open_storage` reads any `clawhdf5_format::storage::Storage` through the whole read API, -every format-crate read path works through `Storage::read_at`/`read_ranges`, and `clawhdf5-remote` serves HTTP(S) and object-store files through a block cache, but: @@ -1008,7 +237,8 @@ cache, but: range; coalescing is the backend's (or the cache's) job. - A group lookup by name in a version-1 (symbol-table) group lists the whole group (dense groups use their name index). Over a range backend that is - one read per symbol-table node and name, per lookup. + one read per symbol-table node and name, per lookup. (The wasm lazy + reader walks the group's B-tree instead; see its entry.) - External virtual-dataset source files are loaded whole through the resolver (`File::set_vds_resolver`), as bytes; they are not read through a `Storage`. @@ -1016,41 +246,29 @@ cache, but: `read_*_zerocopy`) need the file in memory and answer `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` panics for such a file (`File::contiguous_bytes()` is the fallible form). - `LazyFile`, `MmapFile` and the wasm bindings still read a whole file - (`h5rs` and the Python bindings read through `File::storage`, and take - URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`). - - `LazyFile`, `MmapFile` and the Python bindings still read a whole file - (`h5rs` reads through `File::storage`, and takes URLs with its `remote` - feature; the wasm reader's `openUrl` reads by range requests since - 2026-09-27, its `open(bytes)` takes a whole file). -- The file's length is read once, at open: a growing file (SWMR) is not - followed (milestone M5). A remote file is pinned at open, so one that - grows is `RemoteError::FileChanged`. -- Not new, but visible through the equivalence tests: a full read through - the file's chunk cache (`read_raw_data_cached`, `read_raw_data_indexed`, - and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map - order, so which failing chunk it reports can differ from one `File` to - the next (`cve-2025-2310.h5`); the values of a dataset that reads are - not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in - the order the index lists them, as the uncached readers do - (`several_damaged_chunks_report_the_same_chunk_every_time`). +- `LazyFile` and `MmapFile` still read a whole local file. `h5rs` and the + Python bindings read through `File::storage` and take URLs (`h5rs` with + its `remote` feature, Python with `clawhdf5.File(url)`); the wasm + reader's `openUrl` reads by range requests, its `open(bytes)` takes a + whole file. +- A growing local file is followed only when opened with + `File::open_swmr` (milestone M5, 2026-09-27; see `docs/design/swmr.md`): + `File::open` and `open_storage` read the length once, at open. SWMR + reading is not available for remote files (a remote file is pinned at + open, so one that grows is `RemoteError::FileChanged`), `MmapFile` or + `LazyFile`, and clawhdf5 has no SWMR writer. ## Remote files (`clawhdf5-remote`) limits -**Status:** open (added 2026-09-26, milestone M3 of -`docs/design/range-reads.md`). +**Status:** open (added 2026-09-26 with milestone M3 of +`docs/design/range-reads.md`). Python (`clawhdf5.File(url)`) and the +browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27. -- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is - milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but - the default wheel reads plain `http://` only: `https://` needs a wheel - built with `--features https` (rustls with ring, which compiles C), and - `s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs). - The Python tests run against an in-process `http.server` only. - -- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through - `File::as_bytes`, which a remote file does not have. (The browser can - since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.) +- **Default builds read plain `http://` only.** `https://` needs the + `https` feature (rustls with ring, which compiles C) and `s3://`, + `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs); the + default Python wheel has none of them. The Python tests run against an + in-process `http.server` only. - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). The design's policy of using a paged file's page size as the block size is not implemented, and only the first block is read ahead. @@ -1089,7 +307,8 @@ cache, but: ## `clawhdf5-wasm` (browser) limits -**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27). +**Status:** open (by design for now; added 2026-09-26, `openUrl` +2026-09-27). - `open()` holds the whole file in memory (it takes its bytes), so a multi-GB local file does not fit a browser tab. A file on a web server @@ -1098,33 +317,27 @@ cache, but: with these limits: - **Round trips:** a call runs as passes over the blocks fetched so far and is re-run after each wave of misses, so a call costs one round - trip per wave, not one for everything: a chunk index is walked a - level (or a node) per round trip, while the chunks of a read are - fetched together. Listing a group asks for every child's object - header, and every node of a level of the group's index, in one pass - (since 2026-09-27; it was one round trip per header block): 3000 - datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a - `libver="latest"` file (dense links). **Since 2026-09-27 (later):** - 4 and 5 passes (5 and 6 at 64 KiB, from 8 and 11): the index walks go - on past a missing node, and parsers hint what they read next - (`Storage::hint`: node bodies, the heap's blocks, each child's - header), which the lazy reader fetches with a pass's misses. That is - the depth of the chain (index levels, then symbol table nodes or - heap objects, then headers) plus the pass that finishes; it cannot - go lower without reading structures before their addresses are - known. Opening one dataset of a v1 group looks its name up down the - group's B-tree (it read every entry: 74 requests, 193 MB at 1 MiB - blocks for one 64 KiB dataset of the 3000; now 5 requests, 5 MB). - Each pass re-parses what the call reads (CPU, not network). With - headers spread through the file (h5py writes each next to its data) - a listing still fetches most of the file at 1 MiB blocks (192 of - 198 MB; 35 MB in 530 requests at 64 KiB); a smaller `blockSize` - fetches less. Merging nearby requests does not help such a file: the - blocks a listing needs are five or six apart at 64 KiB, so fewer requests would - mean fetching most of the file. Listing it a second time is free at - 64 KiB blocks, but at 1 MiB its metadata blocks (192 MB) exceed the - 64 MiB `cacheSize`, so they are fetched again (the earliest file: 4 - passes, 50 requests); a larger `cacheSize` keeps them. A file's paged + trip per wave: a chunk index is walked a level (or a node) per round + trip, while the chunks of a read are fetched together. Since + 2026-09-27 listing 3000 datasets takes 4 passes for an h5py file and + 5 for a `libver="latest"` file at 1 MiB blocks (5 and 6 at 64 KiB): + index walks go on past a missing node, and parsers hint what they read + next (`Storage::hint`), which the lazy reader fetches with a pass's + misses. That is the depth of the chain (index levels, then symbol + table nodes or heap objects, then headers) plus the pass that + finishes; it cannot go lower without reading structures before their + addresses are known. Opening one dataset of a v1 group looks its name + up down the group's B-tree (5 requests, 5 MB at 1 MiB blocks for one + 64 KiB dataset of the 3000). Each pass re-parses what the call reads + (CPU, not network). + - **Scattered metadata:** with headers spread through the file (h5py + writes each next to its data) a listing still fetches most of the file + at 1 MiB blocks (192 of 198 MB; 35 MB in 530 requests at 64 KiB); a + smaller `blockSize` fetches less. Merging nearby requests does not + help such a file (the blocks a listing needs are five or six apart at + 64 KiB). Listing it a second time is free at 64 KiB blocks, but at + 1 MiB its metadata blocks (192 MB) exceed the 64 MiB `cacheSize`, so + they are fetched again; a larger `cacheSize` keeps them. A file's paged aggregation (metadata in pages) is not used to fetch its metadata in one request. - **Memory:** a call keeps every block it reads until it finishes (the @@ -1139,11 +352,9 @@ cache, but: whole module (every open file on the page), which these limits keep from happening; before 2026-09-27 both did abort it. - **File size:** at most 4 GiB - 1 bytes; a larger file is refused at - open. The format code turns file offsets into `usize` to use them - (with a clean error past it), which is 32 bits on wasm32, so nothing - at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are - tested (with a mock server); files above 200 MB have not been served - for real. + open (file offsets become `usize`, 32 bits on wasm32). Offsets between + 2 and 4 GiB are tested with a mock server; files above 200 MB have not + been served for real. - **Cross-origin servers** must allow CORS for the page's origin and either expose `Content-Range` (`Access-Control-Expose-Headers`) or answer `HEAD` with `Content-Length`. The file is pinned at open by its @@ -1151,48 +362,33 @@ cache, but: either header, only a change of length is detected. - **A server without range support** (it answers `200`) costs a whole download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error - with `fallback: "error"`. Every body, this one and each `206`, is read - as it arrives and cut off past its limit (the range asked for, or - `maxDownload`): a server cannot make the page buffer more. + with `fallback: "error"`. Every body is read as it arrives and cut off + past its limit: a server cannot make the page buffer more. - Fixed block size (`blockSize`, 1 MiB by default); a paged file's page size is not used. No retries: a failed request fails the call (calling again retries it; what was fetched stays cached). - Tested under Node 22 and headless Chromium (Playwright's build) against - a local server, cross-origin included (a page on 127.0.0.1 reading a - file from localhost, with and without exposed headers); not in Firefox - or Safari. - - The native corpus comparison (`tests/lazy.rs` with - `CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file, - `cve-2025-2310.h5`: two of its datasets have more than one bad chunk, - and which chunk's error is reported depends on the iteration order of - the chunk index (a `HashMap`, seeded per process), so the lazy and - the range-storage reads can name different errors. Both are errors; - not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the - chunk cache keeps the index's chunk order, so every read path names the - same (first) damaged chunk. + a local server, cross-origin included; not in Firefox or Safari. - Compound, reference, opaque, bitfield, time and VL-sequence datasets are refused with an error naming the type; attributes of those types come back - as `value: null` with their `dtype`. + as `value: null` with their `dtype`. (VL strings read, with h5py's + values, through the same `VlResolver` as `File` and `h5rs`.) - No Zstd or SZIP (both link C): such datasets fail with `unsupported filter: 32015` / `: 4`. pcodec is not enabled either. - External links and virtual-dataset sources in other files cannot be followed (no file system). -- Variable-length string datasets are read by decoding `read_selection`'s - bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still - cannot (see the audit gaps above). (`File` can since 2026-09-26. Since - 2026-09-26 the wasm crate resolves them with the same `VlResolver` as - `File` and `h5rs`, so all three return h5py's values.) ## The Node.js package (`packages/clawhdf5-node`) does not work -**Status:** open (found 2026-09-25). Unpublished; not built or tested in CI. +**Status:** open (found 2026-09-25; re-checked 2026-09-28, unchanged). +Unpublished; not built or tested in CI. The TypeScript wrapper over `crates/clawhdf5-napi` has never run successfully: - napi-rs converts `#[napi(object)]` fields to camelCase, but the wrapper reads snake_case (`r.line_range`, `s.total_records`, `s.working_count`, …), so every stats and consolidation field comes back `undefined` - (`src/index.ts:76-120`). + (`src/index.ts`). - It loads `../clawhdf5.node`, but `napi build --platform` produces `clawhdf5..node`; `main` points at `index.js` while `tsc` writes to `dist/`; `napi prepublish` expects per-platform packages that are not @@ -1205,3 +401,387 @@ The TypeScript wrapper over `crates/clawhdf5-napi` has never run successfully: It was written for an OpenClaw integration that is not being pursued (see `docs/openclaw.md`). Fix and add CI, or remove it, before anyone depends on it. + +--- + +# Fixed (history) + +Newest first. "Before any release" means no tagged release (v2.7.0 and +earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date +given. + +## `ObjectHeader::parse` 4% slower after range-read M2/M3 + +**Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release. +Speed only; values were always correct. + +Measured on an idle tank (`BENCHMARKS.md`, "Local metadata and data reads +after range-read M2/M3"): `object_header_parse_x401` 23.86 → 24.86 µs +median (+4.2%) from `8f59b2e` to `7a8fae0`, after the parser began reading +continuation chunks from a bounded queue (`a69c5be`). The remaining cost was +the call to the version-1 message loop, kept out of line by +`#[inline(never)]`; inlined, it is 1.0% and 2.6% below `8f59b2e` in two idle +A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's speed"). +A reported 5.6% slowdown of contiguous hyperslab reads, measured under load, +was noise and was withdrawn. + +## Conformance: the last non-ok files + +**Status:** classified 2026-09-27 (PR #21) — none is a clawhdf5 bug. +Result (`conformance/run.sh --no-fetch`, tank): 602 of 697 ok (was 600), +0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read. + +- **2 mismatches were h5py's big-endian VL bug** + (`NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64`, + `hdf5/tools/test/testfiles/tcomplex_be.h5` + `/VariableLengthDatasetFloatComplex`): h5py returns the file's + big-endian bytes under a little-endian dtype. `conformance/ref.py` now + detects the bug in the installed h5py and relabels such elements, so the + values are compared; both files are *ok*. +- **3 our-errors are objects HDF5 2.0 reads only by over-reading memory** + (`cve-2025-2308.h5` `/Scale_offset_long_long_data_le`, + `cve-2025-44904.h5` `/Scale_offset_float_data_le`, + `bad_nbit_parms_walk.h5` `/Nbit_int_data_le`). libhdf5's output for them + is not determined by the file (it changes from run to run and with + `MALLOC_PERTURB_`), and libhdf5's develop branch refuses all three. We + keep refusing them; `conformance/ref_bugs.py` re-checks every run and + classifies such a file *ref-bug* only while its values keep changing. + +## Damaged chunked datasets reported a different failing chunk per open + +**Status:** fixed 2026-09-27 (`4ad8073`, PR #19). Errors only; the values of +readable datasets were never affected. + +A full read through the file's chunk cache listed a damaged dataset's +chunks in hash-map order, seeded per `File`, so two opens of +`cve-2025-2310.h5` could name different failing chunks, and the storage and +wasm corpus comparisons failed now and then. The cache now keeps the chunk +index's order, as the uncached readers do +(`several_damaged_chunks_report_the_same_chunk_every_time`). + +## Files a SWMR writer had open could not be read past a stale end of file + +**Status:** fixed 2026-09-27 (PR #19), before any release: reads have been +bounded by the recorded end of file only since `7d7a7e7` (2026-09-26). + +A libhdf5 SWMR writer does not keep the superblock's end-of-file address up +to date (a mid-write copy records 715 in a 6 030-byte file), so every reader +bounded such a file there: it listed, but chunked reads failed and `h5rs +check` reported chunk indexes past the end. Never wrong data. For a v3 +superblock with the SWMR-write flag the data now ends at the end of the +file, as libhdf5's SWMR reader reads it. Test: +`crates/clawhdf5/tests/swmr_interop.rs` (fixture +`tests/fixtures/swmr_mid_write.h5`). A file still being written is read +with `File::open_swmr` (`docs/design/swmr.md`). + +## Shrinking a chunked dataset with no recorded maximum scrambled it + +**Status:** fixed 2026-09-27 (PR #19), before any release +(`FileEditor::resize` shipped on main in PR #18). + +clawhdf5's writer stored no maximum dimensions for a chunked dataset +created without a `maxshape`; a Fixed Array index then places chunks by the +current dimensions, and `FileEditor::resize` changed only those, so a +shrink moved every chunk (wrong values in every reader) and the dataset +could not grow back. The editor now records the maximum libhdf5 would have +written before the first resize, and the writer records it for every +chunked dataset. Test: `crates/clawhdf5/tests/edit_resize_interop.rs`. + +**What users must do:** datasets written before the fix (v2.7.0 and +earlier, and `main` before 2026-09-27) have no recorded maximum. The fixed +editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`) +does not** — it scrambles them the same way and lets them grow past their +Fixed Array. Resize them once with a fixed `FileEditor` (a resize to the +same shape changes nothing; shrink and grow back) before letting libhdf5 +resize them. A dataset already shrunk by the unfixed editor holds misplaced +chunks; rewrite it from a good copy. + +## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 + +**Status:** fixed 2026-09-26 (PR #18), after v2.7.0. **Every release +(v2.1.0 to v2.7.0) is affected**, in both directions. + +Our Fletcher-32 reduced its sums with `% 65535`; libhdf5 folds them with +`(s & 0xffff) + (s >> 16)`, which differs where a sum is a non-zero +multiple of 65535. Chunks we wrote with such a sum are refused by h5py and +libhdf5 ("filter returned failure during read"); chunks libhdf5 wrote with +one were refused here with `Fletcher32Mismatch` (the data itself was never +wrong). The fix, `clawhdf5_format::checksum::fletcher32`, ports +`H5_checksum_fletcher32` and also accepts the byte-swapped (libhdf5 ≤ 1.6.2) +and the old clawhdf5 forms. Test: `crates/clawhdf5/tests/fletcher32_interop.rs`. + +**What users must do:** a Fletcher-32 dataset written by v2.7.0 or earlier +may hold chunks libhdf5 cannot read; a fixed build reads them. Rewrite such +datasets with a fixed build before handing the file to libhdf5 or h5py. + +## LZF/Blosc chunks written with a stale filter mask + +**Status:** fixed 2026-09-26 (PR #17), before any release (v2.7.0 and +earlier write neither filter). + +`FileBuilder` (and `FileEditor`) stored every LZF or Blosc chunk with +filter mask 0, where libhdf5 stores a chunk whose output is no smaller than +the chunk raw with the filter's mask bit set; after libhdf5 rewrote such a +chunk h5py could no longer read the dataset. Both now use +`clawhdf5_format::filters::compress_chunk_masked`. **What users must do:** +files written before the fix read correctly; rewrite them before letting +libhdf5 modify them. + +## Concurrent and contiguous read performance + +**Status:** fixed 2026-09-26 (PRs #15 and #16). Speed only. + +Measured on tank against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, "First run, +before the read fixes"): full reads of chunked data from 16 threads through +one `File` stopped scaling at about 4 threads (880 MB/s vs 4424 MB/s for 16 +h5py processes), and contiguous datasets read 4x slower than h5py on one +thread. Causes: every full read queued on a one-thread decode pool, fresh +buffers per chunk and a second copy of the output, 4 KiB page faults on the +output buffer, and element-by-element hyperslab copies. Re-measured at +`c5334b1` (`BENCHMARKS.md`, "Results after in-place chunk decoding"): +16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py +processes (1.58x), contiguous reads 1.29x h5py on one thread. Tests: +`single_thread_decode_pool.rs`, `busy_decode_pool.rs`. + +## Scale-offset data read back wrong values + +**Status:** fixed 2026-09-26 (PR #16), after v2.7.0. **Every release that +decoded the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on +ordinary files h5py writes, with no error. + +Of 1480 scale-offset datasets h5py wrote (every integer type, `f4`/`f8`, +both byte orders, many fill values and scale factors), v2.7.0 read 332 +differently from h5py: **151 returned wrong values with no error**, 181 +failed. In every case libhdf5 had stored a chunk at full width (`minbits` +equal to the type's width), which was decoded as offsets from `minval`; +two smaller differences (codes start at byte 21 whatever `minval`'s +recorded size; `minbits` 0 with a fill value is all fill) were fixed with +it. Test: `crates/clawhdf5/tests/scaleoffset_interop.rs`. + +**What users must do:** nothing to the files — they were always right; re-read +them with a fixed build. + +## Crafted global heaps exhaust the variable-length reader's memory + +**Status:** fixed 2026-09-26 (PR #15). Every earlier release is affected +through `read_vl_strings`. + +Global heap collections nested inside one another's object data made +retained memory O(elements × file size) (a 744 KB file reached 1.58 GB) and +parse time O(elements × objects). `VlResolver` now caches object locations +within a 32 MiB budget and refuses overlapping collections, and collections +or objects running past their bounds are refused. Test: +`crates/clawhdf5-format/tests/vl_heap_bounds.rs`. The remaining +one-object-many-elements case is listed under +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +## Gaps found by the 2026-09-25 HDF5 audit (fixed parts) + +**Status:** fixed 2026-09-25 and 2026-09-26 (PRs #11 to #17). Each was an +error unless marked **wrong data**; every release up to v2.7.0 has them. +What remains open is under +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +- **Layout message versions 1 and 2** (HDF5 1.6-era files; 84 of the 686 + sweep files) and **compound datatype version 1 array members** (**wrong + data**: an array member read as one scalar) — fixed 2026-09-25. +- **Virtual datasets:** unmapped regions read as 0 instead of the fill + value (**wrong data**); printf-style and unlimited mappings; hyperslab + selection versions 1-3; the version-1 mapping list with a 2.0 low bound — + fixed 2026-09-25. +- **User blocks**, **old-style shared messages**, **user-defined link + types**, **dense groups over about 22 000 links**, **soft links in + `datasets()`**, a local-heap free list outside the heap (**wrong data**: + garbage names) — fixed 2026-09-25. +- **Dense attributes stored as huge/tiny/filtered fractal-heap objects** + (real NetCDF files, `issue671.nc`); an unreadable attribute no longer + hides the others (`attrs_with_errors()`) — fixed 2026-09-25. +- **VL strings through `File`** (`read_string`, `read_vlen::()`, + `File::decode_strings`/`decode_vlen`), **VL data with 4-byte offsets**, + **metadata cache images** — fixed 2026-09-26. +- **Filters:** LZF, bitshuffle, bzip2 and Blosc 1 (read and write), + Blosc2 and ZFP (read) — fixed 2026-09-26, pure Rust. A chunk whose + filters decode to fewer bytes than the chunk read with zeros for the rest + (**wrong data**); now an error, and unfiltered chunks of the wrong stored + size are refused. A hostile Blosc chunk panicked with overflow checks. +- **Header checks:** 18 CVE objects libhdf5 refuses were read (some as + wrong data); object headers, datatypes, chunk dimensions and chunk-index + offsets are now checked as libhdf5 checks them, a dataspace whose storage + size overflows is refused at open, and the superblock extension is + decoded at open with libhdf5's checks — fixed 2026-09-26. +- **N-Bit/scale-offset on corrupt files** (`cve-2025-2308`, + `bad_nbit_parms_walk.h5`): not our bug — see + [the last non-ok files](#conformance-the-last-non-ok-files). +- **Writer:** groups nest to any depth with soft, hard and external links; + attribute creation order (`track_order`); dense link/attribute indexes of + any size (v2 B-trees of any depth); dense storage past 512 KiB was + written unreadable (it affected v2.7.0); libhdf5 could not add a link to + a group we wrote (no Group Info message); B-tree v2 chunk indexes larger + than one leaf — fixed 2026-09-26. Tests: `writer_groups_interop.rs`, + `deep_btree_interop.rs`. + +## Silent wrong data found by the 2026-09-25 HDF5 audit + +**Status:** fixed 2026-09-25 (PR #11), after v2.7.0. **Every release up to +and including v2.7.0 is affected.** + +The audit (686 public files, 567 read cases and 96 write cases against h5py +3.16 / HDF5 1.10-2.0 and h5dump 1.14.6) found these values returned wrong +**without an error**: + +| Area | What happened | Who is affected | +|---|---|---| +| Chunk index (read) | Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place | any file with a max shape larger than its shape and `libver='latest'` (h5py `maxshape=(10, None)`, `(20, 10)`) | +| Chunk index (write) | Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled | files we wrote with one unlimited dimension and > 244 chunks, or e.g. `maxshape=(20, None)` | +| 4-byte offsets | unfiltered chunked datasets read as zeros | files created with `sizeof_addr = 4` | +| Filter mask | any skipped filter skipped the whole pipeline | files with partially filtered chunks (optional filters, direct chunk writes) | +| Numeric reads | float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half | `read_i32`/`read_i64`/`read_u64` callers on float or wider data; HDF5 2.0 bf16 data | +| SZIP | garbage or zeros | every libhdf5-written SZIP dataset | +| Scale-offset | float values 1 ULP off | libhdf5 D-scale float data | +| Shared fill value | read as zero fill | fill values stored as shared messages | +| VL sequences | `read_vl_bytes` truncated non-byte base types | VL int/float sequences | +| Chunk cache | two threads reading two chunked datasets through one `File` could get each other's chunks | multi-threaded readers, including Python with the GIL released | + +It also found files we wrote that libhdf5 **refuses**, fixed with it: Fixed +Array datasets with more than 1 024 chunks, header messages over 64 KiB, +Reference/Opaque/BitField/Time datatypes, files written with +`with_page_size`, several unlimited dimensions, a finite max shape larger +than the shape, an empty-string attribute (which broke every attribute on +its object), and rotated `FillTime` codes. Our LZ4 and Zstd output could not +be read by libhdf5's registered plugins, and our pcodec filter used Granular +BitRound's ID. Details: `CHANGELOG.md`, Correctness and Interop. + +**What users must do:** re-read affected files with a fixed build; rewrite +files clawhdf5 wrote in the affected cases (one unlimited dimension with more +than 244 chunks, an unlimited dimension not first, the refused cases) before +handing them to libhdf5. + +The sweep became `conformance/run.sh` (corpora pinned by commit; nightly by +`.gitea/workflows/conformance.yml`); current numbers are in +`CONFORMANCE.md`. No sweep run found a panic, hang or crash, including on +the 147 CVE and fuzzer files on some of which h5dump 1.14.6 and h5py/HDF5 +2.0 segfault or abort. + +## Every `f32` dataset we wrote was unreadable by h5py / libhdf5 + +**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. **Every release up to +and including v2.7.0 is affected** (the encoder was already wrong in +v2.1.0). + +The floating-point datatype message's sign-bit position was written as 63 +for every float; libhdf5 validates it, so every `f32` dataset — including +every agent store's `/memory/embeddings`, `norms` and `activation_weights` +— failed to open ("sign bit position out of bounds"). It is now +`bit_offset + bit_precision - 1`. Tests: +`float_sign_location_is_the_top_bit_of_the_value`, +`clawhdf5_writes_f32_h5py_reads`, the agent's +`h5py_reads_every_dataset_of_an_agent_store`. + +**What users must do:** an agent store is rewritten at every checkpoint, so +it becomes readable by h5py at its next checkpoint with a fixed build. Other +files with `f32` datasets need to be rewritten. + +## Empty datasets we wrote were unreadable by h5py / libhdf5 + +**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. Every release up to and +including v2.7.0 is affected. + +An empty dataset was written with a real address and size 0, which +libhdf5's overflow check rejects ("invalid dataset size, likely file +corruption") — every agent store without sessions or a knowledge graph. An +empty contiguous dataset now gets the undefined address, as libhdf5 writes. +**What users must do:** as for `f32` above (agent stores heal at their next +checkpoint; rewrite other files). + +## Extensible Array chunk indexes read back wrong data past the inline elements + +**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including +v2.6.0 is affected.** + +A dataset with exactly one unlimited dimension is indexed by an Extensible +Array, whose data and super block layout was computed wrongly: with more +than 36 chunks values came back from the wrong chunks **with no error** (37 +chunks: 1 element wrong; 400: 364 wrong), and from about 1000 chunks the +read failed. Fixed with interop tests against HDF5 2.0 across every +boundary. **What users must do:** re-read with v2.7.0 or later. (The +writer's own 244-chunk bug is under the +[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit).) + +## Crafted B-tree v2 structures crash or exhaust the reader + +**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including +v2.6.0 is affected** (for untrusted files). + +A self-referencing node under a header claiming 65 535 levels overflowed the +stack and aborted the process (under 100 bytes), and shared children made +traversal visit a node fan-out^depth times (memory exhaustion from ~5 KB). +Depth is now capped at 64 and traversal stops past the records the file +could hold. + +## Python interop suites skip silently when no interpreter has h5py + +**Status:** fixed 2026-09-19 (`a29c1b2`), in v2.6.0. + +On a PEP 668 system every h5py/netCDF4 interop suite skipped and CI stayed +green. The probes now read `CLAWHDF5_PYTHON`, `scripts/ci-test.sh` picks up +`.venv/bin/python`, and `CLAWHDF5_REQUIRE_INTEROP=1` (set in CI) makes a +missing interpreter a failure. To restore coverage on a fresh checkout: + +```bash +python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 +``` + +## B-tree v2 chunk index (layout v4, index type 5) is not supported + +**Status:** fixed 2026-09-19, in v2.5.0. Datasets with two or more +unlimited dimensions written with `libver='latest'` failed to read in +earlier releases; record types 10 and 11 are now decoded on every read path. + +## Attributes with unsupported datatypes are silently dropped + +**Status:** fixed 2026-09-19. `attrs()` omitted booleans, complex, compound +and reference attributes without notice, and cast `u64` arrays to `i64`. +Booleans now decode as 0/1, `AttrValue::U64Array` keeps unsigned arrays, and +`AttrValue::Raw` carries anything else. (Multi-dimensional numeric +attributes are still flattened; see +[HDF5 features still unsupported](#hdf5-features-still-unsupported).) + +## Compound datatype versions 1 and 2 are mis-parsed (default libver files) + +**Status:** fixed 2026-09-19. Every compound dataset written with default +libver bounds (plain `h5py.File(path, 'w')`) failed to read: the v1 legacy +array fields are 28 bytes, not 24, and v2 keeps the name padding. + +## `clawhdf5-gpu` `gpu_tests` can hang under the default parallel test runner + +**Status:** fixed 2026-09-19 (`706189c`), in v2.3.0. Tests now hold a +process-wide lock while they own a device, and `GpuAccelerator` readback +times out after 30 s with `GpuError::BufferMap`. + +## Revised reference datatype (class 7, version 4) is not parsed + +**Status:** fixed 2026-09-19 for object references. `H5T_STD_REF` object +references (`ReferenceType::Object2`) decode, tested against a file made +with libhdf5 through ctypes (`tests/fixtures/gen_std_ref.py`). Region and +attribute references are recognised but not decoded — see +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +## Native complex datatype (class 11) is mis-parsed (HDF5 2.0) + +**Status:** fixed 2026-09-18 (`b55b7db`), in v2.2.0. HDF5 2.0 native complex +types were parsed as a compound member list (garbage or `UnexpectedEof`); +class 11 now parses its base type and surfaces as an `{r, i}` compound. +h5py's default complex mapping (a compound) was never affected. + +## Compound datatype message version 5 is not parsed (HDF5 2.0) + +**Status:** fixed on `main` in `a13ff51` (2026-06-03), in v2.2.0; **not in +v2.1.0**, which was cut five commits earlier. Reported by M. Scot +Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0. + +v2.1.0 rejects any compound dataset written by HDF5 2.0 with +`libver='latest'` (`InvalidDatatypeVersion { class: 6, version: 5 }`). +Versions 3-5 are now accepted for compound and array datatypes, and data +layout message version 5 too (every chunked dataset HDF5 2.0 writes). Tests: +`test_compound_v5_from_hdf5_2_0`, `test_array_v5_from_hdf5_2_0`, +`writer_h5py_tests.rs::read_h5py_generated_compound`.