docs: record the 2026-09-25 HDF5 audit fixes and open gaps
CHANGELOG: upgrade notes (changed read results for max-shape files, saturating conversions, new writer errors, format-crate API changes) and the reader/writer correctness fixes. known-issues: the silent-wrong-data table with before/after sweep numbers, the gaps still open, and a correction to the Extensible Array entry, which said files we wrote were unaffected. CLAUDE.md: clawhdf5-gpu is vector distance computation, not I/O, and clawhdf5-filters holds only deflate backends (no Blosc). Also a facade test that libhdf5's 20-bit N-Bit float test data reads as libhdf5's values. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -3,6 +3,38 @@
|
||||
## Unreleased
|
||||
|
||||
### Upgrade Notes
|
||||
- **HDF5 correctness audit (2026-09-25).** A sweep of 686 public files (the
|
||||
libhdf5 test files, the HDF Group's CVE reproducers, pyfive, netcdf-c,
|
||||
netcdf4-python, h5wasm, h5py and xarray corpora), a 567-case read matrix and
|
||||
a 96-case write matrix against HDF5 1.10–2.0 found bugs that returned wrong
|
||||
values with no error, and files we wrote that libhdf5 rejects. The fixes are
|
||||
listed under Correctness and Interop. What changes for callers:
|
||||
- **Chunked datasets whose max shape is larger than their current shape**,
|
||||
or whose unlimited dimension is not the first, were indexed by the current
|
||||
shape instead of the max shape, both when read and when written. Files from
|
||||
libhdf5 now read correctly. Files clawhdf5 wrote with such a max shape were
|
||||
laid out wrongly and now read the way libhdf5 always read them — rewrite
|
||||
them. Agent stores and ClawBrainHub files have no max shape and are
|
||||
unaffected.
|
||||
- Integer reads (`read_i32`/`read_i64`/`read_u64`/...) of float data now
|
||||
convert (truncate toward zero, saturate at the type's range, NaN reads as
|
||||
0) instead of returning the IEEE bit pattern, and out-of-range integers
|
||||
saturate instead of keeping the low bits.
|
||||
- `FileWriter::finish()` now returns an error instead of writing a corrupt
|
||||
file for: a header message over 64 KiB (e.g. an attribute larger than
|
||||
~64 KiB), a group/dataset/link name that is empty, `.` or contains `/`
|
||||
(nested paths were written as one literal link), a max shape smaller than
|
||||
the shape, a page size outside 512 B–1 GiB, and more than 65 535 chunks in
|
||||
a dataset with several unlimited dimensions.
|
||||
- **Breaking (format crate):** `ObjectHeaderWriter::serialize`,
|
||||
`BatchObjectHeaderWriter::compute_sizes`/`serialize_all` and
|
||||
`build_chunked_data_from_precompressed` return `Result`;
|
||||
`read_fixed_array_chunks`/`read_extensible_array_chunks` take `max_dims`;
|
||||
`build_fixed_array_at`/`ea_writer::build_extensible_array_at` take one
|
||||
`Option<WrittenChunk>` per index slot; `fill_value::dataset_fill_value`
|
||||
returns `UnresolvedSharedMessage` for a shared message it cannot resolve
|
||||
instead of `None`. `FillTime::default()` is `IfSet` (libhdf5's default;
|
||||
default files are byte-identical).
|
||||
- **ZeroClaw does not use clawhdf5.** The project described itself as
|
||||
ZeroClaw's memory backend ("imported as a `clawhdf5` Cargo feature"). Checked
|
||||
against ZeroClaw v0.8.5 (the latest release), the `osobh/zeroclaw` fork and
|
||||
@@ -241,6 +273,53 @@
|
||||
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
||||
|
||||
### Correctness
|
||||
- `clawhdf5-format` reader — **values returned wrong with no error:**
|
||||
- Fixed Array and Extensible Array chunk indexes were laid out by the
|
||||
dataset's current shape instead of its max shape (23 libhdf5 test files,
|
||||
and any h5py file with e.g. `maxshape=(10, None)` or `(20, 10)` under
|
||||
`libver='latest'`).
|
||||
- Files with 4-byte offsets: unfiltered chunked datasets read as zeros.
|
||||
Chunk B-tree keys store offsets in 8 bytes whatever the file's offset
|
||||
size.
|
||||
- A chunk's filter mask skipped the whole pipeline when any bit was set;
|
||||
only the flagged filters are skipped now.
|
||||
- Float data read as an integer returned the bit pattern; narrowing integer
|
||||
reads kept the low bits; bfloat16 was decoded as IEEE half. Floats are now
|
||||
decoded from their datatype fields (bf16, FP8 E4M3/E5M2, IEEE half, single
|
||||
and double).
|
||||
- `vl_data::read_vl_bytes` truncated sequences of non-byte base types.
|
||||
- A shared fill-value message read as zero fill; it is resolved now,
|
||||
including from the file's shared-message (SOHM) table, which could never
|
||||
resolve because its index version byte was skipped.
|
||||
- Two threads reading two chunked datasets through one `File` could get each
|
||||
other's chunks (the shared chunk cache was switched between datasets
|
||||
across separate lock acquisitions). The cache is now keyed by dataset.
|
||||
- `clawhdf5-format` reader — errors on valid files: enum and bool datasets
|
||||
through the numeric readers; the "don't filter partial edge chunks" layout
|
||||
flag; Fletcher32 ahead of deflate (NetCDF-4's order). Unknown-message flags
|
||||
follow libhdf5 (`tbogus.h5`): "fail if unknown" is refused, "fail if unknown
|
||||
and writing" is ignored by a reader.
|
||||
- `clawhdf5-format` writer — **files libhdf5 rejects or reads wrong:**
|
||||
- Extensible Array (one unlimited dimension): chunks from index 244 on were
|
||||
written but never indexed and read as 0, by libhdf5 and by us.
|
||||
- Fixed Array: more than 1 024 chunks gave checksum errors (data blocks
|
||||
were never paged).
|
||||
- A finite max shape larger than the shape gave libhdf5 "addr overflow"; an
|
||||
unlimited dimension that is not the first scrambled the data; several
|
||||
unlimited dimensions (`(None, None)`) broke the whole file. These now
|
||||
write the index libhdf5 writes (swizzled Extensible Array, or a B-tree v2
|
||||
index for several unlimited dimensions).
|
||||
- Header messages over 64 KiB (the size field is 16 bits) and compact
|
||||
datasets at 65 534–65 535 bytes produced corrupt files.
|
||||
- Reference, Opaque, BitField and Time datatypes were written as empty
|
||||
messages; they now encode as HDF5 2.0 does.
|
||||
- `with_page_size` wrote a nonexistent superblock version 4; it now writes
|
||||
the v3 superblock and File Space Info message libhdf5 writes.
|
||||
- `FillTime` values were rotated on disk (NEVER was written as ALLOC, and so
|
||||
on). New `DatasetBuilder::with_fill_value`.
|
||||
- An empty-string attribute got a zero-size datatype, which made every
|
||||
attribute on the object unreadable in libhdf5.
|
||||
- `maxshape` equal to the shape no longer forces chunked layout.
|
||||
- `clawhdf5-format`: **a truncated deflate chunk read back short, with no
|
||||
error.** The deflate filter used flate2's streaming reader, which returns the
|
||||
bytes it has when the input runs out before the end-of-stream marker. It now
|
||||
|
||||
Reference in New Issue
Block a user