Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
827 lines
46 KiB
Markdown
827 lines
46 KiB
Markdown
# Known Issues
|
||
|
||
Bugs and limits found during development or downstream use, tracked here
|
||
because this repository's issue tracker is disabled. One entry per bug; when
|
||
an entry is fixed, record the fix in `CHANGELOG.md` and move it to
|
||
[Fixed (history)](#fixed-history) with its date, the releases it affected and
|
||
what users must do, rather than deleting it.
|
||
|
||
Releases referred to below: v2.1.0 (2026-06-03) to v2.7.0 (2026-09-20).
|
||
Everything fixed after v2.7.0 is on `main` and unreleased.
|
||
|
||
## Open issues
|
||
|
||
Checked against `main` at `9b5803f` on 2026-09-28.
|
||
|
||
| Issue | Kind | Since |
|
||
|---|---|---|
|
||
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
|
||
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
|
||
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
|
||
| [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 |
|
||
| [HDF5 features still unsupported](#hdf5-features-still-unsupported) | errors, never wrong data | 2026-09-25 audit |
|
||
| [External links and external raw data are not followed](#external-links-and-external-raw-data-are-not-followed) | error, by design for now | 2026-09-19 |
|
||
| [Range reads (`File::open_storage`) limits](#range-reads-fileopen_storage-limits) | cost; zero-copy APIs need an in-memory file | 2026-09-26 |
|
||
| [Remote files (`clawhdf5-remote`) limits](#remote-files-clawhdf5-remote-limits) | untested backends, fixed block size, timeouts | 2026-09-26 |
|
||
| [`clawhdf5-wasm` (browser) limits](#clawhdf5-wasm-browser-limits) | memory, round trips, unsupported types | 2026-09-26 |
|
||
| [Conformance: `bad_nbit_parms_walk.h5` flips between ref-bug and our-error](#conformance-bad_nbit_parms_walkh5-flips-between-ref-bug-and-our-error) | report noise (ok count unaffected) | 2026-09-28 |
|
||
| [The Node.js package does not work](#the-nodejs-package-packagesclawhdf5-node-does-not-work) | broken, unpublished, not in CI | 2026-09-25 |
|
||
|
||
---
|
||
|
||
## In-place modification (`FileEditor`) limits
|
||
|
||
**Status:** open (documented 2026-09-26, updated when version-2 B-tree
|
||
chunk indexes, shrinking, dense attributes and space reuse were added).
|
||
`clawhdf5::FileEditor` refuses, with `Error::Unsupported` and without
|
||
writing anything:
|
||
- new chunks in an **implicit** index (it has all of its chunks from the
|
||
start; they are written in place, and allocated/filled on growth under
|
||
early allocation as libhdf5 does);
|
||
- variable-length and reference data;
|
||
- chunks through a filter this build cannot encode (scale-offset, N-Bit,
|
||
SZIP, or a plugin filter it lacks), even an optional one: libhdf5 skips
|
||
an optional filter only when its own build lacks it, which none does for
|
||
these;
|
||
- attributes in dense storage when the heap cannot take them the way
|
||
libhdf5 would: replacing the last attribute left in a heap block by one
|
||
of another size (libhdf5 frees the block), a heap with I/O filters or
|
||
child indirect blocks (more than about 512 KiB of attributes), free
|
||
space in child indirect blocks, directly addressed huge objects; and
|
||
shared attribute messages. Measured 2026-09-26 on tank with the review's
|
||
random-edit harness (120 runs of 150 random edits, `earliest`/`v110`/
|
||
`latest`, about 5600 `set_attr` calls of 8 bytes to 6 KiB): 2.2% of
|
||
`set_attr` calls are refused, every one the last-attribute-in-a-block
|
||
replacement (24% before heap blocks could be skipped);
|
||
- version-1 object headers asked for an attribute larger than a header
|
||
message (they have no dense storage);
|
||
- partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external
|
||
raw data files, virtual datasets;
|
||
- files with a metadata cache image, paged or persistent free-space
|
||
management, a driver info block, or version-3 consistency flags set.
|
||
|
||
**Space is reused only within one editor.** Space an edit frees (a filtered
|
||
chunk that moves, chunks a shrink removes, B-tree nodes merged away, a
|
||
heap's replaced blocks) is reused by later edits of the same `FileEditor`;
|
||
what is left when it is dropped is leaked, as libhdf5 leaks it without a
|
||
persistent free-space manager (`h5repack` reclaims it). A chunk that is the
|
||
last thing in the file grows in place, which covers the usual append.
|
||
Measured 2026-09-26 on tank (`cargo test -p clawhdf5-tools --test
|
||
edit_interop -- --ignored --nocapture measure_append_waste`, one editor for
|
||
the whole workload; sizes are deterministic):
|
||
|
||
| Workload | `FileEditor` | libhdf5 | `h5repack` of ours |
|
||
|---|---|---|---|
|
||
| 1000 appends of 100 `f8`, 1024-element chunks, unfiltered | 810 504 B | 810 504 B | 810 360 B |
|
||
| same, gzip | 306 780 B | 306 058 B | 306 104 B |
|
||
| 2000 appends of 10 values, 4096-element gzip chunks | 79 829 B | 50 292 B | 49 930 B |
|
||
|
||
In the last case the chunk being appended to is followed by new index
|
||
blocks and moves each time it grows, and the space it leaves is too small
|
||
for its next, larger version.
|
||
|
||
**No journal.** A crash while an edit patches existing structures can leave
|
||
the file inconsistent; see the `FileEditor` documentation.
|
||
|
||
**Renamed files outside Linux.** Edits always go to the file the editor
|
||
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
|
||
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
|
||
the path, and after the path was renamed or replaced it fails (on Unix;
|
||
Windows cannot tell and would read whatever the path names).
|
||
|
||
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||
|
||
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||
before anything is written. On top of them:
|
||
|
||
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||
`NotImplementedError` (the editor changes values, shapes and
|
||
attributes only). Mode `'a'` works on an existing file only.
|
||
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||
whole elements of the same structured dtype are), HDF5 array-type
|
||
elements, variable-length data, strings padded with spaces or
|
||
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||
ones, h5py's, are writable), compounds containing such strings, null
|
||
dataspaces, index-list writes of more than 2²² elements (write them
|
||
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
|
||
h5py supports both).
|
||
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||
variable-length strings: h5py reads them back as `bytes`
|
||
(`numpy.bytes_`), not `str`.
|
||
- **Numeric conversion follows libhdf5's native-order results, not its
|
||
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||
saturate, floats are truncated toward zero and clipped), checked value by
|
||
value against h5py 3.16 / HDF5 2.0 in
|
||
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||
differs from h5py on purpose:
|
||
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||
minimum or 2⁶³ depending on the type);
|
||
- when the dataset or the array is not in native byte order, libhdf5's
|
||
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||
wrap an unsigned value too large for the signed type of the same size
|
||
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||
native order;
|
||
- libhdf5's native casts that are undefined in C: half floats into
|
||
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||
integers (→ minimum), a float equal to the integer maximum rounded up
|
||
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||
reads from other threads wait while an edit is written.
|
||
|
||
## Selection reads that decode more than the selection
|
||
|
||
**Status:** open (documented 2026-09-26; re-checked 2026-09-28 in
|
||
`Dataset::read_selection`, `crates/clawhdf5/src/reader.rs`). A selection
|
||
read (and so the Python `ds[...]`) of contiguous data copies just the
|
||
selected runs; of chunked data it materialises only the selection's
|
||
bounding box when that box covers at most half the dataset
|
||
(`partial_read`). It decodes the whole dataset and extracts the selection
|
||
instead when:
|
||
- the bounding box covers more than half the dataset — which a strided
|
||
selection across a chunked dataset (`ds[::100]`) always does, although
|
||
it may touch few chunks;
|
||
- the dataset is compact or virtual, or has no storage;
|
||
- it is chunked with a non-default fill value (the box path does not fill
|
||
unallocated chunks, so the fill-aware full read is used).
|
||
|
||
Values are correct in every case; this is cost only. Selections other than
|
||
`Selection::All` also bypass the file's chunk cache.
|
||
|
||
## HDF5 features still unsupported
|
||
|
||
**Status:** open. What remains of the gaps the
|
||
[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit)
|
||
found (the rest are [fixed](#gaps-found-by-the-2026-09-25-hdf5-audit-fixed-parts)),
|
||
plus later residuals. Each is an error or a documented difference, never
|
||
wrong data.
|
||
|
||
- **Virtual datasets:** the "first missing" view and a printf gap other
|
||
than 0 (libhdf5 access properties we always read at their defaults),
|
||
source-to-virtual type conversion other than a byte swap, nested virtual
|
||
sources, and source files outside the virtual file's directory (refused
|
||
with an error).
|
||
- **Datatypes:** x87 long double and binary128 are refused. Revised
|
||
(version-4) region and attribute references are recognised but not
|
||
decoded, and external references are an error (object references
|
||
decode). A multi-dimensional numeric attribute is returned as a flat
|
||
array (its shape is not reported; `AttrValue::Raw` carries the shape).
|
||
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
|
||
that: libhdf5 fails only the first metadata read of an image it cannot
|
||
load and then reads the file's own (possibly stale) metadata, where we
|
||
keep failing; an image entry that runs past the end of file is refused
|
||
(libhdf5 checks only its start); a flush-dependency parent flag is
|
||
checked against the child count as libhdf5's debug build checks it (HDF5
|
||
2.0 release builds refuse every entry that has children, even in images
|
||
they wrote); the superblock extension's driver-info and shared-message
|
||
table messages are not decoded at open.
|
||
- **Filters:** Blosc2 and ZFP are read-only (clawhdf5 cannot write them).
|
||
Blosc2 frames using dictionaries, lazy chunks, variable-length blocks,
|
||
user-defined codecs or registered filters (e.g. bytedelta) are refused.
|
||
Other filters (32023 Granular BitRound among them, even with the
|
||
`pcodec` feature) can be plugged in with
|
||
`filter_registry::register_filter`.
|
||
- **Writer:** in dense storage (more than 8 attributes on an object, or
|
||
more than 8 links in a group) one attribute or link message over 65 515
|
||
bytes is an error (no huge fractal-heap objects). The writer does not
|
||
produce output that HDF5 1.8 can read.
|
||
- **Checks we deliberately do not make:**
|
||
- a float sign bit position outside the type, and a size-0 string type:
|
||
clawhdf5 up to v2.7.0 wrote them;
|
||
- a v4 chunked layout whose dimensions are encoded in more bytes than
|
||
they need: current libhdf5 reads it (HDFGroup/hdf5@e124c36,
|
||
2026-06-05) though HDF5 2.0.0 (h5py 3.16) refuses it, and clawhdf5
|
||
wrote such layouts until 2026-09-26;
|
||
- bit-field offset/precision outside the type, an unknown
|
||
variable-length kind, an array type whose stored size is not its
|
||
element count times its base size: HDF5 2.0 reads them though newer
|
||
libhdf5 refuses them.
|
||
|
||
`h5rs check` validates with the library's parsers, so it inherits what
|
||
they accept: of the 150 CVE and fuzzer files, `check --data` passes 15,
|
||
and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26).
|
||
- **VL sequences:** a file may point many elements at one large global
|
||
heap object, and a VL-sequence read then returns that object once per
|
||
element, as h5py would (memory is bounded otherwise; see
|
||
[Crafted global heaps](#crafted-global-heaps-exhaust-the-variable-length-readers-memory)).
|
||
|
||
## External links and external raw data are not followed
|
||
|
||
**Status:** open (by design for now; re-checked 2026-09-28); both are
|
||
explicit errors. A path through an external link returns
|
||
`FormatError::ExternalLinkUnsupported { filename, object_path }`, and a
|
||
dataset created with `external=[...]` storage returns
|
||
`FormatError::ExternalDataFilesUnsupported`. If support is added, file
|
||
names must be confined to the opened file's directory, as the
|
||
virtual-dataset resolver does.
|
||
|
||
## Range reads (`File::open_storage`) limits
|
||
|
||
**Status:** open (added 2026-09-26 with milestone M2 of
|
||
`docs/design/range-reads.md`; updated for M3-M5). `File::open_storage`
|
||
reads any `clawhdf5_format::storage::Storage` through the whole read API,
|
||
and `clawhdf5-remote` serves HTTP(S) and object-store files through a block
|
||
cache, but:
|
||
|
||
- **A `Storage` without a cache is asked for each structure as the parsers
|
||
need it**, several times over for some (an object header is re-read by
|
||
each lookup through it): one pass over the conformance corpus — open,
|
||
list, every attribute, every dataset once — is 176 092 `read_at` calls
|
||
for 621 files, 92 489 of them for the 35 001-group `h5stat_newgrat.h5`
|
||
(2026-09-26, tank, `crates/clawhdf5/tests/storage_equivalence.rs` with
|
||
`CLAWHDF5_STORAGE_CORPUS`). A backend of your own over a network needs a
|
||
cache in front of it: wrap it in `clawhdf5_remote::BlockCache`, as
|
||
`open_url` does. `Storage::read_ranges` defaults to one `read_at` per
|
||
range; coalescing is the backend's (or the cache's) job.
|
||
- A group lookup by name in a version-1 (symbol-table) group lists the whole
|
||
group (dense groups use their name index). Over a range backend that is
|
||
one read per symbol-table node and name, per lookup. (The wasm lazy
|
||
reader walks the group's B-tree instead; see its entry.)
|
||
- External virtual-dataset source files are loaded whole through the
|
||
resolver (`File::set_vds_resolver`), as bytes; they are not read through
|
||
a `Storage`.
|
||
- The zero-copy methods (`Dataset::read_raw_ref`, `read_as_slice`,
|
||
`read_*_zerocopy`) need the file in memory and answer
|
||
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
||
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
||
- `LazyFile` and `MmapFile` still read a whole local file. `h5rs` and the
|
||
Python bindings read through `File::storage` and take URLs (`h5rs` with
|
||
its `remote` feature, Python with `clawhdf5.File(url)`); the wasm
|
||
reader's `openUrl` reads by range requests, its `open(bytes)` takes a
|
||
whole file.
|
||
- A growing local file is followed only when opened with
|
||
`File::open_swmr` (milestone M5, 2026-09-27; see `docs/design/swmr.md`):
|
||
`File::open` and `open_storage` read the length once, at open. SWMR
|
||
reading is not available for remote files (a remote file is pinned at
|
||
open, so one that grows is `RemoteError::FileChanged`), `MmapFile` or
|
||
`LazyFile`, and clawhdf5 has no SWMR writer.
|
||
|
||
## Remote files (`clawhdf5-remote`) limits
|
||
|
||
**Status:** open (added 2026-09-26 with milestone M3 of
|
||
`docs/design/range-reads.md`). Python (`clawhdf5.File(url)`) and the
|
||
browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
|
||
|
||
- **Default builds read plain `http://` only.** `https://` needs the
|
||
`https` feature (rustls with ring, which compiles C) and `s3://`,
|
||
`gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs); the
|
||
default Python wheel has none of them. The Python tests run against an
|
||
in-process `http.server` only.
|
||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||
The design's policy of using a paged file's page size as the block size
|
||
is not implemented, and only the first block is read ahead.
|
||
- **Checked against local servers and one public one.** The tests use an
|
||
in-process HTTP/1.1 server and object_store's in-memory and local-file
|
||
stores. HTTPS was checked by hand against `raw.githubusercontent.com`
|
||
(2026-09-26, tank: `h5rs dump` of h5py's `vlen_string_dset.h5` by URL
|
||
equals the downloaded file's). The `s3`, `gcs` and `azure` backends are
|
||
built and their URL parsing tested, but they have not been run against a
|
||
real bucket.
|
||
- **A server with neither a strong ETag nor Last-Modified** can only be
|
||
checked by length, so a same-length replacement mid-read would go
|
||
unnoticed; `HttpOptions::require_validator` refuses such servers. A weak
|
||
ETag (`W/"…"`) cannot be sent as `If-Match`, so it counts as none.
|
||
- **Credentials:** HTTP takes extra headers (`HttpOptions::headers`, e.g.
|
||
`Authorization`); `h5rs` has no option for them. The cloud stores read
|
||
credentials from the environment only. Messages and `Debug` output show
|
||
URLs through `redact_url` (no userinfo, query values `REDACTED`);
|
||
`HttpStorage::url()` returns the URL as given and must not be logged.
|
||
- **Redirects:** at most `HttpOptions::max_redirects` (5) per request,
|
||
never from `https` to `http`, and the custom headers are not sent once a
|
||
redirect leaves the URL's origin. The redirect target is not remembered:
|
||
every request of a redirected file costs its hops again.
|
||
- **Timeouts:** ureq has no idle timeout, only a total one for the body,
|
||
so the body's budget is `timeout + size / min_speed` (30 s + 16 KiB/s
|
||
by default). A connection that stalls mid-body is detected only when
|
||
that budget runs out (94 s for a 1 MiB block, 9 min for an 8 MiB run).
|
||
- Each `ObjectStoreStorage` owns a tokio runtime with two worker threads.
|
||
A read called from inside another runtime blocks that runtime's thread
|
||
for its duration (it works, but `spawn_blocking` is the better place).
|
||
- `h5rs check` downloads a remote file whole (it validates every byte), up
|
||
to `--max-download` (1 GiB by default), and a URL cannot carry a
|
||
`FILE/OBJECT` suffix; `h5rs` uses the default cache settings.
|
||
- The zero-copy methods and `File::as_bytes` are unavailable on a remote
|
||
file (see the range-read limits above).
|
||
|
||
## `clawhdf5-wasm` (browser) limits
|
||
|
||
**Status:** open (by design for now; added 2026-09-26, `openUrl`
|
||
2026-09-27).
|
||
|
||
- `open()` holds the whole file in memory (it takes its bytes), so a
|
||
multi-GB local file does not fit a browser tab. A file on a web server
|
||
can be opened with `openUrl()` instead, which fetches only the byte
|
||
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
|
||
with these limits:
|
||
- **Round trips:** a call runs as passes over the blocks fetched so far
|
||
and is re-run after each wave of misses, so a call costs one round
|
||
trip per wave: a chunk index is walked a level (or a node) per round
|
||
trip, while the chunks of a read are fetched together. Since
|
||
2026-09-27 listing 3000 datasets takes 4 passes for an h5py file and
|
||
5 for a `libver="latest"` file at 1 MiB blocks (5 and 6 at 64 KiB):
|
||
index walks go on past a missing node, and parsers hint what they read
|
||
next (`Storage::hint`), which the lazy reader fetches with a pass's
|
||
misses. That is the depth of the chain (index levels, then symbol
|
||
table nodes or heap objects, then headers) plus the pass that
|
||
finishes; it cannot go lower without reading structures before their
|
||
addresses are known. Opening one dataset of a v1 group looks its name
|
||
up down the group's B-tree (5 requests, 5 MB at 1 MiB blocks for one
|
||
64 KiB dataset of the 3000). Each pass re-parses what the call reads
|
||
(CPU, not network).
|
||
- **Scattered metadata:** with headers spread through the file (h5py
|
||
writes each next to its data) a listing still fetches most of the file
|
||
at 1 MiB blocks (192 of 198 MB; 35 MB in 530 requests at 64 KiB); a
|
||
smaller `blockSize` fetches less. Merging nearby requests does not
|
||
help such a file (the blocks a listing needs are five or six apart at
|
||
64 KiB). Listing it a second time is free at 64 KiB blocks, but at
|
||
1 MiB its metadata blocks (192 MB) exceed the 64 MiB `cacheSize`, so
|
||
they are fetched again; a larger `cacheSize` keeps them. A file's paged
|
||
aggregation (metadata in pages) is not used to fetch its metadata in
|
||
one request.
|
||
- **Memory:** a call keeps every block it reads until it finishes (the
|
||
cache budget applies between calls). It may fetch at most `maxFetch`
|
||
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||
than that fails before anything is fetched: a hostile server cannot
|
||
make the page fetch or allocate what a file's lengths claim. `read()`
|
||
of a dataset that would take more than 1 GiB while it is decoded
|
||
(stored bytes, the values widened to 64 bits, the result) fails,
|
||
naming `readHyperslab`; read windows of large datasets. On wasm32 a
|
||
buffer past 2 GiB cannot exist and a failed allocation aborts the
|
||
whole module (every open file on the page), which these limits keep
|
||
from happening; before 2026-09-27 both did abort it.
|
||
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
|
||
open (file offsets become `usize`, 32 bits on wasm32). Offsets between
|
||
2 and 4 GiB are tested with a mock server; files above 200 MB have not
|
||
been served for real.
|
||
- **Cross-origin servers** must allow CORS for the page's origin and
|
||
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||
either header, only a change of length is detected.
|
||
- **A server without range support** (it answers `200`) costs a whole
|
||
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
|
||
with `fallback: "error"`. Every body is read as it arrives and cut off
|
||
past its limit: a server cannot make the page buffer more.
|
||
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||
size is not used. No retries: a failed request fails the call (calling
|
||
again retries it; what was fetched stays cached).
|
||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||
a local server, cross-origin included; not in Firefox or Safari.
|
||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||
refused with an error naming the type; attributes of those types come back
|
||
as `value: null` with their `dtype`. (VL strings read, with h5py's
|
||
values, through the same `VlResolver` as `File` and `h5rs`.)
|
||
- No Zstd or SZIP (both link C): such datasets fail with
|
||
`unsupported filter: 32015` / `: 4`. pcodec is not enabled either.
|
||
- External links and virtual-dataset sources in other files cannot be
|
||
followed (no file system).
|
||
|
||
## Conformance: `bad_nbit_parms_walk.h5` flips between ref-bug and our-error
|
||
|
||
**Status:** open (found 2026-09-28). Report noise, not a clawhdf5 bug: the
|
||
ok count (602 of 697) and the gate are unaffected.
|
||
|
||
`hdf5/test/testfiles/bad_nbit_parms_walk.h5` `/Nbit_int_data_le` is an
|
||
object clawhdf5 refuses and HDF5 2.0 reads only by over-reading memory (see
|
||
[the last non-ok files](#conformance-the-last-non-ok-files)).
|
||
`conformance/ref_bugs.py` counts it as *ref-bug* only when h5py's values
|
||
differ across its six differently set-up reads. The committed
|
||
`CONFORMANCE.md` (run of 2026-09-28 04:29 UTC, `bf5a163`) saw one distinct
|
||
result in six reads, so it reports the file as **1 our-error** and 2
|
||
ref-bug; a rerun on tank the same day (16:06 UTC, `conformance/run.sh
|
||
--no-fetch`, docs-only changes on top of `9b5803f`) saw three distinct
|
||
results and reports 0 our-error, 3 ref-bug. Whether libhdf5's over-read
|
||
changes between processes depends on heap layout, so six reads do not
|
||
always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more
|
||
varied reads for this object) or in documenting the file as a known
|
||
refusal; neither is done.
|
||
|
||
## NetCDF-4: an unlimited dimension reports size 0
|
||
|
||
**Status:** open (found 2026-09-28 while verifying the README refresh).
|
||
`clawhdf5-netcdf4`'s `NetCDF4File::dimensions()` reports an unlimited
|
||
dimension's `size` as 0 when variables along it hold records. Reproducer: with
|
||
netCDF4-python, create dimension `time` (unlimited) and `x` (3), a variable
|
||
`t(time, x)`, and write 2 records; netCDF4 reports `time` = 2 and `t` shape
|
||
(2, 3). clawhdf5-netcdf4 reports `dim time size 0 unlimited true`, while
|
||
`variable("t").shape()` correctly gives `[2, 3]`. In NetCDF-4 an unlimited
|
||
dimension's length is the largest extent of the variables that use it (its
|
||
dimension-scale dataset is not extended by netCDF-C), and the size is read
|
||
from the dimension scale instead. Variable shapes and values are correct;
|
||
only `Dimension::size` of unlimited dimensions is wrong. Workaround: use the
|
||
variables' shapes.
|
||
|
||
## The Node.js package (`packages/clawhdf5-node`) does not work
|
||
|
||
**Status:** open (found 2026-09-25; re-checked 2026-09-28, unchanged).
|
||
Unpublished; not built or tested in CI.
|
||
|
||
The TypeScript wrapper over `crates/clawhdf5-napi` has never run successfully:
|
||
|
||
- napi-rs converts `#[napi(object)]` fields to camelCase, but the wrapper reads
|
||
snake_case (`r.line_range`, `s.total_records`, `s.working_count`, …), so
|
||
every stats and consolidation field comes back `undefined`
|
||
(`src/index.ts`).
|
||
- It loads `../clawhdf5.node`, but `napi build --platform` produces
|
||
`clawhdf5.<triple>.node`; `main` points at `index.js` while `tsc` writes to
|
||
`dist/`; `napi prepublish` expects per-platform packages that are not
|
||
defined.
|
||
- `save`/`saveBatch` exist in the napi layer but not in the wrapper, so a
|
||
TypeScript caller cannot store an embedding at all.
|
||
- The WAL for `agent.brain` is `agent.h5.wal` (the store uses
|
||
`with_extension("h5.wal")`), not `agent.brain.wal` as the old docs and the
|
||
test cleanup assume.
|
||
|
||
It was written for an OpenClaw integration that is not being pursued (see
|
||
`docs/openclaw.md`). Fix and add CI, or remove it, before anyone depends on it.
|
||
|
||
---
|
||
|
||
# Fixed (history)
|
||
|
||
Newest first. "Before any release" means no tagged release (v2.7.0 and
|
||
earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date
|
||
given.
|
||
|
||
## `ObjectHeader::parse` 4% slower after range-read M2/M3
|
||
|
||
**Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release.
|
||
Speed only; values were always correct.
|
||
|
||
Measured on an idle tank (`BENCHMARKS.md`, "Local metadata and data reads
|
||
after range-read M2/M3"): `object_header_parse_x401` 23.86 → 24.86 µs
|
||
median (+4.2%) from `8f59b2e` to `7a8fae0`, after the parser began reading
|
||
continuation chunks from a bounded queue (`a69c5be`). The remaining cost was
|
||
the call to the version-1 message loop, kept out of line by
|
||
`#[inline(never)]`; inlined, it is 1.0% and 2.6% below `8f59b2e` in two idle
|
||
A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's speed").
|
||
A reported 5.6% slowdown of contiguous hyperslab reads, measured under load,
|
||
was noise and was withdrawn.
|
||
|
||
## Conformance: the last non-ok files
|
||
|
||
**Status:** classified 2026-09-27 (PR #21) — none is a clawhdf5 bug.
|
||
Result (`conformance/run.sh --no-fetch`, tank): 602 of 697 ok (was 600),
|
||
0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read.
|
||
|
||
- **2 mismatches were h5py's big-endian VL bug**
|
||
(`NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64`,
|
||
`hdf5/tools/test/testfiles/tcomplex_be.h5`
|
||
`/VariableLengthDatasetFloatComplex`): h5py returns the file's
|
||
big-endian bytes under a little-endian dtype. `conformance/ref.py` now
|
||
detects the bug in the installed h5py and relabels such elements, so the
|
||
values are compared; both files are *ok*.
|
||
- **3 our-errors are objects HDF5 2.0 reads only by over-reading memory**
|
||
(`cve-2025-2308.h5` `/Scale_offset_long_long_data_le`,
|
||
`cve-2025-44904.h5` `/Scale_offset_float_data_le`,
|
||
`bad_nbit_parms_walk.h5` `/Nbit_int_data_le`). libhdf5's output for them
|
||
is not determined by the file (it changes from run to run and with
|
||
`MALLOC_PERTURB_`), and libhdf5's develop branch refuses all three. We
|
||
keep refusing them; `conformance/ref_bugs.py` re-checks every run and
|
||
classifies such a file *ref-bug* only while its values keep changing.
|
||
|
||
## Damaged chunked datasets reported a different failing chunk per open
|
||
|
||
**Status:** fixed 2026-09-27 (`4ad8073`, PR #19). Errors only; the values of
|
||
readable datasets were never affected.
|
||
|
||
A full read through the file's chunk cache listed a damaged dataset's
|
||
chunks in hash-map order, seeded per `File`, so two opens of
|
||
`cve-2025-2310.h5` could name different failing chunks, and the storage and
|
||
wasm corpus comparisons failed now and then. The cache now keeps the chunk
|
||
index's order, as the uncached readers do
|
||
(`several_damaged_chunks_report_the_same_chunk_every_time`).
|
||
|
||
## Files a SWMR writer had open could not be read past a stale end of file
|
||
|
||
**Status:** fixed 2026-09-27 (PR #19), before any release: reads have been
|
||
bounded by the recorded end of file only since `7d7a7e7` (2026-09-26).
|
||
|
||
A libhdf5 SWMR writer does not keep the superblock's end-of-file address up
|
||
to date (a mid-write copy records 715 in a 6 030-byte file), so every reader
|
||
bounded such a file there: it listed, but chunked reads failed and `h5rs
|
||
check` reported chunk indexes past the end. Never wrong data. For a v3
|
||
superblock with the SWMR-write flag the data now ends at the end of the
|
||
file, as libhdf5's SWMR reader reads it. Test:
|
||
`crates/clawhdf5/tests/swmr_interop.rs` (fixture
|
||
`tests/fixtures/swmr_mid_write.h5`). A file still being written is read
|
||
with `File::open_swmr` (`docs/design/swmr.md`).
|
||
|
||
## Shrinking a chunked dataset with no recorded maximum scrambled it
|
||
|
||
**Status:** fixed 2026-09-27 (PR #19), before any release
|
||
(`FileEditor::resize` shipped on main in PR #18).
|
||
|
||
clawhdf5's writer stored no maximum dimensions for a chunked dataset
|
||
created without a `maxshape`; a Fixed Array index then places chunks by the
|
||
current dimensions, and `FileEditor::resize` changed only those, so a
|
||
shrink moved every chunk (wrong values in every reader) and the dataset
|
||
could not grow back. The editor now records the maximum libhdf5 would have
|
||
written before the first resize, and the writer records it for every
|
||
chunked dataset. Test: `crates/clawhdf5/tests/edit_resize_interop.rs`.
|
||
|
||
**What users must do:** datasets written before the fix (v2.7.0 and
|
||
earlier, and `main` before 2026-09-27) have no recorded maximum. The fixed
|
||
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
|
||
does not** — it scrambles them the same way and lets them grow past their
|
||
Fixed Array. Resize them once with a fixed `FileEditor` (a resize to the
|
||
same shape changes nothing; shrink and grow back) before letting libhdf5
|
||
resize them. A dataset already shrunk by the unfixed editor holds misplaced
|
||
chunks; rewrite it from a good copy.
|
||
|
||
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
|
||
|
||
**Status:** fixed 2026-09-26 (PR #18), after v2.7.0. **Every release
|
||
(v2.1.0 to v2.7.0) is affected**, in both directions.
|
||
|
||
Our Fletcher-32 reduced its sums with `% 65535`; libhdf5 folds them with
|
||
`(s & 0xffff) + (s >> 16)`, which differs where a sum is a non-zero
|
||
multiple of 65535. Chunks we wrote with such a sum are refused by h5py and
|
||
libhdf5 ("filter returned failure during read"); chunks libhdf5 wrote with
|
||
one were refused here with `Fletcher32Mismatch` (the data itself was never
|
||
wrong). The fix, `clawhdf5_format::checksum::fletcher32`, ports
|
||
`H5_checksum_fletcher32` and also accepts the byte-swapped (libhdf5 ≤ 1.6.2)
|
||
and the old clawhdf5 forms. Test: `crates/clawhdf5/tests/fletcher32_interop.rs`.
|
||
|
||
**What users must do:** a Fletcher-32 dataset written by v2.7.0 or earlier
|
||
may hold chunks libhdf5 cannot read; a fixed build reads them. Rewrite such
|
||
datasets with a fixed build before handing the file to libhdf5 or h5py.
|
||
|
||
## LZF/Blosc chunks written with a stale filter mask
|
||
|
||
**Status:** fixed 2026-09-26 (PR #17), before any release (v2.7.0 and
|
||
earlier write neither filter).
|
||
|
||
`FileBuilder` (and `FileEditor`) stored every LZF or Blosc chunk with
|
||
filter mask 0, where libhdf5 stores a chunk whose output is no smaller than
|
||
the chunk raw with the filter's mask bit set; after libhdf5 rewrote such a
|
||
chunk h5py could no longer read the dataset. Both now use
|
||
`clawhdf5_format::filters::compress_chunk_masked`. **What users must do:**
|
||
files written before the fix read correctly; rewrite them before letting
|
||
libhdf5 modify them.
|
||
|
||
## Concurrent and contiguous read performance
|
||
|
||
**Status:** fixed 2026-09-26 (PRs #15 and #16). Speed only.
|
||
|
||
Measured on tank against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, "First run,
|
||
before the read fixes"): full reads of chunked data from 16 threads through
|
||
one `File` stopped scaling at about 4 threads (880 MB/s vs 4424 MB/s for 16
|
||
h5py processes), and contiguous datasets read 4x slower than h5py on one
|
||
thread. Causes: every full read queued on a one-thread decode pool, fresh
|
||
buffers per chunk and a second copy of the output, 4 KiB page faults on the
|
||
output buffer, and element-by-element hyperslab copies. Re-measured at
|
||
`c5334b1` (`BENCHMARKS.md`, "Results after in-place chunk decoding"):
|
||
16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py
|
||
processes (1.58x), contiguous reads 6718 MB/s against h5py's 5545 on one
|
||
thread (1.21x; 1.29x h5py processes). Tests:
|
||
`single_thread_decode_pool.rs`, `busy_decode_pool.rs`.
|
||
|
||
## Scale-offset data read back wrong values
|
||
|
||
**Status:** fixed 2026-09-26 (PR #16), after v2.7.0. **Every release that
|
||
decoded the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on
|
||
ordinary files h5py writes, with no error.
|
||
|
||
Of 1480 scale-offset datasets h5py wrote (every integer type, `f4`/`f8`,
|
||
both byte orders, many fill values and scale factors), v2.7.0 read 332
|
||
differently from h5py: **151 returned wrong values with no error**, 181
|
||
failed. In every case libhdf5 had stored a chunk at full width (`minbits`
|
||
equal to the type's width), which was decoded as offsets from `minval`;
|
||
two smaller differences (codes start at byte 21 whatever `minval`'s
|
||
recorded size; `minbits` 0 with a fill value is all fill) were fixed with
|
||
it. Test: `crates/clawhdf5/tests/scaleoffset_interop.rs`.
|
||
|
||
**What users must do:** nothing to the files — they were always right; re-read
|
||
them with a fixed build.
|
||
|
||
## Crafted global heaps exhaust the variable-length reader's memory
|
||
|
||
**Status:** fixed 2026-09-26 (PR #15). Every earlier release is affected
|
||
through `read_vl_strings`.
|
||
|
||
Global heap collections nested inside one another's object data made
|
||
retained memory O(elements × file size) (a 744 KB file reached 1.58 GB) and
|
||
parse time O(elements × objects). `VlResolver` now caches object locations
|
||
within a 32 MiB budget and refuses overlapping collections, and collections
|
||
or objects running past their bounds are refused. Test:
|
||
`crates/clawhdf5-format/tests/vl_heap_bounds.rs`. The remaining
|
||
one-object-many-elements case is listed under
|
||
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
|
||
|
||
## Gaps found by the 2026-09-25 HDF5 audit (fixed parts)
|
||
|
||
**Status:** fixed 2026-09-25 and 2026-09-26 (PRs #11 to #17). Each was an
|
||
error unless marked **wrong data**; every release up to v2.7.0 has them.
|
||
What remains open is under
|
||
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
|
||
|
||
- **Layout message versions 1 and 2** (HDF5 1.6-era files; 84 of the 686
|
||
sweep files) and **compound datatype version 1 array members** (**wrong
|
||
data**: an array member read as one scalar) — fixed 2026-09-25.
|
||
- **Virtual datasets:** unmapped regions read as 0 instead of the fill
|
||
value (**wrong data**); printf-style and unlimited mappings; hyperslab
|
||
selection versions 1-3; the version-1 mapping list with a 2.0 low bound —
|
||
fixed 2026-09-25.
|
||
- **User blocks**, **old-style shared messages**, **user-defined link
|
||
types**, **dense groups over about 22 000 links**, **soft links in
|
||
`datasets()`**, a local-heap free list outside the heap (**wrong data**:
|
||
garbage names) — fixed 2026-09-25.
|
||
- **Dense attributes stored as huge/tiny/filtered fractal-heap objects**
|
||
(real NetCDF files, `issue671.nc`); an unreadable attribute no longer
|
||
hides the others (`attrs_with_errors()`) — fixed 2026-09-25.
|
||
- **VL strings through `File`** (`read_string`, `read_vlen::<T>()`,
|
||
`File::decode_strings`/`decode_vlen`), **VL data with 4-byte offsets**,
|
||
**metadata cache images** — fixed 2026-09-26.
|
||
- **Filters:** LZF, bitshuffle, bzip2 and Blosc 1 (read and write),
|
||
Blosc2 and ZFP (read) — fixed 2026-09-26, pure Rust. A chunk whose
|
||
filters decode to fewer bytes than the chunk read with zeros for the rest
|
||
(**wrong data**); now an error, and unfiltered chunks of the wrong stored
|
||
size are refused. A hostile Blosc chunk panicked with overflow checks.
|
||
- **Header checks:** 18 CVE objects libhdf5 refuses were read (some as
|
||
wrong data); object headers, datatypes, chunk dimensions and chunk-index
|
||
offsets are now checked as libhdf5 checks them, a dataspace whose storage
|
||
size overflows is refused at open, and the superblock extension is
|
||
decoded at open with libhdf5's checks — fixed 2026-09-26.
|
||
- **N-Bit/scale-offset on corrupt files** (`cve-2025-2308`,
|
||
`bad_nbit_parms_walk.h5`): not our bug — see
|
||
[the last non-ok files](#conformance-the-last-non-ok-files).
|
||
- **Writer:** groups nest to any depth with soft, hard and external links;
|
||
attribute creation order (`track_order`); dense link/attribute indexes of
|
||
any size (v2 B-trees of any depth); dense storage past 512 KiB was
|
||
written unreadable (it affected v2.7.0); libhdf5 could not add a link to
|
||
a group we wrote (no Group Info message); B-tree v2 chunk indexes larger
|
||
than one leaf — fixed 2026-09-26. Tests: `writer_groups_interop.rs`,
|
||
`deep_btree_interop.rs`.
|
||
|
||
## Silent wrong data found by the 2026-09-25 HDF5 audit
|
||
|
||
**Status:** fixed 2026-09-25 (PR #11), after v2.7.0. **Every release up to
|
||
and including v2.7.0 is affected.**
|
||
|
||
The audit (686 public files, 567 read cases and 96 write cases against h5py
|
||
3.16 / HDF5 1.10-2.0 and h5dump 1.14.6) found these values returned wrong
|
||
**without an error**:
|
||
|
||
| Area | What happened | Who is affected |
|
||
|---|---|---|
|
||
| Chunk index (read) | Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place | any file with a max shape larger than its shape and `libver='latest'` (h5py `maxshape=(10, None)`, `(20, 10)`) |
|
||
| Chunk index (write) | Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled | files we wrote with one unlimited dimension and > 244 chunks, or e.g. `maxshape=(20, None)` |
|
||
| 4-byte offsets | unfiltered chunked datasets read as zeros | files created with `sizeof_addr = 4` |
|
||
| Filter mask | any skipped filter skipped the whole pipeline | files with partially filtered chunks (optional filters, direct chunk writes) |
|
||
| Numeric reads | float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half | `read_i32`/`read_i64`/`read_u64` callers on float or wider data; HDF5 2.0 bf16 data |
|
||
| SZIP | garbage or zeros | every libhdf5-written SZIP dataset |
|
||
| Scale-offset | float values 1 ULP off | libhdf5 D-scale float data |
|
||
| Shared fill value | read as zero fill | fill values stored as shared messages |
|
||
| VL sequences | `read_vl_bytes` truncated non-byte base types | VL int/float sequences |
|
||
| Chunk cache | two threads reading two chunked datasets through one `File` could get each other's chunks | multi-threaded readers, including Python with the GIL released |
|
||
|
||
It also found files we wrote that libhdf5 **refuses**, fixed with it: Fixed
|
||
Array datasets with more than 1 024 chunks, header messages over 64 KiB,
|
||
Reference/Opaque/BitField/Time datatypes, files written with
|
||
`with_page_size`, several unlimited dimensions, a finite max shape larger
|
||
than the shape, an empty-string attribute (which broke every attribute on
|
||
its object), and rotated `FillTime` codes. Our LZ4 and Zstd output could not
|
||
be read by libhdf5's registered plugins, and our pcodec filter used Granular
|
||
BitRound's ID. Details: `CHANGELOG.md`, Correctness and Interop.
|
||
|
||
**What users must do:** re-read affected files with a fixed build; rewrite
|
||
files clawhdf5 wrote in the affected cases (one unlimited dimension with more
|
||
than 244 chunks, an unlimited dimension not first, the refused cases) before
|
||
handing them to libhdf5.
|
||
|
||
The sweep became `conformance/run.sh` (corpora pinned by commit; nightly by
|
||
`.gitea/workflows/conformance.yml`); current numbers are in
|
||
`CONFORMANCE.md`. No sweep run found a panic, hang or crash, including on
|
||
the 147 CVE and fuzzer files on some of which h5dump 1.14.6 and h5py/HDF5
|
||
2.0 segfault or abort.
|
||
|
||
## Every `f32` dataset we wrote was unreadable by h5py / libhdf5
|
||
|
||
**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. **Every release up to
|
||
and including v2.7.0 is affected** (the encoder was already wrong in
|
||
v2.1.0).
|
||
|
||
The floating-point datatype message's sign-bit position was written as 63
|
||
for every float; libhdf5 validates it, so every `f32` dataset — including
|
||
every agent store's `/memory/embeddings`, `norms` and `activation_weights`
|
||
— failed to open ("sign bit position out of bounds"). It is now
|
||
`bit_offset + bit_precision - 1`. Tests:
|
||
`float_sign_location_is_the_top_bit_of_the_value`,
|
||
`clawhdf5_writes_f32_h5py_reads`, the agent's
|
||
`h5py_reads_every_dataset_of_an_agent_store`.
|
||
|
||
**What users must do:** an agent store is rewritten at every checkpoint, so
|
||
it becomes readable by h5py at its next checkpoint with a fixed build. Other
|
||
files with `f32` datasets need to be rewritten.
|
||
|
||
## Empty datasets we wrote were unreadable by h5py / libhdf5
|
||
|
||
**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. Every release up to and
|
||
including v2.7.0 is affected.
|
||
|
||
An empty dataset was written with a real address and size 0, which
|
||
libhdf5's overflow check rejects ("invalid dataset size, likely file
|
||
corruption") — every agent store without sessions or a knowledge graph. An
|
||
empty contiguous dataset now gets the undefined address, as libhdf5 writes.
|
||
**What users must do:** as for `f32` above (agent stores heal at their next
|
||
checkpoint; rewrite other files).
|
||
|
||
## Extensible Array chunk indexes read back wrong data past the inline elements
|
||
|
||
**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including
|
||
v2.6.0 is affected.**
|
||
|
||
A dataset with exactly one unlimited dimension is indexed by an Extensible
|
||
Array, whose data and super block layout was computed wrongly: with more
|
||
than 36 chunks values came back from the wrong chunks **with no error** (37
|
||
chunks: 1 element wrong; 400: 364 wrong), and from about 1000 chunks the
|
||
read failed. Fixed with interop tests against HDF5 2.0 across every
|
||
boundary. **What users must do:** re-read with v2.7.0 or later. (The
|
||
writer's own 244-chunk bug is under the
|
||
[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit).)
|
||
|
||
## Crafted B-tree v2 structures crash or exhaust the reader
|
||
|
||
**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including
|
||
v2.6.0 is affected** (for untrusted files).
|
||
|
||
A self-referencing node under a header claiming 65 535 levels overflowed the
|
||
stack and aborted the process (under 100 bytes), and shared children made
|
||
traversal visit a node fan-out^depth times (memory exhaustion from ~5 KB).
|
||
Depth is now capped at 64 and traversal stops past the records the file
|
||
could hold.
|
||
|
||
## Python interop suites skip silently when no interpreter has h5py
|
||
|
||
**Status:** fixed 2026-09-19 (`a29c1b2`), in v2.6.0.
|
||
|
||
On a PEP 668 system every h5py/netCDF4 interop suite skipped and CI stayed
|
||
green. The probes now read `CLAWHDF5_PYTHON`, `scripts/ci-test.sh` picks up
|
||
`.venv/bin/python`, and `CLAWHDF5_REQUIRE_INTEROP=1` (set in CI) makes a
|
||
missing interpreter a failure. To restore coverage on a fresh checkout:
|
||
|
||
```bash
|
||
python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4
|
||
```
|
||
|
||
## B-tree v2 chunk index (layout v4, index type 5) is not supported
|
||
|
||
**Status:** fixed 2026-09-19, in v2.5.0. Datasets with two or more
|
||
unlimited dimensions written with `libver='latest'` failed to read in
|
||
earlier releases; record types 10 and 11 are now decoded on every read path.
|
||
|
||
## Attributes with unsupported datatypes are silently dropped
|
||
|
||
**Status:** fixed 2026-09-19. `attrs()` omitted booleans, complex, compound
|
||
and reference attributes without notice, and cast `u64` arrays to `i64`.
|
||
Booleans now decode as 0/1, `AttrValue::U64Array` keeps unsigned arrays, and
|
||
`AttrValue::Raw` carries anything else. (Multi-dimensional numeric
|
||
attributes are still flattened; see
|
||
[HDF5 features still unsupported](#hdf5-features-still-unsupported).)
|
||
|
||
## Compound datatype versions 1 and 2 are mis-parsed (default libver files)
|
||
|
||
**Status:** fixed 2026-09-19. Every compound dataset written with default
|
||
libver bounds (plain `h5py.File(path, 'w')`) failed to read: the v1 legacy
|
||
array fields are 28 bytes, not 24, and v2 keeps the name padding.
|
||
|
||
## `clawhdf5-gpu` `gpu_tests` can hang under the default parallel test runner
|
||
|
||
**Status:** fixed 2026-09-19 (`706189c`), in v2.3.0. Tests now hold a
|
||
process-wide lock while they own a device, and `GpuAccelerator` readback
|
||
times out after 30 s with `GpuError::BufferMap`.
|
||
|
||
## Revised reference datatype (class 7, version 4) is not parsed
|
||
|
||
**Status:** fixed 2026-09-19 for object references. `H5T_STD_REF` object
|
||
references (`ReferenceType::Object2`) decode, tested against a file made
|
||
with libhdf5 through ctypes (`tests/fixtures/gen_std_ref.py`). Region and
|
||
attribute references are recognised but not decoded — see
|
||
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
|
||
|
||
## Native complex datatype (class 11) is mis-parsed (HDF5 2.0)
|
||
|
||
**Status:** fixed 2026-09-18 (`b55b7db`), in v2.2.0. HDF5 2.0 native complex
|
||
types were parsed as a compound member list (garbage or `UnexpectedEof`);
|
||
class 11 now parses its base type and surfaces as an `{r, i}` compound.
|
||
h5py's default complex mapping (a compound) was never affected.
|
||
|
||
## Compound datatype message version 5 is not parsed (HDF5 2.0)
|
||
|
||
**Status:** fixed on `main` in `a13ff51` (2026-06-03), in v2.2.0; **not in
|
||
v2.1.0**, which was cut five commits earlier. Reported by M. Scot
|
||
Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0.
|
||
|
||
v2.1.0 rejects any compound dataset written by HDF5 2.0 with
|
||
`libver='latest'` (`InvalidDatatypeVersion { class: 6, version: 5 }`).
|
||
Versions 3-5 are now accepted for compound and array datatypes, and data
|
||
layout message version 5 too (every chunked dataset HDF5 2.0 writes). Tests:
|
||
`test_compound_v5_from_hdf5_2_0`, `test_array_v5_from_hdf5_2_0`,
|
||
`writer_h5py_tests.rs::read_h5py_generated_compound`.
|