Files
clawhdf5/docs/known-issues.md
T
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00

46 KiB
Raw Blame History

Known Issues

Bugs and limits found during development or downstream use, tracked here because this repository's issue tracker is disabled. One entry per bug; when an entry is fixed, record the fix in CHANGELOG.md and move it to Fixed (history) with its date, the releases it affected and what users must do, rather than deleting it.

Releases referred to below: v2.1.0 (2026-06-03) to v2.7.0 (2026-09-20). Everything fixed after v2.7.0 is on main and unreleased.

Open issues

Checked against main at 9b5803f on 2026-09-28.

Issue Kind Since
NetCDF-4: an unlimited dimension reports size 0 wrong metadata (dimension size; variable shapes and values are right) 2026-09-28
In-place modification (FileEditor) limits refused edits (Error::Unsupported), space reuse per editor, no journal 2026-09-26
Python in-place editing limits refused writes (NotImplementedError), deliberate conversion differences 2026-09-27
Selection reads that decode more than the selection speed only 2026-09-26
HDF5 features still unsupported errors, never wrong data 2026-09-25 audit
External links and external raw data are not followed error, by design for now 2026-09-19
Range reads (File::open_storage) limits cost; zero-copy APIs need an in-memory file 2026-09-26
Remote files (clawhdf5-remote) limits untested backends, fixed block size, timeouts 2026-09-26
clawhdf5-wasm (browser) limits memory, round trips, unsupported types 2026-09-26
Conformance: bad_nbit_parms_walk.h5 flips between ref-bug and our-error report noise (ok count unaffected) 2026-09-28
The Node.js package does not work broken, unpublished, not in CI 2026-09-25

In-place modification (FileEditor) limits

Status: open (documented 2026-09-26, updated when version-2 B-tree chunk indexes, shrinking, dense attributes and space reuse were added). clawhdf5::FileEditor refuses, with Error::Unsupported and without writing anything:

  • new chunks in an implicit index (it has all of its chunks from the start; they are written in place, and allocated/filled on growth under early allocation as libhdf5 does);
  • variable-length and reference data;
  • chunks through a filter this build cannot encode (scale-offset, N-Bit, SZIP, or a plugin filter it lacks), even an optional one: libhdf5 skips an optional filter only when its own build lacks it, which none does for these;
  • attributes in dense storage when the heap cannot take them the way libhdf5 would: replacing the last attribute left in a heap block by one of another size (libhdf5 frees the block), a heap with I/O filters or child indirect blocks (more than about 512 KiB of attributes), free space in child indirect blocks, directly addressed huge objects; and shared attribute messages. Measured 2026-09-26 on tank with the review's random-edit harness (120 runs of 150 random edits, earliest/v110/ latest, about 5600 set_attr calls of 8 bytes to 6 KiB): 2.2% of set_attr calls are refused, every one the last-attribute-in-a-block replacement (24% before heap blocks could be skipped);
  • version-1 object headers asked for an attribute larger than a header message (they have no dense storage);
  • partial edge chunks stored unfiltered (H5Pset_chunk_opts), external raw data files, virtual datasets;
  • files with a metadata cache image, paged or persistent free-space management, a driver info block, or version-3 consistency flags set.

Space is reused only within one editor. Space an edit frees (a filtered chunk that moves, chunks a shrink removes, B-tree nodes merged away, a heap's replaced blocks) is reused by later edits of the same FileEditor; what is left when it is dropped is leaked, as libhdf5 leaks it without a persistent free-space manager (h5repack reclaims it). A chunk that is the last thing in the file grows in place, which covers the usual append. Measured 2026-09-26 on tank (cargo test -p clawhdf5-tools --test edit_interop -- --ignored --nocapture measure_append_waste, one editor for the whole workload; sizes are deterministic):

Workload FileEditor libhdf5 h5repack of ours
1000 appends of 100 f8, 1024-element chunks, unfiltered 810 504 B 810 504 B 810 360 B
same, gzip 306 780 B 306 058 B 306 104 B
2000 appends of 10 values, 4096-element gzip chunks 79 829 B 50 292 B 49 930 B

In the last case the chunk being appended to is followed by new index blocks and moves each time it grows, and the space it leaves is too small for its next, larger version.

No journal. A crash while an edit patches existing structures can leave the file inconsistent; see the FileEditor documentation.

Renamed files outside Linux. Edits always go to the file the editor opened. FileEditor::reader() (which the Python 'r+' handle reads through) reopens it through /proc/self/fd on Linux; elsewhere it reopens the path, and after the path was renamed or replaced it fails (on Unix; Windows cannot tell and would read whatever the path names).

Python in-place editing (clawhdf5.File(path, 'r+')) limits

Status: open (added 2026-09-27). The Python bindings edit through FileEditor, so its limits above apply, raised as NotImplementedError before anything is written. On top of them:

  • No new or deleted objects: create_dataset/create_group in an 'r+' file, del f[name] and del obj.attrs[name] raise NotImplementedError (the editor changes values, shapes and attributes only). Mode 'a' works on an existing file only.
  • Not writable from Python: compound fields by name (ds['x'] = …; whole elements of the same structured dtype are), HDF5 array-type elements, variable-length data, strings padded with spaces or NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded ones, h5py's, are writable), compounds containing such strings, null dataspaces, index-list writes of more than 2²² elements (write them in slices), and boolean-mask keys (ds[mask] = v, and mask reads; h5py supports both).
  • str attributes are fixed-length UTF-8, where h5py writes variable-length strings: h5py reads them back as bytes (numpy.bytes_), not str.
  • Numeric conversion follows libhdf5's native-order results, not its bugs. Arrays are converted as libhdf5 converts them (integers saturate, floats are truncated toward zero and clipped), checked value by value against h5py 3.16 / HDF5 2.0 in crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py (2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5 differs from h5py on purpose:
    • NaN into an integer dataset is a ValueError (libhdf5 stores 0, the minimum or 2⁶³ depending on the type);
    • when the dataset or the array is not in native byte order, libhdf5's "soft" conversions store a float in (-1, 0) as the integer minimum and wrap an unsigned value too large for the signed type of the same size (65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in native order;
    • libhdf5's native casts that are undefined in C: half floats into unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed integers (→ minimum), a float equal to the integer maximum rounded up in its precision (float32(2**31 - 1) into int32, float64(2**64 - 1) into uint64: → minimum or 0); clawhdf5 saturates;
    • a double in (65504, 65520) into a half float: libhdf5 stores infinity, clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
  • Each edit reopens the file (a new memory map) so that reads see it; reads from other threads wait while an edit is written.

Selection reads that decode more than the selection

Status: open (documented 2026-09-26; re-checked 2026-09-28 in Dataset::read_selection, crates/clawhdf5/src/reader.rs). A selection read (and so the Python ds[...]) of contiguous data copies just the selected runs; of chunked data it materialises only the selection's bounding box when that box covers at most half the dataset (partial_read). It decodes the whole dataset and extracts the selection instead when:

  • the bounding box covers more than half the dataset — which a strided selection across a chunked dataset (ds[::100]) always does, although it may touch few chunks;
  • the dataset is compact or virtual, or has no storage;
  • it is chunked with a non-default fill value (the box path does not fill unallocated chunks, so the fill-aware full read is used).

Values are correct in every case; this is cost only. Selections other than Selection::All also bypass the file's chunk cache.

HDF5 features still unsupported

Status: open. What remains of the gaps the 2026-09-25 audit found (the rest are fixed), plus later residuals. Each is an error or a documented difference, never wrong data.

  • Virtual datasets: the "first missing" view and a printf gap other than 0 (libhdf5 access properties we always read at their defaults), source-to-virtual type conversion other than a byte swap, nested virtual sources, and source files outside the virtual file's directory (refused with an error).

  • Datatypes: x87 long double and binary128 are refused. Revised (version-4) region and attribute references are recognised but not decoded, and external references are an error (object references decode). A multi-dimensional numeric attribute is returned as a flat array (its shape is not reported; AttrValue::Raw carries the shape).

  • Metadata cache images (read since 2026-09-26) differ from libhdf5 in that: libhdf5 fails only the first metadata read of an image it cannot load and then reads the file's own (possibly stale) metadata, where we keep failing; an image entry that runs past the end of file is refused (libhdf5 checks only its start); a flush-dependency parent flag is checked against the child count as libhdf5's debug build checks it (HDF5 2.0 release builds refuse every entry that has children, even in images they wrote); the superblock extension's driver-info and shared-message table messages are not decoded at open.

  • Filters: Blosc2 and ZFP are read-only (clawhdf5 cannot write them). Blosc2 frames using dictionaries, lazy chunks, variable-length blocks, user-defined codecs or registered filters (e.g. bytedelta) are refused. Other filters (32023 Granular BitRound among them, even with the pcodec feature) can be plugged in with filter_registry::register_filter.

  • Writer: in dense storage (more than 8 attributes on an object, or more than 8 links in a group) one attribute or link message over 65 515 bytes is an error (no huge fractal-heap objects). The writer does not produce output that HDF5 1.8 can read.

  • Checks we deliberately do not make:

    • a float sign bit position outside the type, and a size-0 string type: clawhdf5 up to v2.7.0 wrote them;
    • a v4 chunked layout whose dimensions are encoded in more bytes than they need: current libhdf5 reads it (HDFGroup/hdf5@e124c36, 2026-06-05) though HDF5 2.0.0 (h5py 3.16) refuses it, and clawhdf5 wrote such layouts until 2026-09-26;
    • bit-field offset/precision outside the type, an unknown variable-length kind, an array type whose stored size is not its element count times its base size: HDF5 2.0 reads them though newer libhdf5 refuses them.

    h5rs check validates with the library's parsers, so it inherits what they accept: of the 150 CVE and fuzzer files, check --data passes 15, and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26).

  • VL sequences: a file may point many elements at one large global heap object, and a VL-sequence read then returns that object once per element, as h5py would (memory is bounded otherwise; see Crafted global heaps).

Status: open (by design for now; re-checked 2026-09-28); both are explicit errors. A path through an external link returns FormatError::ExternalLinkUnsupported { filename, object_path }, and a dataset created with external=[...] storage returns FormatError::ExternalDataFilesUnsupported. If support is added, file names must be confined to the opened file's directory, as the virtual-dataset resolver does.

Range reads (File::open_storage) limits

Status: open (added 2026-09-26 with milestone M2 of docs/design/range-reads.md; updated for M3-M5). File::open_storage reads any clawhdf5_format::storage::Storage through the whole read API, and clawhdf5-remote serves HTTP(S) and object-store files through a block cache, but:

  • A Storage without a cache is asked for each structure as the parsers need it, several times over for some (an object header is re-read by each lookup through it): one pass over the conformance corpus — open, list, every attribute, every dataset once — is 176 092 read_at calls for 621 files, 92 489 of them for the 35 001-group h5stat_newgrat.h5 (2026-09-26, tank, crates/clawhdf5/tests/storage_equivalence.rs with CLAWHDF5_STORAGE_CORPUS). A backend of your own over a network needs a cache in front of it: wrap it in clawhdf5_remote::BlockCache, as open_url does. Storage::read_ranges defaults to one read_at per range; coalescing is the backend's (or the cache's) job.
  • A group lookup by name in a version-1 (symbol-table) group lists the whole group (dense groups use their name index). Over a range backend that is one read per symbol-table node and name, per lookup. (The wasm lazy reader walks the group's B-tree instead; see its entry.)
  • External virtual-dataset source files are loaded whole through the resolver (File::set_vds_resolver), as bytes; they are not read through a Storage.
  • The zero-copy methods (Dataset::read_raw_ref, read_as_slice, read_*_zerocopy) need the file in memory and answer FormatError::ContiguousStorageRequired otherwise; File::as_bytes() panics for such a file (File::contiguous_bytes() is the fallible form).
  • LazyFile and MmapFile still read a whole local file. h5rs and the Python bindings read through File::storage and take URLs (h5rs with its remote feature, Python with clawhdf5.File(url)); the wasm reader's openUrl reads by range requests, its open(bytes) takes a whole file.
  • A growing local file is followed only when opened with File::open_swmr (milestone M5, 2026-09-27; see docs/design/swmr.md): File::open and open_storage read the length once, at open. SWMR reading is not available for remote files (a remote file is pinned at open, so one that grows is RemoteError::FileChanged), MmapFile or LazyFile, and clawhdf5 has no SWMR writer.

Remote files (clawhdf5-remote) limits

Status: open (added 2026-09-26 with milestone M3 of docs/design/range-reads.md). Python (clawhdf5.File(url)) and the browser (clawhdf5-wasm's openUrl) open URLs since 2026-09-27.

  • Default builds read plain http:// only. https:// needs the https feature (rustls with ring, which compiles C) and s3://, gs://, az:// the s3, gcs, azure features (aws-lc-rs); the default Python wheel has none of them. The Python tests run against an in-process http.server only.
  • The block size is fixed (1 MiB unless CacheConfig says otherwise). The design's policy of using a paged file's page size as the block size is not implemented, and only the first block is read ahead.
  • Checked against local servers and one public one. The tests use an in-process HTTP/1.1 server and object_store's in-memory and local-file stores. HTTPS was checked by hand against raw.githubusercontent.com (2026-09-26, tank: h5rs dump of h5py's vlen_string_dset.h5 by URL equals the downloaded file's). The s3, gcs and azure backends are built and their URL parsing tested, but they have not been run against a real bucket.
  • A server with neither a strong ETag nor Last-Modified can only be checked by length, so a same-length replacement mid-read would go unnoticed; HttpOptions::require_validator refuses such servers. A weak ETag (W/"…") cannot be sent as If-Match, so it counts as none.
  • Credentials: HTTP takes extra headers (HttpOptions::headers, e.g. Authorization); h5rs has no option for them. The cloud stores read credentials from the environment only. Messages and Debug output show URLs through redact_url (no userinfo, query values REDACTED); HttpStorage::url() returns the URL as given and must not be logged.
  • Redirects: at most HttpOptions::max_redirects (5) per request, never from https to http, and the custom headers are not sent once a redirect leaves the URL's origin. The redirect target is not remembered: every request of a redirected file costs its hops again.
  • Timeouts: ureq has no idle timeout, only a total one for the body, so the body's budget is timeout + size / min_speed (30 s + 16 KiB/s by default). A connection that stalls mid-body is detected only when that budget runs out (94 s for a 1 MiB block, 9 min for an 8 MiB run).
  • Each ObjectStoreStorage owns a tokio runtime with two worker threads. A read called from inside another runtime blocks that runtime's thread for its duration (it works, but spawn_blocking is the better place).
  • h5rs check downloads a remote file whole (it validates every byte), up to --max-download (1 GiB by default), and a URL cannot carry a FILE/OBJECT suffix; h5rs uses the default cache settings.
  • The zero-copy methods and File::as_bytes are unavailable on a remote file (see the range-read limits above).

clawhdf5-wasm (browser) limits

Status: open (by design for now; added 2026-09-26, openUrl 2026-09-27).

  • open() holds the whole file in memory (it takes its bytes), so a multi-GB local file does not fit a browser tab. A file on a web server can be opened with openUrl() instead, which fetches only the byte ranges each call needs (milestone M4 of docs/design/range-reads.md), with these limits:
    • Round trips: a call runs as passes over the blocks fetched so far and is re-run after each wave of misses, so a call costs one round trip per wave: a chunk index is walked a level (or a node) per round trip, while the chunks of a read are fetched together. Since 2026-09-27 listing 3000 datasets takes 4 passes for an h5py file and 5 for a libver="latest" file at 1 MiB blocks (5 and 6 at 64 KiB): index walks go on past a missing node, and parsers hint what they read next (Storage::hint), which the lazy reader fetches with a pass's misses. That is the depth of the chain (index levels, then symbol table nodes or heap objects, then headers) plus the pass that finishes; it cannot go lower without reading structures before their addresses are known. Opening one dataset of a v1 group looks its name up down the group's B-tree (5 requests, 5 MB at 1 MiB blocks for one 64 KiB dataset of the 3000). Each pass re-parses what the call reads (CPU, not network).
    • Scattered metadata: with headers spread through the file (h5py writes each next to its data) a listing still fetches most of the file at 1 MiB blocks (192 of 198 MB; 35 MB in 530 requests at 64 KiB); a smaller blockSize fetches less. Merging nearby requests does not help such a file (the blocks a listing needs are five or six apart at 64 KiB). Listing it a second time is free at 64 KiB blocks, but at 1 MiB its metadata blocks (192 MB) exceed the 64 MiB cacheSize, so they are fetched again; a larger cacheSize keeps them. A file's paged aggregation (metadata in pages) is not used to fetch its metadata in one request.
    • Memory: a call keeps every block it reads until it finishes (the cache budget applies between calls). It may fetch at most maxFetch bytes (512 MiB by default, at most 1 GiB), and a single read longer than that fails before anything is fetched: a hostile server cannot make the page fetch or allocate what a file's lengths claim. read() of a dataset that would take more than 1 GiB while it is decoded (stored bytes, the values widened to 64 bits, the result) fails, naming readHyperslab; read windows of large datasets. On wasm32 a buffer past 2 GiB cannot exist and a failed allocation aborts the whole module (every open file on the page), which these limits keep from happening; before 2026-09-27 both did abort it.
    • File size: at most 4 GiB - 1 bytes; a larger file is refused at open (file offsets become usize, 32 bits on wasm32). Offsets between 2 and 4 GiB are tested with a mock server; files above 200 MB have not been served for real.
    • Cross-origin servers must allow CORS for the page's origin and either expose Content-Range (Access-Control-Expose-Headers) or answer HEAD with Content-Length. The file is pinned at open by its ETag or Last-Modified (and its length): when the page cannot see either header, only a change of length is detected.
    • A server without range support (it answers 200) costs a whole download, up to maxDownload (512 MiB, at most 1 GiB), or an error with fallback: "error". Every body is read as it arrives and cut off past its limit: a server cannot make the page buffer more.
    • Fixed block size (blockSize, 1 MiB by default); a paged file's page size is not used. No retries: a failed request fails the call (calling again retries it; what was fetched stays cached).
    • Tested under Node 22 and headless Chromium (Playwright's build) against a local server, cross-origin included; not in Firefox or Safari.
  • Compound, reference, opaque, bitfield, time and VL-sequence datasets are refused with an error naming the type; attributes of those types come back as value: null with their dtype. (VL strings read, with h5py's values, through the same VlResolver as File and h5rs.)
  • No Zstd or SZIP (both link C): such datasets fail with unsupported filter: 32015 / : 4. pcodec is not enabled either.
  • External links and virtual-dataset sources in other files cannot be followed (no file system).

Conformance: bad_nbit_parms_walk.h5 flips between ref-bug and our-error

Status: open (found 2026-09-28). Report noise, not a clawhdf5 bug: the ok count (602 of 697) and the gate are unaffected.

hdf5/test/testfiles/bad_nbit_parms_walk.h5 /Nbit_int_data_le is an object clawhdf5 refuses and HDF5 2.0 reads only by over-reading memory (see the last non-ok files). conformance/ref_bugs.py counts it as ref-bug only when h5py's values differ across its six differently set-up reads. The committed CONFORMANCE.md (run of 2026-09-28 04:29 UTC, bf5a163) saw one distinct result in six reads, so it reports the file as 1 our-error and 2 ref-bug; a rerun on tank the same day (16:06 UTC, conformance/run.sh --no-fetch, docs-only changes on top of 9b5803f) saw three distinct results and reports 0 our-error, 3 ref-bug. Whether libhdf5's over-read changes between processes depends on heap layout, so six reads do not always expose it. A fix belongs in conformance/ref_bugs.py (more or more varied reads for this object) or in documenting the file as a known refusal; neither is done.

NetCDF-4: an unlimited dimension reports size 0

Status: open (found 2026-09-28 while verifying the README refresh). clawhdf5-netcdf4's NetCDF4File::dimensions() reports an unlimited dimension's size as 0 when variables along it hold records. Reproducer: with netCDF4-python, create dimension time (unlimited) and x (3), a variable t(time, x), and write 2 records; netCDF4 reports time = 2 and t shape (2, 3). clawhdf5-netcdf4 reports dim time size 0 unlimited true, while variable("t").shape() correctly gives [2, 3]. In NetCDF-4 an unlimited dimension's length is the largest extent of the variables that use it (its dimension-scale dataset is not extended by netCDF-C), and the size is read from the dimension scale instead. Variable shapes and values are correct; only Dimension::size of unlimited dimensions is wrong. Workaround: use the variables' shapes.

The Node.js package (packages/clawhdf5-node) does not work

Status: open (found 2026-09-25; re-checked 2026-09-28, unchanged). Unpublished; not built or tested in CI.

The TypeScript wrapper over crates/clawhdf5-napi has never run successfully:

  • napi-rs converts #[napi(object)] fields to camelCase, but the wrapper reads snake_case (r.line_range, s.total_records, s.working_count, …), so every stats and consolidation field comes back undefined (src/index.ts).
  • It loads ../clawhdf5.node, but napi build --platform produces clawhdf5.<triple>.node; main points at index.js while tsc writes to dist/; napi prepublish expects per-platform packages that are not defined.
  • save/saveBatch exist in the napi layer but not in the wrapper, so a TypeScript caller cannot store an embedding at all.
  • The WAL for agent.brain is agent.h5.wal (the store uses with_extension("h5.wal")), not agent.brain.wal as the old docs and the test cleanup assume.

It was written for an OpenClaw integration that is not being pursued (see docs/openclaw.md). Fix and add CI, or remove it, before anyone depends on it.


Fixed (history)

Newest first. "Before any release" means no tagged release (v2.7.0 and earlier) contains the bug. Full detail is in CHANGELOG.md under the date given.

ObjectHeader::parse 4% slower after range-read M2/M3

Status: fixed 2026-09-27 (96086ad, PR #21), before any release. Speed only; values were always correct.

Measured on an idle tank (BENCHMARKS.md, "Local metadata and data reads after range-read M2/M3"): object_header_parse_x401 23.86 → 24.86 µs median (+4.2%) from 8f59b2e to 7a8fae0, after the parser began reading continuation chunks from a bounded queue (a69c5be). The remaining cost was the call to the version-1 message loop, kept out of line by #[inline(never)]; inlined, it is 1.0% and 2.6% below 8f59b2e in two idle A/B runs (BENCHMARKS.md, "ObjectHeader::parse back at 8f59b2e's speed"). A reported 5.6% slowdown of contiguous hyperslab reads, measured under load, was noise and was withdrawn.

Conformance: the last non-ok files

Status: classified 2026-09-27 (PR #21) — none is a clawhdf5 bug. Result (conformance/run.sh --no-fetch, tank): 602 of 697 ok (was 600), 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read.

  • 2 mismatches were h5py's big-endian VL bug (NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5 /@vlen_uint64, hdf5/tools/test/testfiles/tcomplex_be.h5 /VariableLengthDatasetFloatComplex): h5py returns the file's big-endian bytes under a little-endian dtype. conformance/ref.py now detects the bug in the installed h5py and relabels such elements, so the values are compared; both files are ok.
  • 3 our-errors are objects HDF5 2.0 reads only by over-reading memory (cve-2025-2308.h5 /Scale_offset_long_long_data_le, cve-2025-44904.h5 /Scale_offset_float_data_le, bad_nbit_parms_walk.h5 /Nbit_int_data_le). libhdf5's output for them is not determined by the file (it changes from run to run and with MALLOC_PERTURB_), and libhdf5's develop branch refuses all three. We keep refusing them; conformance/ref_bugs.py re-checks every run and classifies such a file ref-bug only while its values keep changing.

Damaged chunked datasets reported a different failing chunk per open

Status: fixed 2026-09-27 (4ad8073, PR #19). Errors only; the values of readable datasets were never affected.

A full read through the file's chunk cache listed a damaged dataset's chunks in hash-map order, seeded per File, so two opens of cve-2025-2310.h5 could name different failing chunks, and the storage and wasm corpus comparisons failed now and then. The cache now keeps the chunk index's order, as the uncached readers do (several_damaged_chunks_report_the_same_chunk_every_time).

Files a SWMR writer had open could not be read past a stale end of file

Status: fixed 2026-09-27 (PR #19), before any release: reads have been bounded by the recorded end of file only since 7d7a7e7 (2026-09-26).

A libhdf5 SWMR writer does not keep the superblock's end-of-file address up to date (a mid-write copy records 715 in a 6 030-byte file), so every reader bounded such a file there: it listed, but chunked reads failed and h5rs check reported chunk indexes past the end. Never wrong data. For a v3 superblock with the SWMR-write flag the data now ends at the end of the file, as libhdf5's SWMR reader reads it. Test: crates/clawhdf5/tests/swmr_interop.rs (fixture tests/fixtures/swmr_mid_write.h5). A file still being written is read with File::open_swmr (docs/design/swmr.md).

Shrinking a chunked dataset with no recorded maximum scrambled it

Status: fixed 2026-09-27 (PR #19), before any release (FileEditor::resize shipped on main in PR #18).

clawhdf5's writer stored no maximum dimensions for a chunked dataset created without a maxshape; a Fixed Array index then places chunks by the current dimensions, and FileEditor::resize changed only those, so a shrink moved every chunk (wrong values in every reader) and the dataset could not grow back. The editor now records the maximum libhdf5 would have written before the first resize, and the writer records it for every chunked dataset. Test: crates/clawhdf5/tests/edit_resize_interop.rs.

What users must do: datasets written before the fix (v2.7.0 and earlier, and main before 2026-09-27) have no recorded maximum. The fixed editor handles them; libhdf5 (h5py Dataset.resize, H5Dset_extent) does not — it scrambles them the same way and lets them grow past their Fixed Array. Resize them once with a fixed FileEditor (a resize to the same shape changes nothing; shrink and grow back) before letting libhdf5 resize them. A dataset already shrunk by the unfixed editor holds misplaced chunks; rewrite it from a good copy.

Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768

Status: fixed 2026-09-26 (PR #18), after v2.7.0. Every release (v2.1.0 to v2.7.0) is affected, in both directions.

Our Fletcher-32 reduced its sums with % 65535; libhdf5 folds them with (s & 0xffff) + (s >> 16), which differs where a sum is a non-zero multiple of 65535. Chunks we wrote with such a sum are refused by h5py and libhdf5 ("filter returned failure during read"); chunks libhdf5 wrote with one were refused here with Fletcher32Mismatch (the data itself was never wrong). The fix, clawhdf5_format::checksum::fletcher32, ports H5_checksum_fletcher32 and also accepts the byte-swapped (libhdf5 ≤ 1.6.2) and the old clawhdf5 forms. Test: crates/clawhdf5/tests/fletcher32_interop.rs.

What users must do: a Fletcher-32 dataset written by v2.7.0 or earlier may hold chunks libhdf5 cannot read; a fixed build reads them. Rewrite such datasets with a fixed build before handing the file to libhdf5 or h5py.

LZF/Blosc chunks written with a stale filter mask

Status: fixed 2026-09-26 (PR #17), before any release (v2.7.0 and earlier write neither filter).

FileBuilder (and FileEditor) stored every LZF or Blosc chunk with filter mask 0, where libhdf5 stores a chunk whose output is no smaller than the chunk raw with the filter's mask bit set; after libhdf5 rewrote such a chunk h5py could no longer read the dataset. Both now use clawhdf5_format::filters::compress_chunk_masked. What users must do: files written before the fix read correctly; rewrite them before letting libhdf5 modify them.

Concurrent and contiguous read performance

Status: fixed 2026-09-26 (PRs #15 and #16). Speed only.

Measured on tank against h5py 3.16 / HDF5 2.0 (BENCHMARKS.md, "First run, before the read fixes"): full reads of chunked data from 16 threads through one File stopped scaling at about 4 threads (880 MB/s vs 4424 MB/s for 16 h5py processes), and contiguous datasets read 4x slower than h5py on one thread. Causes: every full read queued on a one-thread decode pool, fresh buffers per chunk and a second copy of the output, 4 KiB page faults on the output buffer, and element-by-element hyperslab copies. Re-measured at c5334b1 (BENCHMARKS.md, "Results after in-place chunk decoding"): 16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py processes (1.58x), contiguous reads 6718 MB/s against h5py's 5545 on one thread (1.21x; 1.29x h5py processes). Tests: single_thread_decode_pool.rs, busy_decode_pool.rs.

Scale-offset data read back wrong values

Status: fixed 2026-09-26 (PR #16), after v2.7.0. Every release that decoded the scale-offset filter (v2.2.0 to v2.7.0) is affected, on ordinary files h5py writes, with no error.

Of 1480 scale-offset datasets h5py wrote (every integer type, f4/f8, both byte orders, many fill values and scale factors), v2.7.0 read 332 differently from h5py: 151 returned wrong values with no error, 181 failed. In every case libhdf5 had stored a chunk at full width (minbits equal to the type's width), which was decoded as offsets from minval; two smaller differences (codes start at byte 21 whatever minval's recorded size; minbits 0 with a fill value is all fill) were fixed with it. Test: crates/clawhdf5/tests/scaleoffset_interop.rs.

What users must do: nothing to the files — they were always right; re-read them with a fixed build.

Crafted global heaps exhaust the variable-length reader's memory

Status: fixed 2026-09-26 (PR #15). Every earlier release is affected through read_vl_strings.

Global heap collections nested inside one another's object data made retained memory O(elements × file size) (a 744 KB file reached 1.58 GB) and parse time O(elements × objects). VlResolver now caches object locations within a 32 MiB budget and refuses overlapping collections, and collections or objects running past their bounds are refused. Test: crates/clawhdf5-format/tests/vl_heap_bounds.rs. The remaining one-object-many-elements case is listed under HDF5 features still unsupported.

Gaps found by the 2026-09-25 HDF5 audit (fixed parts)

Status: fixed 2026-09-25 and 2026-09-26 (PRs #11 to #17). Each was an error unless marked wrong data; every release up to v2.7.0 has them. What remains open is under HDF5 features still unsupported.

  • Layout message versions 1 and 2 (HDF5 1.6-era files; 84 of the 686 sweep files) and compound datatype version 1 array members (wrong data: an array member read as one scalar) — fixed 2026-09-25.
  • Virtual datasets: unmapped regions read as 0 instead of the fill value (wrong data); printf-style and unlimited mappings; hyperslab selection versions 1-3; the version-1 mapping list with a 2.0 low bound — fixed 2026-09-25.
  • User blocks, old-style shared messages, user-defined link types, dense groups over about 22 000 links, soft links in datasets(), a local-heap free list outside the heap (wrong data: garbage names) — fixed 2026-09-25.
  • Dense attributes stored as huge/tiny/filtered fractal-heap objects (real NetCDF files, issue671.nc); an unreadable attribute no longer hides the others (attrs_with_errors()) — fixed 2026-09-25.
  • VL strings through File (read_string, read_vlen::<T>(), File::decode_strings/decode_vlen), VL data with 4-byte offsets, metadata cache images — fixed 2026-09-26.
  • Filters: LZF, bitshuffle, bzip2 and Blosc 1 (read and write), Blosc2 and ZFP (read) — fixed 2026-09-26, pure Rust. A chunk whose filters decode to fewer bytes than the chunk read with zeros for the rest (wrong data); now an error, and unfiltered chunks of the wrong stored size are refused. A hostile Blosc chunk panicked with overflow checks.
  • Header checks: 18 CVE objects libhdf5 refuses were read (some as wrong data); object headers, datatypes, chunk dimensions and chunk-index offsets are now checked as libhdf5 checks them, a dataspace whose storage size overflows is refused at open, and the superblock extension is decoded at open with libhdf5's checks — fixed 2026-09-26.
  • N-Bit/scale-offset on corrupt files (cve-2025-2308, bad_nbit_parms_walk.h5): not our bug — see the last non-ok files.
  • Writer: groups nest to any depth with soft, hard and external links; attribute creation order (track_order); dense link/attribute indexes of any size (v2 B-trees of any depth); dense storage past 512 KiB was written unreadable (it affected v2.7.0); libhdf5 could not add a link to a group we wrote (no Group Info message); B-tree v2 chunk indexes larger than one leaf — fixed 2026-09-26. Tests: writer_groups_interop.rs, deep_btree_interop.rs.

Silent wrong data found by the 2026-09-25 HDF5 audit

Status: fixed 2026-09-25 (PR #11), after v2.7.0. Every release up to and including v2.7.0 is affected.

The audit (686 public files, 567 read cases and 96 write cases against h5py 3.16 / HDF5 1.10-2.0 and h5dump 1.14.6) found these values returned wrong without an error:

Area What happened Who is affected
Chunk index (read) Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place any file with a max shape larger than its shape and libver='latest' (h5py maxshape=(10, None), (20, 10))
Chunk index (write) Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled files we wrote with one unlimited dimension and > 244 chunks, or e.g. maxshape=(20, None)
4-byte offsets unfiltered chunked datasets read as zeros files created with sizeof_addr = 4
Filter mask any skipped filter skipped the whole pipeline files with partially filtered chunks (optional filters, direct chunk writes)
Numeric reads float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half read_i32/read_i64/read_u64 callers on float or wider data; HDF5 2.0 bf16 data
SZIP garbage or zeros every libhdf5-written SZIP dataset
Scale-offset float values 1 ULP off libhdf5 D-scale float data
Shared fill value read as zero fill fill values stored as shared messages
VL sequences read_vl_bytes truncated non-byte base types VL int/float sequences
Chunk cache two threads reading two chunked datasets through one File could get each other's chunks multi-threaded readers, including Python with the GIL released

It also found files we wrote that libhdf5 refuses, fixed with it: Fixed Array datasets with more than 1 024 chunks, header messages over 64 KiB, Reference/Opaque/BitField/Time datatypes, files written with with_page_size, several unlimited dimensions, a finite max shape larger than the shape, an empty-string attribute (which broke every attribute on its object), and rotated FillTime codes. Our LZ4 and Zstd output could not be read by libhdf5's registered plugins, and our pcodec filter used Granular BitRound's ID. Details: CHANGELOG.md, Correctness and Interop.

What users must do: re-read affected files with a fixed build; rewrite files clawhdf5 wrote in the affected cases (one unlimited dimension with more than 244 chunks, an unlimited dimension not first, the refused cases) before handing them to libhdf5.

The sweep became conformance/run.sh (corpora pinned by commit; nightly by .gitea/workflows/conformance.yml); current numbers are in CONFORMANCE.md. No sweep run found a panic, hang or crash, including on the 147 CVE and fuzzer files on some of which h5dump 1.14.6 and h5py/HDF5 2.0 segfault or abort.

Every f32 dataset we wrote was unreadable by h5py / libhdf5

Status: fixed 2026-09-23 (PR #4), after v2.7.0. Every release up to and including v2.7.0 is affected (the encoder was already wrong in v2.1.0).

The floating-point datatype message's sign-bit position was written as 63 for every float; libhdf5 validates it, so every f32 dataset — including every agent store's /memory/embeddings, norms and activation_weights — failed to open ("sign bit position out of bounds"). It is now bit_offset + bit_precision - 1. Tests: float_sign_location_is_the_top_bit_of_the_value, clawhdf5_writes_f32_h5py_reads, the agent's h5py_reads_every_dataset_of_an_agent_store.

What users must do: an agent store is rewritten at every checkpoint, so it becomes readable by h5py at its next checkpoint with a fixed build. Other files with f32 datasets need to be rewritten.

Empty datasets we wrote were unreadable by h5py / libhdf5

Status: fixed 2026-09-23 (PR #4), after v2.7.0. Every release up to and including v2.7.0 is affected.

An empty dataset was written with a real address and size 0, which libhdf5's overflow check rejects ("invalid dataset size, likely file corruption") — every agent store without sessions or a knowledge graph. An empty contiguous dataset now gets the undefined address, as libhdf5 writes. What users must do: as for f32 above (agent stores heal at their next checkpoint; rewrite other files).

Extensible Array chunk indexes read back wrong data past the inline elements

Status: fixed 2026-09-20, in v2.7.0. Every release up to and including v2.6.0 is affected.

A dataset with exactly one unlimited dimension is indexed by an Extensible Array, whose data and super block layout was computed wrongly: with more than 36 chunks values came back from the wrong chunks with no error (37 chunks: 1 element wrong; 400: 364 wrong), and from about 1000 chunks the read failed. Fixed with interop tests against HDF5 2.0 across every boundary. What users must do: re-read with v2.7.0 or later. (The writer's own 244-chunk bug is under the 2026-09-25 audit.)

Crafted B-tree v2 structures crash or exhaust the reader

Status: fixed 2026-09-20, in v2.7.0. Every release up to and including v2.6.0 is affected (for untrusted files).

A self-referencing node under a header claiming 65 535 levels overflowed the stack and aborted the process (under 100 bytes), and shared children made traversal visit a node fan-out^depth times (memory exhaustion from ~5 KB). Depth is now capped at 64 and traversal stops past the records the file could hold.

Python interop suites skip silently when no interpreter has h5py

Status: fixed 2026-09-19 (a29c1b2), in v2.6.0.

On a PEP 668 system every h5py/netCDF4 interop suite skipped and CI stayed green. The probes now read CLAWHDF5_PYTHON, scripts/ci-test.sh picks up .venv/bin/python, and CLAWHDF5_REQUIRE_INTEROP=1 (set in CI) makes a missing interpreter a failure. To restore coverage on a fresh checkout:

python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4

B-tree v2 chunk index (layout v4, index type 5) is not supported

Status: fixed 2026-09-19, in v2.5.0. Datasets with two or more unlimited dimensions written with libver='latest' failed to read in earlier releases; record types 10 and 11 are now decoded on every read path.

Attributes with unsupported datatypes are silently dropped

Status: fixed 2026-09-19. attrs() omitted booleans, complex, compound and reference attributes without notice, and cast u64 arrays to i64. Booleans now decode as 0/1, AttrValue::U64Array keeps unsigned arrays, and AttrValue::Raw carries anything else. (Multi-dimensional numeric attributes are still flattened; see HDF5 features still unsupported.)

Compound datatype versions 1 and 2 are mis-parsed (default libver files)

Status: fixed 2026-09-19. Every compound dataset written with default libver bounds (plain h5py.File(path, 'w')) failed to read: the v1 legacy array fields are 28 bytes, not 24, and v2 keeps the name padding.

clawhdf5-gpu gpu_tests can hang under the default parallel test runner

Status: fixed 2026-09-19 (706189c), in v2.3.0. Tests now hold a process-wide lock while they own a device, and GpuAccelerator readback times out after 30 s with GpuError::BufferMap.

Revised reference datatype (class 7, version 4) is not parsed

Status: fixed 2026-09-19 for object references. H5T_STD_REF object references (ReferenceType::Object2) decode, tested against a file made with libhdf5 through ctypes (tests/fixtures/gen_std_ref.py). Region and attribute references are recognised but not decoded — see HDF5 features still unsupported.

Native complex datatype (class 11) is mis-parsed (HDF5 2.0)

Status: fixed 2026-09-18 (b55b7db), in v2.2.0. HDF5 2.0 native complex types were parsed as a compound member list (garbage or UnexpectedEof); class 11 now parses its base type and surfaces as an {r, i} compound. h5py's default complex mapping (a compound) was never affected.

Compound datatype message version 5 is not parsed (HDF5 2.0)

Status: fixed on main in a13ff51 (2026-06-03), in v2.2.0; not in v2.1.0, which was cut five commits earlier. Reported by M. Scot Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0.

v2.1.0 rejects any compound dataset written by HDF5 2.0 with libver='latest' (InvalidDatatypeVersion { class: 6, version: 5 }). Versions 3-5 are now accepted for compound and array datatypes, and data layout message version 5 too (every chunked dataset HDF5 2.0 writes). Tests: test_compound_v5_from_hdf5_2_0, test_array_v5_from_hdf5_2_0, writer_h5py_tests.rs::read_h5py_generated_compound.