HDF5 1.12 revised the reference datatype (class 7) in datatype message version 4: reference types 2-4 are the new H5T_STD_REF object / dataset-region / attribute references. Datatype::parse rejected them with InvalidReferenceType, so any dataset of that type was unreadable. h5py cannot write this type, which is why it had never been tested. A real file was produced by calling the libhdf5 bundled in the h5py wheel through ctypes (H5T_STD_REF_g, H5Rcreate_object, H5Dwrite); the 2 KB result is committed as tests/fixtures/std_ref_hdf5_2_0.h5 with its generator, gen_std_ref.py. - ReferenceType gains Object2, DatasetRegion2 and Attribute, accepted only from datatype version 4. - read_object_references decodes Object2 elements: type(1) flags(1) token_size(1) token, zero-padded to the element size; the token is the target's object header address. A null reference decodes to the undefined address; an external reference, a wrong type byte or a token that doesn't fit is an error. The fixture test follows both references and checks they resolve to the objects they were created from. Co-Authored-By: Claude Fable 5.1 <[email protected]>
6.9 KiB
Known Issues
Bugs found during development or downstream use, tracked here because this
repository's issue tracker is disabled. One entry per bug; when an entry is
fixed, record the fix in CHANGELOG.md and update its status here rather than
deleting it.
Compound datatype message version 5 is not parsed (HDF5 2.0)
Status: fixed on main in a13ff51 (2026-06-03); not in the v2.1.0
tag, which was cut five commits earlier. Ships in the next release.
Reported by: M. Scot Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0.
Summary: clawhdf5-format v2.1.0 rejects any dataset with a compound
(struct) datatype written by an HDF5 2.0 library in libver='latest' mode:
InvalidDatatypeVersion { class: 6, version: 5 }.
Reproduction (h5py 3.16.0 / HDF5 2.0.0):
import h5py, numpy as np
dt = np.dtype([('x', 'f8'), ('y', 'f8'), ('id', 'i4')])
data = np.array([(1.0, 2.0, 10), (3.0, 4.0, 20)], dtype=dt)
f = h5py.File('compound.h5', 'w', libver='latest')
f.create_dataset('particles', data=data)
f.close()
Committed as crates/clawhdf5-format/tests/writer_h5py_tests.rs::read_h5py_generated_compound
(#[ignore]d; needs python3 with h5py on PATH). Run with
cargo test -p clawhdf5-format --test writer_h5py_tests -- --include-ignored:
v2.1.0 gives 25 passed / 1 failed; main passes everything.
Root cause: the compound (class 6) branch of Datatype::parse
(crates/clawhdf5-format/src/datatype.rs) accepted only versions 1–4. Datatype
message versions 4 and 5 changed only the Reference and Complex classes, so a
v5-tagged compound uses the unchanged v3 member-list layout.
Fix: versions 3–5 are accepted for compound (class 6) and array (class 10)
datatypes, and data layout message version 5 is accepted too (needed for every
chunked dataset written by HDF5 2.0). Byte-level regression tests:
test_compound_v5_from_hdf5_2_0, test_array_v5_from_hdf5_2_0.
Native complex datatype (class 11) is mis-parsed (HDF5 2.0)
Status: fixed 2026-09-18. Found while validating the report above.
Summary: HDF5 2.0 native complex types (H5T_COMPLEX_IEEE_F64LE etc.)
were parsed as if they carried a compound-style member list. The properties are
actually a single base floating-point datatype, so the parser produced a garbage
datatype, or UnexpectedEof when the complex type was a compound member. h5py's
default numpy-complex mapping is unaffected (it writes a {r, i} compound);
only files using the native type through the C API / h5py low-level API hit this.
Fix: class 11 parses its base type and is surfaced as the equivalent
{r, i} compound. Tests: test_complex_v5_from_hdf5_2_0,
test_compound_with_complex_member_from_hdf5_2_0,
writer_h5py_tests.rs::read_h5py_generated_native_complex.
Revised reference datatype (class 7, version 4) is not parsed
Status: fixed 2026-09-19 for object references; region and attribute references are recognised but not decoded.
Summary: HDF5 1.12+ H5T_STD_REF references use datatype message version 4
with reference types 2-4 (object / region / attribute), which Datatype::parse
rejected with InvalidReferenceType. h5py still writes the legacy references,
so no file had been available to test against.
Fix: a real file was produced by driving the libhdf5 bundled in the h5py
wheel through ctypes (tests/fixtures/gen_std_ref.py ->
std_ref_hdf5_2_0.h5). The three new types parse as
ReferenceType::{Object2, DatasetRegion2, Attribute}, and
read_object_references decodes Object2 elements (type, flags, token size,
token = target object header address). External references (flag bit 0) and
the region/attribute payloads are errors rather than misreads.
clawhdf5-gpu gpu_tests can hang under the default parallel test runner
Status: fixed 2026-09-19.
Summary: during cargo test --workspace the gpu_tests binary sat idle for
25+ minutes. Every test created its own wgpu::Instance + device (requesting
adapter-maximum limits) concurrently, and readback used an unbounded
device.poll(Wait).
Fix: tests hold a process-wide lock while they own a device, and
GpuAccelerator readback waits time out after 30 s with GpuError::BufferMap.
Compound datatype versions 1 and 2 are mis-parsed (default libver files)
Status: fixed 2026-09-19. Found by adding a default-libver axis to the h5py interop tests.
Summary: any compound dataset written with default libver bounds (plain
h5py.File(path, 'w'), datatype message version 1) failed to read, typically
with Overflow("compound member 'x': byte_offset(0) + field_size(4136977) ...").
Only libver='latest' files (version 3+) and files written by clawhdf5 itself
worked, which is why the existing tests never caught it.
Root cause: Datatype::parse skipped 24 bytes of legacy per-member array
fields for v1 where the format has 28 (dimensionality 1 + reserved 3 +
permutation 4 + reserved 4 + 4 dimension sizes 16), and treated v2 like v1 minus
name padding, whereas v2 keeps the 8-byte name padding and has no array fields.
Attributes with unsupported datatypes are silently dropped
Status: fixed 2026-09-19.
Summary: Dataset::attrs() / Group::attrs() returned only attributes
convertible to AttrValue and omitted the rest without any indication — every
Python bool (an HDF5 enum), complex, compound and reference attributes. Unsigned
64-bit arrays were also cast to I64Array, turning values above i64::MAX
negative.
Fix: booleans decode as 0/1 integers, AttrValue::U64Array keeps unsigned
arrays unsigned, and AttrValue::Raw { datatype, shape, data } carries any other
attribute verbatim. Both new variants are writable. Still lossy: a
multi-dimensional numeric attribute is returned as a flat array (its shape is
not reported).
B-tree v2 chunk index (layout v4, index type 5) is not supported
Status: fixed 2026-09-19.
Summary: a chunked dataset with two or more unlimited dimensions written
with libver='latest' indexes its chunks with a version-2 B-tree, and reading it
failed with unsupported chunked layout version=4, index_type=Some(5).
Fix: record types 10 (unfiltered) and 11 (filtered) are decoded — address, stored size, filter mask, scaled offsets — through the shared chunk-listing function, so full reads, cached reads, partial reads and fill-value handling all work. Covered by an h5py interop test (plain, gzip+shuffle, a 2500-chunk tree with internal nodes, a sparse dataset with a fill value, a hyperslab).
External links and external raw data are not followed
Status: open (by design for now); both are explicit errors.
Summary: a path through an external link returns
FormatError::ExternalLinkUnsupported { filename, object_path }, and a dataset
created with external=[...] storage returns
FormatError::ExternalDataFilesUnsupported. Neither is resolved. If support is
added, file names must be confined to the opened file's directory, as the
virtual-dataset resolver now does.