Files
clawhdf5/docs/known-issues.md
T
osobhandClaude Fable 5.1 97ab658c11 feat: attrs() reports every attribute; unsigned arrays stay unsigned
attrs() silently omitted any attribute whose datatype had no AttrValue variant
— including every Python bool, which h5py stores as an enum — plus complex,
compound and reference attributes, and cast unsigned 64-bit arrays to
I64Array so values above i64::MAX came back negative.

- numpy/h5py-style booleans (an enum of exactly FALSE=0 / TRUE=1 over an
  integer base) decode as I64 / I64Array of 0/1.
- AttrValue::U64Array keeps unsigned arrays unsigned. Behaviour change: an
  unsigned array attribute no longer arrives as I64Array; the netCDF-4 CF
  helpers (_FillValue, valid_range) and the Python bindings handle it.
- AttrValue::Raw { datatype, shape, data } carries any other attribute
  verbatim (also used when a value fails to decode as its declared type), so
  the attribute list is always complete. Decodable with data_read against the
  datatype; Python receives {"dtype", "shape", "data"}.
- Both new variants are writable, so attributes round-trip between files.
  h5py interop tests cover reading 13 attribute kinds and h5py reading back a
  compound and a u64 attribute written by clawhdf5.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-19 07:05:34 -07:00

6.3 KiB
Raw Permalink Blame History

Known Issues

Bugs found during development or downstream use, tracked here because this repository's issue tracker is disabled. One entry per bug; when an entry is fixed, record the fix in CHANGELOG.md and update its status here rather than deleting it.


Compound datatype message version 5 is not parsed (HDF5 2.0)

Status: fixed on main in a13ff51 (2026-06-03); not in the v2.1.0 tag, which was cut five commits earlier. Ships in the next release.

Reported by: M. Scot Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0.

Summary: clawhdf5-format v2.1.0 rejects any dataset with a compound (struct) datatype written by an HDF5 2.0 library in libver='latest' mode: InvalidDatatypeVersion { class: 6, version: 5 }.

Reproduction (h5py 3.16.0 / HDF5 2.0.0):

import h5py, numpy as np
dt = np.dtype([('x', 'f8'), ('y', 'f8'), ('id', 'i4')])
data = np.array([(1.0, 2.0, 10), (3.0, 4.0, 20)], dtype=dt)
f = h5py.File('compound.h5', 'w', libver='latest')
f.create_dataset('particles', data=data)
f.close()

Committed as crates/clawhdf5-format/tests/writer_h5py_tests.rs::read_h5py_generated_compound (#[ignore]d; needs python3 with h5py on PATH). Run with cargo test -p clawhdf5-format --test writer_h5py_tests -- --include-ignored: v2.1.0 gives 25 passed / 1 failed; main passes everything.

Root cause: the compound (class 6) branch of Datatype::parse (crates/clawhdf5-format/src/datatype.rs) accepted only versions 14. Datatype message versions 4 and 5 changed only the Reference and Complex classes, so a v5-tagged compound uses the unchanged v3 member-list layout.

Fix: versions 35 are accepted for compound (class 6) and array (class 10) datatypes, and data layout message version 5 is accepted too (needed for every chunked dataset written by HDF5 2.0). Byte-level regression tests: test_compound_v5_from_hdf5_2_0, test_array_v5_from_hdf5_2_0.

Native complex datatype (class 11) is mis-parsed (HDF5 2.0)

Status: fixed 2026-09-18. Found while validating the report above.

Summary: HDF5 2.0 native complex types (H5T_COMPLEX_IEEE_F64LE etc.) were parsed as if they carried a compound-style member list. The properties are actually a single base floating-point datatype, so the parser produced a garbage datatype, or UnexpectedEof when the complex type was a compound member. h5py's default numpy-complex mapping is unaffected (it writes a {r, i} compound); only files using the native type through the C API / h5py low-level API hit this.

Fix: class 11 parses its base type and is surfaced as the equivalent {r, i} compound. Tests: test_complex_v5_from_hdf5_2_0, test_compound_with_complex_member_from_hdf5_2_0, writer_h5py_tests.rs::read_h5py_generated_native_complex.

Revised reference datatype (class 7, version 4) is not parsed

Status: open, unconfirmed against a real file.

Summary: HDF5 1.12+ H5T_STD_REF references use datatype version 4 with reference types 24 (object2 / region2 / attribute), which Datatype::parse rejects with InvalidReferenceType. h5py still writes the legacy v1 object/region references, which read correctly, so no reproducing file has been generated yet; one written with the C API (H5T_STD_REF) is needed.

clawhdf5-gpu gpu_tests can hang under the default parallel test runner

Status: fixed 2026-09-19.

Summary: during cargo test --workspace the gpu_tests binary sat idle for 25+ minutes. Every test created its own wgpu::Instance + device (requesting adapter-maximum limits) concurrently, and readback used an unbounded device.poll(Wait).

Fix: tests hold a process-wide lock while they own a device, and GpuAccelerator readback waits time out after 30 s with GpuError::BufferMap.

Compound datatype versions 1 and 2 are mis-parsed (default libver files)

Status: fixed 2026-09-19. Found by adding a default-libver axis to the h5py interop tests.

Summary: any compound dataset written with default libver bounds (plain h5py.File(path, 'w'), datatype message version 1) failed to read, typically with Overflow("compound member 'x': byte_offset(0) + field_size(4136977) ..."). Only libver='latest' files (version 3+) and files written by clawhdf5 itself worked, which is why the existing tests never caught it.

Root cause: Datatype::parse skipped 24 bytes of legacy per-member array fields for v1 where the format has 28 (dimensionality 1 + reserved 3 + permutation 4 + reserved 4 + 4 dimension sizes 16), and treated v2 like v1 minus name padding, whereas v2 keeps the 8-byte name padding and has no array fields.

Attributes with unsupported datatypes are silently dropped

Status: fixed 2026-09-19.

Summary: Dataset::attrs() / Group::attrs() returned only attributes convertible to AttrValue and omitted the rest without any indication — every Python bool (an HDF5 enum), complex, compound and reference attributes. Unsigned 64-bit arrays were also cast to I64Array, turning values above i64::MAX negative.

Fix: booleans decode as 0/1 integers, AttrValue::U64Array keeps unsigned arrays unsigned, and AttrValue::Raw { datatype, shape, data } carries any other attribute verbatim. Both new variants are writable. Still lossy: a multi-dimensional numeric attribute is returned as a flat array (its shape is not reported).

B-tree v2 chunk index (layout v4, index type 5) is not supported

Status: open.

Summary: a chunked dataset with two or more unlimited dimensions written with libver='latest' indexes its chunks with a version-2 B-tree. Reading it fails with ChunkedReadError("unsupported chunked layout version=4, index_type=Some(5)"). Single-chunk, implicit, fixed-array and extensible-array indexes (and the v3 B-tree v1) are supported.

Repro: f.create_dataset("d", shape=(5, 7), chunks=(2, 3), maxshape=(None, None)) with h5py.File(..., libver='latest').

Status: open (by design for now); both are explicit errors.

Summary: a path through an external link returns FormatError::ExternalLinkUnsupported { filename, object_path }, and a dataset created with external=[...] storage returns FormatError::ExternalDataFilesUnsupported. Neither is resolved. If support is added, file names must be confined to the opened file's directory, as the virtual-dataset resolver now does.