fix(format): read huge, tiny and filtered fractal heap objects
A heap ID's type is in bits 4-5 of its first byte (H5HF_ID_TYPE_MASK 0x30); bits 6-7 are the ID version. The reader took the type from bits 6-7, so every huge object ID (0x10) was decoded as a managed one and failed — and since dense attributes are read all at once, one attribute over the heap's 4 KiB managed limit made every attribute on its object unreadable (netcdf4-python's issue671.nc / issue672.nc). - Huge objects (type 1): located directly from the ID when address and length fit in it, otherwise through the huge-object v2 B-tree (record types 1 and 2); filtered huge objects are decoded with the heap's pipeline and their filter mask. - Tiny objects (type 2): read from the ID itself. - Filtered heaps: the header's pipeline is parsed (it was skipped short, so the header checksum was read from the wrong place), indirect-block entries for direct blocks carry their filtered size and mask, and direct blocks are decoded before objects are read from them. - An unknown ID version is an error. FractalHeapHeader gains huge_btree_address, filter_pipeline, root_direct_block_filtered_size, root_direct_block_filter_mask, offset_size and length_size; read_managed_object now accepts any ID type. Regression tests (h5py-written, compared with h5py): dense_attribute_stored_as_a_huge_heap_object, dense_group_with_a_huge_link, dense_group_with_a_filtered_link_heap; unit tests tiny_object_is_read_from_the_id, huge_object_with_a_direct_id, unknown_heap_id_version_is_refused. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -310,6 +310,18 @@
|
||||
estimate instead of libhdf5's per-depth record capacities, and the
|
||||
listing failed. The same B-tree code indexes dense attributes, shared
|
||||
messages and chunks.
|
||||
- Fractal-heap "huge" objects (larger than the heap's managed-object
|
||||
limit, 4 KiB by default — e.g. an 8 KiB dense attribute or a link with a
|
||||
very long name) and "tiny" objects are now read; the ID type was taken
|
||||
from the wrong bits (6-7, the version, instead of 4-5), so a huge object
|
||||
failed and took every attribute on its object down with it (NetCDF-4
|
||||
files such as netcdf4-python's `issue671.nc`). Huge objects are found
|
||||
directly from the ID or through the huge-object v2 B-tree, filtered or
|
||||
not.
|
||||
- Heaps with an I/O filter pipeline (a group created with a filter on its
|
||||
creation property list compresses its link heap) are now read: the
|
||||
header's pipeline was skipped with the wrong size, so its checksum was
|
||||
looked for in the wrong place, and filtered direct blocks were read raw.
|
||||
- `clawhdf5-format` writer — **files libhdf5 rejects or reads wrong:**
|
||||
- Extensible Array (one unlimited dimension): chunks from index 244 on were
|
||||
written but never indexed and read as 0, by libhdf5 and by us.
|
||||
|
||||
Reference in New Issue
Block a user