The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.
The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB
decode limit, the 4 GiB file limit on wasm32, bodies cut off at their
length, listing passes, the cross-origin tests, and a pre-existing
nondeterministic error choice on cve-2025-2310.h5 that can fail the
native corpus comparison), the viewer README (options, how listing
costs, tests) and the M4 status in docs/design/range-reads.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.
Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor re-opened its path to plan each edit but wrote through the file
it held open, and the Python 'r+' handle re-opened the path after every
edit to read. When the path came to name another file between edits (a
rename or replacement, or a relative path after os.chdir), an edit was laid
out from the other file's metadata and written into the held one,
corrupting it, and later reads came from the other file (the review's
repro: h5py then reports "invalid dataset size, likely file corruption").
The editor now plans from a mapping of its own file (a clone of the held
descriptor, dropped before the edit writes) and canonicalises its path at
open. New FileEditor::reader() opens the held file anew for reading,
without sharing the editor's flock (a mapping of a cloned descriptor holds
the lock until unmapped): through /proc/self/fd on Linux, which follows a
renamed file; elsewhere by path, refused on Unix when the path no longer
names the held file. The Python handle reads through it and keeps no path;
a 'w' file is written at the absolute path it was opened with.
Tests: edit_tests.rs edits_go_to_the_file_held_not_the_path; test_edit.py
test_relative_path_and_chdir and test_path_replaced_between_edits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A Fixed or Extensible Array index whose maximum extent (the current one
when no maximum is recorded) is 0 along a dimension has a zero stride for
every dimension before it; ChunkGrid::offsets divided by it. The unfixed
editor made such files by resizing a clawhdf5-written dataset to a zero
extent: `h5rs check` panicked and the next resize raised an internal
error (12 of the reviewer's random-edit seeds 10..39). Such an index has
no slot for any chunk of the dataset; offsets now returns None.
Tests: chunk_grid::zero_extent_has_no_chunks; edit_interop's
zero_extent_resizes_without_a_recorded_maximum on a file the unfixed
editor left (fixture) and on a 2.7.0-written file taken through zero
extents with `h5rs check --data` and h5dump at every step; test_edit.py
random edits on seeds 10..39 of a clawhdf5-written file.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.
libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).
Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
read_string, read_string_bytes, read_string_selection, read_vlen and
read_vlen_selection now retry as a whole, the global-heap decoding after the
read included; so do File::decode_strings / decode_string_bytes /
decode_vlen, Group::datasets / groups / attrs / attr, Dataset::attrs / attr,
the typed full reads (read_f64 ...), the header messages behind shape() and
dtype() (a shared message is read from another header), and
verify_provenance. Before, a transient failure on those reached the caller.
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as AttrValue::Raw) instead of failing, which hid
a transient error as a missing or raw attribute: on a live file such an
error of a retried kind now runs the read again too, and after the last
attempt the last result is returned as before. The format crate gains
find_attribute_reporting_in, which returns the errors find_attribute_in
skips (dense name-index lookups dropped them). The zero-copy reads need the
file in memory, which a live file never is, so they have nothing to retry.
Test: a storage that fails one read with a checksum mismatch; for 14 read
paths over a new fixture (tests/fixtures/swmr_strings_attrs.h5, an h5py
copy with the SWMR-write flag: vlen strings, vlen int32, dense attributes),
each read the path makes fails once in turn and the path must return the
same result with exactly one retry. It fails on the previous commit
(root.attrs, read 0). Dataset::attr on dense attributes returned None
before find_attribute_reporting_in.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
is_transient_format counted every format error but a handful as transient,
so on an open_swmr handle a permanent failure (a file that is not HDF5, an
unsupported version or message, a truncated file) was retried 100 times,
about 0.9 s of pauses per failing operation. Now only these are retried:
a checksum mismatch; a read past the file's current end (UnexpectedEof;
libhdf5 reads zeros there, which fail the checksum); and an object header
prefix whose signature or version does not decode, which libhdf5's
H5C__load_entry also retries (a header garbled whole fails there before its
checksum). Everything else is returned at once.
Tests: a unit test that every permanent kind returns after one call within
50 ms and every transient kind is retried to the limit; open_storage_swmr
of a non-HDF5 buffer returns SignatureNotFound within 100 ms (0.87 s before)
and a missing name on a live file fails without retries. The torn-read and
live h5py-writer tests still pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A file opened with open_swmr whose superblock does not have the SWMR-write
flag (its writer has closed it) is now read exactly as File::open reads it:
bounded by its recorded end of file, through the chunk cache, each
operation tried once, and is_swmr_read() is false. Before, every file opened
with open_swmr ignored its recorded end of file, so a closed file whose end
of file is below its length (h5clear_fsm_persist_less.h5, or the mid-write
fixture with its flags cleared) listed and read objects that File::open and
libhdf5's plain reader refuse. Opening such a file is not retried once the
superblock has been read.
libhdf5's SWMR reader is looser than either (it reads the fixture's chunk
index past the end of file even with the flags cleared); a test pins what
h5py does and docs/design/swmr.md says why we follow the plain reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():
- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
slices and integers with extra leading 1-axes allowed; the exact shape
for an index list, a scalar only where h5py expands it). Arrays are
converted as libhdf5 converts them in native byte order (integers
saturate, floats truncate toward zero and clip, integers go into h5py's
bool enum by value); other values through
numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
dataset is a ValueError instead of libhdf5's arbitrary value. The value
preparation is a small Python module compiled into the extension
(src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
attrs.modify: numeric, bool, complex, bytes and str data of any shape,
with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.
Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.
Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.
Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased), known-issues (the browser can open URLs; the
limits of openUrl: round trips per wave of misses, a call holds what it
reads, CORS and validator visibility, the download fallback), the M4
status in docs/design/range-reads.md (why NeedBytes rather than a
Worker, why not clawhdf5-remote's BlockCache, what was tested; the
status paragraph at the top lost a garbled duplicate), the viewer's
README (API, options, how it works, tests; the size table is marked as
predating openUrl), CLAUDE.md and the README crate list.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.
- Every object of a File now shares one handle (src/handle.rs) that
runs all file access, metadata included, with the GIL released and
parses through File::storage() and the clawhdf5_format *_in functions.
Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
clawhdf5_remote::storage_for_url (read-only; another mode is a
ValueError). File.open_url(url, **options) takes the cache and HTTP
options (block_size, cache_size, headers, retries, timeout,
allow_full_download, max_full_download, require_validator,
max_redirects, max_parallel); File.remote_stats gives the block
cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
OSError, never KeyError/ValueError and never data; `key in group`
raises it instead of answering False.
Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Since the M2 merge the facade handed in-memory files to the generic
`*_in` parsers as `&[u8]` (`with_bytes!`), which instantiates them in
the facade crate, where the format crate's private helpers do not
inline without LTO: listing a 400-group v1 file through `File::open`
was 7-10% slower than main. `ObjectHeader::parse_in`,
`group_v2::{resolve_child_in, resolve_group_children_in,
resolve_path_any_in}` and `attribute::{extract_attributes_tolerant_in,
find_attribute_in}` now pass a storage with `as_contiguous()` to their
non-generic slice entry point, compiled once in the format crate; other
storages reach the same generic core as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Reading continuation chunks from a queue (7e5e920, a69c5be) allocated a
queue Vec and a BTreeSet of chunk starts for every header, and inlined
the per-chunk message loop into the generic parser: ObjectHeader::parse
over 401 version-1 headers went from 24.8 to 45.7 us.
ChunkSpans now keeps the first 8 chunks in an inline array (cycle check
by scan) and is also the read queue; only a header of more chunks
allocates (a boxed spill list and start set). The message loop of one
version-1 chunk is its own non-generic function. Same checks as before:
any number of chunks up to 65,536, cycles refused, chunks bounded by the
file size, one chunk buffer alive at a time, libhdf5 message order,
overlap allowed. The cycle test now also covers spilled chunk lists.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The facade equivalence harness turned every data-read error into "Err", so
it could not see File::open_storage failing differently from File::open
(a Storage or ContiguousStorageRequired error where the mmap path gives a
decode error, say).
- value() keeps the whole error. The only allowance is for a line on which
File::open itself varies between opens — the chunk cache lists a damaged
dataset's chunks in hash-map order, so which failing chunk a full read
reports varies (cve-2025-2310.h5, the one corpus file where this shows):
both sides must fail there, and a fresh File::open (up to 64) must
reproduce the storage's exact error. Open errors were already compared
in full; they still agree.
- The storage transcript may not contain ContiguousStorageRequired.
- More selections: a strided hyperslab (every third row) through
read_f64_selection, and out-of-order points through read_selection and
read_i64_selection.
- harness_compares_errors_not_just_failures checks the harness itself:
two different errors are different values, and a difference File::open
does not produce is reported.
With full errors the harness passes on the 61 fixtures and on the corpus
(701 files, 621 open).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An attribute needing a heap block larger than the next one was refused
("skipping blocks too small for an object", "a first object too large for
the starting block"); once an object's move to dense storage was refused
it refused every new attribute, so 24% of set_attr calls in the review's
random workload failed.
Following H5HF__hdr_update_iter, H5HF__man_iblock_root_create/_double and
H5HF__hdr_skip_blocks, the smaller blocks are now skipped: the iterator
moves past them and they become an indirect free section with a first
row section (serialized, class 1, as H5HF__sect_indirect_serialize writes
it) and ghost normal rows, added as returned space so it merges with a
range skipped just before it (H5HF__sect_indirect_merge_row). Later
objects that best-fit a row section get a block created there
(H5HF__man_iblock_alloc_row / H5HF__sect_indirect_reduce_row: from the
start or end of the range, or from its middle, which splits it, with
libhdf5's span bookkeeping). Heaps with such sections, as libhdf5 writes
them, are now read too (they were refused at open).
dense_skipped_blocks_match_libhdf5 drives every path (merge, split, end,
last entry, row wrap) on earliest/v110/latest files against libhdf5
doing the same edits one session each; heaps, free sections and index
B-trees are equal after every phase. The refusal test now checks the
skip against libhdf5 and keeps a real refusal (last object in a block);
clawhdf5-written heaps get 1-4 KiB attributes too. The three tests fail
on the previous fheap.rs. Random workload refusals: 24% -> 2.2%, all the
documented last-object-in-a-block case.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The version-1 chunk walk nested continuation chunks depth-first and kept
every enclosing chunk's buffer alive, up to 65 536 chunks. With storage
that hands out owned buffers (CountingStorage, the Storage trait, remote
storage) a crafted chain of chunks nested in each other read and held the
square of the file's size (a 192 KB file read 768 MB).
Chunks are now read from a FIFO queue of (address, length) pairs in the
order their continuation messages are found, as H5O_protect does and as
the editor's header walker already did, each buffer released before the
next read. In both header versions a chunk starting at an address seen
before (cycle) is refused, and so are chunks adding up to more than the
file, which bounds a header's reads by the file's size. Overlap itself is
allowed: libhdf5 reads cve-2025-7067.h5, whose continuation chunk overlaps
chunk 0 (refusing overlap cost that conformance file).
Tests: the nested chain is refused having read at most the file (it read
n^2 bytes before); a 3000-chunk chain reads each chunk once; a chunk's
messages follow the whole previous chunk (they were inserted at the
continuation message); an overlapping continuation chunk is read.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
prune_plan stored one Vec<u64> for every chunk coordinate of the region a
shrink cuts off, existing or not, so a sparse dataset exhausted memory
(about 62 bytes per coordinate; (4, 2e7) with chunks (1, 1) took 2.5 GB,
larger extents never finished). It now places each existing chunk in
H5D__chunk_prune_by_extent's walk (its pass, then its coordinates) and
sorts, which gives the same chunks, order and actions in memory and time
proportional to the chunks that exist.
A unit test checks the plan against the full walk (kept as the test's
reference) for 3000 random extents and chunk subsets. The interop test
shrinks a (4, 10^12) dataset with chunks (1, 1) and 9 chunks (v1 and v2
B-tree): 0.56 s and 43 MB peak; the old code aborted on allocation under an
8 GB limit.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Our checksum reduced its sums with `% 65535`; libhdf5's
H5_checksum_fletcher32 folds them with `(s & 0xffff) + (s >> 16)`, which
leaves 0xffff where the modulo leaves 0. On about one chunk in 32768
libhdf5 refused the chunks we wrote and we refused the chunks it wrote.
Every release since v2.1.0 is affected.
clawhdf5_format::checksum::fletcher32 is a port of H5_checksum_fletcher32
and the filter's only implementation. Verification also accepts the
byte-swapped form libhdf5 accepts (1.6.2 and earlier) and the `% 65535`
form earlier releases wrote, so their files stay readable.
The new interop test compares the checksum with libhdf5's own function
(ctypes) on every 1- and 2-byte input and 40 000 random and fold-heavy
inputs, and moves fold-case chunks between h5py and FileBuilder/FileEditor
in both directions; with the old filters.rs the three file tests fail.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ExtentBytes and read_exact_at/read_upto rejected short results but passed
longer-than-asked ones through, and FileData forwarded them too, so a
Storage that broke read_at's contract by returning extra bytes had them
decoded or returned as data (a contiguous dataset read gained 37 junk
bytes). gather_storage alone trimmed.
- storage::exact_len (new, pub): a read of len bytes as exactly len — cut
when longer, an error when short. read_exact_at, read_upto and
ExtentBytes (so chunk fetches and selection gathers) go through it.
- FileData cuts a backend's answer to what it asked for before laying the
cache image over it.
- Tests: over a storage that appends 37 junk bytes to every read, every
format-crate fixture reads exactly as from the slice
(overlong_reads_are_cut_to_the_range_asked_for), and every facade
fixture opens and reads through File::open_storage as through File::open
(overlong_storage_reads_identically). Both failed before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG, the clawhdf5-remote and h5rs READMEs and the remote-files
known issues: redirect rules, scaled timeouts (min_speed), URL redaction,
claimed lengths never allocated (download, --max-download), a 200 for a
small file accepted, and ObjectStoreStorage from any thread.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It refused whenever Handle::try_current() was Ok, which is also the case
inside spawn_blocking threads — so the workaround its own error message
recommended failed the same way, and the backend could only be used from
a bare std::thread in a tokio application.
Reads are now spawned on the storage's own runtime and the caller waits on
a channel: the future never runs on the caller's thread, so neither a
spawn_blocking thread nor a current-thread runtime can deadlock or panic
(a read inside a runtime blocks that thread, like any blocking call; the
docs still recommend spawn_blocking there).
Tests: a read in spawn_blocking of a multi-thread runtime and a read inside
a current-thread runtime's task give File::open's values (both errors
before).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gather_storage merged only runs that touch, so a strided selection of a
contiguous dataset over a Storage became one range (and one owned Vec) per
element: a stride-2 read of 32M f32 through File::open_storage made
16,777,232 read_at calls, took 2.0 s and peaked at 2.09 GB.
The selection is now walked twice. The first walk checks the runs and plans
spans: runs in increasing order at most 4 KiB apart (GATHER_GAP_BYTES) are
read as one span up to 8 MiB (GATHER_SPAN_BYTES; a longer run is split), so
nothing is stored per run. The spans are fetched in RAW_BATCH_BYTES batches
while the second walk copies each run out of its span. Same checks and
errors as before.
The same read is now 32 reads and 0.31 s (File::open: 0.08 s).
contiguous_read_interop: every h5py-checked selection is also read through
File::open_storage and must give libhdf5's bytes; a new test bounds the
range reads of strided, blocked, column and point selections (stride 2: at
most 1 data read; 563,200 before).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the full read split its chunk fetches into 64 MiB batches. The
selection path, the indexed read and the parallel_read decoders fetched
every chunk's stored bytes in one read_ranges call, each extent bounded only
by the file length, so a crafted chunk index pointing many chunks at one
large extent made File::open_storage hold chunks x extent bytes (3.3 GB from
a 16.8 MB file) before the first decode error.
- storage::for_each_extent_batch is now the one way raw-data reads fetch
chunk bytes: batches of at most RAW_BATCH_BYTES (now pub), each decoded
before the next is fetched. Used by the full, cached, indexed, selection
and parallel_read paths; the sweep read uses read_extent per chunk.
- ExtentReq carries each chunk's claimed extent (bounds-checked as before,
same errors) and the prefix actually fetched:
filters::stored_chunk_limit — the chunk size if unfiltered, else each
applied filter's worst-case growth (n + n/4 + 4096 per codec; unbounded
only for an application-registered codec). The in-memory path cuts the
slice it decodes the same way, so both paths still agree.
- tests/raw_fetch_bounds.rs: a crafted chunked_large.h5 (ten chunks all
claiming 20 MiB at one padding blob) read through every path over a
storage that records the largest single fetch; and 160 MiB of legitimate
unfiltered chunks fetched batch by batch. Before: one 80 MiB fetch
(selection) and one 160 MiB fetch; after: within the budget.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
README: "Reading remote files" (open_url, the range_server and read_url
examples with their real output against the fixtures, h5rs on URLs), the
crate in the crate map and the unreleased highlights. CHANGELOG: the
clawhdf5-remote crate, h5rs URLs, File::storage and
VlResolver::element_in, with the request counts over the conformance
corpus (tank, 2026-09-26, the command given). known-issues: the M2
range-read entry updated (the cache now exists; h5rs reads through
storage) and a new entry for the remote backends' limits (no Python or
browser URLs yet, fixed block size, cloud stores not run against a real
bucket, validators, credentials). Design doc: M3 status with the choices
that differ from the plan (a crate rather than a clawhdf5-io feature,
ureq for HTTP so the default build has no C) and the corpus counts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased): the new FileEditor operations, space reuse, and
the two reader fixes (implicit index grid, object-header continuation
chains). known-issues: the editor's remaining refusals (skipped heap
blocks, heaps with filters or child indirect blocks, freeing a heap
block, implicit-index insertions, ...) and the append-waste sizes before
and after reuse (measure_append_waste, tank 2026-09-26; file sizes are
deterministic). range-reads design: status note on the reader changes.
README and CLAUDE.md: what the editor covers and how to test it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased): File::open_storage, raw data and v2 B-trees over
Storage, the tests and their corpus results (2026-09-26, tank; conformance
600 of 697, results.json identical to 8f59b2e). known-issues: what
open_storage does not do yet (no remote backend or block cache, read_at
counts of a one-pass read, v1 group lookups, whole-file VDS sources,
zero-copy methods, SWMR growth, hash-order error choice on damaged chunked
datasets). Design: M2 status and the choices made.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
+6-7% (about 4 ns per header) is real; symbol-table nodes -17%, group
B-tree walk -16%, facade listing -2.4%: local metadata reads are net
slightly faster than main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileBuilder stored every chunk of an LZF or Blosc dataset through the
filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller
than the chunk as a failure of the optional filter and stores the chunk raw
with the filter's mask bit set. For an LZF chunk whose stream was exactly
the chunk's size, the first libhdf5 rewrite stored raw data at the same
size and kept our stale mask 0 in the index, and h5py could no longer read
the dataset.
precompress_chunks now runs chunks through compress_chunk_masked (as
FileEditor does since f7e2ab1), sequentially and on the parallel path, and
build_chunked_data_from_precompressed records each chunk's real mask in
every index the writer builds: single chunk (layout field), Fixed Array and
Extensible Array filtered elements, and version-2 B-tree type 11 records
(create_datasets_parallel goes through the same path). The writer builds
no version-1 B-tree or implicit index. PrecompressedChunks::chunks gains
the mask. Files whose chunks all compress are byte-identical.
Latent only in the unreleased LZF/Blosc writer (added 2026-09-26); no
tagged release writes either filter.
Tests:
- plugin_filters_interop skipped_optional_filters_are_masked_as_libhdf5_masks_them:
LZF, shuffle+LZF+fletcher32 and Blosc over random, compressible and
alternating chunks in every index; masks equal an h5py-written twin's;
h5py r+ rewrites and extends them; h5py, h5dump and our reader read
every value. Before: 20 of 24 datasets had masks other than h5py's, and
with that check disabled h5py failed to read the rewritten datasets
("filter returned failure during read").
- plugin_filters_interop files_whose_chunks_all_compress_are_unchanged:
pins the pre-fix bytes of five all-compressing files.
- chunked_write skipped_lzf_chunks_are_masked_in_every_index (fails before:
mask 0, want 2).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Per-file probe output is identical for 696 of 697 files, not all:
cve-2025-2310.h5's error string depends on which parallel chunk decode
fails first, at f2ff2c4 as on this branch.
- The parser cores are generic (S: Storage + ?Sized); provisional A/B
numbers against f2ff2c4, including the one bench that still shows
ObjectHeader::parse slower when old and new are separate binaries.
- Reads sized by untrusted fields are bounded; the harness only accepts
the known whole-file fallbacks.
- range-reads.md records why M1 went generic rather than &dyn.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A dataset whose filter this build cannot encode (scale-offset, N-Bit, SZIP;
a plugin filter the build lacks) failed with Error::Format("unsupported
filter: 6"), although the editor documents every refused edit as
Error::Unsupported, and the Python bindings raised ValueError rather than
NotImplementedError. Every edit now maps FormatError::UnsupportedFilter to
Error::Unsupported; the file is left untouched as before.
Test: edit_interop unencodable_filters_are_unsupported — h5py scale-offset
datasets (integer with chunks, integer never written, float D-scale):
Error::Unsupported naming the filter, and the file byte for byte unchanged.
Fails on the previous editor (Format(UnsupportedFilter(6))).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
libhdf5 counts a version-2 object header's attributes through its Attribute
Info message (0x15) and reports none when the header has none. set_attr gave
v110/latest groups, the root group and datasets without attributes an
attribute message only, so h5py listed the attribute but len(obj.attrs) and
H5Oget_info's num_attrs said 0, and stayed wrong after h5py r+ added more.
Like H5O__attr_create, the edit now adds the message when a version-2 header
lacks it, in the same planned edit: version 0, the header's creation-order
track/index flags, maximum creation index 0, undefined fractal heap and
B-tree addresses, message flag DONTSHARE — byte for byte what libhdf5
writes. It goes before the attribute (libhdf5's order) when free space
holds both, else after it, so a continuation chunk made for the attribute
also takes it.
Test: edit_interop attribute_count_in_version_2_headers — v110 and latest
files, attributes set on the root group, groups and datasets with and
without existing attributes: h5py's len/num_attrs/list/values, h5dump -A
and our reader agree, also after h5py r+ adds attributes up to and past the
compact limit. Fails on the previous editor (h5py len 0).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A valid group has one link per name, but a damaged or hand-made one can
have two. resolve_child followed the first soft link of the name, the
listing skipped a dangling one and listed the name via a later link, and
path resolution followed the last symbolic link: three answers. All now
take the first link of the name (header message order in a compact group,
name index order in a dense one) and ignore the rest, even if the first
dangles. That is libhdf5's rule for compact groups (H5G__compact_lookup
stops at the first Link message); h5py opens nothing for a dangling first
link although a later one resolves. For a dense group libhdf5
binary-searches the index and may land on another of several exact
duplicates; documented on first_link_named. find_symbolic_link's v2 branch
was dead (only v1 groups reach it) and is now v1-only.
Test: an h5py compact group with soft links dup_A (dangling, or to /d) and
dup_B (the other), dup_B renamed to dup_A in the header and re-checksummed.
Lookup, path and listing through all three readers match h5py for both
orders. With the old group_v2.rs the path lookup returned 42 where h5py
opens nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The editor stored every chunk through the whole pipeline with filter mask 0.
For LZF that did not shrink a chunk, h5py instead stores it raw with the
filter's mask bit set. A chunk the editor stored LZF-encoded at exactly the
raw size was then rewritten raw by libhdf5 at the same size; libhdf5 does not
touch the index entry when the size is unchanged, so the stale mask 0 stayed
and h5py (and h5dump) could no longer read the dataset.
clawhdf5_format::filters::compress_chunk_masked runs the pipeline as
H5Z_pipeline does: an optional filter (H5Z_FLAG_OPTIONAL) that fails is
skipped and its bit set, a mandatory one fails the write, and LZF/Blosc
output no smaller than the input counts as failure, as in the reference
filters (their output buffer is the input's size). Deflate, LZ4, Zstd,
bitshuffle and bzip2 never fail on size in libhdf5 and are kept as before.
Test: edit_interop optional_filters_that_fail_are_skipped — the reviewer's
repro at every libver: the editor stores the chunk exactly as h5py does
(mask 1, size 5; shuffle+LZF+fletcher32 mask 2), h5py r+ rewrites and
extends the datasets, and h5py, h5dump and our reader read every value.
Fails on the previous editor (mask 0; h5dump cannot read /u8).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Three `chunk_info.address as usize` casts behind the `parallel` feature
survived the conversion, because check-32bit-casts.sh linted only default
features plus plugin-filters. On a 32-bit target with rayon a chunk address
past 4 GiB still wrapped onto another part of the file. They go through
addr::to_usize now, and the lane index (h % n, always < n) through
saturating_usize.
The script now lints no default features, default features, and every
optional feature but szip (wasm32; the set with zstd, which does not build
for wasm32, on the host, where the lint reports the same casts). With the
old parallel_read.rs/lane_partition.rs it fails listing the four casts; the
old script passed them. CHANGELOG and the design note give the exact count
(119) and what is not covered.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
known-issues records what the editor refuses, that freed space is never
reused (append-workload file sizes measured 2026-09-26 on tank with the
ignored measure_append_waste test; sizes are deterministic), and that there
is no journal.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Lists the parsers now reading through Storage, what still needs the
whole file (v2 B-tree-indexed structures: a clean error; raw data: M2),
the equivalence harness, and the evidence that nothing changed: existing
tests, a byte-identical conformance results.json and per-file probe
output against f2ff2c4, and identical slice-API transcripts over 748
files between the two builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>