A selection of a chunked dataset with a fill value is compared with the
full read in every index, through a map and through positioned reads. An
LZ4 chunk larger than 256 MiB is bounded by the chunk size, not refused.
The wasm package test reads the 4 GiB-chunk fixture and gets a clean
error in every index.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The writer gives a chunk of more than u32::MAX bytes layout message
version 5 and, filtered, index elements whose stored size takes the file's
size of lengths, as libhdf5 2.x does (the Fixed and Extensible Array
structures match libhdf5's byte for byte). Chunk dimensions of 2^32 or more
and filters that cannot take such a chunk (LZF, bitshuffle, bzip2, Blosc,
pcodec) are refused instead of truncated. Chunks are extracted row by row
and one at a time; deflate no longer cuts input at 4 GiB - 1 bytes, nor
holds the worst-case bound of a large chunk; an LZ4 chunk of 4 GiB or more
is read as the registered framing.
FileEditor refuses writing values into, or pruning/allocating, chunks of
4 GiB or more before anything is written; growing the extent and
attributes still work.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkInfo::chunk_size (and ChunkMapping::file_size) are u64: sizes past
u32 were truncated for Single Chunk, Implicit, Fixed and Extensible Array
indexes, and a v2 B-tree index refused them. A selection of a chunked
dataset with a non-default fill value is read over a box of fill values
instead of a full read, an unfiltered chunk of a file that is not in memory
is read row by row, and an intermediate deflate stage no longer reserves the
chunk's whole bound.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.
With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).
The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.
Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.
variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.
Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Datatype::Complex serializes class 11 version 5 byte-identically to
libhdf5 2.2.0; containers holding it are written as version 5.
- DatasetBuilder::with_complex_f32/f64_data (h5py's {r, i} compound,
default) and with_native_complex_f32/f64_data (class 11, opt-in);
make_(native_)complex_f32/f64_type for attributes.
- Dataset::read_complex_f64/f32 read either form.
- Python create_dataset accepts complex64/complex128 (compound form).
- Parsing unchanged: class 11 still surfaces as {r, i}.
- Tests vs h5py 3.16 / libhdf5 2.0.0 and h5dump 2.2.0; docs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Fixture written by libhdf5 2.2.0 (built from tag 2.2.0) through ctypes:
every bit pattern of FP4 E2M1, FP6 E2M3/E3M2, FP8 E4M3/E5M2 and a
bfloat16 LE/BE set, as datasets and attributes, with what H5Dread/H5Aread
return into double and float and the conversion exceptions libhdf5
raises. clawhdf5 already decoded every value as libhdf5 does, including
an all-ones exponent as inf/NaN in the OCP formats that have none
(documented as a deliberate match in known-issues).
- data_read: NaNs of non-native float layouts get libhdf5's bits (sign
kept, every mantissa bit set) in f64 and f32.
- h5rs dump/ls name these types as h5dump/h5ls 2.x do
(H5T_FLOAT_F4E2M1, "FP4 E2M1 4-bit float", float4-e2m1 ...), checked
against h5dump 2.2.0's output of the fixture.
- Python bindings read them as h5py 3.16 does (float32 for bfloat16,
float16 for the 1-byte formats, file byte order, same bytes as h5py);
writing them in 'r+' is refused.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Dimension::size of an unlimited dimension was its dimension scale's
extent, which netCDF-C leaves at 0, so it read 0 for a dimension holding
records. It is now what netCDF-C reports (nc4_find_dim_len): the largest
current extent of the variables using it in any group, found through the
scale's REFERENCE_LIST, a coordinate variable's own extent included; 0
when nothing has been written.
interop_tests::unlimited_dimension_lengths_match_netcdf4_python compares
with netCDF4-python (variables of different lengths, one in a subgroup,
an unwritten dimension, a coordinate variable shorter than another
variable on its dimension, a subgroup's own dimension); before the fix it
got time 0/6, rec 3/5, srec 0/1.
The known-issues entry moves to Fixed; the crate README's warning goes.
A new open entry records a related bug found meanwhile: variables'
dimensions are matched by size, not DIMENSION_LIST.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor's flock belongs to the open file description. When another
thread forks to spawn a process, the child shares the locked descriptor
until it execs, so a reopen right after the drop could be refused with
Error::Locked (a one-off failure of edit_interop::editor_locks_the_file in
a parallel test run). Drop now unlocks before closing, which releases the
lock for every descriptor sharing it.
Reproducer edit_tests::drop_releases_the_lock_while_other_threads_spawn_processes
(4 threads running `true`, 2000 open/drop rounds): 1483 of 2000 reopens
refused before, 0 in 30 runs after (tank). An OFD lock would not help: it
is inherited across fork the same way and does not conflict with
libhdf5's flock. The agent store's lock file unlocks on drop too.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.
- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
them (2026-09-19/20, machine not recorded, not re-run) instead of none;
the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
time datasets too; zlib-rs byte-identity scoped to what was measured;
macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
dated figures, fast-math is not BLAS.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- docs/README.md links the improvement logs and June plans where the
refresh archived them (docs/archive/).
- clawhdf5-py README: 'r+' creates and replaces attributes (compact or
dense); only deleting them is unsupported.
- scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format
and overwrote BENCHMARKS.md; nothing referenced it. Removed.
- Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu
says it is not HDF5 I/O.
- benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer
exists.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.
Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.
The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).
`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.
Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB
earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
-> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.
- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
this operation". Default: nothing (every backend that reads when
asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
(shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
that misses nothing ignores them (a hint never adds a round trip); a
pass that misses also asks for them, in file order, while the pass
stays within what is left of the operation's `maxFetch` budget (a
hint never makes a call fail). At most 1 MiB (or a block) per hint
and 65536 blocks per pass are recorded, whatever a file makes a
parser hint. `Operation::attempt` follows hints; the plain
`LazyStorage::attempt` does not. `run_blocking` and the browser's
driver use the former. `LazyStats::hinted_blocks` counts them.
No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.
Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkCache::chunks_for returned the chunk index's values in hash-map
order, which differs between File instances (a new HashMap per open).
Readers that stop at the first failing chunk therefore named different
chunks on different opens of the same damaged file: whole-dataset
selections here, and the indexed reader behind the storage harness's
intermittent cve-2025-2310.h5 failure. The cache now also keeps the
chunks in the order the index lists them (fetched or built under one
lock) and returns that order, as the uncached readers use.
Regression: several_damaged_chunks_report_the_same_chunk_every_time
(four chunks inflating to different short lengths; before the fix two
opens reported chunk [40, 0] and chunk [320, 0]).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The editor branch added File::from_std_file and the SWMR branch added the
swmr field to File; merged, the constructor did not set it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A boolean array of the axis's length, or of the dataset's whole shape, is
a mask h5py would apply (NotImplementedError here); one of any other shape
(np.array(True), a wrong length) is a key h5py itself refuses with
TypeError, which test_errors_match_h5py requires us to match.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- opts.headers is read as fetch reads it (a Headers, [name, value] pairs
or a plain object); it was spread as an object, which silently dropped
a Headers instance (a common way to pass Authorization). A caller's
Range is not sent.
- When one range request of a batch fails, the others in flight are
aborted (one AbortController per batch, its signal passed to fetch)
and no new ones start; the first failure is the error. The workers
used to go on issuing requests nobody waited for.
- parallel must be an integer from 1 to 1024 (openUrl) or a positive
integer (fetchRanges): a non-number gave NaN workers, so none ran and
fetchRanges returned nothing.
Tests (test.mjs): headers as object, Headers and pairs reach the fetch;
bad parallel values are option errors; fetchRanges with a 500 on the
third of 20 ranges at parallel 3 starts 3 requests and aborts the 2 in
flight. Before: the Headers case sent no header, 17 requests started
after the failure with none aborted, and parallel "x" returned.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Planning from a mapping of a clone of the locked descriptor shared its
flock: a process forked by another thread meanwhile (any Command) kept the
lock alive for a moment after the editor was dropped, and a test that
reopened the file at once saw Error::Locked (once in a full run). On Linux
plan through /proc/self/fd (a new open file description of the same
file, which still follows a rename); elsewhere keep the clone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.
Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A thread appending timestamps in a loop runs while a 1024x1024 gzip
dataset is rewritten through 'r+': the largest gap between its stamps
during the edit must be under half the edit's duration (an edit holding
the GIL stalls it for the whole edit; checked with a GIL-holding regex
standing in for the edit: one 0.20 s gap in a 0.21 s call). Edits already
ran detached; nothing tested it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor re-opened its path to plan each edit but wrote through the file
it held open, and the Python 'r+' handle re-opened the path after every
edit to read. When the path came to name another file between edits (a
rename or replacement, or a relative path after os.chdir), an edit was laid
out from the other file's metadata and written into the held one,
corrupting it, and later reads came from the other file (the review's
repro: h5py then reports "invalid dataset size, likely file corruption").
The editor now plans from a mapping of its own file (a clone of the held
descriptor, dropped before the edit writes) and canonicalises its path at
open. New FileEditor::reader() opens the held file anew for reading,
without sharing the editor's flock (a mapping of a cloned descriptor holds
the lock until unmapped): through /proc/self/fd on Linux, which follows a
renamed file; elsewhere by path, refused on Unix when the path no longer
names the held file. The Python handle reads through it and keeps no path;
a 'w' file is written at the absolute path it was opened with.
Tests: edit_tests.rs edits_go_to_the_file_held_not_the_path; test_edit.py
test_relative_path_and_chdir and test_path_replaced_between_edits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A Fixed or Extensible Array index whose maximum extent (the current one
when no maximum is recorded) is 0 along a dimension has a zero stride for
every dimension before it; ChunkGrid::offsets divided by it. The unfixed
editor made such files by resizing a clawhdf5-written dataset to a zero
extent: `h5rs check` panicked and the next resize raised an internal
error (12 of the reviewer's random-edit seeds 10..39). Such an index has
no slot for any chunk of the dataset; offsets now returns None.
Tests: chunk_grid::zero_extent_has_no_chunks; edit_interop's
zero_extent_resizes_without_a_recorded_maximum on a file the unfixed
editor left (fixture) and on a 2.7.0-written file taken through zero
extents with `h5rs check --data` and h5dump at every step; test_edit.py
random edits on seeds 10..39 of a clawhdf5-written file.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.
libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).
Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
remote.js read every 206 body (the probe and each range) with
resp.arrayBuffer() and checked its length afterwards, so a hostile
server could make the page buffer gigabytes before the check failed.
Every body is now piped through a TransformStream that errors as soon as
the count passes the limit, which cancels the body (and the request):
the requested range length for a 206, maxDownload for the 200 fallback.
A declared Content-Length past the limit is refused before reading. The
200 path now always streams (it read a declared length at once, because
a reader loop stalls on small bodies in headless Chromium under
--virtual-time-budget; a pipe does not).
Test (test.mjs): a probe and a range answered with a 64 MiB body read at
most the range + one 64 KiB piece; before, all 64 MiB were read ("asked
for the first 1048576 bytes of 2000000, got 67108864").
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
read_string, read_string_bytes, read_string_selection, read_vlen and
read_vlen_selection now retry as a whole, the global-heap decoding after the
read included; so do File::decode_strings / decode_string_bytes /
decode_vlen, Group::datasets / groups / attrs / attr, Dataset::attrs / attr,
the typed full reads (read_f64 ...), the header messages behind shape() and
dtype() (a shared message is read from another header), and
verify_provenance. Before, a transient failure on those reached the caller.
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as AttrValue::Raw) instead of failing, which hid
a transient error as a missing or raw attribute: on a live file such an
error of a retried kind now runs the read again too, and after the last
attempt the last result is returned as before. The format crate gains
find_attribute_reporting_in, which returns the errors find_attribute_in
skips (dense name-index lookups dropped them). The zero-copy reads need the
file in memory, which a live file never is, so they have nothing to retry.
Test: a storage that fails one read with a checksum mismatch; for 14 read
paths over a new fixture (tests/fixtures/swmr_strings_attrs.h5, an h5py
copy with the SWMR-write flag: vlen strings, vlen int32, dense attributes),
each read the path makes fails once in turn and the path must return the
same result with exactly one retry. It fails on the previous commit
(root.attrs, read 0). Dataset::attr on dense attributes returned None
before find_attribute_reporting_in.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.
- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
most 1 GiB): a read longer than it fails at once, before anything is
fetched, and an operation whose passes would fetch more than it fails
before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
code turns offsets into usize, so nothing past 4 GiB can be read there
(shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
maxDownload is bounded to 1 GiB.
Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The README example looped while swmr_writer_active(), which never ends when
the writer crashed or was killed: libhdf5 clears the SWMR-write flag only on
close (the mid-write fixture keeps it set for good). The loop now also stops
after a minute without growth, and the README, the swmr_writer_active docs
(with the same loop as a compiled no_run doctest) and the design say why.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
is_transient_format counted every format error but a handful as transient,
so on an open_swmr handle a permanent failure (a file that is not HDF5, an
unsupported version or message, a truncated file) was retried 100 times,
about 0.9 s of pauses per failing operation. Now only these are retried:
a checksum mismatch; a read past the file's current end (UnexpectedEof;
libhdf5 reads zeros there, which fail the checksum); and an object header
prefix whose signature or version does not decode, which libhdf5's
H5C__load_entry also retries (a header garbled whole fails there before its
checksum). Everything else is returned at once.
Tests: a unit test that every permanent kind returns after one call within
50 ms and every transient kind is retried to the limit; open_storage_swmr
of a non-HDF5 buffer returns SignatureNotFound within 100 ms (0.87 s before)
and a missing name on a live file fails without retries. The torn-read and
live h5py-writer tests still pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A file opened with open_swmr whose superblock does not have the SWMR-write
flag (its writer has closed it) is now read exactly as File::open reads it:
bounded by its recorded end of file, through the chunk cache, each
operation tried once, and is_swmr_read() is false. Before, every file opened
with open_swmr ignored its recorded end of file, so a closed file whose end
of file is below its length (h5clear_fsm_persist_less.h5, or the mid-write
fixture with its flags cleared) listed and read objects that File::open and
libhdf5's plain reader refuse. Opening such a file is not retried once the
superblock has been read.
libhdf5's SWMR reader is looser than either (it reads the fixture's chunk
index past the end of file even with the flags cleared); a test pins what
h5py does and docs/design/swmr.md says why we follow the plain reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():
- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
slices and integers with extra leading 1-axes allowed; the exact shape
for an index list, a scalar only where h5py expands it). Arrays are
converted as libhdf5 converts them in native byte order (integers
saturate, floats truncate toward zero and clip, integers go into h5py's
bool enum by value); other values through
numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
dataset is a ValueError instead of libhdf5's arbitrary value. The value
preparation is a small Python module compiled into the extension
(src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
attrs.modify: numeric, bool, complex, bytes and str data of any shape,
with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.
Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.
Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.
Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Range-read milestone M5 (docs/design/swmr.md): read a file a libhdf5
SWMR writer (h5py f.swmr_mode = True) is still appending to, as h5py's
File(path, 'r', swmr=True) does.
- File::open_swmr / open_storage_swmr: the file is read through the new
FileStorage (pread / seek_read, never mapped, len() the current length),
reads are bounded by the length at each read, and the chunk cache is not
used (a cached index hides new chunks; a cached edge chunk reads as fill
where the writer has since written). A file open for writing without
SWMR is refused with Error::Locked, as libhdf5 refuses it.
- Dataset::refresh re-reads the object header (H5Drefresh).
- Operations that fail with an error a racing write can cause (every
format error but a wrong path/selection, an unsupported feature or a bad
argument) are run again, up to 100 attempts (libhdf5's default metadata
read attempts for SWMR readers), 1 us doubling to 10 ms apart; nested
operations retry as a whole. swmr_retries() counts them,
swmr_writer_active() re-reads the superblock flags.
Tests: FileStorage growth and retry policy (unit); the mid-write copy
through open_swmr; a storage that garbles reads (retried, given up after
the attempts, never returned); and a live h5py writer appending to
Extensible-Array (plain and gzip) and v2-B-tree datasets for 2500 steps
while two Rust reader threads and h5py's SWMR reader check every value,
then the closed file read equal to h5py. Leaving the chunk cache on in
live mode fails the live test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
openUrl(url, opts) returns a RemoteFile with the methods of H5File
(kind, list, info, attrs, attrErrors, read, readHyperslab), each a
promise, and stats(). It runs every call through the restartable
LazyStorage: a pass that misses reports the byte ranges, js/remote.js
fetches them with fetch() and Range headers (six at a time), and the
pass is re-run. This keeps the main thread free without a Worker or
synchronous XHR (h5wasm's lazy files need both), as the design doc
recommends; the cost is re-running a pass per wave of misses.
Every answer is checked: a 206 with exactly the bytes asked for, and
the same ETag/Last-Modified and length as at open, else an error (never
data). A server that ignores Range (200) is downloaded whole, up to
maxDownload (512 MiB), unless fallback: "error". Options: blockSize,
cacheSize, headers, credentials, parallel, fetch.
test/serve.py is a range-capable static server with request counting
(and /norange/ for a server without range support). test.mjs repeats
every fixture check on files opened by URL (1 MiB and 512 B blocks),
checks the request budget on a 200 MB h5py file (list, three small
reads and a window of the big dataset: 5 requests, 6 MiB), the
download fallback, and HTTP errors, changed files and wrong answers;
with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.
- Every object of a File now shares one handle (src/handle.rs) that
runs all file access, metadata included, with the GIL released and
parses through File::storage() and the clawhdf5_format *_in functions.
Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
clawhdf5_remote::storage_for_url (read-only; another mode is a
ValueError). File.open_url(url, **options) takes the cache and HTTP
options (block_size, cache_size, headers, retries, timeout,
allow_full_download, max_full_download, require_validator,
max_redirects, max_parallel); File.remote_stats gives the block
cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
OSError, never KeyError/ValueError and never data; `key in group`
raises it instead of answering False.
Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
LazyStorage holds the blocks of a remote file fetched so far. An
operation runs as passes over it: a read that misses records the
missing blocks and fails, the pass's result is dropped whatever it is
(a parser may have caught the error and carried on), and the caller
fetches the reported ranges and re-runs the pass. No block is evicted
while an operation is in flight, so every pass that does not finish
asks for at least one new block and the operation ends. Blocks sit in
an LRU with a byte budget, trimmed between operations, bulk (raw data)
blocks first.
Reader::open_storage opens a file through any Storage, and variable-
length strings resolve through the file's storage instead of
File::as_bytes, which panics for a file not in memory.
tests/lazy.rs compares, file by file, what the viewer can show (kinds,
listings, attributes, info, whole reads and hyperslabs) read lazily
with the facade's range-storage path and the in-memory reader: files
built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with
CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A
listing plus a small read of a 48 MB file fetches 3 ranges.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A SWMR writer does not keep the superblock's end-of-file address up to
date: a copy h5py made of its own file mid-write records 715 in a
6 030-byte file. Superblock::data_end only ignored the recorded end when
it lay past the end of the file, so every open path bounded reads at
715: the file listed, but every chunked read failed ("unexpected EOF:
need 787 bytes, have 715") and h5rs check reported the chunk indexes
past the end of the file. libhdf5's SWMR reader skips the
end-of-allocation check for every read (H5FD_read); for a v3 superblock
with the SWMR-write flag the data now ends at the end of the file.
Test: tests/swmr_interop.rs over the mid-write copy (fixture), through
File::open, open_buffered and from_bytes, and against h5py's SWMR reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Since the M2 merge the facade handed in-memory files to the generic
`*_in` parsers as `&[u8]` (`with_bytes!`), which instantiates them in
the facade crate, where the format crate's private helpers do not
inline without LTO: listing a 400-group v1 file through `File::open`
was 7-10% slower than main. `ObjectHeader::parse_in`,
`group_v2::{resolve_child_in, resolve_group_children_in,
resolve_path_any_in}` and `attribute::{extract_attributes_tolerant_in,
find_attribute_in}` now pass a storage with `as_contiguous()` to their
non-generic slice entry point, compiled once in the format crate; other
storages reach the same generic core as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Reading continuation chunks from a queue (7e5e920, a69c5be) allocated a
queue Vec and a BTreeSet of chunk starts for every header, and inlined
the per-chunk message loop into the generic parser: ObjectHeader::parse
over 401 version-1 headers went from 24.8 to 45.7 us.
ChunkSpans now keeps the first 8 chunks in an inline array (cycle check
by scan) and is also the read queue; only a header of more chunks
allocates (a boxed spill list and start set). The message loop of one
version-1 chunk is its own non-generic function. Same checks as before:
any number of chunks up to 65,536, cycles refused, chunks bounded by the
file size, one chunk buffer alive at a time, libhdf5 message order,
overlap allowed. The cycle test now also covers spilled chunk lists.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The M2 fetch bound gave every codec n + n/4 + 4096 stored bytes, but
ZFP's fixed-rate mode stores up to 64 bits per value, doubling 4-byte
types: 60 of the 2205 zfp_interop datasets (rate 64, f4/i4) failed with
'compressed stream ends early' once M2 and ZFP were merged.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
check-32bit-casts.sh flagged two casts added by the previous commits; both
values are bounded (checked by gather_storage's first walk, and built from
a usize fetch length), so they go through addr::saturating_usize.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The facade equivalence harness turned every data-read error into "Err", so
it could not see File::open_storage failing differently from File::open
(a Storage or ContiguousStorageRequired error where the mmap path gives a
decode error, say).
- value() keeps the whole error. The only allowance is for a line on which
File::open itself varies between opens — the chunk cache lists a damaged
dataset's chunks in hash-map order, so which failing chunk a full read
reports varies (cve-2025-2310.h5, the one corpus file where this shows):
both sides must fail there, and a fresh File::open (up to 64) must
reproduce the storage's exact error. Open errors were already compared
in full; they still agree.
- The storage transcript may not contain ContiguousStorageRequired.
- More selections: a strided hyperslab (every third row) through
read_f64_selection, and out-of-order points through read_selection and
read_i64_selection.
- harness_compares_errors_not_just_failures checks the harness itself:
two different errors are different values, and a difference File::open
does not produce is reported.
With full errors the harness passes on the 61 fixtures and on the corpus
(701 files, 621 open).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>