Standing rules (no C by default, h5py must read what we write, one
float16 implementation, claims need evidence, OpenClaw/ZeroClaw
withdrawn, the known-issues rule) are gathered in one place; library and
agent-memory invariants are split; the CI section lists what
ci-test.sh and the conformance workflow run now instead of a dated
"two jobs, green" line. Adds the conformance, remote, wasm, Python and
benchmark workflows (idle load below 2, dated records), and warns that
scripts/run-benchmarks.sh is stale and overwrites BENCHMARKS.md. Crate
roles corrected (clawhdf5-io is I/O adapters; codecs live in
clawhdf5-format). Benchmark figures now live in BENCHMARKS.md only.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The committed CONFORMANCE.md counts bad_nbit_parms_walk.h5 as an
our-error because its six h5py reads agreed in that run; a rerun the
same day confirmed the over-read and counted it ref-bug. The ok count
(602 of 697) is the same either way.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
range-reads.md opens with a table of milestones M0-M5 and the PRs that
merged them (#17-#21), replacing a header left garbled by earlier
merges, and each milestone's status names its PR. swmr.md says the
reader is merged (PR #19) and the writer does not exist. openclaw.md
links the Node package's known-issues entry.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A "Current headline numbers" table gives each figure's newest dated
measurement with its machine, command and section. Sections a later run
replaced are marked superseded with a link to the newer one; stale
"still open" notes and cross-references now point at the fixes. No
measured value changed.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An "Open issues" table at the top links each open entry; fixed entries
move to "Fixed (history)", newest first, keeping the date, PR, affected
releases and what users must do. Open entries re-checked against main
(9b5803f): the remaining audit gaps are gathered into one entry, the
range-read and remote limits no longer contradict themselves (Python and
the browser open URLs; SWMR reading is File::open_swmr), and the
nondeterministic damaged-chunk error, fixed in PR #19, has its own
history entry.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Regenerated on tank: ok 600 -> 602 (attr_datatypes.hdf5 and
tcomplex_be.h5, compared against h5py's big-endian VL values corrected),
mismatch 2 -> 0; cve-2025-2308.h5 and cve-2025-44904.h5 are ref-bug
(h5py's values varied across six heaps in this run);
bad_nbit_parms_walk.h5 read the same in all six this run, so it stays
our-error, as the classification rule requires. Baseline raised.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).
`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.
Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB
earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
-> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.
- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
this operation". Default: nothing (every backend that reads when
asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
(shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
that misses nothing ignores them (a hint never adds a round trip); a
pass that misses also asks for them, in file order, while the pass
stays within what is left of the operation's `maxFetch` budget (a
hint never makes a call fail). At most 1 MiB (or a block) per hint
and 65536 blocks per pass are recorded, whatever a file makes a
parser hint. `Operation::attempt` follows hints; the plain
`LazyStorage::attempt` does not. `run_blocking` and the browser's
driver use the former. `LazyStats::hinted_blocks` counts them.
No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A ref-bug file's refused objects were still listed as our-error root
causes. The summary line now names every non-ok class and its count, and
an empty root-cause table says None.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.
Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.
- ref.py checks that the installed h5py returns big-endian VL elements
with the file's bytes under a little-endian dtype (writing and reading
a vlen('>f4') in memory) and, if so, relabels them with the file's byte
order before hashing, marking the object `ref_fix`. The values are now
compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
/VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
in six processes with different heaps (import order, MALLOC_PERTURB_).
Values the file determines are the same every time; these three change
(6, 6 and 3 distinct results), so they are over-read memory, not data
clawhdf5 could match. compare.py classifies a file `ref-bug` only when
every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
objects; test_ref.py covers both (run in the nightly job).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep
(both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX
5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load
at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated.
The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim
reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the
0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's
0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or
by one or two questions, most likely from the index tie-break of 3ed0489.
RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one
question between identical runs, so the float16 "flips" note is corrected.
ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.
The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkCache::chunks_for returned the chunk index's values in hash-map
order, which differs between File instances (a new HashMap per open).
Readers that stop at the first failing chunk therefore named different
chunks on different opens of the same damaged file: whole-dataset
selections here, and the indexed reader behind the storage harness's
intermittent cve-2025-2310.h5 failure. The cache now also keeps the
chunks in the order the index lists them (fetched or built under one
lock) and returns that order, as the uncached readers use.
Regression: several_damaged_chunks_report_the_same_chunk_every_time
(four chunks inflating to different short lengths; before the fix two
opens reported chunk [40, 0] and chunk [320, 0]).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The editor branch added File::from_std_file and the SWMR branch added the
swmr field to File; merged, the constructor did not set it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A boolean array of the axis's length, or of the dataset's whole shape, is
a mask h5py would apply (NotImplementedError here); one of any other shape
(np.array(True), a wrong length) is a key h5py itself refuses with
TypeError, which test_errors_match_h5py requires us to match.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB
decode limit, the 4 GiB file limit on wasm32, bodies cut off at their
length, listing passes, the cross-origin tests, and a pre-existing
nondeterministic error choice on cve-2025-2310.h5 that can fail the
native corpus comparison), the viewer README (options, how listing
costs, tests) and the M4 status in docs/design/range-reads.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- opts.headers is read as fetch reads it (a Headers, [name, value] pairs
or a plain object); it was spread as an object, which silently dropped
a Headers instance (a common way to pass Authorization). A caller's
Range is not sent.
- When one range request of a batch fails, the others in flight are
aborted (one AbortController per batch, its signal passed to fetch)
and no new ones start; the first failure is the error. The workers
used to go on issuing requests nobody waited for.
- parallel must be an integer from 1 to 1024 (openUrl) or a positive
integer (fetchRanges): a non-number gave NaN workers, so none ran and
fetchRanges returned nothing.
Tests (test.mjs): headers as object, Headers and pairs reach the fetch;
bad parallel values are option errors; fetchRanges with a 500 on the
third of 20 ranges at parallel 3 starts 3 requests and aborts the 2 in
flight. Before: the Headers case sent no header, 17 requests started
after the failure with none aborted, and parallel "x" returned.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
serve.py always exposed Content-Range and ETag, and Node has no CORS, so
openUrl's documented cross-origin path (length from a HEAD request,
answers checked by body length alone, no validator) was never run.
- serve.py /noexpose/ serves ranges without Content-Range, ETag,
Last-Modified or Accept-Ranges (what a page sees of a server that
does not expose them); /unexposed/ sends them but exposes none, for a
real browser. HEAD requests are counted (0 bytes).
- test.mjs: every fixture check through /noexpose/ at 1 MiB and 512-byte
blocks (one HEAD each, requests and bytes as the server counted them),
concurrent reads with cacheSize 0, a short answer still caught, and a
server without a HEAD length a clear error.
- browser.sh: the page on 127.0.0.1 opens the file from localhost, once
with Content-Range exposed and once through /unexposed/, where the
server's log must show the HEAD.
Checked by breaking the HEAD length in remote.js: the new checks fail.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Planning from a mapping of a clone of the locked descriptor shared its
flock: a process forked by another thread meanwhile (any Command) kept the
lock alive for a moment after the editor was dropped, and a test that
reopened the file at once saw Error::Locked (once in a full run). On Linux
plan through /proc/self/fd (a new open file description of the same
file, which still follows a rename); elsewhere keep the clone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.
Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A thread appending timestamps in a loop runs while a 1024x1024 gzip
dataset is rewritten through 'r+': the largest gap between its stamps
during the edit must be under half the edit's duration (an edit holding
the GIL stalls it for the whole edit; checked with a GIL-holding regex
standing in for the edit: one 0.20 s gap in a 0.21 s call). Edits already
ran detached; nothing tested it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor re-opened its path to plan each edit but wrote through the file
it held open, and the Python 'r+' handle re-opened the path after every
edit to read. When the path came to name another file between edits (a
rename or replacement, or a relative path after os.chdir), an edit was laid
out from the other file's metadata and written into the held one,
corrupting it, and later reads came from the other file (the review's
repro: h5py then reports "invalid dataset size, likely file corruption").
The editor now plans from a mapping of its own file (a clone of the held
descriptor, dropped before the edit writes) and canonicalises its path at
open. New FileEditor::reader() opens the held file anew for reading,
without sharing the editor's flock (a mapping of a cloned descriptor holds
the lock until unmapped): through /proc/self/fd on Linux, which follows a
renamed file; elsewhere by path, refused on Unix when the path no longer
names the held file. The Python handle reads through it and keeps no path;
a 'w' file is written at the absolute path it was opened with.
Tests: edit_tests.rs edits_go_to_the_file_held_not_the_path; test_edit.py
test_relative_path_and_chdir and test_path_replaced_between_edits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A Fixed or Extensible Array index whose maximum extent (the current one
when no maximum is recorded) is 0 along a dimension has a zero stride for
every dimension before it; ChunkGrid::offsets divided by it. The unfixed
editor made such files by resizing a clawhdf5-written dataset to a zero
extent: `h5rs check` panicked and the next resize raised an internal
error (12 of the reviewer's random-edit seeds 10..39). Such an index has
no slot for any chunk of the dataset; offsets now returns None.
Tests: chunk_grid::zero_extent_has_no_chunks; edit_interop's
zero_extent_resizes_without_a_recorded_maximum on a file the unfixed
editor left (fixture) and on a 2.7.0-written file taken through zero
extents with `h5rs check --data` and h5dump at every step; test_edit.py
random edits on seeds 10..39 of a clawhdf5-written file.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.
libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).
Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
remote.js read every 206 body (the probe and each range) with
resp.arrayBuffer() and checked its length afterwards, so a hostile
server could make the page buffer gigabytes before the check failed.
Every body is now piped through a TransformStream that errors as soon as
the count passes the limit, which cancels the body (and the request):
the requested range length for a 206, maxDownload for the 200 fallback.
A declared Content-Length past the limit is refused before reading. The
200 path now always streams (it read a declared length at once, because
a reader loop stalls on small bodies in headless Chromium under
--virtual-time-budget; a pipe does not).
Test (test.mjs): a probe and a range answered with a 64 MiB body read at
most the range + one 64 KiB piece; before, all 64 MiB were read ("asked
for the first 1048576 bytes of 2000000, got 67108864").
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
read_string, read_string_bytes, read_string_selection, read_vlen and
read_vlen_selection now retry as a whole, the global-heap decoding after the
read included; so do File::decode_strings / decode_string_bytes /
decode_vlen, Group::datasets / groups / attrs / attr, Dataset::attrs / attr,
the typed full reads (read_f64 ...), the header messages behind shape() and
dtype() (a shared message is read from another header), and
verify_provenance. Before, a transient failure on those reached the caller.
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as AttrValue::Raw) instead of failing, which hid
a transient error as a missing or raw attribute: on a live file such an
error of a retried kind now runs the read again too, and after the last
attempt the last result is returned as before. The format crate gains
find_attribute_reporting_in, which returns the errors find_attribute_in
skips (dense name-index lookups dropped them). The zero-copy reads need the
file in memory, which a live file never is, so they have nothing to retry.
Test: a storage that fails one read with a checksum mismatch; for 14 read
paths over a new fixture (tests/fixtures/swmr_strings_attrs.h5, an h5py
copy with the SWMR-write flag: vlen strings, vlen int32, dense attributes),
each read the path makes fails once in turn and the path must return the
same result with exactly one retry. It fails on the previous commit
(root.attrs, read 0). Dataset::attr on dense attributes returned None
before find_attribute_reporting_in.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.
- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
most 1 GiB): a read longer than it fails at once, before anything is
fetched, and an operation whose passes would fetch more than it fails
before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
code turns offsets into usize, so nothing past 4 GiB can be read there
(shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
maxDownload is bounded to 1 GiB.
Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The README example looped while swmr_writer_active(), which never ends when
the writer crashed or was killed: libhdf5 clears the SWMR-write flag only on
close (the mid-write fixture keeps it set for good). The loop now also stops
after a minute without growth, and the README, the swmr_writer_active docs
(with the same loop as a compiled no_run doctest) and the design say why.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
is_transient_format counted every format error but a handful as transient,
so on an open_swmr handle a permanent failure (a file that is not HDF5, an
unsupported version or message, a truncated file) was retried 100 times,
about 0.9 s of pauses per failing operation. Now only these are retried:
a checksum mismatch; a read past the file's current end (UnexpectedEof;
libhdf5 reads zeros there, which fail the checksum); and an object header
prefix whose signature or version does not decode, which libhdf5's
H5C__load_entry also retries (a header garbled whole fails there before its
checksum). Everything else is returned at once.
Tests: a unit test that every permanent kind returns after one call within
50 ms and every transient kind is retried to the limit; open_storage_swmr
of a non-HDF5 buffer returns SignatureNotFound within 100 ms (0.87 s before)
and a missing name on a live file fails without retries. The torn-read and
live h5py-writer tests still pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A file opened with open_swmr whose superblock does not have the SWMR-write
flag (its writer has closed it) is now read exactly as File::open reads it:
bounded by its recorded end of file, through the chunk cache, each
operation tried once, and is_swmr_read() is false. Before, every file opened
with open_swmr ignored its recorded end of file, so a closed file whose end
of file is below its length (h5clear_fsm_persist_less.h5, or the mid-write
fixture with its flags cleared) listed and read objects that File::open and
libhdf5's plain reader refuse. Opening such a file is not retried once the
superblock has been read.
libhdf5's SWMR reader is looser than either (it reads the fixture's chunk
index past the end of file even with the flags cleared); a test pins what
h5py does and docs/design/swmr.md says why we follow the plain reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():
- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
slices and integers with extra leading 1-axes allowed; the exact shape
for an index list, a scalar only where h5py expands it). Arrays are
converted as libhdf5 converts them in native byte order (integers
saturate, floats truncate toward zero and clip, integers go into h5py's
bool enum by value); other values through
numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
dataset is a ValueError instead of libhdf5's arbitrary value. The value
preparation is a small Python module compiled into the extension
(src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
attrs.modify: numeric, bool, complex, bytes and str data of any shape,
with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.
Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.
Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.
Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Range-read milestone M5 (docs/design/swmr.md): read a file a libhdf5
SWMR writer (h5py f.swmr_mode = True) is still appending to, as h5py's
File(path, 'r', swmr=True) does.
- File::open_swmr / open_storage_swmr: the file is read through the new
FileStorage (pread / seek_read, never mapped, len() the current length),
reads are bounded by the length at each read, and the chunk cache is not
used (a cached index hides new chunks; a cached edge chunk reads as fill
where the writer has since written). A file open for writing without
SWMR is refused with Error::Locked, as libhdf5 refuses it.
- Dataset::refresh re-reads the object header (H5Drefresh).
- Operations that fail with an error a racing write can cause (every
format error but a wrong path/selection, an unsupported feature or a bad
argument) are run again, up to 100 attempts (libhdf5's default metadata
read attempts for SWMR readers), 1 us doubling to 10 ms apart; nested
operations retry as a whole. swmr_retries() counts them,
swmr_writer_active() re-reads the superblock flags.
Tests: FileStorage growth and retry policy (unit); the mid-write copy
through open_swmr; a storage that garbles reads (retried, given up after
the attempts, never returned); and a live h5py writer appending to
Extensible-Array (plain and gzip) and v2-B-tree datasets for 2500 steps
while two Rust reader threads and h5py's SWMR reader check every value,
then the closed file read equal to h5py. Leaving the chunk cache on in
live mode fails the live test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>