Phase 3, third batch: lazy remote loading in the browser (range-read M4), reading files a SWMR writer is still appending to (M5), and Python bindings for remote reads and in-place editing. Each branch was reviewed adversarially, and every finding is fixed with a test that fails without the fix.
The Python review also found a wrong-data bug that shipped in #18: FileEditor::resize scrambled data when shrinking a chunked dataset that clawhdf5 wrote. It is fixed here.
Conformance: 600 of 697, with 0 panics, hangs or crashes.
M4: lazy remote files in the browser (clawhdf5-wasm)
openUrl(url, opts) returns a RemoteFile with the same methods as open(bytes), each returning a promise. It fetches only the byte ranges it needs over HTTP Range.
How it works: a restartable "need bytes" loop over a block cache. A pass that misses blocks records every missing block, the JS side fetches them in parallel, and the pass runs again. There's no blocking in the main thread and no Web Worker.
Response checks: every response is validated: status 206, an exact Content-Range and length, and the ETag / Last-Modified seen at open.
Servers without Range: the file is downloaded whole, up to a limit, or refused, depending on the option.
Cross-origin servers: these often hide their range headers; the length then comes from a HEAD request.
Viewer: a URL box and a live request/byte counter.
Tests: 1,894 openUrl checks under Node plus headless Chromium, including cross-origin. 622 corpus files read over HTTP match open(bytes). The 200 MB test file opens in 5 requests and 6 MiB.
Fixes from review:
Crash: a hostile server claiming a huge length aborted the wasm instance, and with it every open file on the page. Allocations are now bounded (maxFetch, a 1 GiB decode limit, try_reserve). Files of 4 GiB or more are refused on wasm32.
Oversized responses: every response body is streamed with a hard cap.
Listing speed: listing a large group used to need one round trip per header. It now takes 6 passes instead of 185; the format crate's group walkers now keep collecting after the first missing block.
Options:Headers instances are accepted, requests still in flight are aborted when one fails, and parallel is validated.
Size: the wasm package grew about 11% gzipped (provisional; the viewer README's size table is marked stale).
M5: SWMR reader (docs/design/swmr.md)
New APIs:
File::open_swmr / open_storage_swmr read live through a growing FileStorage (pread, never mmap).
Dataset::refresh() works like H5Drefresh.
swmr_writer_active() reports whether the writer is still attached.
Live reading applies only to files whose superblock has the SWMR-write flag. Any other file reads exactly as File::open reads it.
Bounded retries on live files: only failures a writer that's mid-flush can cause are retried: checksum mismatches, reads past the current end, and header-prefix errors. Permanent errors return immediately.
Every read path retries, including strings, VL data and attributes. Wrapping them exposed that Dataset::attr on dense attributes returned None on a transient failure.
Bug fixed:Superblock::data_end cut live SWMR files off at their stale recorded end of file.
Test: a live h5py SWMR writer appends while two Rust reader threads and h5py's own SWMR reader follow it, for thousands of iterations (10,000 steps in release) with no torn or wrong values.
Limit (documented): HDF5 has no checksum on global-heap data, so a torn VL read is caught only by its signature and bounds checks.
Python bindings: remote reads and editing
Remote files:clawhdf5.File("http://…") and File.open_url(url, …) read through clawhdf5-remote with the GIL released. Every read now goes through File::storage(), so storage-backed files work.
The default wheel reads plain HTTP only and builds no C. https, s3, gcs and azure are opt-in features.
Editing:clawhdf5.File(path, 'r+') writes through FileEditor.
ds[key] = value follows h5py's key and broadcast rules, and numeric conversions follow libhdf5, clamping out-of-range values.
resize and attrs[...] = ... are supported.
Anything unsupported raises NotImplementedError.
Tests: 163 pytest cases, including random edit sequences checked against both h5py and a numpy model, and a test that edits and remote reads release the GIL.
Fixes from review:
FileEditor::resize wrong data, shipped in #18. Our writer stored no maximum dims, and libhdf5 then takes the current dims as the maximum, so a shrink changed the Fixed Array's layout under existing chunks. The editor now records the original dims as the maximum before the first resize. The writer now records a maximum for every chunked dataset, so h5py can resize files we write: 8 bytes per dim, and max_dimensions() now returns Some. Tests cover a v2.7.0 fixture, FileBuilder files and h5py files.
Edits could land in the wrong file: after a chdir or a rename, an edit could be planned from one file and written into another. FileEditor now plans every edit from the file it holds (reader()), and paths are canonicalised.
Zero-size extents no longer divide by zero in the chunk grid, which had panicked h5rs check.
Boolean-mask writes now raise NotImplementedError.
Found while merging
Build break: the SWMR branch added a swmr field to File, and the Python branch added File::from_std_file without setting it.
Non-deterministic errors on damaged files: the chunk cache returned chunks in hash-map order, seeded per File, so two opens of a damaged file could name different failing chunks. This was the intermittent cve-2025-2310.h5 failure. The cache now keeps the index's own order. The new test fails on the old code, where two opens reported chunks [40, 0] and [320, 0].
Conformance: 600 of 697, gate passes, no class changes.
h5rs check --data flags 0 of the 436 fully-read files.
ClawBrainHub builds and passes 204/204. It now runs from a scratch clone with an absolute path, not through a symlink.
Behaviour changes (in the CHANGELOG)
FileBuilder records max dims for chunked datasets, so max_dimensions() returns Some(dims).
Files clawhdf5 wrote before this change have no stored maximum, and h5py still scrambles them if it resizes them. Resize them once with the fixed FileEditor first; this is written up in known-issues.
New JS API openUrl / RemoteFile; new Rust APIs File::open_swmr, Dataset::refresh, FileEditor::reader; clawhdf5.File(url) and mode 'r+' in Python.
Phase 3, third batch: lazy remote loading in the browser (range-read M4), reading files a SWMR writer is still appending to (M5), and Python bindings for remote reads and in-place editing. Each branch was reviewed adversarially, and every finding is fixed with a test that fails without the fix.
The Python review also found a **wrong-data bug that shipped in #18**: `FileEditor::resize` scrambled data when shrinking a chunked dataset that clawhdf5 wrote. It is fixed here.
**Conformance: 600 of 697**, with 0 panics, hangs or crashes.
## M4: lazy remote files in the browser (`clawhdf5-wasm`)
- `openUrl(url, opts)` returns a `RemoteFile` with the same methods as `open(bytes)`, each returning a promise. It fetches only the byte ranges it needs over HTTP `Range`.
- **How it works:** a restartable "need bytes" loop over a block cache. A pass that misses blocks records every missing block, the JS side fetches them in parallel, and the pass runs again. There's no blocking in the main thread and no Web Worker.
- **Response checks:** every response is validated: status 206, an exact `Content-Range` and length, and the ETag / Last-Modified seen at open.
- **Servers without `Range`:** the file is downloaded whole, up to a limit, or refused, depending on the option.
- **Cross-origin servers:** these often hide their range headers; the length then comes from a HEAD request.
- **Viewer:** a URL box and a live request/byte counter.
- **Tests:** 1,894 `openUrl` checks under Node plus headless Chromium, including cross-origin. 622 corpus files read over HTTP match `open(bytes)`. The 200 MB test file opens in 5 requests and 6 MiB.
- **Fixes from review:**
- **Crash:** a hostile server claiming a huge length aborted the wasm instance, and with it every open file on the page. Allocations are now bounded (`maxFetch`, a 1 GiB decode limit, `try_reserve`). Files of 4 GiB or more are refused on wasm32.
- **Oversized responses:** every response body is streamed with a hard cap.
- **Listing speed:** listing a large group used to need one round trip per header. It now takes 6 passes instead of 185; the format crate's group walkers now keep collecting after the first missing block.
- **Options:** `Headers` instances are accepted, requests still in flight are aborted when one fails, and `parallel` is validated.
- **Size:** the wasm package grew about 11% gzipped (provisional; the viewer README's size table is marked stale).
## M5: SWMR reader (`docs/design/swmr.md`)
- New APIs:
- `File::open_swmr` / `open_storage_swmr` read live through a growing `FileStorage` (pread, never mmap).
- `Dataset::refresh()` works like `H5Drefresh`.
- `swmr_writer_active()` reports whether the writer is still attached.
- **Live reading applies only to files whose superblock has the SWMR-write flag.** Any other file reads exactly as `File::open` reads it.
- **Bounded retries on live files:** only failures a writer that's mid-flush can cause are retried: checksum mismatches, reads past the current end, and header-prefix errors. Permanent errors return immediately.
- **Every read path retries**, including strings, VL data and attributes. Wrapping them exposed that `Dataset::attr` on dense attributes returned `None` on a transient failure.
- **Bug fixed:** `Superblock::data_end` cut live SWMR files off at their stale recorded end of file.
- **Test:** a live h5py SWMR writer appends while two Rust reader threads and h5py's own SWMR reader follow it, for thousands of iterations (10,000 steps in release) with no torn or wrong values.
- **Limit (documented):** HDF5 has no checksum on global-heap data, so a torn VL read is caught only by its signature and bounds checks.
## Python bindings: remote reads and editing
- **Remote files:** `clawhdf5.File("http://…")` and `File.open_url(url, …)` read through clawhdf5-remote with the GIL released. Every read now goes through `File::storage()`, so storage-backed files work.
- The default wheel reads plain HTTP only and builds no C. `https`, `s3`, `gcs` and `azure` are opt-in features.
- **Editing:** `clawhdf5.File(path, 'r+')` writes through `FileEditor`.
- `ds[key] = value` follows h5py's key and broadcast rules, and numeric conversions follow libhdf5, clamping out-of-range values.
- `resize` and `attrs[...] = ...` are supported.
- Anything unsupported raises `NotImplementedError`.
- **Tests:** 163 pytest cases, including random edit sequences checked against both h5py and a numpy model, and a test that edits and remote reads release the GIL.
- **Fixes from review:**
- **`FileEditor::resize` wrong data, shipped in #18.** Our writer stored no maximum dims, and libhdf5 then takes the current dims as the maximum, so a shrink changed the Fixed Array's layout under existing chunks. The editor now records the original dims as the maximum before the first resize. The writer now records a maximum for every chunked dataset, so h5py can resize files we write: 8 bytes per dim, and `max_dimensions()` now returns `Some`. Tests cover a v2.7.0 fixture, FileBuilder files and h5py files.
- **Edits could land in the wrong file:** after a `chdir` or a rename, an edit could be planned from one file and written into another. `FileEditor` now plans every edit from the file it holds (`reader()`), and paths are canonicalised.
- **Zero-size extents** no longer divide by zero in the chunk grid, which had panicked `h5rs check`.
- Boolean-mask writes now raise `NotImplementedError`.
## Found while merging
- **Build break:** the SWMR branch added a `swmr` field to `File`, and the Python branch added `File::from_std_file` without setting it.
- **Non-deterministic errors on damaged files:** the chunk cache returned chunks in hash-map order, seeded per `File`, so two opens of a damaged file could name different failing chunks. This was the intermittent `cve-2025-2310.h5` failure. The cache now keeps the index's own order. The new test fails on the old code, where two opens reported chunks `[40, 0]` and `[320, 0]`.
## Verification (tank, 2026-09-27)
- `CLAWHDF5_REQUIRE_INTEROP=1 scripts/ci-test.sh`: **27/27 pass**.
- Conformance: 600 of 697, gate passes, no class changes.
- `h5rs check --data` flags 0 of the 436 fully-read files.
- ClawBrainHub builds and passes 204/204. It now runs from a scratch clone with an absolute path, not through a symlink.
## Behaviour changes (in the CHANGELOG)
- `FileBuilder` records max dims for chunked datasets, so `max_dimensions()` returns `Some(dims)`.
- Files clawhdf5 wrote before this change have no stored maximum, and h5py still scrambles them if it resizes them. Resize them once with the fixed `FileEditor` first; this is written up in known-issues.
- New JS API `openUrl` / `RemoteFile`; new Rust APIs `File::open_swmr`, `Dataset::refresh`, `FileEditor::reader`; `clawhdf5.File(url)` and mode `'r+'` in Python.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
libhdf5's SWMR semantics as they matter to a reader (superblock v3 with
the SWMR-write flag and a stale EOF, append-only writer, flush
dependencies, refresh, 100 metadata read attempts), what clawhdf5 did
with such files, and the plan: data_end bounded by the file length for
SWMR-flagged files, a live File::open_swmr over positioned reads without
the chunk cache, Dataset::refresh, and bounded whole-operation retries.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A SWMR writer does not keep the superblock's end-of-file address up to
date: a copy h5py made of its own file mid-write records 715 in a
6 030-byte file. Superblock::data_end only ignored the recorded end when
it lay past the end of the file, so every open path bounded reads at
715: the file listed, but every chunked read failed ("unexpected EOF:
need 787 bytes, have 715") and h5rs check reported the chunk indexes
past the end of the file. libhdf5's SWMR reader skips the
end-of-allocation check for every read (H5FD_read); for a v3 superblock
with the SWMR-write flag the data now ends at the end of the file.
Test: tests/swmr_interop.rs over the mid-write copy (fixture), through
File::open, open_buffered and from_bytes, and against h5py's SWMR reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
LazyStorage holds the blocks of a remote file fetched so far. An
operation runs as passes over it: a read that misses records the
missing blocks and fails, the pass's result is dropped whatever it is
(a parser may have caught the error and carried on), and the caller
fetches the reported ranges and re-runs the pass. No block is evicted
while an operation is in flight, so every pass that does not finish
asks for at least one new block and the operation ends. Blocks sit in
an LRU with a byte budget, trimmed between operations, bulk (raw data)
blocks first.
Reader::open_storage opens a file through any Storage, and variable-
length strings resolve through the file's storage instead of
File::as_bytes, which panics for a file not in memory.
tests/lazy.rs compares, file by file, what the viewer can show (kinds,
listings, attributes, info, whole reads and hyperslabs) read lazily
with the facade's range-storage path and the in-memory reader: files
built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with
CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A
listing plus a small read of a 48 MB file fetches 3 ranges.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.
- Every object of a File now shares one handle (src/handle.rs) that
runs all file access, metadata included, with the GIL released and
parses through File::storage() and the clawhdf5_format *_in functions.
Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
clawhdf5_remote::storage_for_url (read-only; another mode is a
ValueError). File.open_url(url, **options) takes the cache and HTTP
options (block_size, cache_size, headers, retries, timeout,
allow_full_download, max_full_download, require_validator,
max_redirects, max_parallel); File.remote_stats gives the block
cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
OSError, never KeyError/ValueError and never data; `key in group`
raises it instead of answering False.
Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
openUrl(url, opts) returns a RemoteFile with the methods of H5File
(kind, list, info, attrs, attrErrors, read, readHyperslab), each a
promise, and stats(). It runs every call through the restartable
LazyStorage: a pass that misses reports the byte ranges, js/remote.js
fetches them with fetch() and Range headers (six at a time), and the
pass is re-run. This keeps the main thread free without a Worker or
synchronous XHR (h5wasm's lazy files need both), as the design doc
recommends; the cost is re-running a pass per wave of misses.
Every answer is checked: a 206 with exactly the bytes asked for, and
the same ETag/Last-Modified and length as at open, else an error (never
data). A server that ignores Range (200) is downloaded whole, up to
maxDownload (512 MiB), unless fallback: "error". Options: blockSize,
cacheSize, headers, credentials, parallel, fetch.
test/serve.py is a range-capable static server with request counting
(and /norange/ for a server without range support). test.mjs repeats
every fixture check on files opened by URL (1 MiB and 512 B blocks),
checks the request budget on a 200 MB h5py file (list, three small
reads and a window of the big dataset: 5 requests, 6 MiB), the
download fallback, and HTTP errors, changed files and wrong answers;
with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The page gets a URL box; ?file=<url> now opens with openUrl instead of
downloading the file, and the header shows the requests made and bytes
fetched so far ("5 requests, 5 MiB of 191 MiB fetched (2.6%)"). Every
file call is awaited, so local files (open) and remote ones share the
code; a selection that finishes after another was made is not shown.
browser.sh serves the page and fixtures with test/serve.py (no symlinks
into the repository any more) and also checks a server without range
support and, on the 200 MB file, that a small dataset and a window of
the big one render with a single-digit percentage fetched.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased), known-issues (the browser can open URLs; the
limits of openUrl: round trips per wave of misses, a call holds what it
reads, CORS and validator visibility, the download fallback), the M4
status in docs/design/range-reads.md (why NeedBytes rather than a
Worker, why not clawhdf5-remote's BlockCache, what was tested; the
status paragraph at the top lost a garbled duplicate), the viewer's
README (API, options, how it works, tests; the size table is marked as
predating openUrl), CLAUDE.md and the README crate list.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Range-read milestone M5 (docs/design/swmr.md): read a file a libhdf5
SWMR writer (h5py f.swmr_mode = True) is still appending to, as h5py's
File(path, 'r', swmr=True) does.
- File::open_swmr / open_storage_swmr: the file is read through the new
FileStorage (pread / seek_read, never mapped, len() the current length),
reads are bounded by the length at each read, and the chunk cache is not
used (a cached index hides new chunks; a cached edge chunk reads as fill
where the writer has since written). A file open for writing without
SWMR is refused with Error::Locked, as libhdf5 refuses it.
- Dataset::refresh re-reads the object header (H5Drefresh).
- Operations that fail with an error a racing write can cause (every
format error but a wrong path/selection, an unsupported feature or a bad
argument) are run again, up to 100 attempts (libhdf5's default metadata
read attempts for SWMR readers), 1 us doubling to 10 ms apart; nested
operations retry as a whole. swmr_retries() counts them,
swmr_writer_active() re-reads the superblock flags.
Tests: FileStorage growth and retry policy (unit); the mid-write copy
through open_swmr; a storage that garbles reads (retried, given up after
the attempts, never returned); and a live h5py writer appending to
Extensible-Array (plain and gzip) and v2-B-tree datasets for 2500 steps
while two Rust reader threads and h5py's SWMR reader check every value,
then the closed file read equal to h5py. Leaving the chunk cache on in
live mode fails the live test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():
- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
slices and integers with extra leading 1-axes allowed; the exact shape
for an index list, a scalar only where h5py expands it). Arrays are
converted as libhdf5 converts them in native byte order (integers
saturate, floats truncate toward zero and clip, integers go into h5py's
bool enum by value); other values through
numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
dataset is a ValueError instead of libhdf5's arbitrary value. The value
preparation is a small Python module compiled into the extension
(src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
attrs.modify: numeric, bool, complex, bytes and str data of any shape,
with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.
Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.
Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.
Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A file opened with open_swmr whose superblock does not have the SWMR-write
flag (its writer has closed it) is now read exactly as File::open reads it:
bounded by its recorded end of file, through the chunk cache, each
operation tried once, and is_swmr_read() is false. Before, every file opened
with open_swmr ignored its recorded end of file, so a closed file whose end
of file is below its length (h5clear_fsm_persist_less.h5, or the mid-write
fixture with its flags cleared) listed and read objects that File::open and
libhdf5's plain reader refuse. Opening such a file is not retried once the
superblock has been read.
libhdf5's SWMR reader is looser than either (it reads the fixture's chunk
index past the end of file even with the flags cleared); a test pins what
h5py does and docs/design/swmr.md says why we follow the plain reader.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
is_transient_format counted every format error but a handful as transient,
so on an open_swmr handle a permanent failure (a file that is not HDF5, an
unsupported version or message, a truncated file) was retried 100 times,
about 0.9 s of pauses per failing operation. Now only these are retried:
a checksum mismatch; a read past the file's current end (UnexpectedEof;
libhdf5 reads zeros there, which fail the checksum); and an object header
prefix whose signature or version does not decode, which libhdf5's
H5C__load_entry also retries (a header garbled whole fails there before its
checksum). Everything else is returned at once.
Tests: a unit test that every permanent kind returns after one call within
50 ms and every transient kind is retried to the limit; open_storage_swmr
of a non-HDF5 buffer returns SignatureNotFound within 100 ms (0.87 s before)
and a missing name on a live file fails without retries. The torn-read and
live h5py-writer tests still pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The README example looped while swmr_writer_active(), which never ends when
the writer crashed or was killed: libhdf5 clears the SWMR-write flag only on
close (the mid-write fixture keeps it set for good). The loop now also stops
after a minute without growth, and the README, the swmr_writer_active docs
(with the same loop as a compiled no_run doctest) and the design say why.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.
- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
most 1 GiB): a read longer than it fails at once, before anything is
fetched, and an operation whose passes would fetch more than it fails
before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
code turns offsets into usize, so nothing past 4 GiB can be read there
(shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
maxDownload is bounded to 1 GiB.
Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
read_string, read_string_bytes, read_string_selection, read_vlen and
read_vlen_selection now retry as a whole, the global-heap decoding after the
read included; so do File::decode_strings / decode_string_bytes /
decode_vlen, Group::datasets / groups / attrs / attr, Dataset::attrs / attr,
the typed full reads (read_f64 ...), the header messages behind shape() and
dtype() (a shared message is read from another header), and
verify_provenance. Before, a transient failure on those reached the caller.
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as AttrValue::Raw) instead of failing, which hid
a transient error as a missing or raw attribute: on a live file such an
error of a retried kind now runs the read again too, and after the last
attempt the last result is returned as before. The format crate gains
find_attribute_reporting_in, which returns the errors find_attribute_in
skips (dense name-index lookups dropped them). The zero-copy reads need the
file in memory, which a live file never is, so they have nothing to retry.
Test: a storage that fails one read with a checksum mismatch; for 14 read
paths over a new fixture (tests/fixtures/swmr_strings_attrs.h5, an h5py
copy with the SWMR-write flag: vlen strings, vlen int32, dense attributes),
each read the path makes fails once in turn and the path must return the
same result with exactly one retry. It fails on the previous commit
(root.attrs, read 0). Dataset::attr on dense attributes returned None
before find_attribute_reporting_in.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
remote.js read every 206 body (the probe and each range) with
resp.arrayBuffer() and checked its length afterwards, so a hostile
server could make the page buffer gigabytes before the check failed.
Every body is now piped through a TransformStream that errors as soon as
the count passes the limit, which cancels the body (and the request):
the requested range length for a 206, maxDownload for the 200 fallback.
A declared Content-Length past the limit is refused before reading. The
200 path now always streams (it read a declared length at once, because
a reader loop stalls on small bodies in headless Chromium under
--virtual-time-budget; a pipe does not).
Test (test.mjs): a probe and a range answered with a 64 MiB body read at
most the range + one 64 KiB piece; before, all 64 MiB were read ("asked
for the first 1048576 bytes of 2000000, got 67108864").
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.
libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).
Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A Fixed or Extensible Array index whose maximum extent (the current one
when no maximum is recorded) is 0 along a dimension has a zero stride for
every dimension before it; ChunkGrid::offsets divided by it. The unfixed
editor made such files by resizing a clawhdf5-written dataset to a zero
extent: `h5rs check` panicked and the next resize raised an internal
error (12 of the reviewer's random-edit seeds 10..39). Such an index has
no slot for any chunk of the dataset; offsets now returns None.
Tests: chunk_grid::zero_extent_has_no_chunks; edit_interop's
zero_extent_resizes_without_a_recorded_maximum on a file the unfixed
editor left (fixture) and on a 2.7.0-written file taken through zero
extents with `h5rs check --data` and h5dump at every step; test_edit.py
random edits on seeds 10..39 of a clawhdf5-written file.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FileEditor re-opened its path to plan each edit but wrote through the file
it held open, and the Python 'r+' handle re-opened the path after every
edit to read. When the path came to name another file between edits (a
rename or replacement, or a relative path after os.chdir), an edit was laid
out from the other file's metadata and written into the held one,
corrupting it, and later reads came from the other file (the review's
repro: h5py then reports "invalid dataset size, likely file corruption").
The editor now plans from a mapping of its own file (a clone of the held
descriptor, dropped before the edit writes) and canonicalises its path at
open. New FileEditor::reader() opens the held file anew for reading,
without sharing the editor's flock (a mapping of a cloned descriptor holds
the lock until unmapped): through /proc/self/fd on Linux, which follows a
renamed file; elsewhere by path, refused on Unix when the path no longer
names the held file. The Python handle reads through it and keeps no path;
a 'w' file is written at the absolute path it was opened with.
Tests: edit_tests.rs edits_go_to_the_file_held_not_the_path; test_edit.py
test_relative_path_and_chdir and test_path_replaced_between_edits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A thread appending timestamps in a loop runs while a 1024x1024 gzip
dataset is rewritten through 'r+': the largest gap between its stamps
during the edit must be under half the edit's duration (an edit holding
the GIL stalls it for the whole edit; checked with a GIL-holding regex
standing in for the edit: one 0.20 s gap in a 0.21 s call). Edits already
ran detached; nothing tested it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.
Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Planning from a mapping of a clone of the locked descriptor shared its
flock: a process forked by another thread meanwhile (any Command) kept the
lock alive for a moment after the editor was dropped, and a test that
reopened the file at once saw Error::Locked (once in a full run). On Linux
plan through /proc/self/fd (a new open file description of the same
file, which still follows a rename); elsewhere keep the clone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
serve.py always exposed Content-Range and ETag, and Node has no CORS, so
openUrl's documented cross-origin path (length from a HEAD request,
answers checked by body length alone, no validator) was never run.
- serve.py /noexpose/ serves ranges without Content-Range, ETag,
Last-Modified or Accept-Ranges (what a page sees of a server that
does not expose them); /unexposed/ sends them but exposes none, for a
real browser. HEAD requests are counted (0 bytes).
- test.mjs: every fixture check through /noexpose/ at 1 MiB and 512-byte
blocks (one HEAD each, requests and bytes as the server counted them),
concurrent reads with cacheSize 0, a short answer still caught, and a
server without a HEAD length a clear error.
- browser.sh: the page on 127.0.0.1 opens the file from localhost, once
with Content-Range exposed and once through /unexposed/, where the
server's log must show the HEAD.
Checked by breaking the HEAD length in remote.js: the new checks fail.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- opts.headers is read as fetch reads it (a Headers, [name, value] pairs
or a plain object); it was spread as an object, which silently dropped
a Headers instance (a common way to pass Authorization). A caller's
Range is not sent.
- When one range request of a batch fails, the others in flight are
aborted (one AbortController per batch, its signal passed to fetch)
and no new ones start; the first failure is the error. The workers
used to go on issuing requests nobody waited for.
- parallel must be an integer from 1 to 1024 (openUrl) or a positive
integer (fetchRanges): a non-number gave NaN workers, so none ran and
fetchRanges returned nothing.
Tests (test.mjs): headers as object, Headers and pairs reach the fetch;
bad parallel values are option errors; fetchRanges with a 500 on the
third of 20 ranges at parallel 3 starts 3 requests and aborts the 2 in
flight. Before: the Headers case sent no header, 17 requests started
after the failure with none aborted, and parallel "x" returned.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB
decode limit, the 4 GiB file limit on wasm32, bodies cut off at their
length, listing passes, the cross-origin tests, and a pre-existing
nondeterministic error choice on cve-2025-2310.h5 that can fail the
native corpus comparison), the viewer README (options, how listing
costs, tests) and the M4 status in docs/design/range-reads.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A boolean array of the axis's length, or of the dataset's whole shape, is
a mask h5py would apply (NotImplementedError here); one of any other shape
(np.array(True), a wrong length) is a key h5py itself refuses with
TypeError, which test_errors_match_h5py requires us to match.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The editor branch added File::from_std_file and the SWMR branch added the
swmr field to File; merged, the constructor did not set it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkCache::chunks_for returned the chunk index's values in hash-map
order, which differs between File instances (a new HashMap per open).
Readers that stop at the first failing chunk therefore named different
chunks on different opens of the same damaged file: whole-dataset
selections here, and the indexed reader behind the storage harness's
intermittent cve-2025-2310.h5 failure. The cache now also keeps the
chunks in the order the index lists them (fetched or built under one
lock) and returns that order, as the uncached readers use.
Regression: several_damaged_chunks_report_the_same_chunk_every_time
(four chunks inflating to different short lengths; before the fix two
opens reported chunk [40, 0] and chunk [320, 0]).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Phase 3, third batch: lazy remote loading in the browser (range-read M4), reading files a SWMR writer is still appending to (M5), and Python bindings for remote reads and in-place editing. Each branch was reviewed adversarially, and every finding is fixed with a test that fails without the fix.
The Python review also found a wrong-data bug that shipped in #18:
FileEditor::resizescrambled data when shrinking a chunked dataset that clawhdf5 wrote. It is fixed here.Conformance: 600 of 697, with 0 panics, hangs or crashes.
M4: lazy remote files in the browser (
clawhdf5-wasm)openUrl(url, opts)returns aRemoteFilewith the same methods asopen(bytes), each returning a promise. It fetches only the byte ranges it needs over HTTPRange.Content-Rangeand length, and the ETag / Last-Modified seen at open.Range: the file is downloaded whole, up to a limit, or refused, depending on the option.openUrlchecks under Node plus headless Chromium, including cross-origin. 622 corpus files read over HTTP matchopen(bytes). The 200 MB test file opens in 5 requests and 6 MiB.maxFetch, a 1 GiB decode limit,try_reserve). Files of 4 GiB or more are refused on wasm32.Headersinstances are accepted, requests still in flight are aborted when one fails, andparallelis validated.M5: SWMR reader (
docs/design/swmr.md)File::open_swmr/open_storage_swmrread live through a growingFileStorage(pread, never mmap).Dataset::refresh()works likeH5Drefresh.swmr_writer_active()reports whether the writer is still attached.File::openreads it.Dataset::attron dense attributes returnedNoneon a transient failure.Superblock::data_endcut live SWMR files off at their stale recorded end of file.Python bindings: remote reads and editing
clawhdf5.File("http://…")andFile.open_url(url, …)read through clawhdf5-remote with the GIL released. Every read now goes throughFile::storage(), so storage-backed files work.https,s3,gcsandazureare opt-in features.clawhdf5.File(path, 'r+')writes throughFileEditor.ds[key] = valuefollows h5py's key and broadcast rules, and numeric conversions follow libhdf5, clamping out-of-range values.resizeandattrs[...] = ...are supported.NotImplementedError.FileEditor::resizewrong data, shipped in #18. Our writer stored no maximum dims, and libhdf5 then takes the current dims as the maximum, so a shrink changed the Fixed Array's layout under existing chunks. The editor now records the original dims as the maximum before the first resize. The writer now records a maximum for every chunked dataset, so h5py can resize files we write: 8 bytes per dim, andmax_dimensions()now returnsSome. Tests cover a v2.7.0 fixture, FileBuilder files and h5py files.chdiror a rename, an edit could be planned from one file and written into another.FileEditornow plans every edit from the file it holds (reader()), and paths are canonicalised.h5rs check.NotImplementedError.Found while merging
swmrfield toFile, and the Python branch addedFile::from_std_filewithout setting it.File, so two opens of a damaged file could name different failing chunks. This was the intermittentcve-2025-2310.h5failure. The cache now keeps the index's own order. The new test fails on the old code, where two opens reported chunks[40, 0]and[320, 0].Verification (tank, 2026-09-27)
CLAWHDF5_REQUIRE_INTEROP=1 scripts/ci-test.sh: 27/27 pass.h5rs check --dataflags 0 of the 436 fully-read files.Behaviour changes (in the CHANGELOG)
FileBuilderrecords max dims for chunked datasets, somax_dimensions()returnsSome(dims).FileEditorfirst; this is written up in known-issues.openUrl/RemoteFile; new Rust APIsFile::open_swmr,Dataset::refresh,FileEditor::reader;clawhdf5.File(url)and mode'r+'in Python.🤖 Generated with Claude Code
A SWMR writer does not keep the superblock's end-of-file address up to date: a copy h5py made of its own file mid-write records 715 in a 6 030-byte file. Superblock::data_end only ignored the recorded end when it lay past the end of the file, so every open path bounded reads at 715: the file listed, but every chunked read failed ("unexpected EOF: need 787 bytes, have 715") and h5rs check reported the chunk indexes past the end of the file. libhdf5's SWMR reader skips the end-of-allocation check for every read (H5FD_read); for a v3 superblock with the SWMR-write flag the data now ends at the end of the file. Test: tests/swmr_interop.rs over the mid-write copy (fixture), through File::open, open_buffered and from_bytes, and against h5py's SWMR reader. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>The page gets a URL box; ?file=<url> now opens with openUrl instead of downloading the file, and the header shows the requests made and bytes fetched so far ("5 requests, 5 MiB of 191 MiB fetched (2.6%)"). Every file call is awaited, so local files (open) and remote ones share the code; a selection that finishes after another was made is not shown. browser.sh serves the page and fixtures with test/serve.py (no symlinks into the repository any more) and also checks a server without range support and, on the 200 MB file, that a small dataset and a window of the big one render with a single-digit percentage fetched. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>remote.js read every 206 body (the probe and each range) with resp.arrayBuffer() and checked its length afterwards, so a hostile server could make the page buffer gigabytes before the check failed. Every body is now piped through a TransformStream that errors as soon as the count passes the limit, which cancels the body (and the request): the requested range length for a 206, maxDownload for the 200 fallback. A declared Content-Length past the limit is refused before reading. The 200 path now always streams (it read a declared length at once, because a reader loop stalls on small bodies in headless Chromium under --virtual-time-budget; a pipe does not). Test (test.mjs): a probe and a range answered with a 64 MiB body read at most the range + one 64 KiB piece; before, all 64 MiB were read ("asked for the first 1048576 bytes of 2000000, got 67108864"). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>Listing a group read every child's object header and stopped at the first that was not fetched yet, and so did the traversals of the group's index (v1 B-tree and symbol table nodes, the local heap's names, v2 B-tree nodes and fractal heap objects). Over openUrl's restartable reader each block cost its own pass and round trip: 184 serial requests to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB. - core::Reader::list reads every child's header before returning the first error (the same error, in listing order, Group::groups/datasets return), classifying them as those do. - clawhdf5-format: after the first sibling that fails, the B-tree v1 and v2 collectors, the symbol table node loop and the dense-link loop go on reading (not using) the remaining siblings, then return that first error: results and errors are unchanged, only failing traversals read more, and in memory that is free (storage::touch). A v1 group's local heap segment (names) is read at once, up to 1 MiB. - LazyStorage no longer fills a one-block hole that is already cached (it was fetched again: 215 MB fetched from a 198 MB file). Measured with tests/lazy.rs listing_cost_of_a_given_file on the reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'): libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73 libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight) libver latest, 1 MiB: 189/188 -> 9/98 libver latest, 64 KiB: 453/452 -> 11/452 Bytes fetched are unchanged (the headers are spread through the file). New test listing_a_large_group_takes_a_few_passes (512-byte blocks): FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000 children 8 and 11 passes. Conformance 600 of 697 (baseline 600); check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>