docs: SWMR reader (range-read M5) in CHANGELOG, README, known issues and designs

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:49:04 -05:00
co-authored by Claude Opus 5.5
parent 21f104bb77
commit 3f45755d58
5 changed files with 157 additions and 13 deletions
+60
View File
@@ -2,6 +2,66 @@
## Unreleased ## Unreleased
### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27)
Design: `docs/design/swmr.md`.
- **Fix: files with the SWMR-write flag were bounded by a stale end of
file** (since the end-of-file check of 2026-09-26, unreleased). A
libhdf5 SWMR writer (h5py `f.swmr_mode = True`) does not keep
the superblock's end-of-file address up to date; a copy h5py made of its
own file mid-write records 715 in a 6 030-byte file.
`Superblock::data_end` only ignored the recorded end when it lay past the
end of the file, so such a file listed, but every chunked read failed
("unexpected EOF: need 787 bytes, have 715") and `h5rs check` reported
its chunk indexes past the end of the file. For a v3 superblock with the
SWMR-write flag the data now ends at the end of the file, as libhdf5's
SWMR reader reads it (it skips the end-of-allocation check). Every open
path (`File::open`, `open_buffered`, `from_bytes`, `open_storage`,
`MmapFile`, `LazyFile`, `h5rs`) reads such a copy now.
- **`File::open_swmr(path)` / `File::open_storage_swmr(storage)`**: open a
file a SWMR writer may still be appending to, as libhdf5's SWMR reader
(h5py `File(path, "r", swmr=True)`) does. The file is read with
positioned reads (the new `clawhdf5::FileStorage`: `pread`/`seek_read`,
never mapped, `len()` the file's current length), reads are bounded by
the file's length at the time of each read, and the chunk cache is not
used (a cached chunk index would hide new chunks; a cached edge chunk
would read as fill where the writer has since written). A file open for
writing without SWMR is refused with `Error::Locked`, as libhdf5 refuses
it.
- **`Dataset::refresh()`** reads the dataset's object header again
(`H5Drefresh`, h5py `Dataset.refresh()`), so `shape()` and later reads
see the writer's appends; a handle keeps its extent until refreshed.
- **Bounded retries.** On a SWMR-read file, an operation (open, lookups,
refresh, reads) that fails with an error a concurrent write can cause —
any format error except a wrong name or selection, an unsupported
feature or a bad argument — is run again from the start, up to
`File::swmr_read_attempts()` times (`SWMR_READ_ATTEMPTS` = 100,
libhdf5's default metadata read attempts for SWMR readers,
`H5Pset_metadata_read_attempts`; `set_swmr_read_attempts` changes it),
pausing 1 µs doubling to 10 ms between attempts. Results come only from
an attempt in which every structure verified, so a torn read is at worst
an error, never data. `File::swmr_retries()` counts the retries.
`File::swmr_writer_active()` reads the superblock flags again to tell
when the writer has closed the file.
- Tests (`crates/clawhdf5/tests/swmr_interop.rs`): the mid-write copy
(fixture `tests/fixtures/swmr_mid_write.h5`) through every open path and
against h5py's SWMR reader; garbled reads (a storage that corrupts the
next reads) retried, never returned, and given up after the attempts;
and a live test: an h5py writer appends to a 1-D and a 2-D dataset with
one unlimited dimension (Extensible Array index, one of them gzip) and a
2-D dataset with two (v2 B-tree index) for 2 500 steps
(`CLAWHDF5_SWMR_STEPS`), flushing after each, while two Rust reader
threads refresh and read all three in a loop — every value must be the
one the writer wrote at its position, extents never shrink — beside
h5py's own SWMR reader doing the same checks; after the writer closes,
the live handle and a new `File::open` read exactly what h5py reads.
Run on tank, 2026-09-27, with h5py 3.16 / HDF5 2.0:
`CLAWHDF5_REQUIRE_INTEROP=1 CLAWHDF5_PYTHON=.venv/bin/python cargo test
-p clawhdf5 --test swmr_interop`; also at 20 000 steps in a release
build. A test with the chunk cache left on in live mode fails it.
- Not covered: SWMR writing, remote SWMR (`BlockCache` caches blocks and
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
groups or attributes (a SWMR writer cannot add objects or attributes).
### Range reads, milestone M3: remote files (2026-09-26) ### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+27
View File
@@ -101,6 +101,10 @@ breaking change, are in [CHANGELOG.md](CHANGELOG.md).
HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a
block cache, without downloading it; `h5rs` takes URLs with its `remote` block cache, without downloading it; `h5rs` takes URLs with its `remote`
feature. See [Reading remote files](#reading-remote-files). feature. See [Reading remote files](#reading-remote-files).
- `File::open_swmr` follows a file an h5py/libhdf5 SWMR writer is still
appending to (`Dataset::refresh`, bounded retries); copies of such files
taken mid-write read with every open path. See
[Following a file a SWMR writer is appending to](#following-a-file-a-swmr-writer-is-appending-to).
**Tooling** **Tooling**
- CI now runs the h5py/netCDF4 interop suites for real (they had been skipping - CI now runs the h5py/netCDF4 interop suites for real (they had been skipping
@@ -504,6 +508,29 @@ built with `--features remote` takes the same URLs:
`https://` is the `https` feature (rustls with ring, which compiles C). `https://` is the `https` feature (rustls with ring, which compiles C).
Limits are in [known issues](docs/known-issues.md). Limits are in [known issues](docs/known-issues.md).
### Following a file a SWMR writer is appending to
`File::open_swmr` reads a file that a libhdf5 writer in SWMR mode (h5py
`f.swmr_mode = True`) is still appending to, as h5py's
`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new
extent, every read reads the chunk index as it is now, and a read that
races the writer (a checksum that fails mid-flush) is retried, up to 100
attempts as in libhdf5, and never returned torn.
```rust
let file = clawhdf5::File::open_swmr("live.h5")?;
let mut ds = file.dataset("samples")?;
while file.swmr_writer_active()? {
ds.refresh()?;
let n = ds.shape()?[0];
// read the new rows ...
std::thread::sleep(std::time::Duration::from_millis(100));
}
ds.refresh()?; // the final extent
```
Design and limits: [docs/design/swmr.md](docs/design/swmr.md).
### Python ### Python
`crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with `crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with
+16 -1
View File
@@ -7,7 +7,9 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see `File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch `CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S), `feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is
done on branch `feat/p3-m5-swmr-reader`, with its own design in
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
next. Every count in §1–§2 was next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata change. Progress: M1, first part (the `Storage` trait and the metadata
@@ -523,6 +525,19 @@ fast path within benchmark noise.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow; **M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
blocks past the old end. Needs libhdf5 SWMR semantics research first. blocks past the old end. Needs libhdf5 SWMR semantics research first.
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and
libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch
above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's
`H5Drefresh`), not per file — a SWMR writer only grows datasets, and the
superblock's EOF is not kept up to date by it, so there is nothing to
re-read there (`File::swmr_writer_active` re-reads its flags). A live
file (`File::open_swmr`) reads through `FileStorage` (positioned reads,
`len()` the current length) with no block or chunk cache, rather than
invalidating cached blocks: `BlockCache`/`HttpStorage` stay snapshot
readers. Operations that fail with an error a racing write can cause are
retried up to 100 times (libhdf5's default for SWMR readers).
Tested against a live h5py writer and h5py's SWMR reader
(`crates/clawhdf5/tests/swmr_interop.rs`).
Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured). Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured).
+31 -12
View File
@@ -86,16 +86,25 @@ writer keeps appending. What the format and the library guarantee:
elements inside that extent. elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, dataset lookup, 4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
refresh, every read) that fails with an error a concurrent write can refresh, every read) that fails with an error a concurrent write can
cause — a checksum or Fletcher-32 mismatch, a short read, a bad cause is run again from the start, up to `File::swmr_read_attempts()`
signature, an undecodable chunk or chunk index — is run again from the times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, … in all). libhdf5 retries the one structure whose checksum failed; we
up to 10 ms between attempts (under a second in all). libhdf5 retries retry the whole operation, because the parsers are pure functions of the
the one structure whose checksum failed; we retry the whole operation, bytes they read. Which errors: a garbled read can fail whichever check
because the parsers are pure functions of the bytes they read. Results the parser makes first — a checksum, a signature, a version byte, a
are only returned from a run where every structure verified, so an size (a test that garbles every byte of a header got
error is returned instead of torn data. Other errors are returned at `InvalidObjectHeaderVersion`, not a checksum error) — so every format
once, and non-live files never retry. error counts except those later bytes cannot change: a path or selection
the caller got wrong, an unsupported filter or external storage, a bad
argument. Those, and a live file's permanent errors of the other kinds
after the last attempt, are returned; non-live files never retry.
Results are only returned from a run where every structure verified, so
a torn metadata read is an error, never data. `File::swmr_retries()`
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
as here: correctness rests on the writer's ordering (chunk data before
the index entry, only appends), as for libhdf5's reader.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock 5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file. flags again, so a reader can tell when the writer has closed the file.
@@ -120,5 +129,15 @@ writer cannot add them), and `MmapFile`/`LazyFile`.
## Status ## Status
See the end of this file's history in `CHANGELOG.md` ("Range reads, Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
milestone M5"). above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
+23
View File
@@ -7,6 +7,29 @@ deleting it.
--- ---
## Files a SWMR writer had open could not be read past a stale end of file
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
release: reads have been bounded by the recorded end of file since
`7d7a7e7` (2026-09-26), which no release contains.
A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the
superblock's SWMR-write flag and does not keep its end-of-file address up to
date: a copy h5py made of its own file mid-write records 715 in a
6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only
ignored the recorded end when it lay *past* the end of the file, so every
reader bounded such a file at 715 bytes: it listed, but every chunked read
failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check`
reported the chunk indexes past the end of the file. Never wrong data.
**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the
end of the file, as libhdf5's SWMR reader reads it. **Test:**
`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy,
`tests/fixtures/swmr_mid_write.h5`, through every open path and against
h5py's SWMR reader). A file still being written is read with
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
at its length at open and is not meant for files that change while open.
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 ## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to **Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to