diff --git a/CHANGELOG.md b/CHANGELOG.md index 664a831..6d3275a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,66 @@ ## Unreleased +### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27) +Design: `docs/design/swmr.md`. +- **Fix: files with the SWMR-write flag were bounded by a stale end of + file** (since the end-of-file check of 2026-09-26, unreleased). A + libhdf5 SWMR writer (h5py `f.swmr_mode = True`) does not keep + the superblock's end-of-file address up to date; a copy h5py made of its + own file mid-write records 715 in a 6 030-byte file. + `Superblock::data_end` only ignored the recorded end when it lay past the + end of the file, so such a file listed, but every chunked read failed + ("unexpected EOF: need 787 bytes, have 715") and `h5rs check` reported + its chunk indexes past the end of the file. For a v3 superblock with the + SWMR-write flag the data now ends at the end of the file, as libhdf5's + SWMR reader reads it (it skips the end-of-allocation check). Every open + path (`File::open`, `open_buffered`, `from_bytes`, `open_storage`, + `MmapFile`, `LazyFile`, `h5rs`) reads such a copy now. +- **`File::open_swmr(path)` / `File::open_storage_swmr(storage)`**: open a + file a SWMR writer may still be appending to, as libhdf5's SWMR reader + (h5py `File(path, "r", swmr=True)`) does. The file is read with + positioned reads (the new `clawhdf5::FileStorage`: `pread`/`seek_read`, + never mapped, `len()` the file's current length), reads are bounded by + the file's length at the time of each read, and the chunk cache is not + used (a cached chunk index would hide new chunks; a cached edge chunk + would read as fill where the writer has since written). A file open for + writing without SWMR is refused with `Error::Locked`, as libhdf5 refuses + it. +- **`Dataset::refresh()`** reads the dataset's object header again + (`H5Drefresh`, h5py `Dataset.refresh()`), so `shape()` and later reads + see the writer's appends; a handle keeps its extent until refreshed. +- **Bounded retries.** On a SWMR-read file, an operation (open, lookups, + refresh, reads) that fails with an error a concurrent write can cause — + any format error except a wrong name or selection, an unsupported + feature or a bad argument — is run again from the start, up to + `File::swmr_read_attempts()` times (`SWMR_READ_ATTEMPTS` = 100, + libhdf5's default metadata read attempts for SWMR readers, + `H5Pset_metadata_read_attempts`; `set_swmr_read_attempts` changes it), + pausing 1 µs doubling to 10 ms between attempts. Results come only from + an attempt in which every structure verified, so a torn read is at worst + an error, never data. `File::swmr_retries()` counts the retries. + `File::swmr_writer_active()` reads the superblock flags again to tell + when the writer has closed the file. +- Tests (`crates/clawhdf5/tests/swmr_interop.rs`): the mid-write copy + (fixture `tests/fixtures/swmr_mid_write.h5`) through every open path and + against h5py's SWMR reader; garbled reads (a storage that corrupts the + next reads) retried, never returned, and given up after the attempts; + and a live test: an h5py writer appends to a 1-D and a 2-D dataset with + one unlimited dimension (Extensible Array index, one of them gzip) and a + 2-D dataset with two (v2 B-tree index) for 2 500 steps + (`CLAWHDF5_SWMR_STEPS`), flushing after each, while two Rust reader + threads refresh and read all three in a loop — every value must be the + one the writer wrote at its position, extents never shrink — beside + h5py's own SWMR reader doing the same checks; after the writer closes, + the live handle and a new `File::open` read exactly what h5py reads. + Run on tank, 2026-09-27, with h5py 3.16 / HDF5 2.0: + `CLAWHDF5_REQUIRE_INTEROP=1 CLAWHDF5_PYTHON=.venv/bin/python cargo test + -p clawhdf5 --test swmr_interop`; also at 20 000 steps in a release + build. A test with the chunk cache left on in live mode fails it. +- Not covered: SWMR writing, remote SWMR (`BlockCache` caches blocks and + `HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing + groups or attributes (a SWMR writer cannot add objects or attributes). + ### Range reads, milestone M3: remote files (2026-09-26) - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives a `clawhdf5::File` (through `File::open_storage`) that reads the file by diff --git a/README.md b/README.md index 3301b09..925b6fc 100644 --- a/README.md +++ b/README.md @@ -101,6 +101,10 @@ breaking change, are in [CHANGELOG.md](CHANGELOG.md). HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a block cache, without downloading it; `h5rs` takes URLs with its `remote` feature. See [Reading remote files](#reading-remote-files). +- `File::open_swmr` follows a file an h5py/libhdf5 SWMR writer is still + appending to (`Dataset::refresh`, bounded retries); copies of such files + taken mid-write read with every open path. See + [Following a file a SWMR writer is appending to](#following-a-file-a-swmr-writer-is-appending-to). **Tooling** - CI now runs the h5py/netCDF4 interop suites for real (they had been skipping @@ -504,6 +508,29 @@ built with `--features remote` takes the same URLs: `https://` is the `https` feature (rustls with ring, which compiles C). Limits are in [known issues](docs/known-issues.md). +### Following a file a SWMR writer is appending to + +`File::open_swmr` reads a file that a libhdf5 writer in SWMR mode (h5py +`f.swmr_mode = True`) is still appending to, as h5py's +`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new +extent, every read reads the chunk index as it is now, and a read that +races the writer (a checksum that fails mid-flush) is retried, up to 100 +attempts as in libhdf5, and never returned torn. + +```rust +let file = clawhdf5::File::open_swmr("live.h5")?; +let mut ds = file.dataset("samples")?; +while file.swmr_writer_active()? { + ds.refresh()?; + let n = ds.shape()?[0]; + // read the new rows ... + std::thread::sleep(std::time::Duration::from_millis(100)); +} +ds.refresh()?; // the final extent +``` + +Design and limits: [docs/design/swmr.md](docs/design/swmr.md). + ### Python `crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with diff --git a/docs/design/range-reads.md b/docs/design/range-reads.md index af44ae2..e0d10ef 100644 --- a/docs/design/range-reads.md +++ b/docs/design/range-reads.md @@ -7,7 +7,9 @@ change. Progress: M0 and M1 are done, and so is M2 (branch `File::open_storage` gives the facade's read API over any `Storage` (see `CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch `feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S), -object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is +object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is +done on branch `feat/p3-m5-swmr-reader`, with its own design in +[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is next. Every count in §1–§2 was change. Progress: M1, first part (the `Storage` trait and the metadata @@ -523,6 +525,19 @@ fast path within benchmark noise. **M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow; add `File::refresh()` that re-reads the superblock/EOF and invalidates cached blocks past the old end. Needs libhdf5 SWMR semantics research first. +- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and + libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch + above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's + `H5Drefresh`), not per file — a SWMR writer only grows datasets, and the + superblock's EOF is not kept up to date by it, so there is nothing to + re-read there (`File::swmr_writer_active` re-reads its flags). A live + file (`File::open_swmr`) reads through `FileStorage` (positioned reads, + `len()` the current length) with no block or chunk cache, rather than + invalidating cached blocks: `BlockCache`/`HttpStorage` stay snapshot + readers. Operations that fail with an error a racing write can cause are + retried up to 100 times (libhdf5's default for SWMR readers). + Tested against a live h5py writer and h5py's SWMR reader + (`crates/clawhdf5/tests/swmr_interop.rs`). Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured). diff --git a/docs/design/swmr.md b/docs/design/swmr.md index 339c2d8..c69fd86 100644 --- a/docs/design/swmr.md +++ b/docs/design/swmr.md @@ -86,16 +86,25 @@ writer keeps appending. What the format and the library guarantee: elements inside that extent. 4. **Bounded retries.** In a live file, an operation (open, dataset lookup, refresh, every read) that fails with an error a concurrent write can - cause — a checksum or Fletcher-32 mismatch, a short read, a bad - signature, an undecodable chunk or chunk index — is run again from the - start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's - default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, … - up to 10 ms between attempts (under a second in all). libhdf5 retries - the one structure whose checksum failed; we retry the whole operation, - because the parsers are pure functions of the bytes they read. Results - are only returned from a run where every structure verified, so an - error is returned instead of torn data. Other errors are returned at - once, and non-live files never retry. + cause is run again from the start, up to `File::swmr_read_attempts()` + times (default 100, libhdf5's default; `set_swmr_read_attempts` changes + it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second + in all). libhdf5 retries the one structure whose checksum failed; we + retry the whole operation, because the parsers are pure functions of the + bytes they read. Which errors: a garbled read can fail whichever check + the parser makes first — a checksum, a signature, a version byte, a + size (a test that garbles every byte of a header got + `InvalidObjectHeaderVersion`, not a checksum error) — so every format + error counts except those later bytes cannot change: a path or selection + the caller got wrong, an unsupported filter or external storage, a bad + argument. Those, and a live file's permanent errors of the other kinds + after the last attempt, are returned; non-live files never retry. + Results are only returned from a run where every structure verified, so + a torn metadata read is an error, never data. `File::swmr_retries()` + counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`). + Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5 + as here: correctness rests on the writer's ordering (chunk data before + the index entry, only appends), as for libhdf5's reader. 5. **Writer state.** `File::swmr_writer_active()` reads the superblock flags again, so a reader can tell when the writer has closed the file. @@ -120,5 +129,15 @@ writer cannot add them), and `MmapFile`/`LazyFile`. ## Status -See the end of this file's history in `CHANGELOG.md` ("Range reads, -milestone M5"). +Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed +above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the +same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test +swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release +build): no read returned a value the writer had not written at that +position, h5py's reader agreed, and retries were needed but rare (the +failures seen were checksum mismatches, each cured by one retry). A +variant of the test with the chunk cache left on in live mode fails it +(stale chunk index / edge chunk), which is why live files do not use it. + +Also found: `File::open` of such a file had been failing since the +end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`). diff --git a/docs/known-issues.md b/docs/known-issues.md index e97f602..e24e076 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -7,6 +7,29 @@ deleting it. --- +## Files a SWMR writer had open could not be read past a stale end of file + +**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any +release: reads have been bounded by the recorded end of file since +`7d7a7e7` (2026-09-26), which no release contains. + +A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the +superblock's SWMR-write flag and does not keep its end-of-file address up to +date: a copy h5py made of its own file mid-write records 715 in a +6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only +ignored the recorded end when it lay *past* the end of the file, so every +reader bounded such a file at 715 bytes: it listed, but every chunked read +failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check` +reported the chunk indexes past the end of the file. Never wrong data. + +**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the +end of the file, as libhdf5's SWMR reader reads it. **Test:** +`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy, +`tests/fixtures/swmr_mid_write.h5`, through every open path and against +h5py's SWMR reader). A file still being written is read with +`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file +at its length at open and is not meant for files that change while open. + ## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 **Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to