diff --git a/docs/design/swmr.md b/docs/design/swmr.md new file mode 100644 index 0000000..339c2d8 --- /dev/null +++ b/docs/design/swmr.md @@ -0,0 +1,124 @@ +# Design: reading files a SWMR writer is still appending to (range-read M5) + +Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader` +(see "Status" at the end). This is milestone M5 of +[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a +refresh". It covers the reader only; clawhdf5 does not write SWMR files. + +## What libhdf5 does + +A SWMR ("single writer, multiple readers") writer is a libhdf5 process that +opened a file with `libver='latest'` and switched to SWMR mode (h5py +`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file +with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the +writer keeps appending. What the format and the library guarantee: + +- **Superblock v3, flags set.** The writer sets the superblock's + file-consistency flags to write access + SWMR write (`0x05`) and clears + them on close. The superblock's end-of-file address is *not* kept up to + date while writing: a copy of a file taken mid-write records an EOF of a + few hundred bytes while the file is tens of kilobytes (checked on tank, + 2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR + reader therefore skips libhdf5's end-of-allocation check for every read + (`H5FD_read`: "allow access to data past the end of the allocated space + … for SWMR read access"), and bounds reads by the file's real length. +- **A non-SWMR open of such a file fails** in libhdf5: "file is already + open for write (may use to clear file consistency flags)". +- **The writer only appends.** New objects and attributes cannot be created + in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked + datasets with one unlimited dimension use an Extensible Array index, with + more than one a version-2 B-tree; both are updated in a SWMR-safe way. + Fixed Array and single-chunk indexes are for datasets that cannot grow. +- **Flush ordering.** Every metadata structure the writer uses is + checksummed, and flush dependencies order the writes: a chunk's data is + written before the index entry that points to it, and index blocks before + the object header whose dataspace announces the new extent. A reader + that reads the object header first and the index after sees an index at + least as new as the extent, so every chunk inside the extent it read is + either in the index or never written (then it reads as the fill value, as + it does for libhdf5's reader). +- **Refresh.** A reader sees a dataset's new extent only when it refreshes + it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's + cached metadata and reads the object header again. +- **Retries.** Reads are not atomic against writes on every system, so a + checksum can fail when a structure is read while the writer rewrites it. + A SWMR reader reads checksummed metadata up to 100 times before failing + (`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1 + otherwise — `H5Ppublic.h` of HDF5 1.14.6). + +## What clawhdf5 did before + +- `File::open` of a SWMR-flagged file bounded every read by the recorded + EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the + file). A file copied or read mid-write therefore listed, but every + chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and + `h5rs check` reported chunk indexes "past the end of the file". +- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a + per-file chunk cache that keeps each dataset's chunk index and decoded + chunks for the life of the `File`. A reader could not see growth at all, + and a mapping of a file that is being rewritten can change under a read. + +## Design + +1. **Bound SWMR files by their length.** `Superblock::data_end` returns the + file's length for a version-3 superblock with the SWMR-write flag, + whatever EOF it records, as libhdf5's SWMR reader does. Every existing + open path (`File::open`, `open_storage`, `h5rs`) then reads a finished + copy of a live file. We keep opening such files without a SWMR flag + (libhdf5 refuses): it is read-only and the alternative is an error. +2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.** + - The file is read with positioned reads (`FileStorage`, `pread` on + Unix, `seek_read` on Windows), never mapped, and `len()` is the file's + current length, so reads past the length seen at open work. + - The facade's view of the file (`FileData`) is *live*: reads are not + clamped to an end fixed at open, only by the storage's current length. + - The chunk cache is not used: every read reads the chunk index and the + chunks it needs again. A cached index would hide new chunks, and a + cached partial edge chunk would read as fill where the writer has since + written data (unfiltered edge chunks are rewritten in place). + - A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would + serve stale bytes; `HttpStorage` also pins a file by ETag and length and + refuses a changed file. Remote SWMR is out of scope. +3. **`Dataset::refresh()`** reads the dataset's object header again (same + address) and replaces the handle's copy, so `shape()` and every later + read use the new extent. Like h5py, a handle that is not refreshed keeps + its extent; its reads still read the index as it is now, and only return + elements inside that extent. +4. **Bounded retries.** In a live file, an operation (open, dataset lookup, + refresh, every read) that fails with an error a concurrent write can + cause — a checksum or Fletcher-32 mismatch, a short read, a bad + signature, an undecodable chunk or chunk index — is run again from the + start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's + default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, … + up to 10 ms between attempts (under a second in all). libhdf5 retries + the one structure whose checksum failed; we retry the whole operation, + because the parsers are pure functions of the bytes they read. Results + are only returned from a run where every structure verified, so an + error is returned instead of torn data. Other errors are returned at + once, and non-live files never retry. +5. **Writer state.** `File::swmr_writer_active()` reads the superblock + flags again, so a reader can tell when the writer has closed the file. + +What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol, +not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR +writer cannot add them), and `MmapFile`/`LazyFile`. + +## Tests + +- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose + EOF is below the file length. +- `crates/clawhdf5/tests/swmr_interop.rs`: + - a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader; + - a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a + 2-D dataset with one unlimited dimension (Extensible Array, one of them + gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step, + while a Rust reader refreshes and reads them in a loop, and an h5py SWMR + reader does the same as the reference. Every value read must be the value + the writer wrote (a deterministic function of its position), extents + never shrink, and after the writer closes both readers must read the + same data as h5py. + +## Status + +See the end of this file's history in `CHANGELOG.md` ("Range reads, +milestone M5").