docs: design for reading files a SWMR writer is appending to (range-read M5)

libhdf5's SWMR semantics as they matter to a reader (superblock v3 with
the SWMR-write flag and a stale EOF, append-only writer, flush
dependencies, refresh, 100 metadata read attempts), what clawhdf5 did
with such files, and the plan: data_end bounded by the file length for
SWMR-flagged files, a live File::open_swmr over positioned reads without
the chunk cache, Dataset::refresh, and bounded whole-operation retries.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:29:48 -05:00
co-authored by Claude Opus 5.5
parent a4c2aced55
commit 13b36e2bd7
+124
View File
@@ -0,0 +1,124 @@
# Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
(see "Status" at the end). This is milestone M5 of
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
## What libhdf5 does
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
opened a file with `libver='latest'` and switched to SWMR mode (h5py
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
writer keeps appending. What the format and the library guarantee:
- **Superblock v3, flags set.** The writer sets the superblock's
file-consistency flags to write access + SWMR write (`0x05`) and clears
them on close. The superblock's end-of-file address is *not* kept up to
date while writing: a copy of a file taken mid-write records an EOF of a
few hundred bytes while the file is tens of kilobytes (checked on tank,
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
reader therefore skips libhdf5's end-of-allocation check for every read
(`H5FD_read`: "allow access to data past the end of the allocated space
… for SWMR read access"), and bounds reads by the file's real length.
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
open for write (may use <h5clear file> to clear file consistency flags)".
- **The writer only appends.** New objects and attributes cannot be created
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
datasets with one unlimited dimension use an Extensible Array index, with
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
Fixed Array and single-chunk indexes are for datasets that cannot grow.
- **Flush ordering.** Every metadata structure the writer uses is
checksummed, and flush dependencies order the writes: a chunk's data is
written before the index entry that points to it, and index blocks before
the object header whose dataspace announces the new extent. A reader
that reads the object header first and the index after sees an index at
least as new as the extent, so every chunk inside the extent it read is
either in the index or never written (then it reads as the fill value, as
it does for libhdf5's reader).
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
cached metadata and reads the object header again.
- **Retries.** Reads are not atomic against writes on every system, so a
checksum can fail when a structure is read while the writer rewrites it.
A SWMR reader reads checksummed metadata up to 100 times before failing
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
## What clawhdf5 did before
- `File::open` of a SWMR-flagged file bounded every read by the recorded
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
file). A file copied or read mid-write therefore listed, but every
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
`h5rs check` reported chunk indexes "past the end of the file".
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
per-file chunk cache that keeps each dataset's chunk index and decoded
chunks for the life of the `File`. A reader could not see growth at all,
and a mapping of a file that is being rewritten can change under a read.
## Design
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
file's length for a version-3 superblock with the SWMR-write flag,
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
copy of a live file. We keep opening such files without a SWMR flag
(libhdf5 refuses): it is read-only and the alternative is an error.
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
- The file is read with positioned reads (`FileStorage`, `pread` on
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
current length, so reads past the length seen at open work.
- The facade's view of the file (`FileData`) is *live*: reads are not
clamped to an end fixed at open, only by the storage's current length.
- The chunk cache is not used: every read reads the chunk index and the
chunks it needs again. A cached index would hide new chunks, and a
cached partial edge chunk would read as fill where the writer has since
written data (unfiltered edge chunks are rewritten in place).
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
refuses a changed file. Remote SWMR is out of scope.
3. **`Dataset::refresh()`** reads the dataset's object header again (same
address) and replaces the handle's copy, so `shape()` and every later
read use the new extent. Like h5py, a handle that is not refreshed keeps
its extent; its reads still read the index as it is now, and only return
elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
refresh, every read) that fails with an error a concurrent write can
cause — a checksum or Fletcher-32 mismatch, a short read, a bad
signature, an undecodable chunk or chunk index — is run again from the
start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's
default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, …
up to 10 ms between attempts (under a second in all). libhdf5 retries
the one structure whose checksum failed; we retry the whole operation,
because the parsers are pure functions of the bytes they read. Results
are only returned from a run where every structure verified, so an
error is returned instead of torn data. Other errors are returned at
once, and non-live files never retry.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file.
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
writer cannot add them), and `MmapFile`/`LazyFile`.
## Tests
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
EOF is below the file length.
- `crates/clawhdf5/tests/swmr_interop.rs`:
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
2-D dataset with one unlimited dimension (Extensible Array, one of them
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
reader does the same as the reference. Every value read must be the value
the writer wrote (a deterministic function of its position), extents
never shrink, and after the writer closes both readers must read the
same data as h5py.
## Status
See the end of this file's history in `CHANGELOG.md` ("Range reads,
milestone M5").