Files
clawhdf5/docs/design/swmr.md
T
osobhandClaude Opus 5.5 ae6b960453 facade: open_swmr reads live only files with the SWMR-write flag
A file opened with open_swmr whose superblock does not have the SWMR-write
flag (its writer has closed it) is now read exactly as File::open reads it:
bounded by its recorded end of file, through the chunk cache, each
operation tried once, and is_swmr_read() is false. Before, every file opened
with open_swmr ignored its recorded end of file, so a closed file whose end
of file is below its length (h5clear_fsm_persist_less.h5, or the mid-write
fixture with its flags cleared) listed and read objects that File::open and
libhdf5's plain reader refuse. Opening such a file is not retried once the
superblock has been read.

libhdf5's SWMR reader is looser than either (it reads the fixture's chunk
index past the end of file even with the flags cleared); a test pins what
h5py does and docs/design/swmr.md says why we follow the plain reader.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:31:57 -05:00

156 lines
9.5 KiB
Markdown

# Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
(see "Status" at the end). This is milestone M5 of
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
## What libhdf5 does
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
opened a file with `libver='latest'` and switched to SWMR mode (h5py
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
writer keeps appending. What the format and the library guarantee:
- **Superblock v3, flags set.** The writer sets the superblock's
file-consistency flags to write access + SWMR write (`0x05`) and clears
them on close. The superblock's end-of-file address is *not* kept up to
date while writing: a copy of a file taken mid-write records an EOF of a
few hundred bytes while the file is tens of kilobytes (checked on tank,
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
reader therefore skips libhdf5's end-of-allocation check for every read
(`H5FD_read`: "allow access to data past the end of the allocated space
… for SWMR read access"), and bounds reads by the file's real length.
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
open for write (may use <h5clear file> to clear file consistency flags)".
- **The writer only appends.** New objects and attributes cannot be created
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
datasets with one unlimited dimension use an Extensible Array index, with
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
Fixed Array and single-chunk indexes are for datasets that cannot grow.
- **Flush ordering.** Every metadata structure the writer uses is
checksummed, and flush dependencies order the writes: a chunk's data is
written before the index entry that points to it, and index blocks before
the object header whose dataspace announces the new extent. A reader
that reads the object header first and the index after sees an index at
least as new as the extent, so every chunk inside the extent it read is
either in the index or never written (then it reads as the fill value, as
it does for libhdf5's reader).
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
cached metadata and reads the object header again.
- **Retries.** Reads are not atomic against writes on every system, so a
checksum can fail when a structure is read while the writer rewrites it.
A SWMR reader reads checksummed metadata up to 100 times before failing
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
## What clawhdf5 did before
- `File::open` of a SWMR-flagged file bounded every read by the recorded
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
file). A file copied or read mid-write therefore listed, but every
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
`h5rs check` reported chunk indexes "past the end of the file".
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
per-file chunk cache that keeps each dataset's chunk index and decoded
chunks for the life of the `File`. A reader could not see growth at all,
and a mapping of a file that is being rewritten can change under a read.
## Design
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
file's length for a version-3 superblock with the SWMR-write flag,
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
copy of a live file. We keep opening such files without a SWMR flag
(libhdf5 refuses): it is read-only and the alternative is an error.
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
- The file is read with positioned reads (`FileStorage`, `pread` on
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
current length, so reads past the length seen at open work.
- The facade's view of the file (`FileData`) is *live*: reads are not
clamped to an end fixed at open, only by the storage's current length.
- The chunk cache is not used: every read reads the chunk index and the
chunks it needs again. A cached index would hide new chunks, and a
cached partial edge chunk would read as fill where the writer has since
written data (unfiltered edge chunks are rewritten in place).
- Only a file whose superblock has the SWMR-write flag when it is opened
is read live. Any other file is read exactly as `File::open` reads it
(bounded by its recorded EOF, through the chunk cache, no retries), and
`is_swmr_read()` is `false`. libhdf5's SWMR reader is looser: it skips
the end-of-allocation check in `H5FD_read` for every file it opens,
flagged or not, yet still refuses an object header past the EOF
(`H5O_protect`, "address of object past end of allocation"). On a
closed file whose EOF is below its length (checked 2026-09-27, h5py
3.16 / HDF5 2.0, the mid-write fixture with its flags cleared) it
therefore reads a dataset whose chunk index lies past the EOF, which
its plain reader refuses. We follow the plain reader there: the EOF of
a file no SWMR writer has open is what the file says it is.
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
refuses a changed file. Remote SWMR is out of scope.
3. **`Dataset::refresh()`** reads the dataset's object header again (same
address) and replaces the handle's copy, so `shape()` and every later
read use the new extent. Like h5py, a handle that is not refreshed keeps
its extent; its reads still read the index as it is now, and only return
elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
refresh, every read) that fails with an error a concurrent write can
cause is run again from the start, up to `File::swmr_read_attempts()`
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
in all). libhdf5 retries the one structure whose checksum failed; we
retry the whole operation, because the parsers are pure functions of the
bytes they read. Which errors: a garbled read can fail whichever check
the parser makes first — a checksum, a signature, a version byte, a
size (a test that garbles every byte of a header got
`InvalidObjectHeaderVersion`, not a checksum error) — so every format
error counts except those later bytes cannot change: a path or selection
the caller got wrong, an unsupported filter or external storage, a bad
argument. Those, and a live file's permanent errors of the other kinds
after the last attempt, are returned; non-live files never retry.
Results are only returned from a run where every structure verified, so
a torn metadata read is an error, never data. `File::swmr_retries()`
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
as here: correctness rests on the writer's ordering (chunk data before
the index entry, only appends), as for libhdf5's reader.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file.
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
writer cannot add them), and `MmapFile`/`LazyFile`.
## Tests
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
EOF is below the file length.
- `crates/clawhdf5/tests/swmr_interop.rs`:
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
2-D dataset with one unlimited dimension (Extensible Array, one of them
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
reader does the same as the reference. Every value read must be the value
the writer wrote (a deterministic function of its position), extents
never shrink, and after the writer closes both readers must read the
same data as h5py.
## Status
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).