8.6 KiB
Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27, implemented on branch feat/p3-m5-swmr-reader
(see "Status" at the end). This is milestone M5 of
range-reads.md: "Storage::len() may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
What libhdf5 does
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
opened a file with libver='latest' and switched to SWMR mode (h5py
f.swmr_mode = True, H5Fstart_swmr_write). Readers open the same file
with H5F_ACC_SWMR_READ (h5py File(path, 'r', swmr=True)) while the
writer keeps appending. What the format and the library guarantee:
- Superblock v3, flags set. The writer sets the superblock's
file-consistency flags to write access + SWMR write (
0x05) and clears them on close. The superblock's end-of-file address is not kept up to date while writing: a copy of a file taken mid-write records an EOF of a few hundred bytes while the file is tens of kilobytes (checked on tank, 2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR reader therefore skips libhdf5's end-of-allocation check for every read (H5FD_read: "allow access to data past the end of the allocated space … for SWMR read access"), and bounds reads by the file's real length. - A non-SWMR open of such a file fails in libhdf5: "file is already open for write (may use to clear file consistency flags)".
- The writer only appends. New objects and attributes cannot be created
in SWMR mode; datasets grow with
H5Dset_extentand are written. Chunked datasets with one unlimited dimension use an Extensible Array index, with more than one a version-2 B-tree; both are updated in a SWMR-safe way. Fixed Array and single-chunk indexes are for datasets that cannot grow. - Flush ordering. Every metadata structure the writer uses is checksummed, and flush dependencies order the writes: a chunk's data is written before the index entry that points to it, and index blocks before the object header whose dataspace announces the new extent. A reader that reads the object header first and the index after sees an index at least as new as the extent, so every chunk inside the extent it read is either in the index or never written (then it reads as the fill value, as it does for libhdf5's reader).
- Refresh. A reader sees a dataset's new extent only when it refreshes
it (
H5Drefresh, h5pyDataset.refresh()), which evicts the dataset's cached metadata and reads the object header again. - Retries. Reads are not atomic against writes on every system, so a
checksum can fail when a structure is read while the writer rewrites it.
A SWMR reader reads checksummed metadata up to 100 times before failing
(
H5Pset_metadata_read_attempts; default 100 for SWMR access, 1 otherwise —H5Ppublic.hof HDF5 1.14.6).
What clawhdf5 did before
File::openof a SWMR-flagged file bounded every read by the recorded EOF (Superblock::data_endonly tolerated an EOF past the end of the file). A file copied or read mid-write therefore listed, but every chunked read failed ("unexpected EOF: need 787 bytes, have 715"), andh5rs checkreported chunk indexes "past the end of the file".- A
Fileis a snapshot: an mmap (or a buffer) of the length at open, and a per-file chunk cache that keeps each dataset's chunk index and decoded chunks for the life of theFile. A reader could not see growth at all, and a mapping of a file that is being rewritten can change under a read.
Design
- Bound SWMR files by their length.
Superblock::data_endreturns the file's length for a version-3 superblock with the SWMR-write flag, whatever EOF it records, as libhdf5's SWMR reader does. Every existing open path (File::open,open_storage,h5rs) then reads a finished copy of a live file. We keep opening such files without a SWMR flag (libhdf5 refuses): it is read-only and the alternative is an error. - A live open:
File::open_swmr(path)/File::open_storage_swmr.- The file is read with positioned reads (
FileStorage,preadon Unix,seek_readon Windows), never mapped, andlen()is the file's current length, so reads past the length seen at open work. - The facade's view of the file (
FileData) is live: reads are not clamped to an end fixed at open, only by the storage's current length. - The chunk cache is not used: every read reads the chunk index and the chunks it needs again. A cached index would hide new chunks, and a cached partial edge chunk would read as fill where the writer has since written data (unfiltered edge chunks are rewritten in place).
- A storage that caches blocks (
clawhdf5-remote'sBlockCache) would serve stale bytes;HttpStoragealso pins a file by ETag and length and refuses a changed file. Remote SWMR is out of scope.
- The file is read with positioned reads (
Dataset::refresh()reads the dataset's object header again (same address) and replaces the handle's copy, soshape()and every later read use the new extent. Like h5py, a handle that is not refreshed keeps its extent; its reads still read the index as it is now, and only return elements inside that extent.- Bounded retries. In a live file, an operation (open, dataset lookup,
refresh, every read) that fails with an error a concurrent write can
cause is run again from the start, up to
File::swmr_read_attempts()times (default 100, libhdf5's default;set_swmr_read_attemptschanges it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second in all). libhdf5 retries the one structure whose checksum failed; we retry the whole operation, because the parsers are pure functions of the bytes they read. Which errors: a garbled read can fail whichever check the parser makes first — a checksum, a signature, a version byte, a size (a test that garbles every byte of a header gotInvalidObjectHeaderVersion, not a checksum error) — so every format error counts except those later bytes cannot change: a path or selection the caller got wrong, an unsupported filter or external storage, a bad argument. Those, and a live file's permanent errors of the other kinds after the last attempt, are returned; non-live files never retry. Results are only returned from a run where every structure verified, so a torn metadata read is an error, never data.File::swmr_retries()counts the retries (libhdf5:H5Fget_metadata_read_retry_info). Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5 as here: correctness rests on the writer's ordering (chunk data before the index entry, only appends), as for libhdf5's reader. - Writer state.
File::swmr_writer_active()reads the superblock flags again, so a reader can tell when the writer has closed the file.
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
writer cannot add them), and MmapFile/LazyFile.
Tests
crates/clawhdf5-format:data_endof a SWMR-flagged v3 superblock whose EOF is below the file length.crates/clawhdf5/tests/swmr_interop.rs:- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
- a live test: an h5py writer (
swmr_mode = True) appends to a 1-D and a 2-D dataset with one unlimited dimension (Extensible Array, one of them gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step, while a Rust reader refreshes and reads them in a loop, and an h5py SWMR reader does the same as the reference. Every value read must be the value the writer wrote (a deterministic function of its position), extents never shrink, and after the writer closes both readers must read the same data as h5py.
Status
Implemented 2026-09-27 on branch feat/p3-m5-swmr-reader as designed
above (CHANGELOG.md, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, cargo test -p clawhdf5 --test swmr_interop, and once with CLAWHDF5_SWMR_STEPS=20000 in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: File::open of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; docs/known-issues.md).