is_transient_format counted every format error but a handful as transient, so on an open_swmr handle a permanent failure (a file that is not HDF5, an unsupported version or message, a truncated file) was retried 100 times, about 0.9 s of pauses per failing operation. Now only these are retried: a checksum mismatch; a read past the file's current end (UnexpectedEof; libhdf5 reads zeros there, which fail the checksum); and an object header prefix whose signature or version does not decode, which libhdf5's H5C__load_entry also retries (a header garbled whole fails there before its checksum). Everything else is returned at once. Tests: a unit test that every permanent kind returns after one call within 50 ms and every transient kind is retried to the limit; open_storage_swmr of a non-HDF5 buffer returns SignatureNotFound within 100 ms (0.87 s before) and a missing name on a live file fails without retries. The torn-read and live h5py-writer tests still pass. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
158 lines
9.6 KiB
Markdown
158 lines
9.6 KiB
Markdown
# Design: reading files a SWMR writer is still appending to (range-read M5)
|
|
|
|
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
|
|
(see "Status" at the end). This is milestone M5 of
|
|
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
|
|
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
|
|
|
|
## What libhdf5 does
|
|
|
|
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
|
|
opened a file with `libver='latest'` and switched to SWMR mode (h5py
|
|
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
|
|
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
|
|
writer keeps appending. What the format and the library guarantee:
|
|
|
|
- **Superblock v3, flags set.** The writer sets the superblock's
|
|
file-consistency flags to write access + SWMR write (`0x05`) and clears
|
|
them on close. The superblock's end-of-file address is *not* kept up to
|
|
date while writing: a copy of a file taken mid-write records an EOF of a
|
|
few hundred bytes while the file is tens of kilobytes (checked on tank,
|
|
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
|
|
reader therefore skips libhdf5's end-of-allocation check for every read
|
|
(`H5FD_read`: "allow access to data past the end of the allocated space
|
|
… for SWMR read access"), and bounds reads by the file's real length.
|
|
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
|
|
open for write (may use <h5clear file> to clear file consistency flags)".
|
|
- **The writer only appends.** New objects and attributes cannot be created
|
|
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
|
|
datasets with one unlimited dimension use an Extensible Array index, with
|
|
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
|
|
Fixed Array and single-chunk indexes are for datasets that cannot grow.
|
|
- **Flush ordering.** Every metadata structure the writer uses is
|
|
checksummed, and flush dependencies order the writes: a chunk's data is
|
|
written before the index entry that points to it, and index blocks before
|
|
the object header whose dataspace announces the new extent. A reader
|
|
that reads the object header first and the index after sees an index at
|
|
least as new as the extent, so every chunk inside the extent it read is
|
|
either in the index or never written (then it reads as the fill value, as
|
|
it does for libhdf5's reader).
|
|
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
|
|
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
|
|
cached metadata and reads the object header again.
|
|
- **Retries.** Reads are not atomic against writes on every system, so a
|
|
checksum can fail when a structure is read while the writer rewrites it.
|
|
A SWMR reader reads checksummed metadata up to 100 times before failing
|
|
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
|
|
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
|
|
|
|
## What clawhdf5 did before
|
|
|
|
- `File::open` of a SWMR-flagged file bounded every read by the recorded
|
|
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
|
|
file). A file copied or read mid-write therefore listed, but every
|
|
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
|
|
`h5rs check` reported chunk indexes "past the end of the file".
|
|
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
|
|
per-file chunk cache that keeps each dataset's chunk index and decoded
|
|
chunks for the life of the `File`. A reader could not see growth at all,
|
|
and a mapping of a file that is being rewritten can change under a read.
|
|
|
|
## Design
|
|
|
|
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
|
|
file's length for a version-3 superblock with the SWMR-write flag,
|
|
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
|
|
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
|
|
copy of a live file. We keep opening such files without a SWMR flag
|
|
(libhdf5 refuses): it is read-only and the alternative is an error.
|
|
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
|
|
- The file is read with positioned reads (`FileStorage`, `pread` on
|
|
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
|
|
current length, so reads past the length seen at open work.
|
|
- The facade's view of the file (`FileData`) is *live*: reads are not
|
|
clamped to an end fixed at open, only by the storage's current length.
|
|
- The chunk cache is not used: every read reads the chunk index and the
|
|
chunks it needs again. A cached index would hide new chunks, and a
|
|
cached partial edge chunk would read as fill where the writer has since
|
|
written data (unfiltered edge chunks are rewritten in place).
|
|
- Only a file whose superblock has the SWMR-write flag when it is opened
|
|
is read live. Any other file is read exactly as `File::open` reads it
|
|
(bounded by its recorded EOF, through the chunk cache, no retries), and
|
|
`is_swmr_read()` is `false`. libhdf5's SWMR reader is looser: it skips
|
|
the end-of-allocation check in `H5FD_read` for every file it opens,
|
|
flagged or not, yet still refuses an object header past the EOF
|
|
(`H5O_protect`, "address of object past end of allocation"). On a
|
|
closed file whose EOF is below its length (checked 2026-09-27, h5py
|
|
3.16 / HDF5 2.0, the mid-write fixture with its flags cleared) it
|
|
therefore reads a dataset whose chunk index lies past the EOF, which
|
|
its plain reader refuses. We follow the plain reader there: the EOF of
|
|
a file no SWMR writer has open is what the file says it is.
|
|
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
|
|
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
|
|
refuses a changed file. Remote SWMR is out of scope.
|
|
3. **`Dataset::refresh()`** reads the dataset's object header again (same
|
|
address) and replaces the handle's copy, so `shape()` and every later
|
|
read use the new extent. Like h5py, a handle that is not refreshed keeps
|
|
its extent; its reads still read the index as it is now, and only return
|
|
elements inside that extent.
|
|
4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
|
|
refresh, every read) that fails with an error a concurrent write can
|
|
cause is run again from the start, up to `File::swmr_read_attempts()`
|
|
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
|
|
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
|
|
in all). libhdf5 retries the one structure whose checksum failed; we
|
|
retry the whole operation, because the parsers are pure functions of the
|
|
bytes they read. Which errors: those libhdf5 retries (`H5C__load_entry`)
|
|
— a checksum mismatch, and a failure to decode the prefix it reads
|
|
before the checksum to size the structure (for an object header its
|
|
signature and version: a test that garbles every byte of a header gets
|
|
`InvalidObjectHeaderVersion`) — and a read past the file's current end
|
|
(short here; libhdf5 reads zeros there, which fail the checksum). Every
|
|
other error (an unsupported version or message, a file that is not HDF5,
|
|
a structure corrupt behind a valid checksum) is returned at once: an
|
|
earlier version retried nearly every format error, so a permanent one
|
|
cost about 0.9 s of pauses. Non-live files never retry.
|
|
Results are only returned from a run where every structure verified, so
|
|
a torn metadata read is an error, never data. `File::swmr_retries()`
|
|
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
|
|
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
|
|
as here: correctness rests on the writer's ordering (chunk data before
|
|
the index entry, only appends), as for libhdf5's reader.
|
|
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
|
|
flags again, so a reader can tell when the writer has closed the file.
|
|
|
|
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
|
|
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
|
|
writer cannot add them), and `MmapFile`/`LazyFile`.
|
|
|
|
## Tests
|
|
|
|
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
|
|
EOF is below the file length.
|
|
- `crates/clawhdf5/tests/swmr_interop.rs`:
|
|
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
|
|
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
|
|
2-D dataset with one unlimited dimension (Extensible Array, one of them
|
|
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
|
|
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
|
|
reader does the same as the reference. Every value read must be the value
|
|
the writer wrote (a deterministic function of its position), extents
|
|
never shrink, and after the writer closes both readers must read the
|
|
same data as h5py.
|
|
|
|
## Status
|
|
|
|
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
|
|
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
|
|
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
|
|
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
|
|
build): no read returned a value the writer had not written at that
|
|
position, h5py's reader agreed, and retries were needed but rare (the
|
|
failures seen were checksum mismatches, each cured by one retry). A
|
|
variant of the test with the chunk cache left on in live mode fails it
|
|
(stale chunk index / edge chunk), which is why live files do not use it.
|
|
|
|
Also found: `File::open` of such a file had been failing since the
|
|
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
|