docs: SWMR reader (range-read M5) in CHANGELOG, README, known issues and designs

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:49:04 -05:00
co-authored by Claude Opus 5.5
parent 21f104bb77
commit 3f45755d58
5 changed files with 157 additions and 13 deletions
+31 -12
View File
@@ -86,16 +86,25 @@ writer keeps appending. What the format and the library guarantee:
elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
refresh, every read) that fails with an error a concurrent write can
cause — a checksum or Fletcher-32 mismatch, a short read, a bad
signature, an undecodable chunk or chunk index — is run again from the
start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's
default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, …
up to 10 ms between attempts (under a second in all). libhdf5 retries
the one structure whose checksum failed; we retry the whole operation,
because the parsers are pure functions of the bytes they read. Results
are only returned from a run where every structure verified, so an
error is returned instead of torn data. Other errors are returned at
once, and non-live files never retry.
cause is run again from the start, up to `File::swmr_read_attempts()`
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
in all). libhdf5 retries the one structure whose checksum failed; we
retry the whole operation, because the parsers are pure functions of the
bytes they read. Which errors: a garbled read can fail whichever check
the parser makes first — a checksum, a signature, a version byte, a
size (a test that garbles every byte of a header got
`InvalidObjectHeaderVersion`, not a checksum error) — so every format
error counts except those later bytes cannot change: a path or selection
the caller got wrong, an unsupported filter or external storage, a bad
argument. Those, and a live file's permanent errors of the other kinds
after the last attempt, are returned; non-live files never retry.
Results are only returned from a run where every structure verified, so
a torn metadata read is an error, never data. `File::swmr_retries()`
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
as here: correctness rests on the writer's ordering (chunk data before
the index entry, only appends), as for libhdf5's reader.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file.
@@ -120,5 +129,15 @@ writer cannot add them), and `MmapFile`/`LazyFile`.
## Status
See the end of this file's history in `CHANGELOG.md` ("Range reads,
milestone M5").
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).