144 lines
8.6 KiB
Markdown
144 lines
8.6 KiB
Markdown
# Design: reading files a SWMR writer is still appending to (range-read M5)
|
|
|
|
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
|
|
(see "Status" at the end). This is milestone M5 of
|
|
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
|
|
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
|
|
|
|
## What libhdf5 does
|
|
|
|
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
|
|
opened a file with `libver='latest'` and switched to SWMR mode (h5py
|
|
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
|
|
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
|
|
writer keeps appending. What the format and the library guarantee:
|
|
|
|
- **Superblock v3, flags set.** The writer sets the superblock's
|
|
file-consistency flags to write access + SWMR write (`0x05`) and clears
|
|
them on close. The superblock's end-of-file address is *not* kept up to
|
|
date while writing: a copy of a file taken mid-write records an EOF of a
|
|
few hundred bytes while the file is tens of kilobytes (checked on tank,
|
|
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
|
|
reader therefore skips libhdf5's end-of-allocation check for every read
|
|
(`H5FD_read`: "allow access to data past the end of the allocated space
|
|
… for SWMR read access"), and bounds reads by the file's real length.
|
|
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
|
|
open for write (may use <h5clear file> to clear file consistency flags)".
|
|
- **The writer only appends.** New objects and attributes cannot be created
|
|
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
|
|
datasets with one unlimited dimension use an Extensible Array index, with
|
|
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
|
|
Fixed Array and single-chunk indexes are for datasets that cannot grow.
|
|
- **Flush ordering.** Every metadata structure the writer uses is
|
|
checksummed, and flush dependencies order the writes: a chunk's data is
|
|
written before the index entry that points to it, and index blocks before
|
|
the object header whose dataspace announces the new extent. A reader
|
|
that reads the object header first and the index after sees an index at
|
|
least as new as the extent, so every chunk inside the extent it read is
|
|
either in the index or never written (then it reads as the fill value, as
|
|
it does for libhdf5's reader).
|
|
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
|
|
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
|
|
cached metadata and reads the object header again.
|
|
- **Retries.** Reads are not atomic against writes on every system, so a
|
|
checksum can fail when a structure is read while the writer rewrites it.
|
|
A SWMR reader reads checksummed metadata up to 100 times before failing
|
|
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
|
|
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
|
|
|
|
## What clawhdf5 did before
|
|
|
|
- `File::open` of a SWMR-flagged file bounded every read by the recorded
|
|
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
|
|
file). A file copied or read mid-write therefore listed, but every
|
|
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
|
|
`h5rs check` reported chunk indexes "past the end of the file".
|
|
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
|
|
per-file chunk cache that keeps each dataset's chunk index and decoded
|
|
chunks for the life of the `File`. A reader could not see growth at all,
|
|
and a mapping of a file that is being rewritten can change under a read.
|
|
|
|
## Design
|
|
|
|
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
|
|
file's length for a version-3 superblock with the SWMR-write flag,
|
|
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
|
|
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
|
|
copy of a live file. We keep opening such files without a SWMR flag
|
|
(libhdf5 refuses): it is read-only and the alternative is an error.
|
|
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
|
|
- The file is read with positioned reads (`FileStorage`, `pread` on
|
|
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
|
|
current length, so reads past the length seen at open work.
|
|
- The facade's view of the file (`FileData`) is *live*: reads are not
|
|
clamped to an end fixed at open, only by the storage's current length.
|
|
- The chunk cache is not used: every read reads the chunk index and the
|
|
chunks it needs again. A cached index would hide new chunks, and a
|
|
cached partial edge chunk would read as fill where the writer has since
|
|
written data (unfiltered edge chunks are rewritten in place).
|
|
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
|
|
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
|
|
refuses a changed file. Remote SWMR is out of scope.
|
|
3. **`Dataset::refresh()`** reads the dataset's object header again (same
|
|
address) and replaces the handle's copy, so `shape()` and every later
|
|
read use the new extent. Like h5py, a handle that is not refreshed keeps
|
|
its extent; its reads still read the index as it is now, and only return
|
|
elements inside that extent.
|
|
4. **Bounded retries.** In a live file, an operation (open, dataset lookup,
|
|
refresh, every read) that fails with an error a concurrent write can
|
|
cause is run again from the start, up to `File::swmr_read_attempts()`
|
|
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
|
|
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
|
|
in all). libhdf5 retries the one structure whose checksum failed; we
|
|
retry the whole operation, because the parsers are pure functions of the
|
|
bytes they read. Which errors: a garbled read can fail whichever check
|
|
the parser makes first — a checksum, a signature, a version byte, a
|
|
size (a test that garbles every byte of a header got
|
|
`InvalidObjectHeaderVersion`, not a checksum error) — so every format
|
|
error counts except those later bytes cannot change: a path or selection
|
|
the caller got wrong, an unsupported filter or external storage, a bad
|
|
argument. Those, and a live file's permanent errors of the other kinds
|
|
after the last attempt, are returned; non-live files never retry.
|
|
Results are only returned from a run where every structure verified, so
|
|
a torn metadata read is an error, never data. `File::swmr_retries()`
|
|
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
|
|
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
|
|
as here: correctness rests on the writer's ordering (chunk data before
|
|
the index entry, only appends), as for libhdf5's reader.
|
|
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
|
|
flags again, so a reader can tell when the writer has closed the file.
|
|
|
|
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
|
|
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
|
|
writer cannot add them), and `MmapFile`/`LazyFile`.
|
|
|
|
## Tests
|
|
|
|
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
|
|
EOF is below the file length.
|
|
- `crates/clawhdf5/tests/swmr_interop.rs`:
|
|
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
|
|
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
|
|
2-D dataset with one unlimited dimension (Extensible Array, one of them
|
|
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
|
|
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
|
|
reader does the same as the reference. Every value read must be the value
|
|
the writer wrote (a deterministic function of its position), extents
|
|
never shrink, and after the writer closes both readers must read the
|
|
same data as h5py.
|
|
|
|
## Status
|
|
|
|
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
|
|
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
|
|
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
|
|
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
|
|
build): no read returned a value the writer had not written at that
|
|
position, h5py's reader agreed, and retries were needed but rare (the
|
|
failures seen were checksum mismatches, each cured by one retry). A
|
|
variant of the test with the chunk cache left on in live mode fails it
|
|
(stale chunk index / edge chunk), which is why live files do not use it.
|
|
|
|
Also found: `File::open` of such a file had been failing since the
|
|
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
|