Files
clawhdf5/docs/design/swmr.md
T
osobhandClaude Opus 5.5 ac0020594b docs: design status notes reflect what is merged
range-reads.md opens with a table of milestones M0-M5 and the PRs that
merged them (#17-#21), replacing a header left garbled by earlier
merges, and each milestone's status names its PR. swmr.md says the
reader is merged (PR #19) and the writer does not exist. openclaw.md
links the Node package's known-issues entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:05:38 -05:00

184 lines
11 KiB
Markdown

# Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27; the reader is implemented and merged (branch
`feat/p3-m5-swmr-reader`, PR #19, `7a8fae0`; see "Status" at the end).
clawhdf5 has no SWMR writer. This is milestone M5 of
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
## What libhdf5 does
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
opened a file with `libver='latest'` and switched to SWMR mode (h5py
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
writer keeps appending. What the format and the library guarantee:
- **Superblock v3, flags set.** The writer sets the superblock's
file-consistency flags to write access + SWMR write (`0x05`) and clears
them on close. The superblock's end-of-file address is *not* kept up to
date while writing: a copy of a file taken mid-write records an EOF of a
few hundred bytes while the file is tens of kilobytes (checked on tank,
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
reader therefore skips libhdf5's end-of-allocation check for every read
(`H5FD_read`: "allow access to data past the end of the allocated space
… for SWMR read access"), and bounds reads by the file's real length.
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
open for write (may use <h5clear file> to clear file consistency flags)".
- **The writer only appends.** New objects and attributes cannot be created
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
datasets with one unlimited dimension use an Extensible Array index, with
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
Fixed Array and single-chunk indexes are for datasets that cannot grow.
- **Flush ordering.** Every metadata structure the writer uses is
checksummed, and flush dependencies order the writes: a chunk's data is
written before the index entry that points to it, and index blocks before
the object header whose dataspace announces the new extent. A reader
that reads the object header first and the index after sees an index at
least as new as the extent, so every chunk inside the extent it read is
either in the index or never written (then it reads as the fill value, as
it does for libhdf5's reader).
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
cached metadata and reads the object header again.
- **Retries.** Reads are not atomic against writes on every system, so a
checksum can fail when a structure is read while the writer rewrites it.
A SWMR reader reads checksummed metadata up to 100 times before failing
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
## What clawhdf5 did before
- `File::open` of a SWMR-flagged file bounded every read by the recorded
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
file). A file copied or read mid-write therefore listed, but every
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
`h5rs check` reported chunk indexes "past the end of the file".
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
per-file chunk cache that keeps each dataset's chunk index and decoded
chunks for the life of the `File`. A reader could not see growth at all,
and a mapping of a file that is being rewritten can change under a read.
## Design
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
file's length for a version-3 superblock with the SWMR-write flag,
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
copy of a live file. We keep opening such files without a SWMR flag
(libhdf5 refuses): it is read-only and the alternative is an error.
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
- The file is read with positioned reads (`FileStorage`, `pread` on
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
current length, so reads past the length seen at open work.
- The facade's view of the file (`FileData`) is *live*: reads are not
clamped to an end fixed at open, only by the storage's current length.
- The chunk cache is not used: every read reads the chunk index and the
chunks it needs again. A cached index would hide new chunks, and a
cached partial edge chunk would read as fill where the writer has since
written data (unfiltered edge chunks are rewritten in place).
- Only a file whose superblock has the SWMR-write flag when it is opened
is read live. Any other file is read exactly as `File::open` reads it
(bounded by its recorded EOF, through the chunk cache, no retries), and
`is_swmr_read()` is `false`. libhdf5's SWMR reader is looser: it skips
the end-of-allocation check in `H5FD_read` for every file it opens,
flagged or not, yet still refuses an object header past the EOF
(`H5O_protect`, "address of object past end of allocation"). On a
closed file whose EOF is below its length (checked 2026-09-27, h5py
3.16 / HDF5 2.0, the mid-write fixture with its flags cleared) it
therefore reads a dataset whose chunk index lies past the EOF, which
its plain reader refuses. We follow the plain reader there: the EOF of
a file no SWMR writer has open is what the file says it is.
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
refuses a changed file. Remote SWMR is out of scope.
3. **`Dataset::refresh()`** reads the dataset's object header again (same
address) and replaces the handle's copy, so `shape()` and every later
read use the new extent. Like h5py, a handle that is not refreshed keeps
its extent; its reads still read the index as it is now, and only return
elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, lookups and
listings, refresh, every dataset read — typed, raw, selections, strings
and variable-length data with their global-heap decoding — attributes,
`File::decode_*`, `verify_provenance`) that fails with an error a
concurrent write can
cause is run again from the start, up to `File::swmr_read_attempts()`
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
in all). libhdf5 retries the one structure whose checksum failed; we
retry the whole operation, because the parsers are pure functions of the
bytes they read. Which errors: those libhdf5 retries (`H5C__load_entry`)
— a checksum mismatch, and a failure to decode the prefix it reads
before the checksum to size the structure (for an object header its
signature and version: a test that garbles every byte of a header gets
`InvalidObjectHeaderVersion`) — and a read past the file's current end
(short here; libhdf5 reads zeros there, which fail the checksum). Every
other error (an unsupported version or message, a file that is not HDF5,
a structure corrupt behind a valid checksum) is returned at once: an
earlier version retried nearly every format error, so a permanent one
cost about 0.9 s of pauses. Non-live files never retry.
Results are only returned from a run where every structure verified, so
a torn metadata read is an error, never data. `File::swmr_retries()`
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as `AttrValue::Raw`) instead of failing;
on a live file such an error of a retried kind runs the read again too,
and after the last attempt the last result is returned as before. The
zero-copy reads (`read_raw_ref`, `read_as_slice`, `read_*_zerocopy`)
need the file in memory, which a live file never is: they report
`None` / `ContiguousStorageRequired` without reading data.
Global heap collections (variable-length data) have no checksum in
HDF5, like raw data, so a torn read of one is only caught when it fails
a check (its signature, a bound).
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
as here: correctness rests on the writer's ordering (chunk data before
the index entry, only appends), as for libhdf5's reader.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file.
A writer that crashed or was killed never clears the flag, so a reader
that follows a file needs a second stop condition (the README example
stops after a minute without growth).
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
writer cannot add them), and `MmapFile`/`LazyFile`.
## Tests
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
EOF is below the file length.
- `crates/clawhdf5/tests/swmr_interop.rs`:
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
- the same bytes with the flags cleared read as `File::open` reads them
(and how h5py's two readers read them);
- a permanent error returns at once; every read path (listings,
attributes, strings, variable-length data) with each of its reads
failing once in turn returns the same result;
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
2-D dataset with one unlimited dimension (Extensible Array, one of them
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
reader does the same as the reference. Every value read must be the value
the writer wrote (a deterministic function of its position), extents
never shrink, and after the writer closes both readers must read the
same data as h5py.
## Status
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
above, merged to `main` in PR #19 (`7a8fae0`) (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; fixed before any release, see
[`docs/known-issues.md`](../known-issues.md#files-a-swmr-writer-had-open-could-not-be-read-past-a-stale-end-of-file)).
Not done (tracked in `docs/known-issues.md`, "Range reads" limits): SWMR
writing, remote SWMR, and live reading through `MmapFile`/`LazyFile`.