Lazy remote files in the browser (M4), SWMR reader (M5), Python remote reads and editing #19

Merged
osobh merged 33 commits from feat/p3-wasm-swmr-python into main 2026-09-27 14:47:26 +00:00
2 changed files with 15 additions and 2 deletions
Showing only changes of commit 175d3a1f50 - Show all commits
+9
View File
@@ -2,6 +2,15 @@
## Unreleased
### Deterministic errors on damaged chunked datasets (2026-09-27)
- A read through the file's chunk cache listed a damaged dataset's chunks in
hash-map order, seeded per `File`, so two opens of the same file could
report different failing chunks (`cve-2025-2310.h5`, and an intermittent
failure of the storage and wasm corpus comparisons). The cache now keeps
the chunk index's own order, as the uncached readers use: every path names
the same first damaged chunk. Values of readable datasets were never
affected.
### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27)
Design: `docs/design/swmr.md`.
- **Fix: files with the SWMR-write flag were bounded by a stale end of
+6 -2
View File
@@ -961,7 +961,9 @@ cache, but:
and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map
order, so which failing chunk it reports can differ from one `File` to
the next (`cve-2025-2310.h5`); the values of a dataset that reads are
not affected.
not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in
the order the index lists them, as the uncached readers do
(`several_damaged_chunks_report_the_same_chunk_every_time`).
## Remote files (`clawhdf5-remote`) limits
@@ -1075,7 +1077,9 @@ cache, but:
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it).
not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the
chunk cache keeps the index's chunk order, so every read path names the
same (first) damaged chunk.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.