From 175d3a1f505ea23655465e59d950a991da841efe Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 08:03:24 -0500 Subject: [PATCH] docs: chunk-cache order fix in the changelog and known issues Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 9 +++++++++ docs/known-issues.md | 8 ++++++-- 2 files changed, 15 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 16070c5..f7d44f1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,15 @@ ## Unreleased +### Deterministic errors on damaged chunked datasets (2026-09-27) +- A read through the file's chunk cache listed a damaged dataset's chunks in + hash-map order, seeded per `File`, so two opens of the same file could + report different failing chunks (`cve-2025-2310.h5`, and an intermittent + failure of the storage and wasm corpus comparisons). The cache now keeps + the chunk index's own order, as the uncached readers use: every path names + the same first damaged chunk. Values of readable datasets were never + affected. + ### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27) Design: `docs/design/swmr.md`. - **Fix: files with the SWMR-write flag were bounded by a stale end of diff --git a/docs/known-issues.md b/docs/known-issues.md index 9eeb4fc..7eea8e7 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -961,7 +961,9 @@ cache, but: and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map order, so which failing chunk it reports can differ from one `File` to the next (`cve-2025-2310.h5`); the values of a dataset that reads are - not affected. + not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in + the order the index lists them, as the uncached readers do + (`several_damaged_chunks_report_the_same_chunk_every_time`). ## Remote files (`clawhdf5-remote`) limits @@ -1075,7 +1077,9 @@ cache, but: and which chunk's error is reported depends on the iteration order of the chunk index (a `HashMap`, seeded per process), so the lazy and the range-storage reads can name different errors. Both are errors; - not specific to `openUrl` (it predates it). + not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the + chunk cache keeps the index's chunk order, so every read path names the + same (first) damaged chunk. - Compound, reference, opaque, bitfield, time and VL-sequence datasets are refused with an error naming the type; attributes of those types come back as `value: null` with their `dtype`.