From 37770f594a2943338df7a7ad0070b28b7e100e21 Mon Sep 17 00:00:00 2001 From: osobh Date: Sat, 26 Sep 2026 08:48:05 -0500 Subject: [PATCH] docs: the rayon fix covers a one-thread pool, not the h5py-process gap The review measured the default pool unchanged (about 2900 MB/s at 16 threads before and after) and still short of 16 h5py processes; small pools still make outside readers wait. Say so instead of marking the scaling issue fixed. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 10 +++++----- CHANGELOG.md | 3 ++- docs/known-issues.md | 10 +++++++--- 3 files changed, 14 insertions(+), 9 deletions(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 5d55961..c4e1705 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -528,11 +528,11 @@ What this shows: (about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab reads, which bypass the `File`'s chunk cache, keep scaling, so the cache (one mutex and one 16 MiB budget per `File`, thrashed by 64 MiB - datasets) is the suspect. **Fixed after these measurements - (2026-09-26); the table above predates the fix and has not been - re-measured.** The cause was not the cache: with `--decode-threads 1` - every full read queued its chunks for the pool's single rayon worker; - see `docs/known-issues.md`. + datasets) is the suspect. The cause of the `--decode-threads 1` + ceiling was not the cache: every full read queued its chunks for the + pool's single rayon worker. That case was fixed after these + measurements (2026-09-26, not yet re-measured here). With the default + pool the gap to h5py processes remains (see `docs/known-issues.md`). - *Contiguous reads are slow*: 2.5 GB/s for a single-threaded full read against h5py's 9.8 GB/s (0.25x), and 0.12x for 256 x 256 hyperslabs. Threads close the gap (about 1.0x h5py at 16), but single-thread diff --git a/CHANGELOG.md b/CHANGELOG.md index 91d56f1..9030c3c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,7 +13,8 @@ (`clawhdf5_format::parallel_read::pool_can_parallelise`). The `File`'s chunk cache, the suspect in `docs/known-issues.md`, was not the cause: reads of datasets larger than its budget already skipped inserting, and - its lookups cost a few percent at 16 threads. + its lookups cost a few percent at 16 threads. Throughput with the default + pool is unchanged, and still short of an h5py process pool. ### Plugin filters (2026-09-26) - **LZF, bitshuffle, bzip2 and Blosc read and write, in pure Rust.** Files diff --git a/docs/known-issues.md b/docs/known-issues.md index a08e914..f8d7798 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -9,10 +9,10 @@ deleting it. ## Concurrent and contiguous read performance (measured 2026-09-26) -**Status:** first bullet fixed (2026-09-26), second open. Measured on +**Status:** open; one cause of the first bullet fixed (2026-09-26). Measured on tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, "Concurrent reads"): -- **Fixed 2026-09-26.** Full reads of chunked datasets from several threads +- **Partly fixed 2026-09-26.** Full reads of chunked datasets from several threads through one `File` stop scaling at about 4 threads (880 MB/s on deflate data vs 4424 MB/s for 16 h5py processes). Hyperslab reads, which skip the chunk cache, scale to 1244 MB/s, so the `File`'s shared chunk cache is the @@ -22,7 +22,11 @@ tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, worker (per-thread CPU time: one thread did all the decoding, the 16 readers almost none). Hyperslab reads touch one chunk each and never used the pool. Reads now decode on the calling thread when the pool has one - thread (`tests/single_thread_decode_pool.rs`). Datasets larger than the + thread (`tests/single_thread_decode_pool.rs`). **Still open:** this + fixes only a one-thread pool. With the default pool, 16 reader threads + ran at about 2900 MB/s before and after the change, still short of 16 + h5py processes (4424 MB/s); with a small pool (2-4 threads) readers + outside it still wait on its workers. Datasets larger than the cache's budget were already read without inserting into it, and skipping its lookups entirely gained only a few percent at 16 threads. Remaining per-read overhead, not yet addressed: each full `read_f32` of a chunked