bench: concurrent-read harness against h5py threads and processes
concurrent_read reads one shared File from 1-16 threads: every dataset in full (distinct datasets per thread) and random hyperslabs of one dataset, over a deflate and a contiguous file it generates (or reuses while manifest.json matches). It reports decoded MB/s and scaling efficiency, warm or --cold (posix_fadvise) page cache, sizes the decode pool with --decode-threads, and writes JSON. scripts/concurrent_read_h5py.py runs the same workload on the same files with h5py threads or spawned processes (same splitmix64 data and slab stream, checked at spot elements), and compare_concurrent_read.py prints one table and refuses runs with different workloads. A smoke test runs all three end to end on tiny files (h5py half honours CLAWHDF5_PYTHON / CLAWHDF5_REQUIRE_INTEROP). BENCHMARKS.md gets a "Concurrent reads" section with the commands, marked not yet measured. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -34,6 +34,10 @@ path = "src/bin/consolidation_efficiency.rs"
|
||||
name = "ephemeral_perf"
|
||||
path = "src/bin/ephemeral_perf.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "concurrent_read"
|
||||
path = "src/bin/concurrent_read.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "mpi_io_bench"
|
||||
path = "src/bin/mpi_io_bench.rs"
|
||||
@@ -64,6 +68,10 @@ clawhdf5-io = { path = "../clawhdf5-io" }
|
||||
mpi = { version = "0.8", optional = true }
|
||||
serde = { workspace = true }
|
||||
serde_json = "1"
|
||||
# concurrent_read: size the decode pool (--decode-threads) and evict files
|
||||
# from the page cache (--cold, posix_fadvise). Both pure Rust / bindings only.
|
||||
rayon = "1"
|
||||
libc = "0.2"
|
||||
tempfile = { workspace = true }
|
||||
# Optional: libhdf5 C wrapper for side-by-side comparison (requires system libhdf5).
|
||||
# Enable with: cargo bench -p clawhdf5-bench --features libhdf5-compare
|
||||
|
||||
Reference in New Issue
Block a user