diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 4024ecb..c81101a 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -536,7 +536,10 @@ The rows and columns of the uncompressed layouts are within 20% (chunked column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not explain the slower windows. -### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank) +### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank, loaded) + +**Superseded** by the idle re-run below; kept as the record the default +was first decided on. Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), branch `feat/libver-v18` at `3c61635`, to decide whether the writer's default should become the @@ -598,6 +601,51 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th same; files grow by about 36 bytes per chunk. The default therefore stays the 1.10 format; the 1.8 format is opt-in. +### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank) + +Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load +average 1.66 to 1.84 at the start of each run. Stacked branch +`feat/huge-chunks` at `b55768f` (contains `feat/libver-v18`). Default and +`--v18` runs alternated, five of each per chunk size; median (min-max), ms. + +> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]` +> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset) + +| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 | +|---|---|---|---:|---:|---:| +| 256 | chunked + deflate | full (first) | 5.9 (5.5-6.2) | 6.2 (6.0-6.7) | 1.05x | +| 256 | chunked + deflate | full (repeat) | 4.4 (4.1-4.9) | 4.3 (4.2-4.4) | 0.98x | +| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 (0.15-0.17) | 0.16 (0.15-0.16) | 1.07x | +| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.23 (1.23-1.36) | 1.22 (1.22-1.25) | 0.99x | +| 256 | chunked + deflate | one row | 0.95 (0.94-1.04) | 0.94 (0.93-0.98) | 0.99x | +| 256 | chunked + deflate | one column | 1.92 (1.91-2.12) | 1.90 (1.90-1.91) | 0.99x | +| 256 | chunked | full (first) | 11.1 (10.7-11.6) | 10.9 (10.6-11.9) | 0.98x | +| 256 | chunked | full (repeat) | 5.1 (4.9-5.2) | 5.2 (4.7-5.6) | 1.02x | +| 256 | chunked | 64 x 64 window (1 chunk) | 0.03 (0.03-0.03) | 0.04 (0.04-0.04) | 1.33x | +| 256 | chunked | 512 x 512 window (4-9 chunks) | 0.36 (0.36-0.38) | 0.37 (0.36-0.38) | 1.03x | +| 256 | chunked | one row | 0.05 (0.04-0.05) | 0.05 (0.05-0.05) | 1.00x | +| 256 | chunked | one column | 0.45 (0.43-0.45) | 0.44 (0.44-0.45) | 0.98x | +| 32 | chunked + deflate | full (first) | 11.7 (11.2-12.4) | 12.5 (11.8-13.2) | 1.07x | +| 32 | chunked + deflate | full (repeat) | 9.5 (9.3-9.8) | 9.4 (9.3-11.2) | 0.99x | +| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.46 (0.45-0.49) | 0.89 (0.89-0.90) | 1.93x | +| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.95 (1.94-2.04) | 2.40 (2.39-2.51) | 1.23x | +| 32 | chunked + deflate | one row | 0.69 (0.67-0.72) | 1.15 (1.13-1.18) | 1.67x | +| 32 | chunked + deflate | one column | 1.26 (1.24-1.28) | 1.72 (1.69-1.73) | 1.37x | +| 32 | chunked | full (first) | 12.5 (11.9-12.7) | 12.6 (12.5-13.4) | 1.01x | +| 32 | chunked | full (repeat) | 5.9 (5.8-6.0) | 6.2 (6.0-6.4) | 1.05x | +| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 (0.35-0.36) | 0.84 (0.83-0.86) | 2.33x | +| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 (0.69-0.74) | 1.17 (1.16-1.17) | 1.67x | +| 32 | chunked | one row | 0.38 (0.38-0.41) | 0.85 (0.84-0.86) | 2.24x | +| 32 | chunked | one column | 1.03 (1.02-1.21) | 1.22 (1.20-1.25) | 1.18x | + +Writes: 377 vs 374 ms (128 chunks), 399 vs 409 ms (8192 chunks); file +bytes as in the loaded run. The loaded run's 0.25x and 0.52x windows at 128 +chunks were the load: idle, every 128-chunk read is within 7% except the +0.03 ms single-chunk window (one timer tick). The 8192-chunk result stands: +a selection on a freshly opened file costs 0.2 to 0.5 ms more through the +version-1 B-tree (1.2x to 2.3x), full reads are within 7%. The default stays +the 1.10 format. + ## Local file speed after range reads ### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)