Merge pull request 'docs(clawhdf5): document DType variants, fix unresolved doc links' (#17) from sdlc-docs/clawhdf5-types-20260514-165210 into main

This commit is contained in:
redclawsystems
2026-05-14 23:54:48 +00:00
commit 3f222f6956
3030 changed files with 89917 additions and 0 deletions
+138
View File
@@ -0,0 +1,138 @@
# Benchmarks: Oracle Server (2026-03-01)
## Hardware
- CPU: Intel Xeon E5-2697 v2 @ 2.70GHz (48 cores)
- RAM: 247 GB
- OS: Ubuntu, Linux 6.14.0-37-generic x86_64
## Software
- rustyhdf5: commit 1b12675 (nightly 1.95.0)
- h5py 3.15.1 / HDF5 1.14.6 / numpy 2.4.2 / Python 3.13.3
## Results (1M float64, 8 MB)
| Benchmark | RustyHDF5 | h5py (C HDF5) | Speedup |
|-----------|-----------|----------------|---------|
| **Write contiguous** | 24.3 ms | 4.3 ms | 0.18x (h5py faster) |
| **Write chunked** | 37.0 ms | 6.4 ms | 0.17x (h5py faster) |
| **Write chunked deflate** | 331.8 ms | 551.9 ms | 1.66x |
| **Read contiguous** | 5.74 ms | 3.31 ms | 0.58x (h5py faster) |
| **Read chunked** | 6.51 ms | 4.14 ms | 0.63x (h5py faster) |
| **Read chunked deflate** | 24.1 ms | 21.8 ms | 0.91x (comparable) |
| **Read 50 attrs** | 16.8 µs | 8.25 ms | 491x |
| **Group nav 100** | 17.9 µs | 1.31 ms | 73x |
## Other RustyHDF5 benchmarks (no h5py equivalent)
- write_dataset_20_attrs_dense: 33.8 µs
- read_dataset_20_attrs_dense: 7.2 µs
- parse_object_header_complex: 367 ns
- write_10K_string_attrs: 161.1 µs
- read_100_string_attrs: 31.1 µs
- write_compound_10K_rows: 76.9 µs
- read_compound_10K_rows: 17.3 µs
- jenkins_lookup3_4MB: 2.80 ms
- sha256_4MB: 18.97 ms
- roundtrip_contiguous: 31.7 ms
- roundtrip_chunked_deflate: 354.9 ms
- write_provenance: 58.7 ms
## Notes
- h5py write includes kernel I/O (tmpfile); RustyHDF5 is in-memory buffer
- h5py read includes file open/close overhead
- Metadata operations (attrs, group nav) show massive RustyHDF5 advantage (in-memory parsing)
- Deflate compression is CPU-bound and comparable between both
- Xeon E5-2697 v2 is older (Ivy Bridge-EP, 2013) — no AVX-512, slower single-thread than M3 Max
## Mmap & I/O Strategy Benchmarks
### RustyHDF5 Mmap vs FileReader (1M f64)
| Benchmark | Time |
|-----------|------|
| filereader contiguous | 16.6 ms |
| **mmapreader contiguous** | **16.5 ms** |
| filereader chunked | 22.1 ms |
| **mmapreader chunked** | **7.3 ms** (3x faster) |
| File::open mmap read | 6.5 ms |
| File::open buffered read | 6.7 ms |
| **Zero-copy raw ref** | **601 ns** |
| **Zero-copy f64 slice** | **627 ns** |
| **Zero-copy f64 mmap** | **622 ns** |
| read_f64 (copy) | 5.8 ms |
| **read_f64 zerocopy** | **618 ns** (~9,400x faster) |
| read_as_slice f64 | 610 ns |
### File Open Overhead (10MB file)
| Method | Time |
|--------|------|
| mmap open only | 22.4 µs |
| buffered open only | 1.09 ms (49x slower) |
### Lazy vs Eager (100 datasets, read 1)
| Method | RustyHDF5 | h5py |
|--------|-----------|------|
| Eager open + read 1 | 40.8 µs | 0.66 ms (16x slower) |
| Lazy/mmap open + read 1 | 52.0 µs | — |
### Prefetch (chunked 1M f64, in-memory)
| Method | Time |
|--------|------|
| No prefetch | 7.58 ms |
| With prefetch | 7.52 ms (marginal — data already in memory) |
### h5py I/O Strategies
| Benchmark | Time |
|-----------|------|
| Default driver read | 4.0 ms |
| Core driver (in-memory) | 10.1 ms (slower — full copy) |
| 10 datasets sequential | 44.2 ms |
| 10 datasets threaded (10 workers) | 62.2 ms (GIL bottleneck) |
| 1-of-100 datasets | 0.66 ms |
### Parallel I/O (File::open vs MmapFile::open)
| Method | Time |
|--------|------|
| File::open 1M f64 | 12.1 ms |
| MmapFile::open 1M f64 | 12.2 ms |
## Key Takeaways (Oracle Xeon)
1. **Zero-copy is the killer feature**: 618ns vs 5.8ms copy vs 4.0ms h5py — ~6,500x faster than h5py
2. **Mmap chunked reads 3x faster** than FileReader (OS page cache does the work)
3. **Mmap file open is 49x faster** than buffered (no data copy, just page table setup)
4. **h5py threading hurts** due to GIL — RustyHDF5's Rust-native parallelism has no such limitation
5. **Eager open at 40.8µs** vs h5py's 0.66ms = 16x faster for selective dataset access
## Rayon Parallel Decompression Scaling (10M f64, 1000 chunks, deflate-6)
80MB uncompressed data, lane-partitioned parallel decompression.
| Threads | Median (ms) | Min (ms) | Speedup vs 1T |
|---------|-------------|----------|---------------|
| 1 | 449.5 | 426.2 | 1.0x |
| 2 | 287.4 | 279.0 | 1.56x |
| 4 | 216.0 | 214.1 | 2.08x |
| 8 | 232.5 | 186.2 | 1.93x (2.29x min) |
| 16 | 220.7 | 177.9 | 2.04x (2.40x min) |
| 24 | 206.7 | 167.4 | 2.17x (2.55x min) |
| 32 | 208.2 | 167.4 | 2.16x (2.55x min) |
| 48 | 208.6 | 165.2 | 2.15x (2.58x min) |
| no-parallel feature | 335.0 | 332.5 | 1.34x (sequential codepath) |
### h5py comparison (GIL-limited)
| Method | Median (ms) |
|--------|-------------|
| h5py sequential deflate read (1M) | 21.8 |
| h5py threaded 10 datasets | 62.2 (slower than sequential!) |
### Analysis
- **Peak scaling ~2.6x** at 48 threads (min times), saturating around 8 cores
- Diminishing returns after 4 threads — deflate decompression is memory-bandwidth limited on this Xeon
- The Xeon E5-2697 v2 has 2 sockets × 12 cores with shared L3; NUMA effects likely cause saturation
- **Sequential no-parallel (335ms) vs parallel-1T (449ms)**: the parallel codepath has ~34% overhead from lane partitioning setup when only using 1 thread
- **vs h5py**: RustyHDF5 parallel at 48T decompresses 10M elements in ~208ms; h5py can't parallelize at all due to GIL
- At scale (100M+ elements), the parallelism advantage would compound further