Compare commits

...
Author SHA1 Message Date
osobhandClaude Opus 5.5 4b075656e7 docs: known-issues — cite PRs #27, #28, #29
CI / test-arm64 (pull_request) Successful in 1m36s
CI / test (pull_request) Successful in 16m57s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 00:00:28 -05:00
osobhandClaude Opus 5.5 23e13b3d9c BENCHMARKS: idle re-run of the HDF5 1.8 format A/B (supersedes the loaded run)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:47:01 -05:00
osobhandClaude Opus 5.5 b55768ff76 chunked writer: chunks of 4 GiB or more never use the v1 B-tree; need high bound V200
Stacking feat/libver-v18 under feat/huge-chunks sent every chunked dataset
of a file with a 1.8 low bound to a version-1 B-tree, whose key holds a
32-bit chunk size. As in libhdf5, such a chunk now takes layout version 5
whatever the low bound, and a high bound below 2.0 is FormatError::LibverBound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:46:04 -05:00
osobhandClaude Opus 5.5 740d1124a8 docs: chunks of 4 GiB or more — CHANGELOG, known-issues, README
Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open
limits entry; the selection-read entry loses its fill-value case; the
FileEditor limits gain the refused rewrites; the README tables name what
is supported and what is not.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:55 -05:00
osobhandClaude Opus 5.5 7a2e61ebd1 Huge chunks: selection tests with fill values, LZ4 over 256 MiB, wasm32 check
A selection of a chunked dataset with a fill value is compared with the
full read in every index, through a map and through positioned reads. An
LZ4 chunk larger than 256 MiB is bounded by the chunk size, not refused.
The wasm package test reads the 4 GiB-chunk fixture and gets a clean
error in every index.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
osobhandClaude Opus 5.5 5a20cf04e8 Write chunks of 4 GiB or more; FileEditor refuses to rewrite them
The writer gives a chunk of more than u32::MAX bytes layout message
version 5 and, filtered, index elements whose stored size takes the file's
size of lengths, as libhdf5 2.x does (the Fixed and Extensible Array
structures match libhdf5's byte for byte). Chunk dimensions of 2^32 or more
and filters that cannot take such a chunk (LZF, bitshuffle, bzip2, Blosc,
pcodec) are refused instead of truncated. Chunks are extracted row by row
and one at a time; deflate no longer cuts input at 4 GiB - 1 bytes, nor
holds the worst-case bound of a large chunk; an LZ4 chunk of 4 GiB or more
is read as the registered framing.

FileEditor refuses writing values into, or pruning/allocating, chunks of
4 GiB or more before anything is written; growing the extent and
attributes still work.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
osobhandClaude Opus 5.5 ac7871fff5 Read chunks of 4 GiB or more in every chunk index
ChunkInfo::chunk_size (and ChunkMapping::file_size) are u64: sizes past
u32 were truncated for Single Chunk, Implicit, Fixed and Extensible Array
indexes, and a v2 B-tree index refused them. A selection of a chunked
dataset with a non-default fill value is read over a box of fill values
instead of a full read, an unfiltered chunk of a file that is not in memory
is read row by row, and an intermediate deflate stage no longer reserves the
chunk's whole bound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
osobhandClaude Opus 5.5 01d2a5dc5d docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
osobhandClaude Opus 5.5 b5a5041655 Write files HDF5 1.8 can read: FileWriter/FileBuilder::libver_bounds
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.

With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).

The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.

Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
osobhandClaude Opus 5.5 00b6f76ee0 clawhdf5-netcdf4: variables' dimensions come from the file
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m24s
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.

variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.

Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 22:52:04 -05:00
osobh f9edf4d6ad Merge pull request 'Write complex numbers, incl. HDF5 2.0 native complex (class 11)' (#26) from feat/write-complex into main
CI / test-arm64 (push) Successful in 1m33s
CI / test (push) Successful in 23m27s
Reviewed-on: #26
2026-09-29 03:11:22 +00:00
osobh 56684ff147 Merge pull request 'HDF5 2.x small floats (FP4/FP6/FP8/bfloat16) checked against libhdf5 2.2.0' (#25) from feat/mx-floats into main
CI / test-arm64 (push) Successful in 1m32s
CI / test (push) Successful in 15m40s
Reviewed-on: #25
2026-09-29 03:11:09 +00:00
osobh b6bbe6b604 Merge pull request 'clawhdf5-netcdf4: unlimited dimensions report their length' (#24) from fix/netcdf-unlimited-dim-size into main
CI / test-arm64 (push) Successful in 1m34s
CI / test (push) Successful in 23m33s
Reviewed-on: #24
2026-09-29 03:10:56 +00:00
osobh d279ee06a2 Merge pull request 'clawhdf5: a dropped FileEditor releases its lock at once' (#23) from fix/editor-lock-fork-race into main
CI / test-arm64 (push) Successful in 1m40s
CI / test (push) Successful in 17m50s
Reviewed-on: #23
2026-09-29 03:10:43 +00:00
osobhandClaude Opus 5.5 c033660d0d docs: known-issues — cite PRs #23 and #24
CI / test-arm64 (pull_request) Successful in 1m37s
CI / test (pull_request) Successful in 18m1s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:38:37 -05:00
osobhandClaude Opus 5.5 5e9160f82c docs: README — one datatypes row for complex writing and small floats
CI / test-arm64 (pull_request) Successful in 1m36s
CI / test (pull_request) Successful in 31m25s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:25:12 -05:00
osobhandClaude Opus 5.5 bce07e9cb9 feat: check HDF5 2.x small floats against libhdf5 2.2.0
CI / test-arm64 (pull_request) Successful in 1m38s
CI / test (pull_request) Successful in 31m22s
Fixture written by libhdf5 2.2.0 (built from tag 2.2.0) through ctypes:
every bit pattern of FP4 E2M1, FP6 E2M3/E3M2, FP8 E4M3/E5M2 and a
bfloat16 LE/BE set, as datasets and attributes, with what H5Dread/H5Aread
return into double and float and the conversion exceptions libhdf5
raises. clawhdf5 already decoded every value as libhdf5 does, including
an all-ones exponent as inf/NaN in the OCP formats that have none
(documented as a deliberate match in known-issues).

- data_read: NaNs of non-native float layouts get libhdf5's bits (sign
  kept, every mantissa bit set) in f64 and f32.
- h5rs dump/ls name these types as h5dump/h5ls 2.x do
  (H5T_FLOAT_F4E2M1, "FP4 E2M1 4-bit float", float4-e2m1 ...), checked
  against h5dump 2.2.0's output of the fixture.
- Python bindings read them as h5py 3.16 does (float32 for bfloat16,
  float16 for the 1-byte formats, file byte order, same bytes as h5py);
  writing them in 'r+' is refused.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:25:03 -05:00
osobhandClaude Opus 5.5 d0d5347cd9 feat: write complex numbers, incl. HDF5 2.0 native complex (class 11)
- Datatype::Complex serializes class 11 version 5 byte-identically to
  libhdf5 2.2.0; containers holding it are written as version 5.
- DatasetBuilder::with_complex_f32/f64_data (h5py's {r, i} compound,
  default) and with_native_complex_f32/f64_data (class 11, opt-in);
  make_(native_)complex_f32/f64_type for attributes.
- Dataset::read_complex_f64/f32 read either form.
- Python create_dataset accepts complex64/complex128 (compound form).
- Parsing unchanged: class 11 still surfaces as {r, i}.
- Tests vs h5py 3.16 / libhdf5 2.0.0 and h5dump 2.2.0; docs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:25:03 -05:00
osobhandClaude Opus 5.5 31efac1ae1 clawhdf5-netcdf4: an unlimited dimension reports its current length
CI / test-arm64 (pull_request) Successful in 1m34s
CI / test (pull_request) Successful in 30m38s
Dimension::size of an unlimited dimension was its dimension scale's
extent, which netCDF-C leaves at 0, so it read 0 for a dimension holding
records. It is now what netCDF-C reports (nc4_find_dim_len): the largest
current extent of the variables using it in any group, found through the
scale's REFERENCE_LIST, a coordinate variable's own extent included; 0
when nothing has been written.

interop_tests::unlimited_dimension_lengths_match_netcdf4_python compares
with netCDF4-python (variables of different lengths, one in a subgroup,
an unwritten dimension, a coordinate variable shorter than another
variable on its dimension, a subgroup's own dimension); before the fix it
got time 0/6, rec 3/5, srec 0/1.

The known-issues entry moves to Fixed; the crate README's warning goes.
A new open entry records a related bug found meanwhile: variables'
dimensions are matched by size, not DIMENSION_LIST.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:24:59 -05:00
osobhandClaude Opus 5.5 3eca5d8334 clawhdf5: a dropped FileEditor releases its lock at once
CI / test-arm64 (pull_request) Successful in 1m38s
CI / test (pull_request) Successful in 28m11s
FileEditor's flock belongs to the open file description. When another
thread forks to spawn a process, the child shares the locked descriptor
until it execs, so a reopen right after the drop could be refused with
Error::Locked (a one-off failure of edit_interop::editor_locks_the_file in
a parallel test run). Drop now unlocks before closing, which releases the
lock for every descriptor sharing it.

Reproducer edit_tests::drop_releases_the_lock_while_other_threads_spawn_processes
(4 threads running `true`, 2000 open/drop rounds): 1483 of 2000 reopens
refused before, 0 in 30 runs after (tank). An OFD lock would not help: it
is inherited across fork the same way and does not conflict with
libhdf5's flock. The agent store's lock file unlocks on drop too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:05:17 -05:00
osobh 5eae9ee60b Merge pull request 'docs: refresh README, crate READMEs and reference docs; fact-check every claim' (#22) from docs/readme-refresh into main
CI / test-arm64 (push) Successful in 1m44s
CI / test (push) Successful in 17m48s
Reviewed-on: #22
2026-09-28 17:24:25 +00:00
osobhandClaude Opus 5.5 d3d8d7ded3 docs: README — szip is a clawhdf5-format feature
CI / test-arm64 (pull_request) Successful in 1m35s
CI / test (pull_request) Successful in 15m45s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:25:03 -05:00
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00
osobhandClaude Opus 5.5 a48cb9f1a4 docs: known issue — NetCDF-4 unlimited dimensions report size 0
Found while checking the README refresh: clawhdf5-netcdf4 reads an
unlimited dimension's size from its dimension-scale dataset, which
netCDF-C never extends, so a dimension with 2 records reports 0.
Variable shapes and values are right. To be fixed separately.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:15:51 -05:00
osobhandClaude Opus 5.5 3100f0143b docs: fix cross-links after the refresh; remove the stale benchmark script
- docs/README.md links the improvement logs and June plans where the
  refresh archived them (docs/archive/).
- clawhdf5-py README: 'r+' creates and replaces attributes (compact or
  dense); only deleting them is unsupported.
- scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format
  and overwrote BENCHMARKS.md; nothing referenced it. Removed.
- Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu
  says it is not HDF5 I/O.
- benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer
  exists.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:15:18 -05:00
osobh 2bffd6b622 Merge branch 'docs/refresh-readme' into docs/readme-refresh 2026-09-28 11:14:06 -05:00
osobh c53c43b14f Merge branch 'docs/refresh-crates' into docs/readme-refresh 2026-09-28 11:14:06 -05:00
osobh 14790487a7 Merge branch 'docs/refresh-reference' into docs/readme-refresh 2026-09-28 11:14:06 -05:00
osobhandClaude Opus 5.5 3c557c9f0a docs: archive the improvement log/scan and the June superpowers plans
Moved to docs/archive/ (kept for their history, not deleted), each with a
one-line header saying it is historical and what supersedes it:

- IMPROVEMENT_LOG.md: three automated-loop PRs from April-May 2026, on
  the earlier quantumclaw PR numbering, which now collides with this
  repo's #12-#15. Superseded by CHANGELOG.md and git log.
- IMPROVEMENT_SCAN.md: one scan's notes (2026-05-04) of changes merged
  long ago. Superseded by CHANGELOG.md and git log.
- docs/superpowers/plans/*.md -> docs/archive/plans/: agent pre-work
  plans for d6c4d4f (2026-06-30), already marked implemented. The MPI plan
  promised collective MPI-IO, which is not what shipped (MpiVol is
  root-read + broadcast); its header says so and points at ROADMAP.md,
  where collective I/O is an open item.

Nothing links to the old paths (research/*.md mention IMPROVEMENT_LOG.md
and ROADMAP.md in prose only).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:40 -05:00
osobhandClaude Opus 5.5 419cb52287 docs: ROADMAP rewritten from CHANGELOG, git log and known-issues
The old file was a tracker for the mid-2026 agent-memory tracks, last
updated 2026-08-05, with an OpenClaw track and "what's next" items that
have since shipped (CI, fuzz target, WAL checksums, HNSW parallel build).

Now: releases v2.0.0-v2.7.0 and every PR merged since (#3-#21, merge
dates from git log), range-read milestones M0-M5, a one-paragraph summary
of the agent-memory work, and what is genuinely next: crates.io and PyPI
publishing (plus the broken Node package), SWMR writing, MPI collective
I/O, paged-metadata single-request reads, Blosc2/ZFP encoders, and the
open items of docs/known-issues.md. OpenClaw and ZeroClaw are listed only
as withdrawn. No dates are given for future work.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:40 -05:00
osobhandClaude Opus 5.5 cab952bb00 docs(conformance): file classes, ref-bug and ref_fix, latest result
Explains each class compare.py assigns, how ref_bugs.py confirms a
ref-bug in every run and how ref.py corrects h5py's big-endian VL values
(ref_fix), and records the latest result (602 of 697 ok, 3 ref-bug,
2026-09-27) with a link to CONFORMANCE.md. Mentions --no-fetch.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00
osobhandClaude Opus 5.5 999cb86071 docs(wasm-viewer): package size re-measured after openUrl; round-trip costs
The size table predated openUrl (it said so). Re-measured on tank at
9b5803f with `bash examples/wasm-viewer/build.sh`, then `wc -c` and
`gzip -9 -n -c`: the wasm is 1,384,607 B (378,485 gzipped), was 627,501
(191,639); the glue 40,711 (8,181), was 21,826 (4,487); remote.js 9,326
(3,448). 476,062 B of the wasm is the function-name section. opt-level z
(CARGO_PROFILE_WASM_RELEASE_OPT_LEVEL=z) is now 1% smaller gzipped than
the profile's s; 3 is larger. h5wasm rows unchanged.

Also adds the measured cost of listing a large group and reading one
dataset by URL (passes / requests / bytes, from CHANGELOG, 2026-09-27).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00
osobhandClaude Opus 5.5 b55b24b7ba docs: crate READMEs describe each crate as it is today
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.

Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
  clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
  on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
  HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
  search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
  backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
  exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
  it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
  nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.

The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00
osobhandClaude Opus 5.5 b660952421 docs: docs/README.md indexes every document
One line per document: user guides, evidence (CONFORMANCE, BENCHMARKS,
the conformance README), design docs, crate and package READMEs, and the
historical notes (roadmap, improvement logs, June plans, research briefs).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:10:22 -05:00
osobhandClaude Opus 5.5 b0b4018919 docs: QUICKSTART and USE_CASES on the current APIs
QUICKSTART used APIs that do not exist (file.dataset_names(),
File::attr, AttrValue::Str, memory.search(&q, 5), MemoryConfig::new with
a &str, consolidation without timestamps), `clawhdf5 = "2.0"` from
crates.io, and "3-45x faster than libhdf5". It now covers HDF5 in Rust
(write, read, strings, in-place append, remote, SWMR), Python (read, r+,
w, URLs), NetCDF-4, h5rs, agent memory and the CLI, every snippet
compiled and run (Python against a wheel built from the tree).

USE_CASES dropped claims with no source (the agent crate adds ~2MB,
IVF-PQ under 1.2 ms on modest hardware, an OpenClaw scenario, a .brain
layout and `clawhub publish` commands) and now covers the HDF5 cases
(no-C builds, threads, remote data, untrusted files, SWMR, in-place
edits), the agent cases with measured numbers, and when to use
something else.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:10:22 -05:00
osobhandClaude Opus 5.5 a90373ca84 docs: README rewritten around the HDF5 library and its evidence
Leads with what clawhdf5 is today for an HDF5 reader: the conformance
result (602 of 697, 0 mismatches, no panic/hang/crash; CONFORMANCE.md of
2026-09-28), the CVE corpus against h5dump and h5py, concurrent reads
against h5py threads and processes (BENCHMARKS.md, 2026-09-26, c5334b1)
and the libhdf5 comparison with its date and caveat; then a feature matrix
(supported / read only / not supported), install from git and maturin,
Rust and Python quick starts, remote files, the browser, SWMR, h5rs, a
short agent-memory section, the crate map and a documentation table.

Removed: the unverifiable "1850+ tests" badge and "~86K lines" footer,
the "What's new v2.2 -> v2.7" list (it is CHANGELOG.md), the agent
comparison table with other products, the Phase 1/2 roadmap checklist,
and the long agent sections (now docs/agent-memory.md). Fixed: crates
listed as C-free, SZIP / N-Bit / scale-offset as read-only, virtual
datasets as written too, Python 'r+' can create attributes (it cannot
create or delete objects). Every code snippet was compiled and run
against the workspace, the Python ones against a wheel built from it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:10:22 -05:00
osobhandClaude Opus 5.5 4d1a43fd7a docs: agent memory guide in docs/agent-memory.md
The agent-memory detail that lived only in README.md (architecture,
modules, performance and footprint tables, LongMemEval, feature flags and
settings, file schema, SQLite migration, research foundation) moves to its
own page, so the README can lead with the HDF5 library. Code examples are
updated to the current API (MemoryConfig::new takes a PathBuf,
HDF5Memory::search with SearchOptions, consolidation with timestamps) and
were compiled and run against the workspace; the CLI section was run
against the built `clawhdf5` binary. New: the CLI's search defaults to
0.7/0.3, not the library's 0.4/0.6; the /integrity group of signed stores.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:10:22 -05:00
osobhandClaude Opus 5.5 38107b90ed docs: CLAUDE.md regrouped into rules, invariants and workflows
Standing rules (no C by default, h5py must read what we write, one
float16 implementation, claims need evidence, OpenClaw/ZeroClaw
withdrawn, the known-issues rule) are gathered in one place; library and
agent-memory invariants are split; the CI section lists what
ci-test.sh and the conformance workflow run now instead of a dated
"two jobs, green" line. Adds the conformance, remote, wasm, Python and
benchmark workflows (idle load below 2, dated records), and warns that
scripts/run-benchmarks.sh is stale and overwrites BENCHMARKS.md. Crate
roles corrected (clawhdf5-io is I/O adapters; codecs live in
clawhdf5-format). Benchmark figures now live in BENCHMARKS.md only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:08:53 -05:00
osobhandClaude Opus 5.5 24cc9d14d8 docs: known issue for the flapping bad_nbit_parms_walk classification
The committed CONFORMANCE.md counts bad_nbit_parms_walk.h5 as an
our-error because its six h5py reads agreed in that run; a rerun the
same day confirmed the over-read and counted it ref-bug. The ok count
(602 of 697) is the same either way.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:07:20 -05:00
osobhandClaude Opus 5.5 ac0020594b docs: design status notes reflect what is merged
range-reads.md opens with a table of milestones M0-M5 and the PRs that
merged them (#17-#21), replacing a header left garbled by earlier
merges, and each milestone's status names its PR. swmr.md says the
reader is merged (PR #19) and the writer does not exist. openclaw.md
links the Node package's known-issues entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:05:38 -05:00
osobhandClaude Opus 5.5 06d8e2ee45 docs: benchmark headline numbers, superseded sections labelled
A "Current headline numbers" table gives each figure's newest dated
measurement with its machine, command and section. Sections a later run
replaced are marked superseded with a link to the newer one; stale
"still open" notes and cross-references now point at the fixes. No
measured value changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:04:38 -05:00
osobhandClaude Opus 5.5 798331ddbc docs: known issues split into open issues and a condensed fixed history
An "Open issues" table at the top links each open entry; fixed entries
move to "Fixed (history)", newest first, keeping the date, PR, affected
releases and what users must do. Open entries re-checked against main
(9b5803f): the remaining audit gaps are gathered into one entry, the
range-read and remote limits no longer contradict themselves (Python and
the browser open URLs; SWMR reading is File::open_swmr), and the
nondeterministic damaged-chunk error, fixed in PR #19, has its own
history entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:04:38 -05:00
osobh 9b5803f587 Merge pull request 'Remote files open in a few requests; ObjectHeader::parse back to speed; last conformance mismatches resolved (602/697)' (#21) from feat/listing-header-last-files into main
CI / test-arm64 (push) Successful in 1m39s
CI / test (push) Successful in 15m37s
Reviewed-on: #21
2026-09-28 11:57:03 +00:00
osobhandClaude Opus 5.5 694ee0a090 docs: conformance report after the last non-ok files were classified (602 of 697 ok)
CI / test-arm64 (pull_request) Successful in 1m22s
CI / test (pull_request) Successful in 15m34s
Regenerated on tank: ok 600 -> 602 (attr_datatypes.hdf5 and
tcomplex_be.h5, compared against h5py's big-endian VL values corrected),
mismatch 2 -> 0; cve-2025-2308.h5 and cve-2025-44904.h5 are ref-bug
(h5py's values varied across six heaps in this run);
bad_nbit_parms_walk.h5 read the same in all six this run, so it stays
our-error, as the classification rule requires. Baseline raised.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 23:39:41 -05:00
osobh bf5a163dcf Merge branch 'perf/wasm-listing-passes' into feat/listing-header-last-files
# Conflicts:
#	CHANGELOG.md
2026-09-27 23:16:27 -05:00
osobh 4f90ce02a7 Merge branch 'perf/header-parse-and-last-files' into feat/listing-header-last-files 2026-09-27 23:16:23 -05:00
osobhandClaude Opus 5.5 5b45c60c9c docs: ObjectHeader::parse A/B after inlining the v1 message loop (at or below 8f59b2e)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 23:15:09 -05:00
osobhandClaude Opus 5.5 2e5b059530 docs: fewer round trips for remote files in the browser, counted
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:48:47 -05:00
osobhandClaude Opus 5.5 761bdbf24f format: look a name up in a v1 group down its B-tree, as libhdf5 does
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).

`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.

Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
  earliest, 1 MiB:  7/74/193.6 MB -> 6/5/5.2 MB
  earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
  latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
  -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
  commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:47 -05:00
osobhandClaude Opus 5.5 6f5d14fd62 format: group walks go on past a failed node and hint what they read next
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.

- The v1 and v2 B-tree collectors descend into every child of a node
  after one fails (they only read the siblings before, so a sibling's
  subtree came a pass later), then return the first error: results and
  errors unchanged. The v2 walk stops once its record budget is spent,
  so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
  a group B-tree node's and a symbol table node's body (read once their
  header gives a length, a round trip later when the body is in the
  next block), an object header's first chunk and its continuation
  chunks, the symbol table nodes a B-tree leaf names, a dense group's
  name index header and the heap's root block (both read right after
  the heap header). A listing also hints every child's object header as
  its entry is read, even after a failure, and every direct block of a
  dense group's heap (reading the indirect blocks, at most 4096 entries
  and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
  n entries end) is one helper used by the object reads and the hints.

Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
  earliest, 1 MiB:  6/73/192.5 MB -> 4/68/192.5 MB
  earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
  latest,   1 MiB:  9/98/196.5 MB -> 5/86/196.5 MB
  latest,   64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:44 -05:00
osobhandClaude Opus 5.5 ca81c3ebfa format, wasm: Storage::hint, fetched by the lazy reader with a pass's misses
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.

- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
  this operation". Default: nothing (every backend that reads when
  asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
  (shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
  that misses nothing ignores them (a hint never adds a round trip); a
  pass that misses also asks for them, in file order, while the pass
  stays within what is left of the operation's `maxFetch` budget (a
  hint never makes a call fail). At most 1 MiB (or a block) per hint
  and 65536 blocks per pass are recorded, whatever a file makes a
  parser hint. `Operation::attempt` follows hints; the plain
  `LazyStorage::attempt` does not. `run_blocking` and the browser's
  driver use the former. `LazyStats::hinted_blocks` counts them.

No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:41:09 -05:00
osobhandClaude Opus 5.5 05e1136027 docs: conformance report with the last non-ok files classified (602 of 697 ok, 0 our-error, 0 mismatch)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:40:03 -05:00
osobhandClaude Opus 5.5 ff0b2f8a4e conformance: keep ref-bug files out of the root-cause tables, name the non-ok classes
A ref-bug file's refused objects were still listed as our-error root
causes. The summary line now names every non-ok class and its count, and
an empty root-cause table says None.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:33:42 -05:00
osobhandClaude Opus 5.5 96086add99 format: inline the version-1 message loop again (ObjectHeader::parse back to 8f59b2e's speed)
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.

Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:32:11 -05:00
osobhandClaude Opus 5.5 5c44630ea2 conformance: compare h5py's big-endian VL values corrected, confirm libhdf5 over-reads per run
The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.

- ref.py checks that the installed h5py returns big-endian VL elements
  with the file's bytes under a little-endian dtype (writing and reading
  a vlen('>f4') in memory) and, if so, relabels them with the file's byte
  order before hashing, marking the object `ref_fix`. The values are now
  compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
  /VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
  prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
  in six processes with different heaps (import order, MALLOC_PERTURB_).
  Values the file determines are the same every time; these three change
  (6, 6 and 3 distinct results), so they are over-read memory, not data
  clawhdf5 could match. compare.py classifies a file `ref-bug` only when
  every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
  objects; test_ref.py covers both (run in the nightly job).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:26:32 -05:00
osobh 425585ee71 Merge pull request 'Benchmarks re-measured: LongMemEval with real embeddings, local reads on an idle machine' (#20) from docs/bench-refresh-2026-09-27 into main
CI / test-arm64 (push) Successful in 1m35s
CI / test (push) Successful in 16m51s
Reviewed-on: #20
2026-09-28 02:38:09 +00:00
134 changed files with 12334 additions and 4077 deletions
+1 -1
View File
@@ -40,7 +40,7 @@ jobs:
python3 -m venv /opt/interop
# maturin + pytest: ci-test.sh builds the Python package
# (crates/clawhdf5-py) and runs its tests against h5py.
/opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray hdf5plugin maturin pytest
/opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray h5netcdf hdf5plugin maturin pytest
echo "/opt/interop/bin" >> "$GITHUB_PATH"
- name: Show interop library versions
# h5dump's version too: the h5rs dump test requires its exact output
+2
View File
@@ -40,6 +40,8 @@ jobs:
run: cargo test --release --manifest-path conformance/probe/Cargo.toml
env:
CARGO_TARGET_DIR: conformance/.cache/target
- name: Reference-side tests
run: /opt/conformance/bin/python conformance/test_ref.py
- name: Sweep
# The corpora come from GitHub (pinned commits, conformance/corpus.txt),
# so this job needs a runner that reaches github.com.
+3
View File
@@ -7,3 +7,6 @@ weights/
.venv
__pycache__/
.pytest_cache/
# Scratch files the heavy tests generate (huge_chunks_interop)
crates/*/tests/scratch/
+251 -12
View File
@@ -51,6 +51,31 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
---
## Current headline numbers
The newest dated measurement of each headline figure, as of 2026-09-28.
Everything below this section is the dated record behind them; sections whose
figures a later run replaced are marked *Superseded*. Machine "tank" is an AMD
Ryzen 7 7800X3D (8C/16T); rows marked idle were run with the 1-minute load
average below 2.
| Figure | Value | Measured | Command | Details |
|---|---|---|---|---|
| Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) |
| LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) |
| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | `f32`: 2026-09-24, tank, `5c8323c`; int8: 2026-09-19 (`c0a9206`), machine not recorded | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) |
| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | x86: 2026-09-20 (`dea02f5`), machine not recorded; Pi 5: 2026-09-21 (`114a2df`); not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) |
| `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) |
| Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) |
| Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) |
| `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) |
| Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) |
| Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) |
| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 35x (1.46 vs 51.4 ms, pure-Rust deflate) / 10.3x / 10.6x | write 2026-09-23, tank; attributes and groups 2026-08-03, tank | `cargo bench -p clawhdf5-bench --bench h5bench_write --features libhdf5-compare -- '^write_2d_chunked/'`; `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng), [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) |
| Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) |
---
## Memory footprint
`cargo run --release -p clawhdf5-bench --bin search_harness -- --footprint --full`,
@@ -64,6 +89,10 @@ change at all. Measured that way a store holding the corpus twice and one
holding it once came out *identical* (1.00x both), which is how the first
attempt at this measurement went.
> *Superseded* by the current figures below (2026-09-24): this table is the
> record of the double-copy fix (commit 2e7e045, undated); the store measured
> 2.72x, not 2.43x, by the time the int8 index landed.
| N | vectors (raw) | reopened, before | reopened, after |
|---:|---:|---:|---:|
| 1 000 | 1 MiB | 5 MiB (3.41x) | 4 MiB (2.39x) |
@@ -379,6 +408,9 @@ point: does a selection cost what the *selection* costs?
### Baseline (v2.4.0): every selection decodes the whole dataset
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24). Kept as the before picture.
4096 x 2048 f64 (64 MB per dataset), chunks 256 x 256, file 129 MB
| layout | read | selected | time ms | MB/s of selection | vs full read |
@@ -404,6 +436,9 @@ point: does a selection cost what the *selection* costs?
### After: partial reads
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24).
Only the rows of a contiguous dataset, or the chunks, that overlap the
selection's bounding box are read/decoded. A 64 x 64 window of the compressed
dataset: **105 -> 0.39 ms**; one row: **106 -> 2.7 ms**; one column:
@@ -435,6 +470,9 @@ because the machine's speed drifted; compare the *vs full read* column.)
### After: parallel cached decode, fewer copies (full reads)
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24).
Full-read times, old and new binaries run alternately at the same moment (this
machine's absolute speed drifts over a long session, so only same-moment
comparisons mean anything):
@@ -498,10 +536,164 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
explain the slower windows.
### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank, loaded)
**Superseded** by the idle re-run below; kept as the record the default
was first decided on.
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), branch `feat/libver-v18`
at `3c61635`, to decide whether the writer's default should become the
HDF5 1.8 format (`libver_bounds(LibVer::V18, LibVer::V18)`: version-1 B-tree
chunk indexes) instead of the 1.10 format (Fixed Array indexes for these
datasets). **Not an idle machine:** two other agents were building; the
1-minute load average was 7.0 to 7.6 throughout (the rule is below 2), so
treat differences under about 20% as noise. Default and `--v18` runs
alternated, three of each per chunk size; medians of the three.
> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]`
> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset)
Write (the whole 3-dataset file) and file size:
| chunks | 1.10 write ms | 1.8 write ms | 1.10 file bytes | 1.8 file bytes | size |
|---|---:|---:|---:|---:|---:|
| 256 x 256 | 467-500 | 467-483 | 134 916 080 | 134 933 848 | +0.013% |
| 32 x 32 | 490-495 | 489-506 | 138 879 336 | 139 465 112 | +0.42% |
Reads (chunked datasets; the contiguous one does not change), ms. The
window labels are the harness's, which count 256 x 256 chunks: with 32 x 32
chunks the 64 x 64 window covers 9 chunks and the 512 x 512 one 289.
| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 |
|---|---|---|---:|---:|---:|
| 256 | chunked + deflate | full (first) | 6.70 | 7.30 | 1.09x |
| 256 | chunked + deflate | full (repeat) | 4.80 | 3.80 | 0.79x |
| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 | 0.16 | 1.07x |
| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 2.40 | 1.25 | 0.52x |
| 256 | chunked + deflate | one row | 0.95 | 0.95 | 1.00x |
| 256 | chunked + deflate | one column | 1.93 | 1.94 | 1.01x |
| 256 | chunked | full (first) | 11.30 | 11.80 | 1.04x |
| 256 | chunked | full (repeat) | 6.20 | 5.60 | 0.90x |
| 256 | chunked | 64 x 64 window (1 chunk) | 0.04 | 0.04 | 1.00x |
| 256 | chunked | 512 x 512 window (4-9 chunks) | 1.50 | 0.37 | 0.25x |
| 256 | chunked | one row | 0.05 | 0.05 | 1.00x |
| 256 | chunked | one column | 0.44 | 0.45 | 1.02x |
| 32 | chunked + deflate | full (first) | 11.50 | 12.20 | 1.06x |
| 32 | chunked + deflate | full (repeat) | 9.30 | 9.30 | 1.00x |
| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.45 | 0.89 | 1.98x |
| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.97 | 2.41 | 1.22x |
| 32 | chunked + deflate | one row | 0.69 | 1.13 | 1.64x |
| 32 | chunked + deflate | one column | 1.26 | 1.70 | 1.35x |
| 32 | chunked | full (first) | 11.80 | 13.00 | 1.10x |
| 32 | chunked | full (repeat) | 6.30 | 6.20 | 0.98x |
| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 | 0.83 | 2.31x |
| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 | 1.16 | 1.66x |
| 32 | chunked | one row | 0.38 | 0.84 | 2.21x |
| 32 | chunked | one column | 1.05 | 1.21 | 1.15x |
With 128 chunks per dataset the two formats read and write alike (the 512 x
512 windows' 0.25x and 0.52x are not explained by the index and are likely
the load). With 8192 chunks, full reads stay within 10%, but a selection on
a freshly opened file costs about 0.4 to 0.5 ms more through the version-1
B-tree (1.2x to 2.3x). Each timed selection opens the file anew, so the
likely cause (not profiled) is walking the B-tree's nodes (2.6 KB each,
about 150 per dataset here) against a Fixed Array's few blocks. Writing costs the
same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in.
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load
average 1.66 to 1.84 at the start of each run. Stacked branch
`feat/huge-chunks` at `b55768f` (contains `feat/libver-v18`). Default and
`--v18` runs alternated, five of each per chunk size; median (min-max), ms.
> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]`
> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset)
| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 |
|---|---|---|---:|---:|---:|
| 256 | chunked + deflate | full (first) | 5.9 (5.5-6.2) | 6.2 (6.0-6.7) | 1.05x |
| 256 | chunked + deflate | full (repeat) | 4.4 (4.1-4.9) | 4.3 (4.2-4.4) | 0.98x |
| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 (0.15-0.17) | 0.16 (0.15-0.16) | 1.07x |
| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.23 (1.23-1.36) | 1.22 (1.22-1.25) | 0.99x |
| 256 | chunked + deflate | one row | 0.95 (0.94-1.04) | 0.94 (0.93-0.98) | 0.99x |
| 256 | chunked + deflate | one column | 1.92 (1.91-2.12) | 1.90 (1.90-1.91) | 0.99x |
| 256 | chunked | full (first) | 11.1 (10.7-11.6) | 10.9 (10.6-11.9) | 0.98x |
| 256 | chunked | full (repeat) | 5.1 (4.9-5.2) | 5.2 (4.7-5.6) | 1.02x |
| 256 | chunked | 64 x 64 window (1 chunk) | 0.03 (0.03-0.03) | 0.04 (0.04-0.04) | 1.33x |
| 256 | chunked | 512 x 512 window (4-9 chunks) | 0.36 (0.36-0.38) | 0.37 (0.36-0.38) | 1.03x |
| 256 | chunked | one row | 0.05 (0.04-0.05) | 0.05 (0.05-0.05) | 1.00x |
| 256 | chunked | one column | 0.45 (0.43-0.45) | 0.44 (0.44-0.45) | 0.98x |
| 32 | chunked + deflate | full (first) | 11.7 (11.2-12.4) | 12.5 (11.8-13.2) | 1.07x |
| 32 | chunked + deflate | full (repeat) | 9.5 (9.3-9.8) | 9.4 (9.3-11.2) | 0.99x |
| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.46 (0.45-0.49) | 0.89 (0.89-0.90) | 1.93x |
| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.95 (1.94-2.04) | 2.40 (2.39-2.51) | 1.23x |
| 32 | chunked + deflate | one row | 0.69 (0.67-0.72) | 1.15 (1.13-1.18) | 1.67x |
| 32 | chunked + deflate | one column | 1.26 (1.24-1.28) | 1.72 (1.69-1.73) | 1.37x |
| 32 | chunked | full (first) | 12.5 (11.9-12.7) | 12.6 (12.5-13.4) | 1.01x |
| 32 | chunked | full (repeat) | 5.9 (5.8-6.0) | 6.2 (6.0-6.4) | 1.05x |
| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 (0.35-0.36) | 0.84 (0.83-0.86) | 2.33x |
| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 (0.69-0.74) | 1.17 (1.16-1.17) | 1.67x |
| 32 | chunked | one row | 0.38 (0.38-0.41) | 0.85 (0.84-0.86) | 2.24x |
| 32 | chunked | one column | 1.03 (1.02-1.21) | 1.22 (1.20-1.25) | 1.18x |
Writes: 377 vs 374 ms (128 chunks), 399 vs 409 ms (8192 chunks); file
bytes as in the loaded run. The loaded run's 0.25x and 0.52x windows at 128
chunks were the load: idle, every 128-chunk read is within 7% except the
0.03 ms single-chunk window (one timer tick). The 8192-chunk result stands:
a selection on a freshly opened file costs 0.2 to 0.5 ms more through the
version-1 B-tree (1.2x to 2.3x), full reads are within 7%. The default stays
the 1.10 format.
## Local file speed after range reads
### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)
The remaining 4% (below) was the call to the version-1 message loop, which
`4313917` kept out of line with `#[inline(never)]`. Found with A/B builds
changing one piece at a time (perf is not available: `perf_event_paranoid`
4): `#[inline]` on `parse_v1_messages` alone brought
`object_header_parse_x401` from about 24.5–24.9 µs to 23.6–24.0 µs against
8f59b2e's 23.6–24.1 µs (short 4-second rounds); no attribute measured like
`#[inline(never)]`;
creating the chunk list only when a continuation is found measured no
faster on top and was not kept.
Same method as below: `8f59b2e` built in its own worktree and target
directory, separate binaries alternating, `taskset -c 5
local_metadata_bench --bench --warm-up-time 3 --measurement-time 10`, every
binary started with the 1-minute load average below 2 (0.19–1.86) and no
`rustc` running. Candidate: `96086ad` (this change). Median (range) of 3
rounds; run 2 also alternated `main` `425585e`.
| function | 8f59b2e | 425585e (main) | 96086ad | vs 8f59b2e |
|---|---:|---:|---:|---:|
| run 1: `object_header_parse_x401` | 23.81 µs (23.76–24.00) | | 23.57 µs (23.23–23.89) | **−1.0%** |
| run 1: `snod_parse_all` | 1.840 µs (1.837–1.854) | | 1.839 µs (1.837–1.873) | 0.0% |
| run 1: `btree_v1_walk` | 343 ns (338–349) | | 352 ns (344–363) | +2.7% |
| run 1: `facade_list_400_groups` | 8.03 ms (8.01–8.10) | | 8.05 ms (8.01–8.15) | +0.2% |
| run 2: `object_header_parse_x401` | 24.15 µs (23.74–24.41) | 24.93 µs (24.72–25.05) | 23.52 µs (23.34–23.68) | **−2.6%** |
| run 2: `snod_parse_all` | 1.843 µs (1.830–1.847) | 1.856 µs (1.850–1.865) | 1.861 µs (1.858–1.869) | +1.0% |
| run 2: `btree_v1_walk` | 346 ns (339–366) | 356 ns (347–356) | 357 ns (350–359) | +3.2% |
| run 2: `facade_list_400_groups` | 8.09 ms (8.02–8.10) | 8.12 ms (7.97–8.19) | 8.08 ms (7.98–8.12) | −0.1% |
- `ObjectHeader::parse` is at or below 8f59b2e (−1.0%, −2.6%) and 5.6%
faster than `main` in the same run.
- `btree_v1_walk` (one walk of a 350 ns B-tree) is 3% above 8f59b2e in
both runs, with overlapping ranges, and is the same on `main` (+0.2%
between `main` and this change): not from this change. The walk's code
changed in `e553153` (after a failed child the siblings are only read,
so the error returns after them; the fixture never takes that path, but
the loop carries the extra state); left as is.
- `snod_parse_all` and the facade listing are within noise.
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
> The `object_header_parse_x401` row (+4.2%) is *superseded* by
> [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank)
> (2026-09-27, `96086ad`); the other rows are current.
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as
separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen
@@ -545,8 +737,8 @@ What this shows:
- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header;
the base and candidate ranges do not overlap). It is the cost of reading
continuation chunks from a bounded queue (the fix for unbounded reads on
crafted headers) and does not show in the listing. Kept open in
`docs/known-issues.md`.
crafted headers) and does not show in the listing. (Fixed later the
same day; see the section above and `docs/known-issues.md`.)
- **Full reads of deflate data got faster** after #18 (in-place chunk
decoding into the typed output and per-thread scratch buffers): +1.7% on
one thread, +36% at 16.
@@ -599,6 +791,10 @@ saturate memory bandwidth (1.05x).
### Results after the read fixes (2026-09-26, tank, `408f69e`)
> *Superseded* by [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
> (2026-09-26, `c5334b1`), which closed the 16-thread gap listed at the end
> of this section.
Same machine, files and commands as the first run below, re-run on an idle
tank (load average 1.60 at the start; the 1-minute figure rose to about 5
during the clawhdf5 runs, mostly their own threads) after two fixes:
@@ -639,11 +835,17 @@ Read with care:
- At 16 threads every tool dropped in this run (h5py threads on contiguous
data from 8002 to 2285 MB/s, processes from 12846 to 6942), so the
16-thread rows are noisier than the others.
- Still behind: full reads of chunked data at 16 threads (0.69x-0.76x h5py
processes). See `docs/known-issues.md`.
- Still behind at this commit: full reads of chunked data at 16 threads
(0.69x-0.76x h5py processes); fixed by `c5334b1` (above), recorded as
fixed in `docs/known-issues.md`.
### First run, before the read fixes (2026-09-26, tank, `91644d8`)
> *Superseded* results: the tables and "What this shows" are the before
> picture for [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
> (2026-09-26). The workload description and the **Run** box below are
> still how every `concurrent_read` figure in this file is produced.
Measured on tank (AMD Ryzen 7 7800X3D, 8 cores / 16 threads, 61 GiB, Linux
7.0) at commit `91644d8`, load average 1.84 when the run started (the
1-minute figure rose to 3.7 during the runs; that is mostly the benchmark's
@@ -680,8 +882,8 @@ What this shows:
- **clawhdf5 threads on one `File` do, for hyperslab reads of compressed
data:** 1244 MB/s at 16 threads, 9.7x h5py threads and 0.89x h5py
processes, without a process pool.
- **Where clawhdf5 is behind** (open performance bugs, see
`docs/known-issues.md`):
- **Where clawhdf5 was behind** at `91644d8` (both since fixed; see
`docs/known-issues.md`, "Concurrent and contiguous read performance"):
- *Full reads of chunked datasets stop scaling at about 4 threads*
(about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab
reads, which bypass the `File`'s chunk cache, keep scaling, so the
@@ -771,6 +973,11 @@ Other flags (both harnesses): `--threads`, `--reps`, `--slab`, `--slabs`,
## Search harness baseline (v2.3.0)
> *Historical.* This baseline and the "After: …" subsections that follow
> record each step of the search work; they are *superseded* by
> [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24),
> the last subsection of this part.
Produced by `cargo run --release -p clawhdf5-bench --bin search_harness -- --full`
on deterministic **clustered** synthetic data (384-dim, unit-normalised; points =
cluster centre + noise — uniform random vectors are nearly equidistant in high
@@ -833,10 +1040,11 @@ build: 9752.6 ms (10254 vectors/s) · exact scan: 40 QPS, p50 24648 µs
| 1000 | 11 | 3.9 | 0.9 | 68.1 | 5.48 | 5.57 | 182.5 |
| 10000 | 114 | 32.2 | 10.9 | 845.0 | 48.56 | 78.65 | 19.8 |
| 100000 | 1486 | 713.0 | 354.5 | 10486.5 | 883.51 | 975.23 | 1.1 |
wrote /tmp/claude-1000/-home-osobh-projects-clawhdf5/422f755e-dd25-4c35-8613-5439087e3aaa/scratchpad/baseline_full.json
### After: HNSW neighbour-selection heuristic
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Same harness, same data, after replacing closest-M neighbour selection with the
HNSW paper's diversity heuristic (Algorithm 4, keeping pruned connections) for
both new links and back-link pruning. Recall@10 at `ef = 64`: **0.87 → 1.00**
@@ -882,6 +1090,8 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs
### After: persistent keyword index, no store rewrite per query
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
`hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every
record) and rewrite the whole `.h5` file on **every query**. The index is now
kept for the life of the store and updated incrementally, and activation boosts
@@ -902,6 +1112,8 @@ index removes that.
### After: vector index persisted with the checkpoint
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
The HNSW graph (not the vectors, which the store already holds) is saved to
`<store>.h5.ann` at each checkpoint and reloaded by `open()`, tied to that
checkpoint by a generation id. The index is now built once per store (the *cold
@@ -919,6 +1131,8 @@ index incrementally.
### After: unit-vector dot product, reusable visited set
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Cosine distance recomputed both vector norms on every evaluation; the index now
stores unit vectors and uses a plain dot product. The per-call `HashSet` of
visited nodes became a reusable epoch-stamped array. Recall is unchanged.
@@ -964,6 +1178,8 @@ build: 21084.6 ms (4743 vectors/s) · exact scan: 39 QPS, p50 24739 µs
### After: unranked keyword scores, top-k merge (rankings unchanged)
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
A fusion study (`search_harness --fusion-study`) showed that capping the
keyword candidate pool is **not** a safe optimisation: against the current
full-corpus normalisation the final top-10 overlap is only 0.83-0.92 and the
@@ -986,6 +1202,8 @@ results.
### After: batched bulk build (optionally parallel); deletions handled in search
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Profiling showed **90% of a build's distance evaluations are in back-link
pruning**. The bulk build now inserts in batches: plan each node's neighbours
against the graph as it stood at the start of the batch, link, then prune every
@@ -1548,6 +1766,11 @@ MRR, or a one-question change in recency, is within this variation.
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
This table is **BM25-only** (zero embeddings). With real embeddings and the
default hybrid 0.4/0.6 the same corpus gives turn Hit@5 **81.4%** (2026-09-27;
see [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500)),
which is the headline figure.
47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence
sessions, so retrieval has to actually discriminate.
@@ -1655,7 +1878,9 @@ over rank-1 precision.
activation. Until now its combined score contained **no relevance term at
all** — `RerankInput` did not carry the retrieval score — so a caller that
re-ranked its candidates threw the retriever's ordering away and returned them
ordered by age. The OpenClaw backend did exactly that on every search.
ordered by age. `ClawhdfBackend` (the `openclaw` module) did exactly that on
every search. (OpenClaw itself never integrated clawhdf5; see
`docs/openclaw.md`.)
Measuring that is unambiguous. "Recency" below is the share of
`knowledge-update` questions where the newest gold session outranked the stale
@@ -1753,7 +1978,7 @@ worth stating plainly rather than hiding: LongMemEval questions share substantia
vocabulary with their evidence turns, which is close to the best case for lexical
matching, and MiniLM at 384 dimensions is a small embedding model.
> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \
> **Run:** `cargo run --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings -- \
> benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2`
> For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on
> `PATH` at *build* time — cudarc's build script shells out to it. The toolkit
@@ -2172,7 +2397,8 @@ The tank row was measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D), commit
### Reproducibility
```bash
rustup override set nightly
# Any stable toolchain at or above the MSRV (1.92) works; the original
# 2026-07-01 run used a nightly, later runs stable.
# Latency benchmarks (Criterion)
cargo bench -p clawhdf5-agent
@@ -2275,7 +2501,9 @@ libhdf5 reads from a temp file including `open` + `read` + `close` overhead.
| clawhdf5 hyperslab (f64, 10% slice) | — | 4.09 µs / **1.8 GiB/s** | 50.1 µs / **1.5 GiB/s** |
libhdf5 f64 comparison excluded — clawhdf5's datatype encoding differs from libhdf5's (known
gap), making cross-format reads unreliable for comparison.
gap), making cross-format reads unreliable for comparison. (That gap was the float sign-bit
bug, fixed 2026-09-23: `docs/known-issues.md`, "Every `f32` dataset we wrote was unreadable by
h5py / libhdf5". The comparison has not been re-run since.)
### Chunked Read Throughput
@@ -2383,6 +2611,13 @@ global file mutex and flushes to disk on every attribute write or group creation
## vs libhdf5 Summary
> Measured on the original i7-12650H (clawhdf5 2026-07-01, libhdf5
> 2026-06-30). The newest run of this table is
> [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03)
> (2026-08-03), which reproduces every row within ~15% except chunked
> write (45.3x on tank; 35x on 2026-09-23 with the pure-Rust deflate, see
> [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng)).
| Workload | clawhdf5 | libhdf5 | Speedup |
|----------|----------|---------|---------|
| Sequential read, 1K f32 | 634 ns | 45.2 µs | **71×** |
@@ -2414,7 +2649,7 @@ to the page cache. There is no algorithmic headroom above ~1.7 GiB/s on this har
### Caveats
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap). f64 results are clawhdf5-only.
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap at the time; fixed 2026-09-23, see [Sequential Read Throughput](#sequential-read-throughput)). f64 results are clawhdf5-only.
- Serial benchmarks. clawhdf5 uses Rayon for chunk compression when > 2 chunks; that parallelism is already reflected in the chunked write numbers.
- clawhdf5 reads from `Vec<u8>` (zero-copy from mmap in production); libhdf5 reads from a temp file. This gives clawhdf5 a structural read advantage that reflects realistic API usage.
@@ -2611,6 +2846,10 @@ Same not-like-for-like caveat as the "Comparison to MemX" section at the top of
file applies — MemX's figure is end-to-end, these are a single component. Ratios are
an order-of-magnitude indication, not a benchmark result.
> The Ratio column below was retracted afterwards: see
> [Comparison to MemX](#comparison-to-memx-arxiv260316171). Kept as recorded
> on 2026-08-05; do not cite it.
| Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio |
|--------|----------------------------|----------------------------------|-------|
| 100K flat search | <90 ms | 6.60 ms | ~14x |
+365
View File
@@ -2,6 +2,371 @@
## Unreleased
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
dimension it had written fewer records of got `dim_<n>`, and dimensions
of one size could be swapped. It now resolves them as netCDF-C does
(`libhdf5/hdf5open.c`): the ids in the variable's `_Netcdf4Coordinates`
(each scale's `_Netcdf4Dimid`), else the scales its `DIMENSION_LIST`
references (object references, also the revised `H5T_STD_REF`; of
several scales attached to one axis, the last, as netCDF-C's
`dimscale_visitor` keeps), looked up in its group and then each parent
group; a coordinate variable is on
its own scale. Only an axis with neither (a file not written by a netCDF
library) is still matched by size. A variable may use one dimension
twice (`(p, p)`).
- `variables()` and `variable_names()` leave out the dimension scales that
are only dimensions (`NAME` "This is a netCDF dimension but not a netCDF
variable."), as netCDF-C does, and `variable()` refuses them
(`VariableNotFound`). A variable stored as `_nc4_non_coord_<name>`
(netCDF-C's name for a variable sharing a dimension's name without being
its coordinate variable) is listed and found as `<name>`.
`NetCDF4File::variable_names` is new; `NetCDF4Group::variable_names` used
to list every dataset.
- A variable along an unlimited dimension has the dimension's length, as in
netCDF: `Variable::shape` is that length and the reads return that many
values, the records it has not written as the fill value (`_FillValue`,
else netCDF's `NC_FILL_*` for the type; `""` for strings; NaN from
`read_f64`). It was the HDF5 extent. `Variable::stored_shape` (new) is the
HDF5 extent. Where the unlimited dimension is not a variable's first,
unwritten values are placed per row, as netCDF-C's element and row reads
return them; a whole-variable read through netCDF-C 4.9.3
(netCDF4-python 1.7.4) instead returns the written values first, then the
fill.
- A variable's attributes are read when it is opened (they were read on
first use).
- Tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3) variable by
variable (dimensions, shape, every value): the known-issues reproducer;
two dimensions of one size in either order, one dimension used twice, a
scalar, a non-coordinate variable named like a dimension, a subgroup and
a sub-subgroup on their ancestors' dimensions; unwritten records of
`i1`/`i4`/`u8`/`f4`/`f8`/string variables, with and without
`_FillValue`; a file of h5py dimension scales (no netCDF attributes);
h5netcdf 1.8.1 and xarray 2026.7.0 files (engines netcdf4 and h5netcdf).
The h5netcdf cases skip when h5netcdf is not installed; CI now installs
it. A one-off comparison over the 78 conformance-corpus files
netCDF4-python opens (tank, 2026-09-28; a throwaway test, not committed)
found 52 files with the same variables, dimensions and shapes, and the
same values in every numeric variable of up to 5000 elements (enum
variables were not compared). The other 26 have no dimension scales
(netCDF-C names their axes `phony_dim_<n>`; this crate still matches by
size or names them `dim_<size>`) or hold datasets of types netCDF-C
skips (opaque, references). Affected v2.1.0 to v2.7.0.
`docs/known-issues.md`.
### Writing files HDF5 1.8 can read (2026-09-28)
- New `LibVer` (`V18`, `V110`, `V112`, `V114`, `V200`, `Latest`; re-exported
as `clawhdf5::LibVer`) and `FileBuilder::libver_bounds(low, high)`
(`FileWriter::libver_bounds` in the format crate), as libhdf5's
`H5Pset_libver_bounds` and h5py's `libver=(low, high)`. The default,
`(V110, Latest)`, writes exactly what was written before (the HDF5 1.10
format, which HDF5 1.8.23 refuses to open).
- A low bound of `V18` writes what libhdf5 2.x writes for h5py's
`libver=('v108', 'latest')`: a version-2 superblock (version 3 only for a
paged file), version-3 layout messages for contiguous, compact and chunked
datasets, and a version-1 B-tree chunk index for every chunked dataset,
resizable and multi-unlimited ones included, instead of the single chunk,
Fixed Array, Extensible Array and version-2 B-tree indexes. The B-tree
writer (`btree_v1_write`) replays `H5B_insert` with the chunk callbacks of
`H5Dbtree.c` for chunks arriving in row-major order: libhdf5's split
ratios, right keys moved exactly when `H5D__btree_cmp3` moves them, the
root kept at its address. Its trees are libhdf5's node for node (levels,
child counts, keys) for 1-D, 2-D and 3-D datasets with two- and
three-level trees, deflated or not, against libhdf5 2.0 writing without a
chunk cache (`chunk_btrees_match_libhdf5`). An empty chunked dataset has
no tree (undefined address), as in libhdf5.
- A high bound refuses, with the new `FormatError::LibverBound` and before
anything is written, what needs a newer format: virtual datasets and the
paged file-space strategy (1.10), the 1.12 reference types (datatype
version 4), native complex numbers (datatype version 5, HDF5 2.0), and a
low bound above the high one. `(V18, V18)` therefore writes a file HDF5
1.8 reads, or fails; `(V18, Latest)` writes such objects in their newer
format, as libhdf5 does. `Datatype::max_encoded_version` reports the
version a type needs.
- Checked against a real HDF5 1.8: `scripts/build-hdf5-1.8.sh` builds
1.8.23 (the last 1.8 release) with its tools. `clawhdf5-tools`'
`tests/libver_v18.rs` writes every writer feature under `(V18, V18)`
(contiguous, empty, scalar, compact, f16/f32/f64/i64, chunked with
deflate + shuffle + Fletcher-32 and edge chunks, resizable with one and
two unlimited dimensions, finite maxshape, 100 000 one-element chunks for
a three-level tree, a 100 x 100 grid of chunks, fill values, fixed-length
strings, compound, enum, array, compact/dense/creation-ordered groups,
soft, hard and external links, dense attributes); HDF5 1.8.23's h5dump
dumps the whole file exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset (`-b LE`), and h5py, clawhdf5 and
`h5rs check --data` read it. Then `FileEditor` appends to the resizable
datasets 300 times (splitting B-tree nodes), grows the three-level one,
rewrites a deflated row, grows a 2-D dataset and sets compact and dense
attributes; h5py appends too; and every reader checks again. The 1.8
checks are skipped where no HDF5 1.8 is found (`CLAWHDF5_H5DUMP18` or
`~/.cache/hdf5-1.8.23`), as in CI.
- The default stays the 1.10 format: with 8192 chunks per dataset, small
selections on a freshly opened file read 1.2x to 2.3x slower through a
version-1 B-tree (read harness, tank, 2026-09-28, under load; full reads
and writes within 10%, files about 36 bytes per chunk larger).
`read_harness` gains `--v18` and `--chunk N`. `BENCHMARKS.md`, "HDF5 1.8
format".
- Not yet: the Python bindings' `'w'` mode has no `libver` argument, and
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
written. `docs/known-issues.md`.
### Chunks of 4 GiB or more (HDF5 2.0's layout message version 5) (2026-09-28)
- **What libhdf5 does** (read in libhdf5 2.2.0's `H5Dchunk.c`, `H5Dfarray.c`,
`H5Dearray.c`, `H5Dbtree2.c`, `H5Dbtree.c`, `H5Olayout.c`): a chunk of
more than 0xFFFFFFFF bytes requires layout message version 5, so a high
bound of `H5F_LIBVER_V200` ("chunk size > 4GB requires H5F_LIBVER_V200"),
whatever the low bound, and version 5 always uses the newer indexes.
Under version 5 a filtered Fixed Array, Extensible Array or v2 B-tree
element stores the chunk's size in "size of lengths" bytes (8); under
version 4 in one byte more than the chunk needs. A filtered Single Chunk
stores it in "size of lengths" bytes under both. Unfiltered elements
store no size, and the Implicit index none at all. A v1 B-tree key keeps
a 32-bit size: libhdf5 never writes a larger chunk there and refuses to
open one ("chunk size must be < 4GB with v1 b-tree index").
- **Reading** such chunks works in every index: `ChunkInfo::chunk_size` and
`ChunkMapping::file_size` are now `u64` (**public API change**). The
Single Chunk, Implicit, Fixed and Extensible Array readers truncated an
unfiltered chunk's size (the read failed), and the v2 B-tree reader
refused any chunk over 4 GiB. Checked against a fixture libhdf5 2.0.0
wrote (`tests/fixtures/huge_chunks_filtered.h5`, 188 KB: a 4 GiB + 8
byte chunk in each filtered index, deflated twice) and a 44 GiB sparse
file h5py writes at test time (every unfiltered index).
- **Selections read what they touch:** a selection of a chunked dataset with
a non-default fill value is read over a box of fill values from the
chunks it touches (new `partial_read::read_selection_filled_in`); it
used to decode the whole dataset. An unfiltered chunk of a file that is
not in memory (`File::open_storage`) is read row by row instead of whole.
A filtered chunk is still decoded whole.
- **Writing:** a chunk of more than `u32::MAX` bytes gets layout message
version 5 and libhdf5's index element widths; the Fixed and Extensible
Array structures match libhdf5 2.0.0's byte for byte. New public helpers
`chunked_write::{layout_version_for, chunk_size_len, MAX_V4_CHUNK_BYTES}`.
Refused before anything is written: chunk dimensions of 2^32 or more
(they were cut to 32 bits) and, for such chunks, LZF, bitshuffle, bzip2,
Blosc and pcodec. Chunks are extracted row by row and one at a time
(compressed in parallel only up to 64 MiB), and a filter pipeline no
longer copies its input first.
- **Deflate** compressed only the first 4 GiB - 1 bytes of a larger chunk
(zlib takes at most that much per call, and `Finish` ended the stream
there; since the one-pass deflate of 2026-09-23, in `clawhdf5-format` and
`clawhdf5-filters`), and it reserved the compressor's worst-case bound,
which flate2's Rust backends zero on every call: 4 GiB kept per 4 GiB
chunk. Inputs over 64 MiB now start at 1/16 of the bound and grow, and a
result more than 64 MiB too large is shrunk. On the read side, an
intermediate deflate stage (one followed by a filter other than shuffle
or Fletcher-32) no longer reserves the chunk's whole bound: reading the
double-deflated fixture peaked at 8.5 GiB, now 4.0 GiB.
- **LZ4:** a chunk decoding to more than 256 MiB was refused ("lz4:
declared size exceeds limit") though the chunk size bounded it; the
ceiling now applies only to a decode of unknown size. A chunk of 4 GiB
or more is read as the registered HDF5 framing, whose 64-bit size no
longer starts with four zero bytes.
- **`FileEditor`** refuses, before writing anything, to write values into
or to prune, fill or allocate chunks of 4 GiB or more; growing the extent
(late allocation) and setting attributes work. It sizes a version-5
filtered index element from the file's size of lengths (it assumed 8).
- **32-bit targets:** reading or writing such a chunk is
`FormatError::Overflow`; `scripts/check-32bit-casts.sh` passes, and the
wasm package test reads the fixture and gets that error in every index.
- Tests: `crates/clawhdf5/tests/huge_chunks_interop.rs` (index listing,
byte-for-byte array indexes and the editor always; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of every index filtered and unfiltered,
and a write of every index read back by clawhdf5, h5py 3.16 and — layout
and storage size — h5dump 2.2.0 via `CLAWHDF5_H5DUMP2`),
`partial_read_equivalence::fill_value_selections_match_full_reads`, unit
tests of the encodings (`chunked_write::tests`) and of LZ4. The opt-in
run (`cargo test --release -p clawhdf5 --test huge_chunks_interop --
--test-threads=1`) peaked at 8.06 GiB resident, h5py included (tank,
2026-09-28, commit `9143737`). libhdf5 2.2.0's h5dump cannot print these
chunks' values (its deflate filter fails on any chunk over 4 GiB,
libhdf5's own included); h5py 3.16 reads them. `docs/known-issues.md`: three fixed
entries and the open "Chunks of 4 GiB or more: limits" (chunk dimensions
of 2^32 or more, the refused filters, memory, the untested unfiltered
write).
### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked
descriptor until it exec'd, so reopening the file right after the drop
was sometimes refused with `Error::Locked` (seen once as a failure of
`edit_interop::editor_locks_the_file` in a parallel test run). The drop
now unlocks before closing, which releases the lock for every descriptor
that shares it. Reproducer
`edit_tests::drop_releases_the_lock_while_other_threads_spawn_processes`
(tank, 2026-09-28): 1483 of 2000 reopens refused before, 0 in 30 runs
after. The agent store's lock file unlocks on drop the same way (its
250 ms retry on open had hidden the race). `docs/known-issues.md`.
### NetCDF-4: unlimited dimensions report their length (2026-09-28)
- `clawhdf5-netcdf4`'s `Dimension::size` of an unlimited dimension was the
extent of its dimension scale, which netCDF-C leaves at 0 (or, for a
coordinate variable, that variable's own length), so it read 0 for a
dimension holding records. It is now what netCDF-C reports
(`nc4_find_dim_len`): the largest current extent of the variables using
the dimension in any group (the scale's `REFERENCE_LIST`), a coordinate
variable included; 0 when nothing has been written. Test:
`interop_tests::unlimited_dimension_lengths_match_netcdf4_python`
(variables of different lengths, one in a subgroup, an unwritten
dimension, a coordinate variable shorter than another variable on its
dimension, a subgroup's own dimension), compared with netCDF4-python.
Affected v2.1.0 to v2.7.0. `docs/known-issues.md` also gains an open
entry found meanwhile: variables' dimensions are matched by size.
### HDF5 2.x small floats checked against libhdf5 2.2.0 (2026-09-28)
- New fixture written by libhdf5 2.2.0 itself (built from tag `2.2.0`,
`crates/clawhdf5/tests/fixtures/gen_mx_floats.py`): a dataset and an
attribute of each of `H5T_FLOAT_F4E2M1`, `H5T_FLOAT_F6E2M3`,
`H5T_FLOAT_F6E3M2` (every bit pattern, also with the unused high bits set),
`H5T_FLOAT_F8E4M3`, `H5T_FLOAT_F8E5M2` (all 256) and
`H5T_FLOAT_BFLOAT16LE`/`BE` (zeros, subnormals, max, ±inf, NaNs), with
what libhdf5 returns for each element into `double` and `float`.
`tests/mx_floats_interop.rs` compares `read_f64`, `read_f32` and the
attributes bit for bit. No value was wrong: clawhdf5 already decoded every
pattern as libhdf5 does, including an all-ones exponent as ±inf/NaN in the
formats the OCP MX specification makes finite (FP4, FP6, FP8 E4M3 — a
deliberate match, recorded in `docs/known-issues.md`).
- NaNs of these formats now have libhdf5's bits: sign kept, every mantissa
bit set (`0x7FFF_FFFF_FFFF_FFFF` as `f64`, `0x7FFF_FFFF` as `f32`); they
were `f64::NAN`'s bits and an `as f32` cast's.
- `h5rs dump` names these types as h5dump 2.x does (`H5T_FLOAT_F4E2M1`,
`H5T_FLOAT_BFLOAT16LE`, ...) instead of printing an `H5T_FLOAT { ... }`
block; `h5rs ls -v` describes them as h5ls 2.x does (`FP4 E2M1 4-bit
float`) and `h5rs ls` lists them as `bfloat16`, `float8-e4m3`,
`float6-e2m3`, `float4-e2m1`, ... instead of `float16`/`float8`. Checked
against h5dump 2.2.0's output of the fixture (`tests/mx_floats_dump.rs`).
- Python bindings: datasets and attributes of these types used to raise
`TypeError`; they now read as h5py 3.16 reads them — `float32` for
bfloat16, `float16` for the 1-byte formats, in the file's byte order — with
the same bytes h5py returns (`tests/test_small_floats.py`). Writing them in
`'r+'` raises `NotImplementedError` before anything is written; inside a
compound or array type they still raise `TypeError`.
### Writing complex numbers, including HDF5 2.0's native complex type (2026-09-28)
- **h5py's form (default):** `DatasetBuilder::with_complex_f32_data` /
`with_complex_f64_data` take `[re, im]` pairs and write the compound
`{r, i}` h5py writes for numpy `complex64`/`complex128` (h5py 3.16 still
writes this by default); `make_complex_f32_type`/`make_complex_f64_type`
give the datatype for `AttrValue::Raw` attributes.
- **Native complex (opt-in):** `with_native_complex_f32_data` /
`with_native_complex_f64_data` and `make_native_complex_f32_type` /
`make_native_complex_f64_type` write `H5T_COMPLEX_IEEE_F32LE`/`F64LE`
(datatype class 11, version 5), through the new `Datatype::Complex`
variant. The encoding is byte-identical to libhdf5 2.2.0's
(`H5Odtype.c`: homogeneous, rectangular, base type follows), and a
compound, array or variable-length type holding one is written as
version 5, as libhdf5 raises it. No file-level version bound is needed:
the superblock and object headers we write already open in libhdf5 2.0.
Only libhdf5 2.0+ reads class 11 (Debian's h5dump 1.14 fails on the
object), hence opt-in. Checked on tank: h5py 3.16 (libhdf5 2.0.0) reads
datasets (contiguous and chunked+deflate), attributes, a compound member
and an array of them as numpy `complex64`/`complex128` with class 11;
h5dump 2.2.0 prints them as `H5T_COMPLEX_IEEE_F*LE`; `h5rs check --data`
finds no problems.
- **Reading:** `Dataset::read_complex_f64`/`read_complex_f32` return
`[re, im]` pairs from either form (and from h5py's files, native or
not). Parsing is unchanged: class 11 still surfaces as the `{r, i}`
compound (`Datatype::complex_as_compound`), so `Datatype::parse` never
returns `Datatype::Complex` (see `docs/known-issues.md`).
- **Breaking for exhaustive matches:** `Datatype` gained the `Complex`
variant; downstream `match`es over `Datatype` without a wildcard need an
arm (`Datatype::complex_as_compound(size, base)` gives the compound view).
- **Python:** `create_dataset` accepts `complex64` and `complex128` arrays,
written as h5py's compound (native class 11 is Rust-only).
- Tests: `datatype.rs` (byte equality with libhdf5's encoding),
`integration_tests::complex_datasets_and_attributes_round_trip`,
`writer_h5py_tests::h5py_reads_our_complex_datasets_and_attributes`
(runs h5dump 2.x when `CLAWHDF5_H5DUMP2` names one),
`h5py_interop_tests::h5py_complex_datasets_read_as_complex`, and
`test_write_read.py::test_roundtrip_complex` /
`test_read_native_complex_from_h5py`.
### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27)
- Parsing a version-1 object header was 4% slower than before range-read
M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk
message loop, kept out of line since the allocation fix; it is inlined
again. `object_header_parse_x401` is now 1.0% and 2.6% below `8f59b2e`
and 5.6% below `main` (tank, idle, separate binaries alternating;
`BENCHMARKS.md`). The chunk queue's checks are unchanged (65,536 chunks,
cycle refusal, file-size budget, one chunk buffer at a time, libhdf5's
order, overlapping chunks allowed).
### Conformance: no our-errors or mismatches left (2026-09-27)
- The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us. Re-checked with evidence
(`docs/known-issues.md`, "Conformance: the last non-ok files"):
- The 2 mismatches were h5py's big-endian VL bug (elements returned with
the file's bytes under a little-endian dtype). `conformance/ref.py` now
checks the installed h5py has the bug and relabels such elements before
hashing, so their values are compared: `attr_datatypes.hdf5` and
`tcomplex_be.h5` are identical to clawhdf5's and now **ok**.
- The 3 our-errors are objects HDF5 2.0 reads by over-reading memory
(scale-offset codes past a chunk, unfiltered chunks shorter than a
chunk, an N-Bit parameter list one value short). h5py's values for them
change between runs, with `MALLOC_PERTURB_` and with import order, and
h5dump 1.14.6 prints others, so they are not the file's data and
clawhdf5 keeps refusing them. `conformance/ref_bugs.py` repeats that
check in every run, and such a file is the new class **ref-bug** only
while its values keep changing.
- Result (`conformance/run.sh --no-fetch`, tank, 2026-09-27): 602 of 697
ok (baseline 600), 0 our-error, 0 mismatch, 3 ref-bug, 92
h5py-cannot-read, no panic/hang/crash/oom.
### Remote files in the browser: fewer round trips to list a group or open a dataset (2026-09-27)
- **Opening one dataset of a v1 (symbol table) group no longer reads the
whole group.** A name is looked up down the group's B-tree, as
libhdf5's `H5G__stab_lookup` does (binary search on the node keys, names
in the local heap compared bytewise, then one symbol table node); only
when that finds no hard link of that name (a soft link, or a B-tree out
of name order, where libhdf5 would report it missing) is every entry
read, as before. Local files benefit too (a lookup read O(entries)).
In a group holding two entries of one name, the B-tree's is now the one
found, as in libhdf5.
- **`Storage::hint(offset, len)`** (clawhdf5-format): a parser says what
it reads next — a B-tree node's or symbol table node's body, an object
header's first chunk and continuation chunks, the symbol table nodes a
B-tree leaf names, a dense group's name index header and heap blocks,
and in a listing every child's object header. Every backend ignores it
but the browser's restartable reader (`clawhdf5_wasm::lazy`), which
fetches the hinted blocks it lacks together with the blocks a pass
missed, within the call's `maxFetch` budget; a pass that misses
nothing ignores them, so a hint never adds a round trip, and results
never depend on hints.
- The v1 and v2 B-tree walks of a listing descend into every child after
one fails (before, the siblings were only read, so their subtrees came
a pass later), then return the first error: same results and errors.
- Counted on tank, 2026-09-27, with `CLAWHDF5_WASM_LIST_FILE=<file>
CLAWHDF5_WASM_READ=/d1500 cargo test --release -p clawhdf5-wasm --test
lazy listing_cost_of_a_given_file -- --nocapture` on an h5py file like
the reviewer's (3000 datasets of 16384 `f32`, 198 MB, h5py 3.16 /
HDF5 2.0), passes / requests / bytes, before -> after:
| file, block size | `list('/')` | open + read one dataset |
|---|---|---|
| earliest, 1 MiB | 6 / 73 / 192.5 MB -> 4 / 68 / 192.5 MB | 7 / 74 / 193.6 MB -> 6 / 5 / 5.2 MB |
| earliest, 64 KiB | 8 / 531 / 35.2 MB -> 5 / 530 / 35.3 MB | 9 / 515 / 34.1 MB -> 8 / 7 / 0.52 MB |
| latest, 1 MiB | 9 / 98 / 196.5 MB -> 5 / 86 / 196.5 MB | 8 / 7 / 6.7 MB -> 7 / 7 / 6.7 MB |
| latest, 64 KiB | 11 / 452 / 29.6 MB -> 6 / 454 / 30.5 MB | 9 / 8 / 0.58 MB -> 8 / 8 / 0.58 MB |
The listing's passes now follow the depth of the group's index (the
chain index levels -> symbol table nodes or heap objects -> child
headers); its bytes are the child headers, which h5py spreads through
the file (at 1 MiB blocks most of it). The whole corpus read lazily
(`CLAWHDF5_WASM_CORPUS`, 656 files, every object listed, described and
read): 27 513 -> 27 496 passes, 3 001 -> 2 962 requests and 194.8 ->
195.3 MB at 64 KiB blocks; 41 342 -> 37 517 passes, 32 578 -> 29 511
requests, 65.4 -> 65.7 MB at 512 B. The Node and Chromium suite
(`examples/wasm-viewer/test/run.sh`) passes unchanged (the 200 MB file
still takes 5 requests, 6 MiB); its corpus comparison fetched 33.54 ->
33.61 MB.
- Tests: the listing budgets (`listing_a_large_group_takes_a_few_passes`,
512-byte blocks) are tightened to the new counts (FileBuilder, 600
children: 5 -> 4 passes; h5py, 2000 children: 8 -> 5 and 11 -> 6); new
`reading_one_dataset_of_a_large_group_fetches_a_few_blocks` (h5py
earliest: 529 requests, 333 kB -> at most 6 requests, 27 kB), v1
lookups against the listing (and with a name moved out of B-tree
order), hints riding only on misses and within the fetch budget.
### Deterministic errors on damaged chunked datasets (2026-09-27)
- A read through the file's chunk cache listed a damaged dataset's chunks in
hash-map order, seeded per `File`, so two opens of the same file could
+233 -220
View File
@@ -1,253 +1,267 @@
# clawhdf5
## Purpose
Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated vector search. A standalone library. Its one verified consumer is ClawBrainHub (`.brain` files); no agent framework integrates it (OpenClaw and ZeroClaw claims were withdrawn on 2026-09-25 — neither was ever true).
Pure-Rust HDF5 implementation (read, write, in-place edit, remote and browser
reads) plus agent memory on top of it: HNSW vector search, a WAL-backed store,
and GPU vector distances. A standalone library. Its one verified consumer is
ClawBrainHub (`.brain` files); no agent framework integrates it (see
*Standing rules*).
## Architecture
Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal FFI bindings crate for the optional `szip` feature):
Cargo workspace, 19 crates under `crates/` (plus `libaec-sys`, the FFI crate
behind the optional `szip` feature). MSRV 1.92 (`rust-version`, checked in CI).
| Crate | Role |
|-------|------|
| `clawhdf5-format` | HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants |
| `clawhdf5-io` | Read/write implementation |
| `clawhdf5-filters` | Deflate backends (zlib-rs, zlib-ng, Apple Compression); the HDF5 filter pipeline, the filter registry (`clawhdf5_format::filter_registry`) and the other codecs (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec, and the pure-Rust plugin filters LZF, bitshuffle, bzip2, Blosc 1, and Blosc2 and ZFP read-only) live in `clawhdf5-format`. |
| `clawhdf5-derive` | Proc-macro derive for HDF5-serializable structs |
| `clawhdf5` | Main facade crate |
| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer |
| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index |
| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage |
| `clawhdf5-gpu` | GPU vector distance computation via wgpu (hand-written WGSL compute shaders) — not dataset I/O |
| `clawhdf5-accel` | CPU SIMD acceleration path |
| `clawhdf5-migrate` | SQLite → HDF5 agent-memory migration |
| `clawhdf5-format` | The HDF5 format: parsers and writer (superblock, headers, B-trees, heaps, chunk indexes), the `Storage` trait, the filter pipeline and registry (`filter_registry`), every codec except deflate (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec; pure-Rust LZF, bitshuffle, bzip2, Blosc 1; Blosc2 and ZFP read-only), `float16`, `checksum` |
| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) |
| `clawhdf5-io` | I/O adapters (buffers, mmap, prefetch) |
| `clawhdf5-derive` | `#[derive(H5Type)]` for compound types |
| `clawhdf5` | Facade: `File`, `FileBuilder`, `Dataset`, `FileEditor` (`src/edit/`), SWMR reading (`src/swmr.rs`) |
| `clawhdf5-netcdf4` | NetCDF-4 read support |
| `clawhdf5-remote` | `open_url`: HTTP(S) range requests and object stores (S3, GCS, Azure) through `BlockCache` |
| `clawhdf5-tools` | `h5rs`: `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` |
| `clawhdf5-py` | PyO3 bindings (h5py-like API, remote files, `'r+'` editing) |
| `clawhdf5-wasm` | wasm-bindgen browser reader (`open(bytes)`, `openUrl(url)`); demo in `examples/wasm-viewer/` |
| `clawhdf5-ann` | HNSW index |
| `clawhdf5-agent` | Agent memory store (`HDF5Memory`), sessions, knowledge graph, BM25 |
| `clawhdf5-accel` | CPU SIMD kernels (AVX2, NEON) |
| `clawhdf5-gpu` | wgpu vector distances (WGSL) — not dataset I/O; HDF5 I/O is CPU-only |
| `clawhdf5-migrate` | SQLite → agent store migration |
| `clawhdf5-cli` | Agent-memory CLI |
| `clawhdf5-napi` | Node.js addon (the `packages/clawhdf5-node` wrapper is broken; `docs/known-issues.md`) |
| `clawhdf5-android` | Android JNI bindings |
| `clawhdf5-cli` | Command-line interface (agent memory) |
| `clawhdf5-tools` | `h5rs`: pure-Rust HDF5 tools — `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` (structural + checksum validator) |
| `clawhdf5-napi` | Node.js native addon bindings |
| `clawhdf5-py` | PyO3 Python bindings |
| `clawhdf5-wasm` | WebAssembly (wasm-bindgen) reader for the browser; demo in `examples/wasm-viewer/` |
| `clawhdf5-remote` | Remote files: `open_url` over HTTP(S) range requests and object stores (`object_store`: S3, GCS, Azure) through a mandatory block cache (`BlockCache`) |
| `clawhdf5-bench` | Benchmark suite |
| `clawhdf5-bench` | Benchmarks and harnesses (`search_harness`, `read_harness`, `concurrent_read`, `longmemeval_bench`, …) |
## Key Features
- Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to
pure-Rust zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake).
`ci-test.sh` fails if a C-building crate enters the core crates' default
tree. flate2 must keep `runtime_detection` with zlib-rs — without it zlib-rs
loses SIMD and inflates 3.5x slower. MSRV is 1.92 (`rust-version`, checked
in CI).
- HNSW vector index for semantic similarity search over agent memories — the
`clawhdf5-agent` `hnsw` feature is **on by default**, so `hybrid_search` uses
the approximate `clawhdf5-ann` index for the vector stage (the index mirrors
the cache and self-heals on drift). Build the agent with
`--no-default-features --features float16` to force the exact linear cosine scan.
The agent's `parallel` feature (also default) builds the index on a thread
pool; the graph is identical with or without it.
The index uses the HNSW paper's diversity heuristic for neighbour selection
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
(tied to the checkpoint by a generation id; stale/damaged sidecars are
ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by
default** for new stores, persisted; stores predating the setting load as
`false` and keep their f32 index — guarded by
`tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`)
stores the index's own copy of the embeddings as `i8`,
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
at 100K); because quantised distances are approximate and `ef` cannot
compensate, the query path then re-scores the candidate pool against the
exact embeddings, which holds recall at the f32 index's level. It is also
faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a
Raspberry Pi 5 (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm since
the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64
code is `cfg`'d out on x86, so x86 CI never compiles or lints it — test it
on real ARM (`rpivision02`, 10.0.2.3, is a Pi 5). `hybrid_search` keeps one incremental BM25
index for the life of the store and never writes the store: Hebbian
activation boosts are persisted by the next checkpoint (or on drop), not per
query. Measure any search-path change with
`cargo run --release -p clawhdf5-bench --bin search_harness` (baselines in
`BENCHMARKS.md`).
- WAL (write-ahead log) for crash-safe persistence, with a chained CRC32
trailer per entry (each entry's CRC folds in the previous entry's CRC) so a
corrupted, reordered, duplicated, or spliced entry stops replay cleanly
instead of loading bad or tampered data. The pre-chaining per-entry-CRC
format (v2) is still fully readable; the oldest no-CRC format (v1) is only
reachable through the one-time migration path in `HDF5Memory::open`, not
through the public `WalFile::read_entries`.
**What the WAL guarantees:** integrity, ordering, and recovery from a
*process* crash at any point — including between a checkpoint and the WAL
truncate (each checkpoint records a `WalMark` in `/meta`, and `open()` skips
the WAL prefix the `.h5` already contains, so entries are never applied
twice). Checkpoints and snapshots are made durable as a unit (temp file
synced, renamed, directory synced). **What it does not guarantee:**
individual WAL appends are *not* fsynced (a deliberate latency trade-off), so
saves made since the last checkpoint can be lost on power failure or kernel
panic. Current header version is 4 (adds the `Update` record used by
`save_or_update`); v3 files are read and upgraded in place.
- A store has a **single writer**: `HDF5Memory::create`/`open` hold an exclusive
advisory lock on `<store>.h5.lock` and a second opener gets
`MemoryError::Locked`. Use `HDF5Memory::open_read_only` for a lock-free,
never-writing point-in-time view (the CLI's `recall`/`stats`/`agents-md`/
`export` do). An unreadable WAL (torn header, bad magic) is quarantined to
`<store>.h5.wal.corrupt-<ts>` rather than blocking `open()`; a WAL with an
unknown *newer* version still fails and is left untouched.
- `MemoryConfig::float16` (**on by default** for new stores, persisted;
existing stores keep their recorded `false` — guarded by the v2.5.0
fixture in `tests/float16_store.rs`; CLI opt-out is `create --f32`) writes
`/memory/embeddings` as IEEE half precision (48% smaller file at 100K;
LongMemEval with real MiniLM embeddings identical to f32).
`MemoryCache::half_precision` rounds each embedding as it enters the cache (push, update, WAL replay, and on load of a store still
`f32` on disk), so memory and file agree bit for bit; the conversions live
in `clawhdf5_format::float16` and must stay the single implementation.
Values beyond ±65504 are `MemoryError::InvalidEntry`. Interop: every file
must open in h5py — `f32` datasets and empty datasets did not until
2026-09-23 (see `docs/known-issues.md`); the agent's `h5py_interop` test
guards a whole store.
- `HDF5Memory::search(query_emb, text, &SearchOptions)` is the full search
path: optional source-channel filter (applied before ranking; exact scan of
the allowed records whenever cheaper than `pool × M` index distance
Reference docs: `docs/known-issues.md` (open issues table first — check it
before calling something a bug or a feature), `BENCHMARKS.md` (headline
numbers first), `CONFORMANCE.md` (generated), `docs/design/range-reads.md`
and `docs/design/swmr.md`, `CHANGELOG.md` (full detail of every fix).
## Standing rules
- **No C in the default build.** No libhdf5; deflate defaults to pure-Rust
zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake). `ci-test.sh`
fails if a C-building crate enters the core crates' default tree. Zstd,
SZIP, `https` (ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are opt-in. flate2
must keep `runtime_detection` with zlib-rs — without it zlib-rs loses SIMD
and inflates 3.5x slower.
- **Every file we write must open in h5py/libhdf5.** Interop tests compare
against h5py and h5dump; `f32` and empty datasets did not open until
2026-09-23.
- **float16 has one implementation:** `clawhdf5_format::float16`.
- **Claims need evidence.** Performance and integration claims in docs must
be measured, dated (with machine and command), or withdrawn. Benchmark
numbers are dated records: never edit a measured value, add a new dated
section and mark the old one superseded.
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not and
never was an OpenClaw memory plugin; the old `memory.backend = "clawhdf5"`
config was never valid. `docs/openclaw.md` records what a real plugin would
need. The `openclaw` module's `ClawhdfBackend` is just `search` with
re-rank + confidence on.
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
v0.8.5 and the `osobh/zeroclaw` fork and their history): its memory
backends are its own; `clawhdf5-migrate`'s default SQLite layout is not
ZeroClaw's schema. Don't reintroduce integration claims without an
integration and a test against the real consumer.
- **known-issues.md:** one entry per bug; when fixed, record it in
`CHANGELOG.md` and move the entry to *Fixed (history)* with date, PR,
affected releases and what users must do — never delete it.
## HDF5 library: invariants and gotchas
- **Remote/range reads** (`docs/design/range-reads.md`, M0-M5 merged in PRs
#17-#19, M4 listing costs cut in #21): every format-crate read path goes through `Storage`
(`read_at`/`read_ranges`/`hint`). `File::open_storage` takes any
`Storage`; `clawhdf5_remote::open_url` wraps HTTP (`HttpStorage`, ureq) or
`ObjectStoreStorage` in `BlockCache` (1 MiB blocks, LRU budget, in-flight
dedup, coalesced runs). Remote files are pinned by ETag/Last-Modified and
length (`RemoteError::FileChanged`). Zero-copy APIs and `File::as_bytes`
need an in-memory file. Parse through `File::storage()` and the `*_in`
functions, not `as_bytes`, in new code (the Python bindings do).
`ObjectStoreStorage` runs reads on its own small tokio runtime, so it
works from any thread.
- **SWMR** (`docs/design/swmr.md`): `File::open_swmr` reads a file a libhdf5
SWMR writer is appending to — positioned reads, no chunk cache, bounded
retries (100), `Dataset::refresh()`. clawhdf5 has no SWMR writer; remote
SWMR is out of scope.
- **Browser** (`clawhdf5-wasm`, read-only, no Zstd/SZIP): `openUrl` reads
through the restartable "NeedBytes" cache (`src/lazy.rs`: a call is re-run
after each wave of misses; no block is evicted while a call runs); the HTTP
is JavaScript (`js/remote.js`).
- **In-place editing** (`clawhdf5::FileEditor`): overwrites values, grows and
shrinks chunked datasets (every chunk index) and sets attributes (compact
and dense) without rewriting the file, changing indexes and heaps as
libhdf5 does; freed space is reused within one editor. Anything it cannot
do safely is `Error::Unsupported` before any write (limits in
`docs/known-issues.md`). The algorithms follow libhdf5 `hdf5_1_14_6`
(github.com/HDFGroup/hdf5). Test changes with `cargo test -p
clawhdf5-tools --test edit_interop --test edit_coverage_interop`.
- **Provenance:** `Dataset::verify_provenance()` (facade `provenance`
feature, default) re-hashes a dataset against its `_provenance_sha256`
attribute (`DatasetBuilder::with_provenance`). Opt-in per call; unkeyed
hash — tamper-evident, not tamper-proof.
## Agent memory: invariants and gotchas
- **Search.** `HDF5Memory::search(query_emb, text, &SearchOptions)` is the
full path: optional source-channel filter (before ranking; exact scan of
the allowed records when cheaper than `pool × M` index distance
evaluations, and as the fallback when the pool comes back short), fusion,
activation scaling, optional re-ranking and confidence rejection.
`hybrid_search`/`hybrid_search_with` are thin wrappers; `ClawhdfBackend`
(the `openclaw` module) is `search` with re-rank + confidence on.
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not an
OpenClaw memory plugin and never was — the old `memory.backend = "clawhdf5"`
config was never valid. Don't reintroduce OpenClaw claims; `docs/openclaw.md`
records what a real plugin would need.
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
v0.8.5 and the `osobh/zeroclaw` fork, and their full history): no
`clawhdf5` feature or backend exists; ZeroClaw's memory backends are
sqlite/lucid/postgres/qdrant/markdown/none behind its own `Memory` trait.
`clawhdf5-migrate`'s default SQLite layout (`memory_chunks`, `sessions`,
`entities`, `relations`) is not ZeroClaw's schema either (ZeroClaw's is a
`memories` table). Don't reintroduce integration claims without an
integration and a test against the real consumer. Measure changes with
`search_harness --options-study`.
- `MemoryConfig::compression` is off by default; when on, embeddings are
deflate-compressed, or Zstd with the agent's `zstd` feature (links libzstd).
- Signed checkpoints (`clawhdf5-agent` `signing` module): with
`HDF5Memory::set_signing_key` every checkpoint stores an Ed25519-signed
manifest (SHA-256 per record in a Merkle tree + settings/sessions/graph
hashes; per-record hashes in `/integrity/record_hashes`);
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover exactly
what the file persists in the form the loader returns it (strings lose
trailing NULs; an empty WAL mark is not written) or untouched stores stop
verifying — `tests/signed_store.rs` round-trips awkward strings. The key is
never persisted; a signed store refuses to checkpoint without it
(`MemoryError::SigningKeyRequired`, and `MemoryError` is `#[non_exhaustive]`).
`hybrid_search`/`hybrid_search_with` are thin wrappers. It keeps one
incremental BM25 index for the life of the store and never writes the
store: Hebbian activation boosts are persisted by the next checkpoint (or
on drop).
- **HNSW** (`hnsw` feature, default): the approximate `clawhdf5-ann` index
mirrors the cache and self-heals on drift; build the agent with
`--no-default-features --features float16` for the exact linear scan.
`parallel` (default) builds it on a thread pool with an identical graph.
Neighbour selection uses the HNSW paper's diversity heuristic (closest-M
capped recall at 0.31 recall@10 at 100K on clustered data). The graph is
saved to `<store>.h5.ann` at each checkpoint, tied to it by a generation
id; a stale or damaged sidecar is ignored and the index rebuilt.
- **`MemoryConfig::quantized_index`** (default on for new stores, persisted;
older stores load as `false` — guarded by `tests/fixtures/store_v2_5_0.h5`;
CLI `create --f32-index`): the index's copy of the embeddings is `i8`, and
the query path re-scores candidates against the exact embeddings. The
aarch64 kernels (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm) are
`cfg`'d out on x86, so x86 CI never compiles them — test on real ARM
(`rpivision02`, 10.0.2.3, a Pi 5) or rely on the `test-arm64` job.
- **`MemoryConfig::float16`** (default on for new stores, persisted; older
stores keep `false` — guarded in `tests/float16_store.rs`; CLI `create
--f32`): `/memory/embeddings` is IEEE half. `MemoryCache::half_precision`
rounds each embedding as it enters the cache (push, update, WAL replay, and
load of a store still `f32` on disk) so memory and file agree bit for bit.
Values beyond ±65504 are `MemoryError::InvalidEntry`. The agent's
`h5py_interop` test guards that a whole store opens in h5py.
- **WAL.** Chained CRC32 per entry (a corrupted, reordered, duplicated or
spliced entry stops replay cleanly). Header version 4 (`Update` record for
`save_or_update`); v3 is upgraded in place, v2 read, v1 only through the
one-time migration in `HDF5Memory::open`. Each checkpoint records a
`WalMark` in `/meta` so `open()` never applies an entry twice; checkpoints
and snapshots are durable as a unit (temp file synced, renamed, directory
synced). Individual WAL appends are **not** fsynced (deliberate): saves
since the last checkpoint can be lost on power failure or kernel panic.
- **Single writer.** `create`/`open` hold an exclusive lock on
`<store>.h5.lock` (`MemoryError::Locked` for a second opener);
`open_read_only` is a lock-free point-in-time view (CLI `recall`/`stats`/
`agents-md`/`export`). An unreadable WAL is quarantined to
`<store>.h5.wal.corrupt-<ts>`; a WAL of an unknown newer version fails and
is left untouched.
- **Signed checkpoints** (`signing` module): with `set_signing_key` each
checkpoint stores an Ed25519-signed manifest (per-record SHA-256 in a
Merkle tree plus settings/sessions/graph hashes; `/integrity/record_hashes`);
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover
exactly what the file persists in the form the loader returns it (strings
lose trailing NULs; an empty WAL mark is not written) —
`tests/signed_store.rs` round-trips awkward strings. The key is never
persisted; a signed store refuses to checkpoint without it
(`MemoryError::SigningKeyRequired`; `MemoryError` is `#[non_exhaustive]`).
WAL entries after the checkpoint are not covered.
- `Dataset::verify_provenance()` (clawhdf5 facade, `provenance` feature, on by
default) recomputes a dataset's SHA-256 and compares it against the
`_provenance_sha256` attribute written automatically on save when
`DatasetBuilder::with_provenance` is used. It's opt-in per call, not run
automatically on open — it decodes and hashes the whole dataset. The hash
is unkeyed (tamper-*evident*, not tamper-*proof*): it detects accidental
corruption, not a deliberate actor able to modify both the data and the
stored hash.
- `clawhdf5-agent`'s `HDF5Memory::save`/`save_batch`/`save_or_update` run every
write through an in-memory (session-scoped, not persisted to disk)
provenance ledger and write-anomaly detector: a content hash per record
(`provenance.rs`) for detecting accidental mid-session corruption, plus
rate-limit/injection-pattern/source-distribution checks (`anomaly.rs`).
Alerts never block a save — drain them with `HDF5Memory::take_anomaly_alerts`.
`MemorySource` for this bookkeeping is inferred from the caller-supplied
`source_channel` string (a heuristic, not an authenticated trust boundary).
- In-place modification: `clawhdf5::FileEditor` (`crates/clawhdf5/src/edit/`)
overwrites values, grows and shrinks chunked datasets (every chunk index,
version-2 B-trees included) and sets attributes (compact and dense
storage) in existing files (h5py- or clawhdf5-written) without rewriting
them, changing indexes and heaps as libhdf5 does (index shapes and heap
bookkeeping are compared with libhdf5's in the tests); space an edit
frees is reused by later edits of the same editor. Anything it cannot do
safely is `Error::Unsupported` before any write (limits in
`docs/known-issues.md`). Test changes with
`cargo test -p clawhdf5-tools --test edit_interop --test
edit_coverage_interop` (h5py, h5dump, `h5rs check`, structure comparisons
with libhdf5; libhdf5 sources for the algorithms are at
github.com/HDFGroup/hdf5, tag `hdf5_1_14_6`).
- Remote files (`clawhdf5-remote`, range-read milestone M3 of
`docs/design/range-reads.md`): `open_url("http://…")` gives a
`clawhdf5::File` over `File::open_storage`, read through `BlockCache`
(1 MiB blocks, LRU byte budget, per-block in-flight dedup across threads,
runs coalesced into parallel requests). `HttpStorage` pins the file by
ETag/Last-Modified and length (a change is `RemoteError::FileChanged`),
refuses servers that ignore `Range` unless a full download is allowed,
and retries transient failures. `ObjectStoreStorage` (feature
`object-store`, pure Rust) runs each read on a small owned tokio
runtime and waits on a channel, so it works from any thread, including
inside `spawn_blocking` or another runtime. Default build is plain HTTP with
no C; `https` (rustls + ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are
opt-in. Tests run a std-only HTTP server
(`tests/common/server.rs`, also the `range_server` example);
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
call is re-run after each wave of misses; no block evicted while a call
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
version) and tests it under Node and headless Chromium (a Playwright
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
(range server with request counts, 200 MB budget file); the CI container
has neither, so CI runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
numbers in the example's README predate `openUrl`.
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
- **Write bookkeeping.** `save`/`save_batch`/`save_or_update` feed an
in-memory, session-scoped provenance ledger and anomaly detector
(`provenance.rs`, `anomaly.rs`); alerts never block a save
(`take_anomaly_alerts`). `MemorySource` is inferred from the caller's
`source_channel` string — a heuristic, not a trust boundary.
- `MemoryConfig::compression` is off by default (deflate, or Zstd with the
agent's `zstd` feature, which links libzstd).
## Workflows
### Build
Put `$HOME/.cargo/bin` on `PATH`. The h5py/netCDF4 interop tests find their
Python through `CLAWHDF5_PYTHON` (or `.venv/bin/python`); create it with
`python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 hdf5plugin`.
Set `CLAWHDF5_REQUIRE_INTEROP=1` to make a missing interpreter a failure.
```bash
cargo build --release
```
### Test
```bash
cargo test --workspace
bash scripts/ci-test.sh # everything CI runs (see below)
```
### CI
`.gitea/workflows/ci.yml` has two jobs, both green as of 2026-09-22:
- **`test`** (`ubuntu-latest`, in `rust:latest`) runs `scripts/ci-test.sh` with
the h5py/netCDF4 interop suites required (`CLAWHDF5_REQUIRE_INTEROP=1`).
Served by the `tank` and `architect` runners.
- **`test-arm64`** (`linux_arm64`) lints and tests the aarch64 code — the NEON
kernels are `cfg`'d out on x86, so this is the only place they are built.
Served by `vision-01` (host mode) and `vision-02` (Docker), so steps must
work in both.
### CI (`.gitea/workflows/`)
- **`ci.yml` `test`** (`ubuntu-latest`, `rust:latest` container; runners
`tank`, `architect`): installs h5py/netCDF4/xarray/hdf5plugin/maturin/pytest,
`hdf5-tools` and `cmake`, then runs `scripts/ci-test.sh` with
`CLAWHDF5_REQUIRE_INTEROP=1`. The script runs: fmt; clippy (workspace, the
format feature matrix, each plugin filter alone, parallel, fast-deflate,
remote with all backends, h5rs remote); "no C in the default build";
wasm32 build and clippy; `check-32bit-casts.sh`; the wasm package under
Node when `node` and `wasm-bindgen` exist (not in CI); the MSRV check;
`cargo test` (workspace plus feature variants: format matrix, parallel,
remote/object_store, h5rs URLs, ann parallel, fast-deflate); the h5py
interop suites (`writer_h5py_tests --include-ignored`, plugin filters,
ZFP); the Python package (clippy, `maturin build`, pytest vs h5py);
`cargo bench --no-run`; `check-nostd.sh`; an optional fuzz smoke run
(`CLAWHDF5_FUZZ_SECONDS`).
- **`ci.yml` `test-arm64`** (`linux_arm64`; `vision-01` host mode,
`vision-02` Docker — steps must work in both): clippy of
`clawhdf5-accel`, tests of `-accel`, `-ann`, `-format`; the only place the NEON kernels build.
- **`conformance.yml`** (nightly 03:17 UTC and manual): probe unit tests,
`conformance/test_ref.py`, then `conformance/run.sh` (gate:
`conformance/check.py` against `baseline.json`).
Keep workflows free of JavaScript actions (`actions/checkout`, `actions/cache`,
…): `rust:latest` has no `node`, and not every runner reaches GitHub, where
they are fetched from. Check out with plain `git` instead. The `test` job
installs `cmake` for the opt-in `fast-deflate` (zlib-ng) steps; the default
build needs no C toolchain, so `test-arm64` does not.
All runners are on `gitea-runner` 3.5.0, from `docker.gitea.com/act_runner`
— `gitea/act_runner:latest` on Docker Hub is frozen at 0.6.1.
Keep workflows free of JavaScript actions (`actions/checkout`,
`actions/cache`, …): `rust:latest` has no `node` and not every runner reaches
GitHub. Check out with plain `git`. Runners are `gitea-runner` 3.5.0 from
`docker.gitea.com/act_runner` (`gitea/act_runner:latest` on Docker Hub is
frozen at 0.6.1).
### CLI
### Conformance
```bash
cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands
CLAWHDF5_PYTHON=.venv/bin/python bash conformance/run.sh --no-fetch # writes CONFORMANCE.md
```
Reads 697 files of eight pinned corpora with clawhdf5 and h5py and compares
them object by object (602 ok in the run of 2026-09-28). `CONFORMANCE.md` is
generated — never hand-edit it (its wording lives in `conformance/report.py`).
Use `--update-baseline` only after an intended change in results.
`CONFORMANCE_CACHE` points at an existing corpus cache (`conformance/.cache`,
about 450 MB). See `conformance/README.md`.
### HDF5 tools (`h5rs`, crate `clawhdf5-tools`)
### HDF5 tools (`h5rs`)
```bash
cargo run -p clawhdf5-tools -- ls -r file.h5 # also dump [--json], stat, diff, check
bash scripts/h5rs-fuzz.sh # every subcommand over the CVE corpus: no panic/crash/hang
bash scripts/h5rs-check-ok-files.sh --data # check passes every fully-read conformance file
```
Its interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian
`hdf5-tools`, installed in CI); `dump` must stay byte-identical to h5dump on
the test files.
Interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian `hdf5-tools`);
`dump` must stay byte-identical to h5dump on the test files.
### Remote and browser tests
- `clawhdf5-remote` tests run a std-only HTTP server
(`tests/common/server.rs`, also the `range_server` example);
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- wasm: `bash examples/wasm-viewer/test/run.sh` builds the package (needs the
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node and
headless Chromium (Playwright's download in `~/.cache/ms-playwright` on
tank) against `test/serve.py` (range server with request counts). CI has
neither, so it runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus).
### Python bindings
```bash
cd crates/clawhdf5-py
maturin develop
python -c "import clawhdf5; print(clawhdf5.__version__)"
cd crates/clawhdf5-py && maturin develop
python -m pytest crates/clawhdf5-py/tests # compares with h5py; editing tests want CLAWHDF5_H5RS=<path to h5rs>
```
### Benchmarks
- Search path: `cargo run --release -p clawhdf5-bench --bin search_harness`
(`--full`, `--options-study`, `--footprint`, …); reads: `read_harness`,
`concurrent_read`; criterion benches with `cargo bench -p <crate>`.
- Run on an idle machine (1-minute load average below 2; wait otherwise),
alternate base and candidate binaries for A/B comparisons, and record date,
machine, commit and command with every number in `BENCHMARKS.md`.
- `BENCHMARKS.md` is written by hand from dated runs; no script regenerates
it (the old `scripts/run-benchmarks.sh`, which benchmarked the pre-rename
`rustyhdf5-format` and overwrote the file, was removed on 2026-09-28).
### CLI
```bash
cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot, keygen, verify
```
## Integration
@@ -255,8 +269,7 @@ python -c "import clawhdf5; print(clawhdf5.__version__)"
verified consumer: `cbh-core` reads and writes `.brain` files through the
facade (`File`, `FileBuilder`, `AttrValue`, `Selection`), `cbh-scanner`
uses the facade, and `cbh-cli` uses `clawhdf5_agent::bm25::BM25Index`. It
depends on this repo by path (`../clawhdf5`), so it builds against whatever
is checked out — changes to those APIs reach it directly. Verified
2026-09-25 against main: builds, and its 204 tests pass.
- OpenClaw and ZeroClaw were both described as consumers; neither integrates
clawhdf5 (see Key Features and `docs/openclaw.md`).
depends on this repo by path (`../clawhdf5`), so changes to those APIs
reach it directly. Verified 2026-09-25 against main: builds, and its 204
tests pass.
- OpenClaw and ZeroClaw integrate nothing (see *Standing rules*).
+47 -33
View File
@@ -13,15 +13,15 @@ fatal. This file is generated by `conformance/run.sh`; do not edit it by hand.
| | |
|---|---|
| date | 2026-09-27 00:34 UTC |
| clawhdf5 commit | `f37e7ae3263277319dba4bc39be5397194eb00c3` |
| date | 2026-09-28 04:29 UTC |
| clawhdf5 commit | `bf5a163dcf7fe28d651ada6545d8136ffeffc825` |
| machine | `tank`: AMD Ryzen 7 7800X3D 8-Core Processor, 16 CPUs, 61 GiB, Linux 7.0.0-34-generic x86_64 |
| command | `conformance/run.sh --no-fetch --update-baseline` |
| rustc | rustc 1.98.1 (48a229cea 2026-09-01) |
| reference | h5py 3.16.0, HDF5 2.0.0, numpy 2.5.3, hdf5plugin 7.1.0, Python 3.14.4 |
| h5dump | Version 1.14.6 (CVE corpus only) |
| limits | 20 s timeout (SIGKILL), 4096 MiB address space, per process; 16 files in parallel |
| runtime | 21 s probing + comparing (0 s fetch/build before it) |
| runtime | 20 s probing + comparing (5 s fetch/build before it) |
## Results
@@ -29,25 +29,24 @@ A file's class is the first that applies:
- **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these.
- **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers.
- **ref-bug** — every difference is an object clawhdf5 refuses that h5py reads only through a libhdf5 bug: the values h5py returns for it change with the reading process's heap, re-checked in every run (see *Reference bugs*).
- **our-error** — clawhdf5 returned an error for something h5py reads.
- **mismatch** — both read it, but the shapes, values, object set or attribute set differ.
- **ok** — every object h5py reads, clawhdf5 reads identically.
| corpus | files | ok | our-error | mismatch | h5py-cannot-read | panic | hang | crash | oom |
|---|---|---|---|---|---|---|---|---|---|
| NCAS-CMS_pyfive | 33 | 32 | 0 | 1 | 0 | 0 | 0 | 0 | 0 |
| cve_hdf5 | 147 | 113 | 2 | 0 | 32 | 0 | 0 | 0 | 0 |
| h5py_data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| hdf5 | 466 | 404 | 1 | 1 | 60 | 0 | 0 | 0 | 0 |
| netcdf-c | 20 | 20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| netcdf4-python | 18 | 18 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| usnistgov_h5wasm | 5 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| xarray-data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| **all** | **697** | **600** | **3** | **2** | **92** | **0** | **0** | **0** | **0** |
| corpus | files | ok | our-error | mismatch | h5py-cannot-read | ref-bug | panic | hang | crash | oom |
|---|---|---|---|---|---|---|---|---|---|---|
| NCAS-CMS_pyfive | 33 | 33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| cve_hdf5 | 147 | 113 | 0 | 0 | 32 | 2 | 0 | 0 | 0 | 0 |
| h5py_data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| hdf5 | 466 | 405 | 1 | 0 | 60 | 0 | 0 | 0 | 0 | 0 |
| netcdf-c | 20 | 20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| netcdf4-python | 18 | 18 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| usnistgov_h5wasm | 5 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| xarray-data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| **all** | **697** | **602** | **1** | **0** | **92** | **2** | **0** | **0** | **0** | **0** |
2 of the 2 mismatches are a known h5py bug, not ours (see *Known not-our-bug*).
3 of the 3 our-errors are corrupt data that HDF5 2.0 reads only through a bug and clawhdf5 refuses (see *Known not-our-bug*).
**Our errors and mismatches: 1.** Files not ok: 1 our-error, 92 h5py-cannot-read, 2 ref-bug. 2 object(s) were compared against h5py's values corrected for a known h5py bug (2 identical to clawhdf5's; see *Reference bugs*).
Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):
@@ -72,14 +71,11 @@ Grouped by normalised error message. *files* counts files whose class this cause
| files | objects | error | examples |
|---:|---:|---|---|
| 3 | 3 | `ChunkedReadError("…")` | `cve_hdf5/cvefiles/cve-2025-2308.h5`, `cve_hdf5/cvefiles/cve-2025-44904.h5`, `hdf5/test/testfiles/bad_nbit_parms_walk.h5` |
| 1 | 1 | `ChunkedReadError("…")` | `hdf5/test/testfiles/bad_nbit_parms_walk.h5` |
## Mismatch root causes
| files | objects | cause | examples |
|---:|---:|---|---|
| 1 | 1 | `attr-values: ours=vlen(>u8) h5py=object layout=- filters=-` | `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` |
| 1 | 1 | `values: ours=vlen({r:>f4,i:>f4}8) h5py=object layout=contiguous filters=-` | `hdf5/tools/test/testfiles/tcomplex_be.h5` |
None.
## CVE corpus: clawhdf5 vs h5dump vs h5py
@@ -203,7 +199,7 @@ columns are.
| cvefiles/cve-2024-33876.h5 | ok | read 3 obj, 1 errors | read 3 obj, 1 errors | ok |
| cvefiles/cve-2024-33877.h5 | error exit | read 8 obj, 1 errors | read 8 obj, 1 errors | ok |
| cvefiles/cve-2025-2153.h5 | error exit | open error | read 1 obj, 1 errors | h5py-cannot-read |
| cvefiles/cve-2025-2308.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | our-error |
| cvefiles/cve-2025-2308.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | ref-bug |
| cvefiles/cve-2025-2309.h5 | ok | read 6 obj, 1 errors | read 6 obj | ok |
| cvefiles/cve-2025-2310.h5 | error exit | read 24 obj, 8 errors | read 24 obj, 8 errors | ok |
| cvefiles/cve-2025-2912.h5 | error exit | open error | open error | h5py-cannot-read |
@@ -214,7 +210,7 @@ columns are.
| cvefiles/cve-2025-2924.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-2925.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-2926.h5 | error exit | open error | open error | h5py-cannot-read |
| cvefiles/cve-2025-44904.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | our-error |
| cvefiles/cve-2025-44904.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | ref-bug |
| cvefiles/cve-2025-44905.h5 | error exit | read 25 obj, 3 errors | read 25 obj, 3 errors | ok |
| cvefiles/cve-2025-6269-1.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-6269-2.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
@@ -249,13 +245,36 @@ columns are.
</details>
## Known not-our-bug
## Reference bugs
### Objects h5py reads only through a libhdf5 bug (*ref-bug*)
clawhdf5 refuses these objects; h5py 3.16 / HDF5 2.0 returns values for them. `conformance/ref_bugs.py`
re-reads each with h5py in six fresh processes whose heaps differ (h5py imported before numpy, three
times and twice more with `MALLOC_PERTURB_`, and numpy imported first). Values the file determines
come out the same every time; these do not, so they are memory libhdf5 over-reads, not the file's
data. A file is *ref-bug* only while every one of its differences is such an object confirmed in
the same run; an object that reads the same every time goes back to *our-error*. Reproducer:
`python conformance/ref_bugs.py conformance/.cache/corpus` (prints every read's outcome).
| file | object | distinct results in 6 reads | confirmed | what goes wrong |
|---|---|---:|---|---|
| `cve_hdf5/cvefiles/cve-2025-2308.h5` | `/Scale_offset_long_long_data_le` | 6 | yes | the first chunk records minbits 11: its 12 values need 17 bytes of codes, and the 26-byte chunk holds 5 after its 21-byte header; libhdf5's scale-offset decoder reads past its buffer, and develop refuses the chunk ("Buffer too short") |
| `cve_hdf5/cvefiles/cve-2025-44904.h5` | `/Scale_offset_float_data_le` | 6 | yes | unfiltered chunks stored as 38 and 37 bytes for 48-byte chunks: 1.14/2.0 read the stored bytes into a buffer of that size and use it as the whole chunk (H5D__chunk_lock), so the rest is heap memory; develop refuses them ("incorrect chunk size returned from index for unfiltered chunk") |
| `hdf5/test/testfiles/bad_nbit_parms_walk.h5` | `/Nbit_int_data_le` | 1 | **no** | the N-Bit parameter list holds 7 values (cd_values[0] = 7) where an integer needs 8: the decoder takes the bit offset from cd_values[7], past the list; libhdf5's own test (`test_filter_bad_params`, test/dsets.c on develop) requires the read to fail |
### Values corrected for a known h5py bug
- **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence
whose base type is big-endian with the file's big-endian bytes but a native (little-endian)
numpy dtype, so the values it reports are byte-swapped garbage; `h5dump` prints the values
clawhdf5 reads. Reproducer: `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]`
reads back in h5py as `[4.6e-41, 9.0e-44]`. Affected here: `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5`, `hdf5/tools/test/testfiles/tcomplex_be.h5`.
numpy dtype: a `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]` reads back as
`[4.6e-41, 9.0e-44]`; `h5dump` prints the file's values. `ref.py` checks that the installed
h5py still does this (by writing and reading exactly that dataset in memory) and, if so,
relabels such elements with the file's byte order before hashing, so the values are still
compared. Corrected objects: `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64` (same as clawhdf5), `hdf5/tools/test/testfiles/tcomplex_be.h5` `/VariableLengthDatasetFloatComplex` (same as clawhdf5).
## Other comparison rules
- **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose
bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a
bit offset / reduced precision into the plain numpy type of the same size. The probe
@@ -264,11 +283,6 @@ columns are.
- **Types h5py widens.** Where h5py reads a type into a numpy type of a different size
(FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not
compared (shape and presence still are): dataset file type size 1 -> numpy float16 (2) (15x), attr file type size 1 -> numpy float16 (2) (15x), dataset file type size 2 -> numpy float32 (4) (2x), dataset file type size 8 -> numpy float128 (16) (1x), dataset file type size 12 -> numpy float128 (16) (1x), attr file type size 2 -> numpy float32 (4) (1x), dataset file type size 2 -> numpy >f4 (4) (1x), attr file type size 2 -> numpy >f4 (4) (1x).
- **Corrupt data HDF5 2.0 reads through a bug.** clawhdf5 refuses these objects; h5py 3.16 /
HDF5 2.0 returns values for them that the file does not hold:
- `cve_hdf5/cvefiles/cve-2025-2308.h5` `/Scale_offset_long_long_data_le`: scale-offset codes run past the end of the chunk: HDF5 2.0 reads past its buffer; libhdf5's develop branch refuses the chunk ("Buffer too short").
- `cve_hdf5/cvefiles/cve-2025-44904.h5` `/Scale_offset_float_data_le`: unfiltered chunks of 38 and 37 bytes for 48-byte chunks: HDF5 2.0 fills the rest with whatever its buffer held; libhdf5's develop branch refuses them ("incorrect chunk size returned from index for unfiltered chunk").
- `hdf5/test/testfiles/bad_nbit_parms_walk.h5` `/Nbit_int_data_le`: an N-Bit parameter list one value short: HDF5 2.0 reads past the list; libhdf5's own test (`test_filter_bad_params`, test/dsets.c) now requires the read to fail.
- **References** are compared by presence only (`R`), not by target.
## Objects h5py fails on but clawhdf5 reads
+389 -1057
View File
File diff suppressed because it is too large Load Diff
+144 -177
View File
@@ -1,193 +1,160 @@
# ClawhDF5 Roadmap — Agent Memory Evolution
# clawhdf5 roadmap
> Making clawhdf5 the defacto agentic memory solution.
> Single file. Pure Rust. Zero dependencies. Trusted everywhere.
What has shipped, and what is genuinely next. Everything here is checked
against `CHANGELOG.md`, `git log` and [`docs/known-issues.md`](docs/known-issues.md);
dates are merge dates on `main`. Nothing after v2.7.0 has been released:
the work since then is on `main` under `CHANGELOG.md` "Unreleased".
_Last updated: 2026-09-28 (at `9b5803f`, PR #21)._
---
## Track 1: Knowledge Graph in HDF5
**Status:** 🟢 Phase 1 Complete
**Priority:** Critical
**Crate:** `clawhdf5-agent`
## Done
- [x] **1.1** Entity storage — entities with properties, embeddings, timestamps (created_at/updated_at)
- [x] **1.2** Relation storage — typed edges with RelationType enum (Temporal/Causal/Associative/Hierarchical/Custom), metadata, timestamps
- [x] **1.3** Entity extraction helpers — rule-based extraction (Person, Org, Location, Date, Technology, Project) with extract_and_store_entities() integration
- [x] **1.4** Entity resolution — fuzzy name matching (Levenshtein distance) via resolve_or_create()
- [x] **1.5** Graph traversal queries — BFS neighbors with depth, subgraph extraction from seeds
- [x] **1.6** Spreading activation — weighted activation propagation with configurable decay
- [x] **1.7** Graph-aware retrieval — get_entity_context() for formatted context injection
- [x] **1.8** Tests — comprehensive tests for all new features
### Releases
**Research:** Graph-Native Cognitive Memory (2026), Graph-based Agent Memory survey (2026), SYNAPSE (2025)
| Version | Date | Headline |
|---|---|---|
| v2.0.0 | 2026-03-19 | rustyhdf5 (11 crates) and edgehdf5 (4 crates) unified into one workspace as `clawhdf5-*` |
| v2.1.0 | 2026-06-03 | HNSW backs the agent's vector search by default; live, mutable HNSW index |
| v2.2.0 – v2.7.0 | 2026-09-18 – 2026-09-20 | bounded decompression and read-path bounds checks, single-writer store locking, WAL v4, HNSW recall fix (0.31 -> 0.98 recall@10 at 100K), fusion weights tuned on LongMemEval, int8 index, Extensible Array read fix and chunk-index checksums |
Details per release: [`CHANGELOG.md`](CHANGELOG.md).
### Since v2.7.0 (unreleased, on `main`)
| PR | Merged | What |
|---|---|---|
| #3 | 2026-09-23 | pure-Rust deflate (zlib-rs) by default, no C in the core crates' default build (checked in CI), MSRV 1.92 |
| #4 | 2026-09-25 | files open in h5py again (every `f32` and every empty dataset clawhdf5 wrote was unreadable by libhdf5); float16 embedding storage |
| #5 | 2026-09-25 | `HDF5Memory::search` with `SearchOptions` (source filters, re-ranking, confidence); float16 on by default |
| #6 | 2026-09-25 | `clawhdf5-migrate` writes real agent stores; knowledge-graph fix; dated benchmark re-run |
| #7 | 2026-09-25 | consolidation benchmark completed (cheaper novelty scoring) |
| #8 | 2026-09-25 | Ed25519-signed checkpoints (`HDF5Memory::verify`) |
| #9, #10 | 2026-09-25 | OpenClaw and ZeroClaw integration claims withdrawn — neither ever integrated clawhdf5 |
| #11 | 2026-09-26 | silent wrong data and libhdf5 interop bugs found by the HDF5 audit fixed |
| #12 | 2026-09-26 | reproducible conformance sweep over eight public corpora, nightly CI job ([`CONFORMANCE.md`](CONFORMANCE.md)) |
| #13 | 2026-09-26 | reads HDF5 1.6-era layouts, user blocks, virtual datasets, dense attributes, very large groups |
| #14 | 2026-09-26 | `h5rs` tools (`ls`, `dump`, `stat`, `diff`, `check`), the browser reader (`clawhdf5-wasm`), libhdf5's header checks, plugin filters (LZF, bitshuffle, bzip2, Blosc), concurrency benchmark |
| #15 | 2026-09-26 | fast contiguous and concurrent reads, variable-length data, nested groups and links in the writer, Python bindings |
| #16 | 2026-09-26 | chunked full reads faster than an h5py process pool, writer B-trees of any size, Blosc2 (read), 599/697 conformance |
| #17 | 2026-09-26 | range reads M0/M1 (indexed name lookups, the `Storage` trait), ZFP (read), in-place editing (`FileEditor`) |
| #18 | 2026-09-27 | range reads M2/M3 (`File::open_storage`; `clawhdf5-remote`: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes |
| #19 | 2026-09-27 | remote files in the browser (`openUrl`, M4), SWMR reader (`File::open_swmr`, M5), Python remote reads and `'r+'` editing |
| #20 | 2026-09-28 | benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine |
| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 mismatch (the run of 2026-09-28 in [`CONFORMANCE.md`](CONFORMANCE.md) still counts 1 our-error, a corrupt N-Bit file libhdf5's own tests refuse) |
### Range reads (design: [`docs/design/range-reads.md`](docs/design/range-reads.md))
- [x] M0 — indexed name lookups (#17)
- [x] M1 — metadata parsed through the `Storage` trait (#17)
- [x] M2 — raw data through `Storage`, `File::open_storage` (#18)
- [x] M3 — `clawhdf5-remote`: HTTP(S) range requests and object stores through a block cache; `h5rs` URLs (#18); Python URLs (#19)
- [x] M4 — `openUrl` in the browser, restartable "NeedBytes" cache (#19; fewer round trips in #21)
- [x] M5 — reading files a SWMR writer is appending to ([`docs/design/swmr.md`](docs/design/swmr.md), #19)
### Agent memory (`clawhdf5-agent`)
Shipped before and during the v2 releases, and kept current since:
knowledge graph with entity extraction and resolution; three-tier
consolidation with decay; hybrid retrieval (HNSW + BM25, weighted or RRF
fusion, re-ranking, confidence rejection, query expansion); temporal index
and session DAG; per-save provenance ledger and write-anomaly detection;
multi-modal embeddings; WAL with chained CRC32; single-writer locking;
signed checkpoints. Retrieval is measured, not claimed: see
[`BENCHMARKS.md`](BENCHMARKS.md) ("LongMemEval Results" reports retrieval
recall, not QA accuracy; earlier headline numbers that compared different
granularities were retracted there).
---
## Track 2: Memory Consolidation Engine
**Status:** 🟢 Phase 1 Complete
**Priority:** Critical
**Crate:** `clawhdf5-agent`
## Next
- [x] **2.1** Importance scoring — surprise (novelty), correction boost, length scoring with configurable weights
- [x] **2.2** Three-tier memory model — Working → Episodic → Semantic with bounded capacities
- [x] **2.3** Time-decay with reactivation — exponential decay with configurable half-life, access resets timestamp
- [x] **2.4** Bounded memory with graceful degradation — evict lowest-decay entries when over capacity
- [x] **2.5** Consolidation cycles — promote/evict across tiers based on importance and access thresholds
- [x] **2.6** Memory statistics — ConsolidationStats with per-tier counts, eviction/promotion tracking
- [x] **2.7** Tests — comprehensive tests for all features
Not scheduled; listed roughly by how much they unblock. None has a date.
**Research:** CraniMem (2026), D-MEM (2026), AI Hippocampus survey (2026)
### Distribution
- [ ] **Publish the crates to crates.io.** Nothing is published; the READMEs
say to depend on git. Before publishing: no `publish` settings exist
(only `clawhdf5-wasm` has `publish = false`).
- [ ] **Publish Python wheels to PyPI.** `crates/clawhdf5-py` builds with
maturin and is tested in CI, but no wheel is published. The default wheel
reads plain `http://` only; `https`/`s3`/`gcs`/`azure` wheels compile C
(ring, aws-lc-rs).
- [ ] **The Node.js package** (`packages/clawhdf5-node` over
`clawhdf5-napi`) has never worked and is not in CI: fix it and add CI, or
remove it ([known issue](docs/known-issues.md)).
### HDF5 features
- [ ] **SWMR writing.** The reader is done (M5); writing a file while
libhdf5 readers follow it is not. Also not covered: remote SWMR (a remote
file is pinned at open), `MmapFile`/`LazyFile` SWMR reads, refreshing
groups or attributes.
- [ ] **MPI collective I/O.** `clawhdf5-io`'s `MpiVol` (`mpi-io`) is
root-read + broadcast and gather-to-root writes, not collective MPI-IO
(`MPI_File_read_at_all`/`write_at_all`).
- [ ] **Paged-metadata single-request reads.** Files written with paged
aggregation (`H5Pset_file_space_strategy(PAGE)`, `h5repack -S PAGE`)
keep their metadata in a few pages; range reads could fetch those in one
request and use the file's page size as the block size. Today the block
size is fixed (1 MiB) and only the first block is read ahead
(range-reads design, option (c) as a policy).
- [ ] **Blosc2 and ZFP encoders.** Both filters are read-only; the other
plugin filters (LZF, bitshuffle, bzip2, Blosc 1) read and write.
- [ ] **External links and external raw data** are explicit errors, not
followed.
- [ ] **Virtual datasets:** the "first missing" view and printf gaps other
than 0, source-to-virtual type conversion other than a byte swap, nested
virtual sources, source files outside the virtual file's directory.
- [ ] **Datatypes:** x87 long double and binary128 are refused.
- [ ] **Writer:** one attribute or link message over 65 515 bytes in dense
storage is an error (huge fractal-heap objects); no option to write
files HDF5 1.8 can read.
- [ ] **`FileEditor`:** new chunks in implicit indexes, variable-length and
reference data, filters it cannot encode (scale-offset, N-Bit, SZIP),
some dense-attribute heap layouts, creating or deleting objects and
attributes (also from Python `'r+'`), and no journal (a crash mid-edit
can leave the file inconsistent). Freed space is reused only within one
editor.
- [ ] **Selection reads** decode the whole dataset when the selection's
bounding box covers more than half of it (a strided `ds[::100]`), and
for compact/virtual datasets or a non-default fill value: correct, but
more work than needed.
- [ ] **Readers:** `LazyFile` and `MmapFile` still need the whole file;
the zero-copy methods need the file in memory.
### Remote and browser
- [ ] Run the `s3`/`gcs`/`azure` backends against real buckets (only built
and URL-parsing-tested so far).
- [ ] `h5rs` options for request headers and cache settings.
- [ ] Browser limits in [`docs/known-issues.md`](docs/known-issues.md)
("`clawhdf5-wasm` (browser) limits"): files of 4 GiB or more (wasm32),
compound/reference/opaque datasets, round trips per index level. The
package doubled in size with `openUrl`
([size table](examples/wasm-viewer/README.md#size)); dropping the
function-name section would take a third off the raw size (13% gzipped).
### Quality
- [ ] Scheduled fuzz campaigns: the cargo-fuzz targets
([`crates/clawhdf5-format/fuzz`](crates/clawhdf5-format/fuzz/README.md),
and the agent's WAL target) run only by hand or with
`CLAWHDF5_FUZZ_SECONDS`.
---
## Track 3: Hybrid Retrieval Pipeline
**Status:** 🟢 Phase 1 Complete
**Priority:** High
**Crate:** `clawhdf5-agent`
## Withdrawn
- [x] **3.1** Reciprocal Rank Fusion (RRF) — rrf_hybrid_search() with k=60 constant
- [x] **3.2** Multi-factor re-ranking — temporal decay, source authority hierarchy, activation scores (reranker.rs)
- [x] **3.3** Low-confidence rejection — min_score threshold, gap filtering, max_results (confidence.rs)
- [x] **3.4** Query expansion — synonyms, acronyms, temporal rewrites, morphological variants, knowledge graph aliases + expanded_search() with RRF merge
- [x] **3.5** Result explanation — ReRankResult with full score breakdown per factor
- [x] **3.6** Configurable pipeline — ReRankConfig + ConfidenceConfig with tunable weights/thresholds
- [x] **3.7** Tests + MemX-comparable benchmarks — 5 integration tests (Hit@1≥90%, search<500ms@100K, BM25<200ms@100K, hybrid<50ms@10K, compact<200ms@10K)
- **OpenClaw integration** (withdrawn 2026-09-25, PR #9). clawhdf5 was
never an OpenClaw memory plugin; the documented
`memory.backend = "clawhdf5"` was never valid. The Rust `ClawhdfBackend`
remains as a library API. [`docs/openclaw.md`](docs/openclaw.md) records
what a real plugin would need.
- **ZeroClaw integration** (withdrawn 2026-09-25, PR #10). ZeroClaw has no
clawhdf5 backend, and `clawhdf5-migrate`'s SQLite layout is not
ZeroClaw's schema.
**Research:** MemX (2026), SwiftMem (2026)
---
## Track 4: Temporal Reasoning
**Status:** 🟢 Phase 1 Complete
**Priority:** High
**Crate:** `clawhdf5-agent`
- [x] **4.1** Temporal index — sorted timestamp index with binary search, insert/remove
- [x] **4.2** Time-range queries — range_query, before, after, latest, earliest
- [x] **4.3** Session DAG — parent/child linking, chain walking, time-range overlap queries
- [x] **4.4** Temporal re-ranking — query hint enum (Latest/Earliest/Around/Between/None) with boost scoring
- [x] **4.5** Temporal entity tracking — EntityTimeline with state change history + point-in-time reconstruction
- [x] **4.6** Tests — comprehensive tests for all features
**Research:** MemX temporal gaps (≤43.6% Hit@5), MemoryArena multi-session tasks (2026)
---
## Track 5: Memory Security & Provenance
**Status:** 🟢 Phase 1 Complete
**Priority:** Medium-High
**Crate:** `clawhdf5-agent`
- [x] **5.1** Source attribution — MemoryProvenance with source, creator, session, FNV-1a content hash
- [x] **5.2** Write anomaly detection — rate limiting, 15 injection patterns, source distribution analysis
- [x] **5.3** Source isolation — per-MemorySource sub-stores preventing cross-contamination
- [x] **5.4** Memory integrity verification — content hash comparison via verify_integrity()
- [x] **5.5** Poisoning resistance — pattern detection for prompt injection attempts
- [x] **5.6** Tests — comprehensive tests including adversarial patterns
**Research:** MemoryGraft (2025), SSGM Framework (2026)
---
## Track 6: Multi-Modal Memory
**Status:** 🟢 Phase 1 Complete
**Priority:** Medium
**Crate:** `clawhdf5-agent`
- [x] **6.1** Image embedding storage — ModalEmbedding with model provenance (CLIP, SigLIP, etc.)
- [x] **6.2** Audio fingerprints — Audio modality with embedding storage
- [x] **6.3** Multi-modal search — search_by_modality (filtered) + search_cross_modal (all embeddings)
- [x] **6.4** Observation records — raw perception vs interpretation with confidence scoring
- [x] **6.5** Media reference storage — MediaRef with Path/Url/Inline, MIME types, FNV-1a checksums
- [x] **6.6** Tests — 35 comprehensive tests
**Research:** Neuro-Symbolic Memory (2026), RAGdb multi-modal RAG (2025)
---
## Track 7: OpenClaw Integration — withdrawn (2026-09-25)
**Status:** ⚪ Withdrawn (the items below were library work; no OpenClaw integration shipped)
**Priority:** Critical (for adoption)
**Crates:** `clawhdf5-agent`, `clawhdf5-napi`
- [x] **7.1** Memory backend trait — MemoryBackend with search/get/write/ingest/export/stats
- [x] **7.2** Hybrid retrieval pipeline — ClawhdfBackend wires RRF → reranker → confidence rejection
- [x] **7.3** Markdown import/export — MarkdownParser + MarkdownExporter with line tracking + metadata
- [x] **7.4** `search()` — backed by the full hybrid retrieval pipeline (a Rust method; no OpenClaw tool was ever registered)
- [x] **7.5** `get()` — read back by path, with a line slice (not an OpenClaw tool either)
- [x] **7.6** Compaction integration — run_compaction() (decay + compact + WAL flush), run_consolidation() (hippocampal engine), tick_session(), flush_wal()
- [ ] **7.7** ~~Config surface — `memory.backend = "clawhdf5"`~~ — never valid OpenClaw config; docs removed
- [ ] **7.8** ~~Documentation + migration guide~~ — removed: they described an integration that never worked
**Node.js bridge:** `clawhdf5-napi` (napi-rs) and a TypeScript wrapper in `packages/clawhdf5-node` exist but are unpublished, untested in CI and known to be broken (docs/known-issues.md).
---
> **Withdrawn.** None of this track produced a working OpenClaw integration: no
> plugin was built, the documented `memory.backend = "clawhdf5"` config was never
> valid in any OpenClaw release, and the Node package was never published. The
> Rust `ClawhdfBackend` remains as a library API. Not pursued for now; see
> [docs/openclaw.md](docs/openclaw.md) for what a plugin would need today.
## Track 8: Benchmarking & Validation
**Status:** 🟢 Complete
**Priority:** High
**Crates:** `clawhdf5-agent`, `clawhdf5-bench`
- [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547
- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results)
- [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal
- [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion
- [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss
- [x] **8.6** Cross-platform benchmarks — x86 measured, ARM estimated, cross_platform.sh script
- [x] **8.7** Published results in BENCHMARKS.md with ephemeral tier Redis comparison (70-140x faster)
---
## Implementation Order
**Phase 1:** ~~Tracks 1, 2, 3 — core memory intelligence~~ 🟢 Complete
**Phase 2:** ~~Track 4 (temporal) + Track 5 (security)~~ 🟢 Complete
**Phase 3:** ~~Track 6 (multi-modal)~~ 🟢 Complete; Track 7 (OpenClaw integration) withdrawn
**Phase 4:** ~~Track 8 (benchmarking + validation)~~ 🟢 Complete
All 8 tracks delivered. 1,650+ tests passing, zero clippy warnings.
---
## What's Next
Verified against current repo state on 2026-08-05 (see also `docs/superpowers/plans/` for the filter-codec/format-write/MPI-IO work, now shipped):
- [ ] TypeScript bridge not wired into CI — `packages/clawhdf5-node/` already has a complete, working napi-rs package (package.json, tsconfig, hand-written TS wrapper matching all 21 `#[napi]` items, Jest test suite, README); it isn't published to npm and has no committed lockfile
- [ ] Publish crates to crates.io — no `publish` config anywhere in the workspace yet
- [ ] Python wheel distribution via maturin — `crates/clawhdf5-py/pyproject.toml` exists (maturin-buildable locally) but wheels aren't published anywhere
- [ ] `chunked_read.rs`/`data_read.rs` full bounds-check audit + scheduled fuzz campaigns (the new `fuzz_dataset_read` target covers the two files' main entry points; a full manual audit of every indexing site is still open) — see Tier 4 below
- [ ] WAL per-entry checksum landed as CRC32 (see below); a stronger per-entry format (explicit length prefix, avoiding the read-then-verify restructuring) could still be revisited if profiling shows it matters
- [ ] HNSW build parallelism is still narrow (only `prune_connections`); the correctness-sensitive outer insert loop needs its own dedicated design pass before parallelizing
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open
- [x] `chunked_read.rs`/`data_read.rs`/`local_heap.rs` bounds-check audit: added `ensure_len` overflow guards, a recursion-depth guard against cyclic B-trees, and a fix for an unguarded compound-datatype byte-offset overrun. Added a new `fuzz_dataset_read` cargo-fuzz target exercising the contiguous/chunked/compact read paths — it found and we fixed 3 real crash bugs (integer-overflow panics) within the first few runs
- [x] `clawhdf5-ann`: optional `parallel` feature (rayon) for HNSW's `prune_connections` neighbor-distance computation
- [x] `[workspace.dependencies]` added for `tempfile`/`criterion`/`half`/`serde`, fixing a real version skew on `half` (2 vs 2.7)
### Recently closed out (2026-08-05 hardening pass)
- [x] CI/CD pipeline — `.gitea/workflows/ci.yml` now runs `scripts/ci-test.sh` (fmt, clippy, tests, no_std check) on push/PR to `main`
- [x] Fixed no_std build breakage in `clawhdf5-format` (missing alloc imports, `AtomicU64` unsupported on thumbv7em, `f64::powi` requiring std/libm)
- [x] Fixed version skew: `clawhdf5-py` (pyproject.toml) and `packages/clawhdf5-node` (package.json) were both behind the actual crate version
### Recently closed out (2026-08-03 cleanup pass)
- [x] Removed `clawhdf5-types` — it was an empty 1-line stub crate; shared type definitions already live in `clawhdf5-format`, so CLAUDE.md and the workspace manifest were corrected instead of filling it in
- [x] Superblock v4 (page-buffer mode) read/write — the only unimplemented task from `docs/superpowers/plans/2026-06-29-format-write-extensions.md`; now done (`Superblock::parse_v4`/`serialize`, `FileWriter::with_page_size`)
- [x] Reconciled the three `docs/superpowers/plans/*.md` docs against actual shipped code — they were pre-work plans for `d6c4d4f` (2026-06-30), committed to git late; checkboxes now reflect reality
---
_Last updated: 2026-08-05_
The old track-by-track tracker this file used to be (agent-memory
Tracks 1–8, mid-2026) is in git history (`git log -- ROADMAP.md`).
+2 -1
View File
@@ -29,7 +29,8 @@
# - Use wasm-pack with a custom bench harness
# - Replace std::time::Instant with web_sys::Performance::now()
# - Replace TempDir/HDF5 I/O with an in-memory backend (separate effort)
# See ROADMAP.md §WASM for the full scope.
# Browser reads are tested (not benchmarked) by
# examples/wasm-viewer/test/run.sh; see examples/wasm-viewer/README.md.
set -euo pipefail
+33 -2
View File
@@ -6,9 +6,38 @@ h5py/libhdf5, compares the two readings object by object, and writes
```sh
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --no-fetch # use the cached corpus as is
conformance/run.sh --update-baseline # after an intended change in results
```
Latest result (tank, 2026-09-28 04:29 UTC, `conformance/run.sh --no-fetch
--update-baseline`): 602 of 697 files ok, 1 our-error, 0 mismatch, 2
ref-bug, 92 h5py-cannot-read, and no panic, hang, crash or out-of-memory.
The our-error file is `bad_nbit_parms_walk.h5`, which flips between ref-bug
and our-error from run to run (see `docs/known-issues.md`). The report with every file is
[`CONFORMANCE.md`](../CONFORMANCE.md).
## Classes
`compare.py` puts each file in one class:
| class | meaning |
|---|---|
| **ok** | clawhdf5 and h5py read the same objects with the same values |
| **our-error** | h5py reads something clawhdf5 refuses |
| **mismatch** | both read it, with different values or structure |
| **h5py-cannot-read** | h5py (libhdf5) cannot read the file; not compared |
| **ref-bug** | h5py reads an object clawhdf5 refuses, but only through a libhdf5 over-read: `ref_bugs.py` re-reads it in six processes with different heaps (import order, `MALLOC_PERTURB_`) and its values change. The file is ref-bug only while that is confirmed in the same run; if the values become stable it counts as our-error again |
| **panic / hang / crash / oom** | a clawhdf5 failure under the timeout and address-space limit; the gate fails on any |
Where h5py itself returns wrong values through a known h5py bug (the
big-endian variable-length bug: elements returned with the file's bytes
under a little-endian dtype), `ref.py` checks that the installed h5py has
the bug, corrects the values before hashing and marks them `ref_fix`, so
those objects are still compared. The evidence for the three remaining
non-ok files (ref-bug or, for one, our-error) is under "Conformance: the last non-ok files" in
[`docs/known-issues.md`](../docs/known-issues.md).
Needs Rust, `git`, `h5dump` (Debian/Ubuntu `hdf5-tools`), `libaec` (for the
probe's `szip` feature; `libaec-dev`), and a Python with the packages in
`requirements.txt`. The first run downloads about 450 MB of sparse checkouts.
@@ -19,9 +48,11 @@ probe's `szip` feature; `libaec-dev`), and a Python with the packages in
| `fetch-corpus.sh` | shallow, sparse, blob-filtered checkout of each pinned commit into `.cache/src/` (gitignored); no-op when already there |
| `list_files.py` | which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
| `probe/` | the clawhdf5 side: a standalone crate (outside the workspace, so `cargo test --workspace` never builds it) that walks a file with `clawhdf5-format` and prints canonical JSON |
| `ref.py` | the h5py side: the same JSON from h5py |
| `ref.py` | the h5py side: the same JSON from h5py (values corrected for a known h5py bug are marked `ref_fix`) |
| `ref_bugs.py` | re-reads the objects h5py reads only through a libhdf5 bug in six differently-set-up processes; an object whose values change is confirmed as a libhdf5 over-read |
| `test_ref.py` | tests of `ref.py`'s correction and `ref_bugs.py`'s confirmation (`python conformance/test_ref.py`) |
| `run_one.sh` | runs both sides on one file (and `h5dump` on the CVE corpus) under a timeout and an address-space limit |
| `compare.py` | classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes |
| `compare.py` | classifies each file (ok / our-error / mismatch / h5py-cannot-read / ref-bug / panic / hang / crash / oom) and groups root causes |
| `report.py` | writes `CONFORMANCE.md` |
| `check.py` | the gate: fails on any panic/hang/crash/oom, on an ok count below `baseline.json`, or on a baseline-ok file that is no longer ok |
| `baseline.json` | the ok files the gate holds the line on |
+11 -11
View File
@@ -1,33 +1,31 @@
{
"comment": "conformance/run.sh fails if the ok count drops below `ok` or a file in `ok_files` stops being ok. Regenerate with `conformance/run.sh --update-baseline` after an intended change.",
"commit": "f37e7ae3263277319dba4bc39be5397194eb00c3",
"date": "2026-09-27 00:34 UTC",
"commit": "bf5a163dcf7fe28d651ada6545d8136ffeffc825",
"date": "2026-09-28 04:29 UTC",
"reference": "h5py 3.16.0 / HDF5 2.0.0",
"files": 697,
"ok": 600,
"ok": 602,
"counts": {
"h5py-cannot-read": 92,
"mismatch": 2,
"ok": 600,
"our-error": 3
"ok": 602,
"our-error": 1,
"ref-bug": 2
},
"per_corpus": {
"NCAS-CMS_pyfive": {
"mismatch": 1,
"ok": 32
"ok": 33
},
"cve_hdf5": {
"h5py-cannot-read": 32,
"ok": 113,
"our-error": 2
"ref-bug": 2
},
"h5py_data": {
"ok": 4
},
"hdf5": {
"h5py-cannot-read": 60,
"mismatch": 1,
"ok": 404,
"ok": 405,
"our-error": 1
},
"netcdf-c": {
@@ -45,6 +43,7 @@
},
"ok_files": [
"NCAS-CMS_pyfive/tests/compact.hdf5",
"NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5",
"NCAS-CMS_pyfive/tests/data/btreev2.hdf5",
"NCAS-CMS_pyfive/tests/data/chunked.hdf5",
"NCAS-CMS_pyfive/tests/data/cmip_bad_eg.nc",
@@ -455,6 +454,7 @@
"hdf5/tools/test/testfiles/tcmpdints.h5",
"hdf5/tools/test/testfiles/tcmpdintsize.h5",
"hdf5/tools/test/testfiles/tcomplex.h5",
"hdf5/tools/test/testfiles/tcomplex_be.h5",
"hdf5/tools/test/testfiles/tcompound.h5",
"hdf5/tools/test/testfiles/tcompound_complex.h5",
"hdf5/tools/test/testfiles/tcompound_complex2.h5",
+30 -3
View File
@@ -5,6 +5,9 @@ Writes <results_dir>/results.csv, results.json and summary.md.
File classes (first match wins):
hang, oom, crash, panic ours: timeout / allocation failure / signal / any panic (caught or not)
h5py-cannot-read libhdf5/h5py failed to open the file (or crashed/hung)
ref-bug every issue is an object we refuse that h5py reads only through a
libhdf5 bug, confirmed in this run by ref_bugs.py (its values
change with the reading process's heap)
our-error we fail to open, list, or read something h5py reads
mismatch we read something with different shape/values, or a different object set
ok
@@ -19,6 +22,20 @@ import sys
R = sys.argv[1]
RUNS = os.path.join(R, "runs")
# Objects ref_bugs.py confirmed in this run: h5py's values for them come from
# libhdf5 reading memory the file does not determine.
try:
REF_BUGS = {(b["file"], b["object"])
for b in json.load(open(os.path.join(R, "ref_bugs.json")))["read_bugs"] if b.get("confirmed")}
except (OSError, ValueError, KeyError):
REF_BUGS = set()
def is_ref_bug(rel, issue):
"""An our-error on reading an object that ref_bugs.py confirmed."""
kind, detail = issue[0], issue[1]
return kind == "our-error" and any(f == rel and detail.startswith(obj + ": error: ") for f, obj in REF_BUGS)
def load(d, name):
rc_p = os.path.join(d, name + ".rc")
@@ -86,6 +103,8 @@ mismatch_causes = collections.defaultdict(lambda: {"files": set(), "count": 0, "
panics = []
ref_only_errors = collections.Counter()
incomparable = collections.Counter()
# Values ref.py corrected for a known h5py bug: (file, object, fixes, same as ours)
ref_fixes = []
def add(bucket, key, file, example):
@@ -169,6 +188,8 @@ for rel in files:
elif a.get("hash") != b.get("hash"):
issues.append(("mismatch", f"{p}: values differ (h5py {a.get('dtype')} vs ours {b.get('dtype')})", "values", b | {"ref_head": a.get("head"), "ref_dtype": a.get("dtype")}))
ok = False
if a.get("ref_fix") and "hash" in b:
ref_fixes.append((rel, p, a["ref_fix"], a.get("hash") == b.get("hash")))
ra, oa = a.get("attrs") or {}, b.get("attrs") or {}
if "attrs_error" not in b and "attrs_error" not in a and not ref_unopened:
for an in sorted(set(ra) | set(oa)):
@@ -187,6 +208,8 @@ for rel in files:
issues.append(("mismatch", f"{p}@{an}: attr shape {x.get('shape')} vs ours {y.get('shape')}", "attr-shape", y | {"ref_dtype": x.get("dtype")}))
elif x.get("hash") != y.get("hash"):
issues.append(("mismatch", f"{p}@{an}: attr values differ (h5py {x.get('dtype')} vs ours {y.get('dtype')})", "attr-values", y | {"ref_head": x.get("head"), "ref_dtype": x.get("dtype")}))
if x.get("ref_fix") and "hash" in y:
ref_fixes.append((rel, f"{p}@{an}", x["ref_fix"], x.get("hash") == y.get("hash")))
if ok:
n_ok += 1
@@ -197,6 +220,8 @@ for rel in files:
cls = "panic"
elif ref_open_fail:
cls = "h5py-cannot-read"
elif issues and all(is_ref_bug(rel, i) for i in issues):
cls = "ref-bug"
elif ours_open_err:
cls = "our-error"
issues.append(("our-error", f"open: {ours_open_err}", ours_open_err, {}))
@@ -214,7 +239,8 @@ for rel in files:
"caught": [(p, w, m[:2500]) for p, w, m in caught_panics[:3]],
"n_caught": len(caught_panics),
})
for kind, detail, key, rec in issues:
# A ref-bug file's differences are listed with the evidence instead.
for kind, detail, key, rec in (issues if cls != "ref-bug" else []):
if kind == "our-error":
add(root_causes, norm(key), rel, detail[:300])
else:
@@ -259,10 +285,11 @@ def ser(b):
json.dump({"rows": rows, "issues": issues_by_file, "root_causes": ser(root_causes), "mismatch_causes": ser(mismatch_causes),
"panics": panics, "incomparable": incomparable.most_common(), "ref_only_errors": ref_only_errors.most_common()},
"panics": panics, "incomparable": incomparable.most_common(), "ref_only_errors": ref_only_errors.most_common(),
"ref_fixes": ref_fixes, "ref_bugs_confirmed": sorted(REF_BUGS)},
open(os.path.join(R, "results.json"), "w"), indent=1)
classes = ["ok", "our-error", "mismatch", "h5py-cannot-read", "hang", "panic", "crash", "oom"]
classes = ["ok", "our-error", "mismatch", "h5py-cannot-read", "ref-bug", "hang", "panic", "crash", "oom"]
by_corpus = collections.defaultdict(collections.Counter)
for r in rows:
by_corpus[r["corpus"]][r["class"]] += 1
+53 -1
View File
@@ -53,6 +53,49 @@ def packed(dt):
return dt
# --- reference corrections ---------------------------------------------------
# Where h5py is known to return values the file does not hold, and the right
# values follow from what it returned, ref.py corrects them and records the
# correction on the object ("ref_fix"), so the comparison is still a real
# comparison and CONFORMANCE.md lists every corrected object. Each correction
# first checks that the installed h5py still has the bug.
# Corrections applied while encoding the current object.
FIXES = set()
_BE_VLEN_BUG = None
def be_vlen_bug():
"""h5py (3.16 / HDF5 2.0 at least) returns the elements of a
variable-length sequence whose base type is big-endian with the file's
big-endian bytes under a native (little-endian) dtype: a
`vlen_dtype('>f4')` dataset holding [1.0, 2.0] reads back as
[4.6e-41, 9.0e-44]. `h5dump` prints the file's values. Checked once per
process by writing and reading exactly that dataset in memory."""
global _BE_VLEN_BUG
if _BE_VLEN_BUG is None:
import io
try:
bio = io.BytesIO()
with h5py.File(bio, "w") as f:
d = f.create_dataset("v", (1,), dtype=h5py.vlen_dtype(np.dtype(">f4")))
d[0] = np.array([1.0, 2.0], dtype=">f4")
with h5py.File(bio, "r") as f:
got = np.asarray(f["v"][0])
_BE_VLEN_BUG = (got.dtype == np.dtype("<f4")
and got.view(">f4").tolist() == [1.0, 2.0]
and got.tolist() != [1.0, 2.0])
except Exception: # noqa: BLE001
_BE_VLEN_BUG = False
return _BE_VLEN_BUG
def unswapped(got, base):
"""`got` is `base` (big-endian somewhere) with every field in native
little-endian order instead: the shape of h5py's big-endian VL bug."""
return base.newbyteorder("<") == got and base != got
def canon_el(dt, val, out):
if dt.fields:
for n in dt.names:
@@ -79,7 +122,13 @@ def canon_el(dt, val, out):
base = h5py.check_vlen_dtype(dt)
if base is None:
raise TypeError(f"unhandled object dtype {dt!r}")
arr = np.asarray(val if val is not None else [], dtype=base).reshape(-1)
arr = np.asarray(val if val is not None else [])
if arr.dtype != base and be_vlen_bug() and unswapped(arr.dtype, base):
# h5py's big-endian VL bug (see be_vlen_bug): the bytes are
# the file's, the dtype label is wrong. Relabel, don't convert.
arr = arr.view(base)
FIXES.add("h5py-be-vlen")
arr = np.asarray(arr, dtype=base).reshape(-1)
out += b"V" + struct.pack("<I", arr.shape[0])
if simple(base):
out += arr.astype(packed(base)).tobytes()
@@ -118,6 +167,7 @@ def hash_values(arr, dt, rec):
while dt.subdtype is not None:
dt = dt.subdtype[0]
arr = np.asarray(arr, dtype=dt)
FIXES.clear()
if simple(dt):
c = np.ascontiguousarray(arr).astype(packed(dt)).tobytes()
else:
@@ -125,6 +175,8 @@ def hash_values(arr, dt, rec):
for x in arr.reshape(-1):
canon_el(dt, x, out)
c = bytes(out)
if FIXES:
rec["ref_fix"] = sorted(FIXES)
rec["hash"] = hashlib.sha256(c).hexdigest()
rec["head"] = c[:48].hex()
+105
View File
@@ -0,0 +1,105 @@
#!/usr/bin/env python3
"""ref_bugs.py <corpus_dir>: re-check the objects h5py reads only through a
libhdf5 bug.
For each object of READ_BUGS (below), h5py reads it in several fresh
processes whose heaps differ: h5py imported before numpy (three runs, plus
two with glibc's MALLOC_PERTURB_, which fills newly allocated and freed heap
blocks with a byte pattern) and numpy imported first. Values the file
determines come out the same every time. An object whose values differ
between those runs is read from memory the file does not determine — an
over-read or an uninitialised buffer in libhdf5 — so the values h5py reports
for it are not the file's, and clawhdf5 refusing the object is not a
clawhdf5 error. compare.py classifies a file as `ref-bug` only on objects
confirmed that way in the same run (`$OUT/ref_bugs.json`); an object whose
reading turns out stable stays an our-error.
Run by conformance/run.sh; on its own it is the reproducer (JSON on stdout).
"""
import concurrent.futures
import json
import os
import subprocess
import sys
# (file, object) -> what goes wrong. Checked 2026-09-27 against HDF5 2.0.0
# (h5py 3.16), h5dump 1.14.6 and the HDFGroup/hdf5 sources (tag hdf5_1_14_6
# and develop); see docs/known-issues.md, "Conformance: the last non-ok files".
READ_BUGS = {
("cve_hdf5/cvefiles/cve-2025-2308.h5", "/Scale_offset_long_long_data_le"):
"the first chunk records minbits 11: its 12 values need 17 bytes of codes, and the "
"26-byte chunk holds 5 after its 21-byte header; libhdf5's scale-offset decoder reads "
"past its buffer, and develop refuses the chunk (\"Buffer too short\")",
("cve_hdf5/cvefiles/cve-2025-44904.h5", "/Scale_offset_float_data_le"):
"unfiltered chunks stored as 38 and 37 bytes for 48-byte chunks: 1.14/2.0 read the "
"stored bytes into a buffer of that size and use it as the whole chunk "
"(H5D__chunk_lock), so the rest is heap memory; develop refuses them (\"incorrect chunk "
"size returned from index for unfiltered chunk\")",
("hdf5/test/testfiles/bad_nbit_parms_walk.h5", "/Nbit_int_data_le"):
"the N-Bit parameter list holds 7 values (cd_values[0] = 7) where an integer needs 8: "
"the decoder takes the bit offset from cd_values[7], past the list; libhdf5's own test "
"(`test_filter_bad_params`, test/dsets.c on develop) requires the read to fail",
}
# (which module is imported first, MALLOC_PERTURB_)
RUNS = [("h5py", None), ("h5py", None), ("h5py", None), ("h5py", "170"), ("h5py", "255"),
("numpy", None)]
READ = r"""
import hashlib, sys
if sys.argv[3] == "h5py":
import h5py, numpy as np
else:
import numpy as np, h5py
try:
import hdf5plugin # noqa: F401
except Exception:
pass
try:
with h5py.File(sys.argv[1], "r") as f:
a = np.ascontiguousarray(f[sys.argv[2]][()])
print("values " + hashlib.sha256(a.tobytes()).hexdigest()[:16])
except Exception as e:
print("error " + (str(e).splitlines() or [type(e).__name__])[0][:120])
"""
def read_once(path, obj, first, perturb):
env = dict(os.environ)
env.pop("MALLOC_PERTURB_", None)
if perturb:
env["MALLOC_PERTURB_"] = perturb
try:
p = subprocess.run([sys.executable, "-c", READ, path, obj, first], env=env,
capture_output=True, text=True, timeout=60)
out = p.stdout.strip().splitlines()
return out[-1] if out else f"exit {p.returncode}"
except subprocess.TimeoutExpired:
return "timeout"
def check(corpus, key):
f, obj = key
path = os.path.join(corpus, f)
rec = {"file": f, "object": obj, "why": READ_BUGS[key]}
if not os.path.exists(path):
return rec | {"missing": True, "confirmed": False}
runs = [{"first": a, "malloc_perturb": p, "outcome": read_once(path, obj, a, p)} for a, p in RUNS]
distinct = sorted({r["outcome"] for r in runs})
return rec | {
"runs": runs,
"distinct": len(distinct),
"confirmed": len(distinct) > 1 and any(o.startswith("values ") for o in distinct),
}
def main():
corpus = sys.argv[1]
keys = list(READ_BUGS)
with concurrent.futures.ThreadPoolExecutor(max_workers=len(keys)) as ex:
out = list(ex.map(lambda k: check(corpus, k), keys))
print(json.dumps({"read_bugs": out}, indent=1))
if __name__ == "__main__":
main()
+62 -69
View File
@@ -26,7 +26,7 @@ except Exception: # noqa: BLE001
R, OUT_MD, CORPUS = sys.argv[1], sys.argv[2], sys.argv[3]
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.dirname(HERE)
CLASSES = ["ok", "our-error", "mismatch", "h5py-cannot-read", "panic", "hang", "crash", "oom"]
CLASSES = ["ok", "our-error", "mismatch", "h5py-cannot-read", "ref-bug", "panic", "hang", "crash", "oom"]
def sh(*cmd, cwd=ROOT):
@@ -89,46 +89,15 @@ def ex_list(files, n=3):
return s + (f" (+{len(files) - n} more)" if len(files) > n else "")
# --- known causes that are not clawhdf5 bugs --------------------------------
def is_h5py_be_vlen(i):
"""h5py returns the elements of a VL sequence of a big-endian base type
with their file (big-endian) bytes but a native-endian dtype."""
return (i["kind"] == "mismatch" and i["key"] in ("values", "attr-values")
and (i.get("ref_dtype") == "object") and (i.get("ours_dtype") or "").startswith("vlen(")
and ">" in (i.get("ours_dtype") or ""))
# Objects the reference (h5py 3.16 / HDF5 2.0) reads only because of an
# HDF5 2.0 bug, and that clawhdf5 refuses: each one reads past a buffer or
# returns bytes the file does not hold, and libhdf5's develop branch refuses all
# three. (file, object) -> why. Checked 2026-09-26 against HDF5 2.0.0
# and HDFGroup/hdf5 develop sources; see docs/known-issues.md.
LIBHDF5_BUGS = {
("cve_hdf5/cvefiles/cve-2025-2308.h5", "/Scale_offset_long_long_data_le"):
"scale-offset codes run past the end of the chunk: HDF5 2.0 reads past its buffer; "
"libhdf5's develop branch refuses the chunk (\"Buffer too short\")",
("cve_hdf5/cvefiles/cve-2025-44904.h5", "/Scale_offset_float_data_le"):
"unfiltered chunks of 38 and 37 bytes for 48-byte chunks: HDF5 2.0 fills the rest with "
"whatever its buffer held; libhdf5's develop branch refuses them (\"incorrect chunk size returned "
"from index for unfiltered chunk\")",
("hdf5/test/testfiles/bad_nbit_parms_walk.h5", "/Nbit_int_data_le"):
"an N-Bit parameter list one value short: HDF5 2.0 reads past the list; libhdf5's own "
"test (`test_filter_bad_params`, test/dsets.c) now requires the read to fail",
}
def is_libhdf5_bug(rel, i):
return i["kind"] == "our-error" and any(
f == rel and i["detail"].startswith(obj + ":") for (f, obj) in LIBHDF5_BUGS)
known = collections.defaultdict(list)
for r in rows:
iss = issues.get(r["file"], [])
if r["class"] == "mismatch" and iss and all(is_h5py_be_vlen(i) for i in iss):
known["h5py-be-vlen"].append(r["file"])
if r["class"] == "our-error" and iss and all(is_libhdf5_bug(r["file"], i) for i in iss):
known["libhdf5-2.0"].append(r["file"])
# --- reference bugs ---------------------------------------------------------
# ref_bugs.py's re-check of the objects h5py reads only through a libhdf5 bug
# (compare.py classifies on the confirmed ones), and the objects whose h5py
# values ref.py corrected (compare.py's ref_fixes).
try:
ref_bugs = json.load(open(os.path.join(R, "ref_bugs.json")))["read_bugs"]
except (OSError, ValueError, KeyError):
ref_bugs = []
ref_fixes = res.get("ref_fixes", [])
# --- the CVE corpus: clawhdf5 vs h5dump vs h5py ------------------------------
@@ -243,6 +212,7 @@ w("A file's class is the first that applies:")
w("")
w("- **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these.")
w("- **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers.")
w("- **ref-bug** — every difference is an object clawhdf5 refuses that h5py reads only through a libhdf5 bug: the values h5py returns for it change with the reading process's heap, re-checked in every run (see *Reference bugs*).")
w("- **our-error** — clawhdf5 returned an error for something h5py reads.")
w("- **mismatch** — both read it, but the shapes, values, object set or attribute set differ.")
w("- **ok** — every object h5py reads, clawhdf5 reads identically.")
@@ -254,14 +224,12 @@ for c in sorted(by_corpus):
w(f"| {c} | {sum(cnt.values())} | " + " | ".join(str(cnt.get(k, 0)) for k in CLASSES) + " |")
w(f"| **all** | **{len(rows)}** | " + " | ".join(f"**{total.get(k, 0)}**" for k in CLASSES) + " |")
w("")
if known["h5py-be-vlen"]:
w(f"{len(known['h5py-be-vlen'])} of the {total.get('mismatch', 0)} mismatches are a known h5py bug, "
"not ours (see *Known not-our-bug*).")
w("")
if known["libhdf5-2.0"]:
w(f"{len(known['libhdf5-2.0'])} of the {total.get('our-error', 0)} our-errors are corrupt data that "
"HDF5 2.0 reads only through a bug and clawhdf5 refuses (see *Known not-our-bug*).")
w("")
nonok = total.get("our-error", 0) + total.get("mismatch", 0)
w(f"**Our errors and mismatches: {nonok}.** Files not ok: "
+ (", ".join(f"{total[c]} {c}" for c in CLASSES if c != "ok" and total.get(c)) or "none") + "."
+ (f" {len(ref_fixes)} object(s) were compared against h5py's values corrected for a known h5py bug"
f" ({sum(1 for x in ref_fixes if x[3])} identical to clawhdf5's; see *Reference bugs*)." if ref_fixes else ""))
w("")
w("Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):")
w("")
w("| corpus | source | commit |")
@@ -281,19 +249,25 @@ w("")
w("## Our-error root causes")
w("")
w("Grouped by normalised error message. *files* counts files whose class this cause affects.")
w("")
w("| files | objects | error | examples |")
w("|---:|---:|---|---|")
for k, v in res["root_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
if res["root_causes"]:
w("Grouped by normalised error message. *files* counts files whose class this cause affects.")
w("")
w("| files | objects | error | examples |")
w("|---:|---:|---|---|")
for k, v in res["root_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
else:
w("None.")
w("")
w("## Mismatch root causes")
w("")
w("| files | objects | cause | examples |")
w("|---:|---:|---|---|")
for k, v in res["mismatch_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
if res["mismatch_causes"]:
w("| files | objects | cause | examples |")
w("|---:|---:|---|---|")
for k, v in res["mismatch_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
else:
w("None.")
w("")
w("## CVE corpus: clawhdf5 vs h5dump vs h5py")
@@ -321,14 +295,38 @@ w("")
w("</details>")
w("")
w("## Known not-our-bug")
w("## Reference bugs")
w("")
w("### Objects h5py reads only through a libhdf5 bug (*ref-bug*)")
w("")
w("clawhdf5 refuses these objects; h5py 3.16 / HDF5 2.0 returns values for them. `conformance/ref_bugs.py`")
w("re-reads each with h5py in six fresh processes whose heaps differ (h5py imported before numpy, three")
w("times and twice more with `MALLOC_PERTURB_`, and numpy imported first). Values the file determines")
w("come out the same every time; these do not, so they are memory libhdf5 over-reads, not the file's")
w("data. A file is *ref-bug* only while every one of its differences is such an object confirmed in")
w("the same run; an object that reads the same every time goes back to *our-error*. Reproducer:")
w("`python conformance/ref_bugs.py conformance/.cache/corpus` (prints every read's outcome).")
w("")
w("| file | object | distinct results in 6 reads | confirmed | what goes wrong |")
w("|---|---|---:|---|---|")
for b in ref_bugs:
n = "missing" if b.get("missing") else b.get("distinct", "?")
w(f"| `{b['file']}` | `{b['object']}` | {n} | {'yes' if b.get('confirmed') else '**no**'} | {b['why']} |")
w("")
w("### Values corrected for a known h5py bug")
w("")
w("- **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence")
w(" whose base type is big-endian with the file's big-endian bytes but a native (little-endian)")
w(" numpy dtype, so the values it reports are byte-swapped garbage; `h5dump` prints the values")
w(" clawhdf5 reads. Reproducer: `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]`")
w(" reads back in h5py as `[4.6e-41, 9.0e-44]`. Affected here: "
+ (ex_list(sorted(known["h5py-be-vlen"]), 10) if known["h5py-be-vlen"] else "none") + ".")
w(" numpy dtype: a `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]` reads back as")
w(" `[4.6e-41, 9.0e-44]`; `h5dump` prints the file's values. `ref.py` checks that the installed")
w(" h5py still does this (by writing and reading exactly that dataset in memory) and, if so,")
w(" relabels such elements with the file's byte order before hashing, so the values are still")
w(" compared. Corrected objects: "
+ (", ".join(f"`{f}` `{p}` ({'same as clawhdf5' if same else '**differs from clawhdf5**'})"
for f, p, _, same in ref_fixes) if ref_fixes else "none") + ".")
w("")
w("## Other comparison rules")
w("")
w("- **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose")
w(" bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a")
w(" bit offset / reduced precision into the plain numpy type of the same size. The probe")
@@ -339,11 +337,6 @@ if res["incomparable"]:
w(" (FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not")
w(" compared (shape and presence still are): "
+ ", ".join(f"{k} ({n}x)" for k, n in res["incomparable"]) + ".")
w("- **Corrupt data HDF5 2.0 reads through a bug.** clawhdf5 refuses these objects; h5py 3.16 /")
w(" HDF5 2.0 returns values for them that the file does not hold:")
for (f, obj), why in sorted(LIBHDF5_BUGS.items()):
here = "" if f in known["libhdf5-2.0"] else " (not an our-error in this run)"
w(f" - `{f}` `{obj}`: {why}{here}.")
w("- **References** are compared by presence only (`R`), not by target.")
w("")
if res.get("ref_only_errors"):
+2
View File
@@ -71,6 +71,8 @@ xargs -a "$OUT/files.txt" -d '\n' -P "$JOBS" -I{} bash -c '
f="$1"; d="$OUT/runs/${f//\//__}"
case "$f" in cve_hdf5/*) export WITH_H5DUMP=1 ;; esac
"$HERE/run_one.sh" "$C/$f" "$d"' _ {} 2>"$OUT/probe.log"
echo "== re-checking the objects h5py reads only through a libhdf5 bug"
"$PY" "$HERE/ref_bugs.py" "$C" > "$OUT/ref_bugs.json" 2> "$OUT/ref_bugs.err" || true
echo "== comparing"
"$PY" "$HERE/compare.py" "$OUT" >/dev/null
t2=$(date +%s)
+92
View File
@@ -0,0 +1,92 @@
#!/usr/bin/env python3
"""Tests of the reference side's corrections: `python conformance/test_ref.py`.
- ref.py compares a big-endian VL sequence by the file's values even though
h5py returns them byte-swapped (and records that it corrected them);
- ref_bugs.py confirms an object only when its reads disagree.
"""
import json
import os
import subprocess
import sys
import tempfile
import unittest
import h5py
import numpy as np
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, HERE)
import ref_bugs # noqa: E402
def ref_objects(path):
out = subprocess.run([sys.executable, os.path.join(HERE, "ref.py"), path],
capture_output=True, text=True, check=True).stdout
return {o["path"]: o for o in json.loads(out)["objects"]}
class BigEndianVlen(unittest.TestCase):
def test_be_vlen_compared_by_file_values(self):
with tempfile.TemporaryDirectory() as d:
path = os.path.join(d, "v.h5")
with h5py.File(path, "w") as f:
for name, order in (("be", ">"), ("le", "<")):
t = np.dtype(order + "f4")
ds = f.create_dataset(name, (2,), dtype=h5py.vlen_dtype(t))
ds[0] = np.array([1.0, 2.0], dtype=t)
ds[1] = np.array([3.0], dtype=t)
u = np.dtype(order + "u8")
f.attrs.create(name, [np.array([1, 2], dtype=u), np.array([42], dtype=u)],
dtype=h5py.vlen_dtype(u))
objs = ref_objects(path)
be, le = objs["/be"], objs["/le"]
# Same values, so the same canonical hash whatever the file's byte order.
self.assertEqual(be["hash"], le["hash"])
self.assertEqual(objs["/"]["attrs"]["be"]["hash"], objs["/"]["attrs"]["le"]["hash"])
self.assertNotIn("ref_fix", le)
# And the correction is recorded wherever h5py needed it.
import ref
if ref.be_vlen_bug():
self.assertEqual(be.get("ref_fix"), ["h5py-be-vlen"])
self.assertEqual(objs["/"]["attrs"]["be"].get("ref_fix"), ["h5py-be-vlen"])
class RefBugsConfirmation(unittest.TestCase):
def run_check(self, outcomes):
seq = iter(outcomes)
saved = ref_bugs.read_once
ref_bugs.read_once = lambda *a: next(seq)
try:
key = next(iter(ref_bugs.READ_BUGS))
with tempfile.TemporaryDirectory() as d:
p = os.path.join(d, key[0])
os.makedirs(os.path.dirname(p))
open(p, "wb").close()
return ref_bugs.check(d, key)
finally:
ref_bugs.read_once = saved
def test_stable_values_are_not_confirmed(self):
r = self.run_check(["values a"] * len(ref_bugs.RUNS))
self.assertFalse(r["confirmed"])
def test_changing_values_are_confirmed(self):
r = self.run_check(["values a"] * (len(ref_bugs.RUNS) - 1) + ["values b"])
self.assertTrue(r["confirmed"])
r = self.run_check(["values a"] * (len(ref_bugs.RUNS) - 1) + ["error filter failed"])
self.assertTrue(r["confirmed"])
def test_errors_only_are_not_confirmed(self):
# h5py cannot read it at all: nothing it reads, nothing to excuse.
r = self.run_check(["error x"] * (len(ref_bugs.RUNS) - 1) + ["error y"])
self.assertFalse(r["confirmed"])
def test_missing_file_is_not_confirmed(self):
key = next(iter(ref_bugs.READ_BUGS))
with tempfile.TemporaryDirectory() as d:
self.assertFalse(ref_bugs.check(d, key)["confirmed"])
if __name__ == "__main__":
unittest.main()
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-accel"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "SIMD-accelerated operations for rustyhdf5"
description = "SIMD kernels (AVX2, NEON) used by clawhdf5 — pure Rust"
license = "MIT"
repository = "https://git.redclaw.dev/quantumclaw/clawhdf5"
readme = "README.md"
+52 -14
View File
@@ -1,24 +1,62 @@
# clawhdf5-accel
[![crates.io](https://img.shields.io/crates/v/clawhdf5-accel.svg)](https://crates.io/crates/clawhdf5-accel)
[![docs.rs](https://docs.rs/clawhdf5-accel/badge.svg)](https://docs.rs/clawhdf5-accel)
CPU SIMD kernels for vector search: dot products, cosine similarity, L2
distance, norms and int8 dot products, dispatched at run time to the best
backend the CPU has, with a portable scalar fallback for every operation.
[`clawhdf5-ann`](../clawhdf5-ann/README.md) and
[`clawhdf5-agent`](../clawhdf5-agent/README.md) use it in their distance
loops; it has nothing to do with HDF5 file I/O.
SIMD-accelerated operations for clawhdf5.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## API
```rust
use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance};
let a = [1.0f32, 2.0, 3.0, 4.0];
let b = [4.0f32, 3.0, 2.0, 1.0];
assert_eq!(dot_product(&a, &b), 20.0);
let _cos = cosine_similarity(&a, &b);
let _l2 = l2_distance(&a, &b);
assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24);
println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64
```
Also `vector_norm`, `batch_norms`, `batch_cosine`, `batch_cosine_prenorm`,
`f16_to_f32_batch`, `checksum_fletcher32` and `align_to_cache_line`.
## Backends
`detect_backend()` picks once per process: `Avx512` (with the `avx512`
feature), `Avx2` (AVX2 + FMA), `Neon` (every aarch64 CPU), or `Scalar`.
`Sse4` and `WasmSimd128` are reported when detected but run the scalar
kernels.
`dot_i8`, used by the agent's quantised (int8) HNSW index, runs on
AVX2 and on NEON — with the `SDOT` instruction (through inline assembly,
since the intrinsic is unstable) on cores that have dotprod, such as the
Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8
index answers 1.63x the queries per second of the f32 one on x86-64
(AVX2; 2026-09-20, machine not recorded, not re-run) and 1.18x on a
Raspberry Pi 5 (2026-09-21) ([`BENCHMARKS.md` § Quantising the index copy](../../BENCHMARKS.md#quantising-the-index-copy-quantized_index)).
The aarch64 code is compiled out on x86, so only the `test-arm64` CI job
builds and tests it.
## Features
- AVX2 and NEON SIMD acceleration
- AVX-512 support (`avx512` feature)
- Float16 conversion (`float16` feature)
- CRC32 checksum acceleration
| Feature | Default | What | Builds C |
|---|---|---|---|
| `avx512` | no | AVX-512F kernels | no |
| `float16` | no | `f16_to_f32_batch` through the `half` crate (a software conversion otherwise) | no |
## Usage
```rust
use clawhdf5_accel::checksum::crc32_simd;
let crc = crc32_simd(&data);
```
The half-precision conversion used for stored embeddings is
`clawhdf5_format::float16`, not this crate's.
## License
+112 -16
View File
@@ -1,28 +1,124 @@
# clawhdf5-agent
[![crates.io](https://img.shields.io/crates/v/clawhdf5-agent.svg)](https://crates.io/crates/clawhdf5-agent)
[![docs.rs](https://img.shields.io/docsrs/clawhdf5-agent)](https://docs.rs/clawhdf5-agent)
Persistent memory for AI agents in a single HDF5 file: text chunks with
embeddings and metadata, hybrid search (HNSW vector search + BM25 keyword
search, fused), sessions, a knowledge graph, a write-ahead log for crash
safety, and optionally Ed25519-signed checkpoints. Stores open in h5py like
any other HDF5 file. Built on [`clawhdf5`](../clawhdf5/README.md),
[`clawhdf5-ann`](../clawhdf5-ann/README.md) and
[`clawhdf5-accel`](../clawhdf5-accel/README.md).
HDF5-backed persistent memory store for on-device AI agents.
It is a library: no agent framework integrates it (OpenClaw and ZeroClaw
integration claims were withdrawn on 2026-09-25; see
[`docs/openclaw.md`](../../docs/openclaw.md)). The command-line front end
is [`clawhdf5-cli`](../clawhdf5-cli/README.md).
Built on [clawhdf5](https://crates.io/crates/clawhdf5), clawhdf5-agent provides a vector-searchable memory backend optimized for edge AI workloads. Store embeddings, text chunks, and metadata in a single HDF5 file with SIMD-accelerated similarity search.
## Features
- Persistent vector store in HDF5 format
- Cosine similarity and L2 distance search
- SIMD-accelerated via clawhdf5-accel (AVX2, NEON)
- Optional GPU acceleration via clawhdf5-gpu
- Memory-mapped access for large stores
- f16 storage support for compact embeddings
## Usage
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-agent = "2.1.0"
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust,no_run
use std::path::PathBuf;
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};
let config = MemoryConfig::new(PathBuf::from("agent.h5"), "my-agent", 384);
let mut mem = HDF5Memory::create(config)?;
mem.save(MemoryEntry {
chunk: "The deploy key rotates every Monday.".into(),
embedding: vec![0.01; 384], // from your embedding model
source_channel: "chat".into(),
timestamp: 1_790_000_000.0,
session_id: "s1".into(),
tags: "ops".into(),
})?;
let query = vec![0.01f32; 384];
let hits = mem.search(&query, "deploy key", &SearchOptions::new(5).with_sources(["chat"]));
for h in &hits {
println!("{:.3} {}", h.score, h.chunk);
}
mem.flush_wal()?; // checkpoint now; otherwise one is made once the WAL holds more than 500 entries (wal_max_entries)
# Ok::<(), clawhdf5_agent::MemoryError>(())
```
## What is in it
- **`HDF5Memory`** — `create`, `open` (single writer: an exclusive lock on
`<store>.h5.lock`, a second opener gets `MemoryError::Locked`),
`open_read_only` (no lock, never writes). Through the `AgentMemory`
trait: `save`, `save_batch`, `delete`, `compact`, `count`, `snapshot`,
sessions; also `save_or_update`, `delete_batch`, `flush_wal`.
- **Search** — `search(query_embedding, text, &SearchOptions)`: optional
source-channel filter applied before ranking, vector + BM25 fusion
(weighted or RRF), Hebbian activation scaling, optional re-ranking
(`reranker::ReRankConfig`) and confidence rejection
(`confidence::ConfidenceConfig`). `hybrid_search` and
`hybrid_search_with` are thin wrappers. The vector stage uses the HNSW
index (`hnsw` feature); its graph is saved to `<store>.h5.ann` at each
checkpoint and reloaded on open (rebuilt if stale or damaged).
- **Storage settings** (`MemoryConfig`, persisted with the store):
`float16` embeddings (on by default for new stores; 48% smaller file at
100K records, same retrieval on LongMemEval), `quantized_index` (int8
copy of the vectors in the index, on by default; re-scored against the
exact embeddings), `compression` (off by default), HNSW `m`/`ef`
parameters, WAL settings (`wal_enabled`, on by default; `wal_max_entries`,
500: the WAL is checkpointed into the `.h5` once it holds more).
- **WAL** (`wal`) — every write is appended to `<store>.h5.wal` with a
chained CRC32 per entry, so a corrupted, reordered or spliced entry stops
replay. Recovers from a process crash at any point, including between a
checkpoint and the WAL truncate. WAL appends are not fsynced: saves since
the last checkpoint can be lost on power failure. An unreadable WAL is
quarantined to `<store>.h5.wal.corrupt-<ts>`.
- **Signed checkpoints** (`signing`) — `set_signing_key` signs a manifest
(SHA-256 Merkle tree over records, plus settings, sessions and graph) at
every checkpoint; `HDF5Memory::verify(path, &public_key)` checks it and
locates edits. WAL entries after the checkpoint are not covered.
- **Knowledge graph** (`knowledge`, `entity_extract`) — `add_entity`,
`add_entity_alias`, `add_relation`, `extract_and_store_entities`,
traversal and spreading activation.
- **Also:** sessions (`session`), temporal index (`temporal`),
consolidation tiers (`consolidation`), an in-memory TTL tier
(`ephemeral`), multi-modal embeddings (`multimodal`), `AGENTS.md`
generation (`agents_md`), query expansion, and a session-scoped
provenance ledger and write-anomaly detector on every save
(`take_anomaly_alerts`; alerts never block a save, and the source is
inferred from `source_channel`, not authenticated).
- `openclaw::ClawhdfBackend` is `search` with re-ranking and confidence
on, plus Markdown import/export. The module name is historical: it is not
an OpenClaw plugin.
## Features
| Feature | Default | What | Builds C |
|---|---|---|---|
| `hnsw` | yes | HNSW vector index (`clawhdf5-ann`); without it the vector stage is an exact linear cosine scan | no |
| `parallel` | yes | build the HNSW index on a rayon pool (same graph either way) | no |
| `float16` | yes | f16 helpers in `vector_search` (`half`). Stores' `MemoryConfig::float16` works without it. | no |
| `fast-math` | no | `matrixmultiply` batch distances in `strategy` | no |
| `accelerate` | no | Apple Accelerate BLAS in `strategy` (macOS) | links a system framework |
| `openblas` | no | OpenBLAS in `strategy` | yes (`openblas-src`) |
| `gpu` | no | `gpu_search` through [`clawhdf5-gpu`](../clawhdf5-gpu/README.md) (wgpu), used by `strategy`, not by `HDF5Memory::search` | no, but needs GPU drivers |
| `zstd` | no | Zstd instead of deflate when `MemoryConfig::compression` is on | yes (libzstd) |
| `async` | no | `async_memory` wrapper on tokio | no |
`--no-default-features --features float16` forces the exact linear scan.
## Measurements and limits
- Search recall and latency, file size, LongMemEval and MemoryArena
retrieval numbers: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with
the `clawhdf5-bench` binaries (`search_harness`, `longmemeval_bench`,
`footprint_bench`, ...).
- Known issues and their history: [`docs/known-issues.md`](../../docs/known-issues.md).
- Migrating a SQLite memory database:
[`clawhdf5-migrate`](../clawhdf5-migrate/README.md).
## License
MIT
+12 -2
View File
@@ -20,7 +20,7 @@ const LOCK_RETRY_DELAY: std::time::Duration = std::time::Duration::from_millis(1
/// never leaves a stale lock behind; the empty lock file itself is harmless).
#[derive(Debug)]
pub(crate) struct StoreLock {
_file: File,
file: File,
}
impl StoreLock {
@@ -42,7 +42,7 @@ impl StoreLock {
let mut attempts_left = LOCK_RETRIES;
loop {
match file.try_lock() {
Ok(()) => return Ok(Self { _file: file }),
Ok(()) => return Ok(Self { file }),
Err(TryLockError::WouldBlock) if attempts_left > 0 => {
attempts_left -= 1;
std::thread::sleep(LOCK_RETRY_DELAY);
@@ -60,6 +60,16 @@ impl StoreLock {
}
}
impl Drop for StoreLock {
/// Unlocks before the file is closed: a process another thread forks
/// inherits the descriptor until it execs, and a `flock` lasts while any
/// descriptor of the open file does, so closing alone could keep the
/// store locked for a moment after the drop (see `FileEditor`'s `Drop`).
fn drop(&mut self) {
let _ = self.file.unlock();
}
}
#[cfg(test)]
mod tests {
use super::*;
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-android"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "Android JNI bridge for edgehdf5-memory HDF5 backend"
description = "Android JNI bindings for clawhdf5 agent memory"
license = "MIT"
[lib]
+44
View File
@@ -0,0 +1,44 @@
# clawhdf5-android
A C ABI over [`clawhdf5-agent`](../clawhdf5-agent/README.md) for Android
apps: a `cdylib` exporting `extern "C"` functions (`edgehdf5_*`, a name
kept from the project's earlier "edgehdf5" days) that manage an
`HDF5Memory` through an opaque handle.
The functions are plain C symbols, not JNI-mangled `Java_...` entry points:
a Kotlin/Java app calls them through a thin JNI shim or JNA of its own. No
such shim, Gradle project or AAR is in this repository, and the crate is
not built for an Android target in CI (only its host-side unit tests run
with the workspace).
## Functions
| Function | What |
|---|---|
| `edgehdf5_create(path, agent_id, embedding_dim)` / `edgehdf5_open(path)` | a handle, or null on failure |
| `edgehdf5_close(handle)` | drop the store; what is not yet checkpointed stays in its WAL, as with any `HDF5Memory` |
| `edgehdf5_save(handle, ...)` | save one entry; the embedding length is checked against the store's dimension before the pointer is read |
| `edgehdf5_delete`, `edgehdf5_count`, `edgehdf5_count_active` | |
| `edgehdf5_hybrid_search(handle, query, len, text, vector_weight, keyword_weight, max_results, out_indices, out_scores, out_chunks)` | results into caller-provided arrays; returns the number written |
| `edgehdf5_add_session`, `edgehdf5_get_session_summary` | sessions |
| `edgehdf5_add_entity`, `edgehdf5_add_relation` | knowledge graph |
| `edgehdf5_free_string` | free a string this library returned |
Every function is `unsafe`: the caller guarantees valid, NUL-terminated
strings and correctly sized buffers (see each function's `# Safety`
section), and serialises access to a handle; separate handles are
independent.
## Build
```bash
cargo build --release -p clawhdf5-android # host build; for a device, add --target aarch64-linux-android with the NDK's linker configured
```
It depends on `clawhdf5-agent` with **default features off**, so there is
no HNSW index (the vector stage is an exact linear scan) and no rayon
pool. No C is compiled.
## License
MIT
+55 -10
View File
@@ -1,25 +1,70 @@
# clawhdf5-ann
[![crates.io](https://img.shields.io/crates/v/clawhdf5-ann.svg)](https://crates.io/crates/clawhdf5-ann)
[![docs.rs](https://docs.rs/clawhdf5-ann/badge.svg)](https://docs.rs/clawhdf5-ann)
An HNSW (Hierarchical Navigable Small World) approximate nearest-neighbour
index in pure Rust, with cosine or L2 distance, optional int8 storage of
the vectors, deletions, and persistence as an HDF5 file. It is the vector
stage of [`clawhdf5-agent`](../clawhdf5-agent/README.md)'s search (the
agent's `hnsw` feature, on by default); distances run on
[`clawhdf5-accel`](../clawhdf5-accel/README.md)'s SIMD kernels.
HNSW approximate nearest neighbor index stored as HDF5.
Neighbours are chosen with the HNSW paper's diversity heuristic, not plain
closest-M (which capped recall on clustered data at 0.31 recall@10 at 100K
vectors).
## Features
Not on crates.io yet; depend on it from git:
- Build and query HNSW indexes persisted in HDF5 format
- Pure Rust, no C dependencies
- Efficient similarity search for high-dimensional vectors
```toml
[dependencies]
clawhdf5-ann = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
use clawhdf5_ann::HnswIndex;
use clawhdf5_ann::{DistanceMetric, HnswIndex, Storage};
let index = HnswIndex::from_hdf5("vectors.h5").unwrap();
let neighbors = index.search(&query, 10);
let vectors: Vec<Vec<f32>> = (0..500)
.map(|i| (0..16).map(|j| ((i * 31 + j * 7) % 97) as f32 / 97.0).collect())
.collect();
// m = 16 connections per node, ef_construction = 200
let mut index = HnswIndex::build_with(&vectors, 16, 200, DistanceMetric::Cosine, Storage::Int8);
let hits = index.search(&vectors[42], 10, 64); // (id, distance), closest first; ef >= k
assert!(hits[0].1 < 1e-3); // vector 42 itself (or an identical one)
let id = index.insert(vec![0.5; 16]);
index.mark_deleted(id);
// Persist as HDF5 (a self-contained file: graph and vectors) and load it back
let bytes = index.to_hdf5_bytes().unwrap();
let loaded = HnswIndex::load_from_hdf5(&bytes).unwrap();
assert_eq!(loaded.len(), index.len());
```
- `HnswIndex::build` (L2), `build_with_metric`, `build_with` (metric and
storage); `new`/`new_with` plus `insert` for an index built
incrementally.
- `Storage::Int8` keeps each vector as `i8`, a quarter of the memory; it
applies to `Cosine` only (an L2 index keeps `Float32`). Distances are then
approximate, so a caller that needs exact ranking re-scores the
candidates, as the agent does.
- `mark_deleted`, `is_deleted`, `deleted_count`, `active_len`, `compact`
(returns the old-to-new id map).
- `save_to_hdf5(&mut writer)` / `to_hdf5_bytes` / `load_from_hdf5` store
the whole index; `graph_to_bytes` / `from_graph_bytes` store only the
graph (with a CRC32) for a caller that keeps the vectors elsewhere — the
agent's `<store>.h5.ann` sidecar.
## Features
| Feature | Default | What | Builds C |
|---|---|---|---|
| `parallel` | no | build the graph on a rayon pool; the graph is identical with or without it | no |
Recall and speed against exact search, for the index alone and in the
agent: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with
`cargo run --release -p clawhdf5-bench --bin search_harness`.
## License
MIT
+49
View File
@@ -0,0 +1,49 @@
# clawhdf5-bench
The measurement harnesses behind [`BENCHMARKS.md`](../../BENCHMARKS.md):
HDF5 read and write speed (against libhdf5 and h5py where noted) and the
agent store's search, footprint and retrieval quality. Not meant for
publishing; nothing else in the workspace depends on it. Run everything with
`--release`, and quote numbers with the machine, date and command, as
`BENCHMARKS.md` does.
## Binaries
| Binary | Measures |
|---|---|
| `read_harness` | full reads vs hyperslab selections of a chunked 2-D dataset (compressed and not) and a contiguous one: does a selection cost scale with the selection or the dataset? (`-- --large` for 512 MB) |
| `concurrent_read` | decoded read throughput vs threads on one open `File`; `scripts/concurrent_read_h5py.py` runs the same workload with h5py (threads and processes) and `scripts/compare_concurrent_read.py` tabulates both |
| `search_harness` | HNSW recall@10 vs exact search, QPS and latency per `ef`, and end-to-end `HDF5Memory` ingest/checkpoint/open/search at 1K–100K (`--full`); studies: `--float16-study`, `--options-study`, `--signing-study`, `--ann-only --uniform` |
| `longmemeval_bench` | LongMemEval retrieval recall (turn and session Hit@k, MRR) — **retrieval, not QA accuracy**. Oracle or full `longmemeval_s` haystack; `--features embeddings` (or `embeddings-cuda`) embeds with MiniLM, otherwise the vector stage is inert and the run is BM25-only |
| `memory_arena` | a deterministic multi-session retrieval benchmark (BM25-only) |
| `footprint_bench` | file size and bytes per record at 100–100K records, float16 or `--f32`, WAL on/off, compressed or not |
| `consolidation_efficiency` | retrieval before and after consolidation on signal + noise records |
| `ephemeral_perf` | the in-memory ephemeral tier's set/get latency |
| `mpi_io_bench` | `clawhdf5-io`'s `MpiVol` (root-read + broadcast, not collective I/O); needs `--features mpi-io` and `mpirun` |
```bash
cargo run --release -p clawhdf5-bench --bin search_harness -- --full
cargo run --release -p clawhdf5-bench --bin read_harness
```
## Criterion benches and example
- `cargo bench -p clawhdf5-bench` runs `h5bench_write`, `h5bench_read` and
`h5bench_meta` (h5bench-style sequential, chunked, strided and metadata
workloads). `--features libhdf5-compare` adds the same workloads through
libhdf5 (the `hdf5-metno` crate; needs a system libhdf5 1.14).
- `examples/worldmodel_sampling.rs`: shuffled per-frame reads of a
`(N, H, W, C)` `uint8` dataset, clawhdf5 against h5py on the same file.
## Features
| Feature | What | Builds C |
|---|---|---|
| `libhdf5-compare` | libhdf5 variants of the Criterion benches | links the system libhdf5 |
| `mpi-io` | `mpi_io_bench` | yes (`mpi-sys`; needs an MPI installation) |
| `embeddings` | MiniLM embeddings for `longmemeval_bench` (candle) | yes (a `cc` build dependency in the candle/tokenizers tree) |
| `embeddings-cuda` | the same on a CUDA GPU (minutes instead of hours on the full haystack) | yes (CUDA) |
## License
MIT
+32 -9
View File
@@ -7,15 +7,19 @@
//! ```text
//! cargo run --release -p clawhdf5-bench --bin read_harness
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --large # 512 MB
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --v18 # HDF5 1.8 format
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk 32 # 32 x 32 chunks
//! ```
//!
//! `--v18` writes the file with `libver_bounds(V18, V18)` (version-1 B-tree
//! chunk indexes) instead of the default 1.10 format (Fixed Array indexes
//! here), to compare the two.
use std::time::{Duration, Instant};
use clawhdf5::{File, FileBuilder};
use clawhdf5::{File, FileBuilder, LibVer};
use clawhdf5_format::selection::Selection;
const CHUNK: u64 = 256;
struct Layout {
name: &'static str,
chunked: bool,
@@ -46,16 +50,19 @@ fn value(row: u64, col: u64) -> f64 {
(row * 100_003 + col) as f64 * 0.5
}
fn write_file(path: &std::path::Path, rows: u64, cols: u64) {
fn write_file(path: &std::path::Path, rows: u64, cols: u64, chunk: u64, v18: bool) {
let data: Vec<f64> = (0..rows)
.flat_map(|r| (0..cols).map(move |c| value(r, c)))
.collect();
let mut builder = FileBuilder::new();
if v18 {
builder.libver_bounds(LibVer::V18, LibVer::V18);
}
for (i, layout) in LAYOUTS.iter().enumerate() {
let ds = builder.create_dataset(&format!("d{i}"));
ds.with_f64_data(&data).with_shape(&[rows, cols]);
if layout.chunked {
ds.with_chunks(&[CHUNK, CHUNK]);
ds.with_chunks(&[chunk, chunk]);
}
if layout.deflate {
ds.with_deflate(4);
@@ -91,7 +98,14 @@ fn slab(start: [u64; 2], count: [u64; 2]) -> Selection {
}
fn main() {
let large = std::env::args().any(|a| a == "--large");
let args: Vec<String> = std::env::args().collect();
let large = args.iter().any(|a| a == "--large");
let v18 = args.iter().any(|a| a == "--v18");
let chunk: u64 = args
.iter()
.position(|a| a == "--chunk")
.and_then(|i| args.get(i + 1))
.map_or(256, |c| c.parse().expect("--chunk N"));
let (rows, cols) = if large { (8192, 8192) } else { (4096, 2048) };
let total_mb = (rows * cols * 8) as f64 / (1 << 20) as f64;
if cfg!(debug_assertions) {
@@ -100,12 +114,21 @@ fn main() {
let dir = tempfile::TempDir::new().unwrap();
let path = dir.path().join("read_harness.h5");
write_file(&path, rows, cols);
let file_mb = std::fs::metadata(&path).unwrap().len() as f64 / (1 << 20) as f64;
let t = Instant::now();
write_file(&path, rows, cols, chunk, v18);
let write_ms = t.elapsed().as_secs_f64() * 1e3;
let file_bytes = std::fs::metadata(&path).unwrap().len();
let file_mb = file_bytes as f64 / (1 << 20) as f64;
println!("## Read harness");
println!(
"\n{rows} x {cols} f64 ({total_mb:.0} MB per dataset), chunks {CHUNK} x {CHUNK}, file {file_mb:.0} MB\n"
"\n{rows} x {cols} f64 ({total_mb:.0} MB per dataset), chunks {chunk} x {chunk}, \
format {}, file {file_mb:.0} MB ({file_bytes} bytes), written in {write_ms:.0} ms\n",
if v18 {
"1.8 (v1 B-tree)"
} else {
"1.10 (default)"
}
);
// (label, selection, elements selected)
+49
View File
@@ -0,0 +1,49 @@
# clawhdf5-cli
The `clawhdf5` command: create, fill, search and inspect a
[`clawhdf5-agent`](../clawhdf5-agent/README.md) memory store from the
shell. Output is JSON. (For general HDF5 files use `h5rs` from
[`clawhdf5-tools`](../clawhdf5-tools/README.md).)
```bash
cargo install --path crates/clawhdf5-cli # installs `clawhdf5`; not on crates.io yet
# or: cargo run -p clawhdf5-cli -- --help
```
No C is compiled.
## Commands
The store is `--path FILE` (or `CLAWHDF5_PATH`) before the subcommand.
| Command | What |
|---|---|
| `create [--agent-id ID] [--dim N] [--wal] [--f32] [--f32-index]` | a new store (dimension 384 by default); float16 embeddings and an int8 index copy unless `--f32` / `--f32-index`. The WAL is off unless `--wal` (the library's default is on), so each save is checkpointed at once |
| `save [--json '{...}']` | save one entry, from `--json` or stdin: `{"chunk", "embedding", "source_channel", "timestamp", "session_id", "tags"}` |
| `search --embedding '[...]' [--query TEXT] [-k N] [--vector-weight W] [--keyword-weight W]` | hybrid search (defaults 5 results, weights 0.7 / 0.3) |
| `recall INDEX` | one entry by index |
| `stats` | counts and configuration |
| `flush-wal` | checkpoint the WAL into the `.h5` |
| `agents-md [--output FILE]` | generate an `AGENTS.md` from the store |
| `export` | every entry as JSON lines |
| `snapshot DEST` | a copy of the store's `.h5` file |
| `keygen --out FILE` | a new Ed25519 signing key (64 hex characters, created owner-only on Unix) |
| `verify --public-key HEX_OR_FILE` | check a signed store; exit status 2 if it does not verify |
`recall`, `stats`, `agents-md` and `export` open the store read-only
(no lock, nothing written), so they work while another process has it
open. `save`, `search` (which records activation boosts) and `flush-wal`
open it for writing and take the store's lock. With
`--signing-key FILE` (or `CLAWHDF5_SIGNING_KEY`) every checkpoint a command
makes is signed; a signed store refuses to checkpoint without the key.
```bash
clawhdf5 --path mem.h5 create --agent-id demo --dim 3
echo '{"chunk":"hello","embedding":[0.1,0.2,0.3],"source_channel":"cli","timestamp":0,"session_id":"s1","tags":""}' \
| clawhdf5 --path mem.h5 save
clawhdf5 --path mem.h5 search --embedding '[0.1,0.2,0.3]' --query hello -k 3
```
## License
MIT
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-derive"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "Derive macros for rustyhdf5 HDF5 traits"
description = "Derive macro (H5Type) for clawhdf5 compound types"
license = "MIT"
repository = "https://git.redclaw.dev/quantumclaw/clawhdf5"
readme = "README.md"
+34 -12
View File
@@ -1,28 +1,50 @@
# clawhdf5-derive
[![crates.io](https://img.shields.io/crates/v/clawhdf5-derive.svg)](https://crates.io/crates/clawhdf5-derive)
[![docs.rs](https://docs.rs/clawhdf5-derive/badge.svg)](https://docs.rs/clawhdf5-derive)
`#[derive(H5Type)]`: maps a Rust struct with named fields to an HDF5
compound datatype. The derive generates three inherent methods:
Derive macros for clawhdf5 HDF5 traits.
- `hdf5_datatype() -> clawhdf5_format::datatype::Datatype` — the
`Datatype::Compound` (members in field order, packed, little-endian);
- `to_bytes(&self) -> Vec<u8>` — one element in that layout;
- `from_bytes(&[u8]) -> Self` — the reverse (panics if the slice is shorter
than the compound).
## Features
Supported field types: `f32`, `f64`, `i8`–`i64`, `u8`–`u64`, `bool`
(stored as `u8`) and fixed-size arrays `[T; N]` of those numeric types.
Tuple structs, enums and nested structs are refused at compile time.
- `#[derive(HDF5Type)]` for automatic HDF5 datatype mapping
- Struct-to-compound-type derivation
The generated code names `clawhdf5_format`, so the crate using the derive
must depend on [`clawhdf5-format`](../clawhdf5-format/README.md) too. Not
on crates.io yet:
## Usage
```toml
[dependencies]
clawhdf5-derive = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Example
```rust
use clawhdf5_derive::HDF5Type;
use clawhdf5_derive::H5Type;
use clawhdf5_format::datatype::Datatype;
#[derive(HDF5Type)]
#[derive(H5Type, Debug, PartialEq)]
struct Point {
x: f64,
y: f64,
z: f64,
id: u32,
pos: [f64; 3],
valid: bool,
}
let p = Point { id: 7, pos: [1.0, 2.0, 3.0], valid: true };
let bytes = p.to_bytes();
assert_eq!(bytes.len(), 4 + 24 + 1);
assert_eq!(Point::from_bytes(&bytes), p);
assert!(matches!(Point::hdf5_datatype(), Datatype::Compound { size: 29, .. }));
```
Tests: `crates/clawhdf5-format/tests/derive_tests.rs`.
## License
MIT
+39 -10
View File
@@ -1,27 +1,56 @@
# clawhdf5-filters
[![crates.io](https://img.shields.io/crates/v/clawhdf5-filters.svg)](https://crates.io/crates/clawhdf5-filters)
[![docs.rs](https://docs.rs/clawhdf5-filters/badge.svg)](https://docs.rs/clawhdf5-filters)
Standalone deflate (zlib) compression and decompression with a choice of
backend: pure-Rust zlib-rs (default), zlib-ng, Apple's Compression
framework, or miniz_oxide.
Filter and compression pipeline for clawhdf5.
This crate holds **deflate backends only**. The HDF5 filter pipeline, the
filter registry and every other codec (shuffle, Fletcher-32, N-Bit,
scale-offset, LZ4, Zstd, SZIP, pcodec, LZF, bitshuffle, bzip2, Blosc,
Blosc2, ZFP) live in [`clawhdf5-format`](../clawhdf5-format/README.md),
which calls flate2 itself and selects its deflate backend with its own
features. No library crate of the workspace depends on this one (the
`clawhdf5` facade uses it only in tests).
## Features
Not on crates.io yet; depend on it from git:
- DEFLATE compression/decompression
- Pure-Rust deflate via zlib-rs (default, `zlib-rs` feature)
- zlib-ng instead, if you want it (`fast-deflate` feature; C, needs cmake)
- Apple Compression framework support (`apple-compression` feature)
```toml
[dependencies]
clawhdf5-filters = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
## API
```rust
use clawhdf5_filters::{deflate_compress, deflate_decompress};
use clawhdf5_filters::{deflate_backend, deflate_compress, deflate_decompress};
let data: Vec<u8> = (0..10_000u32).map(|i| (i % 251) as u8).collect();
let compressed = deflate_compress(&data, 6).unwrap();
// The second argument bounds the output: the expected decompressed size.
let decompressed = deflate_decompress(&compressed, data.len()).unwrap();
assert_eq!(decompressed, data);
println!("backend: {}", deflate_backend()); // "zlib-rs" by default
```
Also `deflate_compress_miniz`/`deflate_decompress_miniz` (always
miniz_oxide) and `fast_deflate::{compress, decompress, active_backend}`.
## Features
Backend priority: `apple-compression` (macOS only) > zlib-ng > zlib-rs >
miniz_oxide (with none enabled).
| Feature | Default | Backend | Builds C |
|---|---|---|---|
| `zlib-rs` | yes | zlib-rs through flate2, with `runtime_detection` (needed for its SIMD) | no |
| `fast-deflate` | no | zlib-ng through flate2 | yes (cmake) |
| `system-zlib` | no | the system zlib through flate2 | yes (`libz-sys`) |
| `apple-compression` | no | Apple Compression framework, macOS only (ignored elsewhere) | no (links a system framework) |
zlib-rs matches zlib-ng on HDF5 reads and writes and produces
byte-identical output: see "Deflate backend" in
[`BENCHMARKS.md`](../../BENCHMARKS.md).
## License
MIT
+32 -3
View File
@@ -327,24 +327,53 @@ fn inflate_bounded(data: &[u8], size_hint: usize, limit: usize) -> Result<Vec<u8
}
}
/// Largest compression output reserved at its worst-case size up front.
const DEFLATE_EXACT_BOUND: usize = 64 << 20;
/// Compress data using flate2 (zlib-ng, zlib-rs or miniz_oxide; see module docs).
pub(crate) fn flate2_compress(data: &[u8], level: u32) -> Result<Vec<u8>, String> {
use flate2::{Compress, Compression, FlushCompress, Status};
// zlib's compressBound, plus the zlib header and trailer.
let bound = data.len() + (data.len() >> 12) + (data.len() >> 14) + (data.len() >> 25) + 13 + 6;
// flate2's Rust backends (zlib-rs, miniz_oxide) zero the whole spare
// capacity on each call, so a large input's worst-case bound would be
// memory held for nothing (4 GiB for a 4 GiB chunk that deflates to a
// few MiB): past 64 MiB the output starts at 1/16 of the bound and
// doubles as needed.
let first = if bound <= DEFLATE_EXACT_BOUND {
bound
} else {
bound / 16
};
let mut out = Vec::new();
out.try_reserve_exact(bound)
out.try_reserve_exact(first)
.map_err(|e| format!("deflate: cannot allocate output: {e}"))?;
let mut deflater = Compress::new(Compression::new(level), true);
loop {
let (in_before, out_before) = (deflater.total_in(), deflater.total_out());
let rest = &data[in_before as usize..];
// zlib takes at most u32::MAX input bytes per call, and `Finish`
// ends the stream after the bytes it took: input of 4 GiB or more
// was cut at 4 GiB - 1. Finish only once the rest fits one call.
let flush = if rest.len() > u32::MAX as usize {
FlushCompress::None
} else {
FlushCompress::Finish
};
let status = deflater
.compress_vec(&data[in_before as usize..], &mut out, FlushCompress::Finish)
.compress_vec(rest, &mut out, flush)
.map_err(|e| format!("deflate: {e}"))?;
match status {
Status::StreamEnd => return Ok(out),
Status::StreamEnd => {
if out.capacity() - out.len() > DEFLATE_EXACT_BOUND {
out.shrink_to_fit();
}
return Ok(out);
}
// Out of room (the bound makes it unreachable below
// `DEFLATE_EXACT_BOUND`): grow rather than fail.
Status::Ok | Status::BufError if out.len() == out.capacity() => out
.try_reserve(out.capacity().max(4096))
.map_err(|e| format!("deflate: cannot allocate output: {e}"))?,
+97 -16
View File
@@ -1,27 +1,108 @@
# clawhdf5-format
[![crates.io](https://img.shields.io/crates/v/clawhdf5-format.svg)](https://crates.io/crates/clawhdf5-format)
[![docs.rs](https://docs.rs/clawhdf5-format/badge.svg)](https://docs.rs/clawhdf5-format)
The HDF5 file format in pure Rust: parsers and writers for every on-disk
structure, the filter pipeline and its codecs, and the shared type
definitions the other crates use. Most users want the
[`clawhdf5`](../clawhdf5/README.md) facade, which wraps this crate in an
h5py-like API; use this one directly for low-level access or in `no_std`
code.
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## What is in it
- **Parsing:** superblock v0–v3 (`superblock`, with the superblock
extension and metadata cache images, `superblock_ext`), object headers v1
and v2 (`object_header`), every header message the readers use
(`datatype`, `dataspace`, `data_layout` v1–v4 including virtual datasets,
`fill_value`, `attribute`, `link_message`, `shared_message`, ...), groups
old and new (`group_v1` symbol tables with local heaps, `group_v2` with
fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single
chunk, implicit, fixed array, extensible array, v2 B-tree).
- **Reading data:** `data_read` (contiguous, compact, chunked),
`partial_read` and `selection` (hyperslabs and points), `vl_data`
(variable-length strings and sequences through the global heap),
`chunk_cache`.
- **Storage:** the `storage::Storage` trait (`read_at`, `read_ranges`,
`len`, `hint`) that every read path goes through, so a file can be read
from memory, a file handle or a remote backend
([`clawhdf5-remote`](../clawhdf5-remote/README.md)).
- **Writing:** `file_writer::FileWriter` and the builders in
`type_builders` (datasets, groups, attributes, compound and enum types,
links, virtual datasets, creation-order tracking); chunk indexes and
dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`,
`ea_writer`, and the version-1 chunk B-tree of `btree_v1_write`). Output
is read by h5py and h5dump; `FileWriter::libver_bounds` (`libver`) picks
the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other
IDs can be registered at run time with `register_filter`). Built in:
deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4,
Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle,
bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
- **Shared pieces:** `float16` (the one IEEE half-precision conversion the
workspace uses), `provenance` (SHA-256 dataset hashes), `checksum`
(Jenkins lookup3 for v2+ structures).
## Example
```rust
use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};
// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3])
.set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();
// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
.unwrap();
assert!(!hdr.messages.is_empty());
```
## Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
- `no_std` support (disable `std` feature)
- Optional parallel reads via Rayon
- SHA-256 provenance tracking
| Feature | Default | What | Builds C |
|---|---|---|---|
| `std` | yes | standard library; without it the crate is `no_std` + `alloc` (CI builds it for `thumbv7em-none-eabihf`) | no |
| `checksum` | yes | verify Jenkins lookup3 checksums | no |
| `deflate` | yes | deflate through flate2 | no |
| `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no |
| `system-zlib-decompress` | yes | macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere | no (links the system libz on macOS) |
| `provenance` | yes | SHA-256 provenance hashes | no |
| `lzf` | yes | LZF (32000) | no |
| `parallel` | no | rayon-parallel chunk decoding | no |
| `fast-checksum` | no | hardware CRC32 through `crc32fast` | no |
| `lz4` | no | LZ4 (32004) | no |
| `pcodec` | no | pcodec | no |
| `bitshuffle`, `bzip2`, `blosc` | no | 32008, 307, 32001, read and write | no |
| `blosc2`, `zfp` | no | 32026, 32013, read only | no |
| `plugin-filters` | no | all six plugin filters above | no |
| `lookup-stats` | no | counters for name-lookup benchmarks | no |
| `zstd` | no | Zstandard (32015) | yes (libzstd) |
| `szip` | no | SZIP (4) decoding | links the system libaec (`libaec-dev`) |
| `fast-deflate` | no | zlib-ng | yes (cmake) |
| `system-zlib` | no | the system zlib | yes (`libz-sys`) |
| `blake3_hash` | no | `provenance::blake3_hash` | yes (`cc`) |
## Usage
## Robustness
```rust
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
```
Every parser is meant to return an error, never panic, on hostile input:
nine cargo-fuzz targets live in [`fuzz/`](fuzz/README.md), the conformance
sweep includes the HDF Group's CVE corpus
([`CONFORMANCE.md`](../../CONFORMANCE.md)), and header checks follow
libhdf5's. Open gaps are in [`docs/known-issues.md`](../../docs/known-issues.md).
## License
+12 -4
View File
@@ -51,10 +51,18 @@ done
## CI
These targets are **not** run in CI (`.gitea/workflows/ci.yml`) — cargo-fuzz
requires nightly and each meaningful run takes minutes, which doesn't fit a
per-PR gate. Run them manually on a schedule (e.g. before a release, or after
touching parser code) instead.
These targets are **not** run by the CI workflows (`.gitea/workflows/ci.yml`)
— cargo-fuzz requires nightly and each meaningful run takes minutes, which
doesn't fit a per-PR gate. Run them by hand before a release or after
touching parser code. `scripts/ci-test.sh` has an opt-in smoke run: with
`CLAWHDF5_FUZZ_SECONDS=N` it runs every target of this crate and of
`crates/clawhdf5-agent/fuzz` (the WAL parser) for N seconds each.
Other robustness checks that do run: the nightly conformance sweep reads
the HDF Group's CVE reproducers and fails on any panic, hang, crash or
out-of-memory ([`conformance/README.md`](../../../conformance/README.md)),
and `scripts/h5rs-fuzz.sh` runs every `h5rs` subcommand over them, optionally
on byte-flipped copies.
## Reproducing Crashes
+30 -12
View File
@@ -92,6 +92,8 @@ impl BTreeV1Node {
// + left_sibling(offset_size) + right_sibling(offset_size)
let os = offset_size as usize;
let header_size = 8 + os * 2;
// The body is read once the header says how long it is.
file.hint(offset, NODE_HINT_LEN);
let header = read_exact_at(file, offset, header_size)?;
let file_data: &[u8] = &header;
// The header's read checked that `offset + header_size` fits.
@@ -156,7 +158,18 @@ impl BTreeV1Node {
}
/// Maximum recursion depth for B-tree traversal (malformed data protection).
const MAX_BTREE_DEPTH: usize = 64;
pub(crate) const MAX_BTREE_DEPTH: usize = 64;
/// What a symbol table node takes with libhdf5's default group leaf K (4):
/// its 8-byte header and 2K entries of 40 bytes (8-byte offsets). Hinted
/// before one is read ([`Storage::hint`]); a node of another size is read
/// all the same.
const SNOD_HINT_LEN: usize = 8 + 8 * 40;
/// What a group B-tree node takes with libhdf5's default internal K (16):
/// its header (24 bytes with 8-byte offsets), 2K + 1 keys and 2K children
/// of 8 bytes. Hinted before one is read.
const NODE_HINT_LEN: usize = 24 + (2 * 16 + 1 + 2 * 16) * 8;
/// Collect all leaf-level child addresses (SNOD addresses) by traversing the B-tree.
pub fn collect_symbol_table_nodes(
@@ -196,20 +209,22 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
}
if node.node_level == 0 {
// Leaf: children are SNOD addresses
// Leaf: children are SNOD addresses, read next (see
// `Storage::hint`).
for &snod in &node.children {
file.hint(snod, SNOD_HINT_LEN);
}
Ok(node.children)
} else {
// Internal: recurse into children. After the first child that
// fails, the others are only read (as `storage::touch` does), not
// descended into; that error is returned.
// Internal: recurse into children. A child that fails does not
// stop the walk: the others are still descended into (reading, not
// using, what they hold), then the first error is returned. The
// result and the error are those of stopping at the first failure;
// a storage that records what it lacks (see `storage::touch`)
// learns every node the walk can reach in one attempt.
let mut result = Vec::new();
let mut failed = None;
for &child_addr in &node.children {
if failed.is_some() {
// Parsing reads the node's header, then its body.
let _ = BTreeV1Node::parse_in(file, child_addr, offset_size, length_size);
continue;
}
match collect_symbol_table_nodes_inner(
file,
child_addr,
@@ -217,8 +232,11 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
length_size,
depth + 1,
) {
Ok(child_snods) => result.extend(child_snods),
Err(e) => failed = Some(e),
Ok(child_snods) if failed.is_none() => result.extend(child_snods),
Ok(_) => {}
Err(e) => {
failed.get_or_insert(e);
}
}
}
match failed {
@@ -0,0 +1,470 @@
//! Writing a version-1 B-tree chunk index (node type 1): the chunk index of
//! layout message versions 1-3, and the only one HDF5 1.8 reads.
//!
//! The tree is built the way libhdf5 builds it when the chunks reach it one
//! after another in row-major order (a whole-dataset `H5Dwrite` of a 1-D
//! dataset, or of any dataset without a chunk cache; with one, libhdf5
//! inserts the small chunks of a multi-dimensional dataset in the order its
//! cache evicts them, which fills the nodes differently): each
//! chunk goes through the same steps as `H5B_insert` (`H5B.c`) with the
//! chunk callbacks of `H5Dbtree.c`, so nodes split where libhdf5's split,
//! with its default split ratios (a full right-most node keeps 90% of its
//! children, a left-most one 10%, any other half), and keys hold what
//! libhdf5's hold:
//!
//! - a chunk's key is its size in the file, its filter mask and its offsets
//! (the element-size coordinate 0);
//! - a node's final key is the zero-size key one chunk past the chunk that
//! last moved it (every scaled coordinate plus one, `H5D__btree_new_node`),
//! which libhdf5 moves only when a new chunk is not below it
//! (`H5D__btree_cmp3`) — so after an even number of appends in one
//! dimension it lies on the last chunk itself;
//! - a full root is copied to a new node and becomes the parent of the copy
//! and its new sibling, so the root's address (the layout message's) never
//! changes.
//!
//! Nodes are laid out in the order libhdf5 allocates them (the root first,
//! then each new node as a split creates it), all of the full node size, the
//! unused slots zero.
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use core::cmp::Ordering;
use crate::error::FormatError;
/// libhdf5's default chunk B-tree K (`HDF5_BTREE_CHUNK_IK_DEF`): nodes hold
/// up to 2K = 64 children. Superblocks of version 2 cannot record another
/// value without a superblock extension, which this writer does not emit.
pub(crate) const CHUNK_BTREE_K: u16 = 32;
/// libhdf5's default split ratios (`H5D_XFER_BTREE_SPLIT_RATIO_DEF`) for a
/// left-most, middle and right-most node.
const SPLIT_RATIOS: [f64; 3] = [0.1, 0.5, 0.9];
/// A chunk to index: scaled coordinates (offset / chunk dimension) in each
/// dataset dimension, stored size, filter mask and address.
pub(crate) struct ChunkEntry {
pub(crate) scaled: Vec<u64>,
pub(crate) nbytes: u64,
pub(crate) filter_mask: u32,
pub(crate) address: u64,
}
#[derive(Debug, Clone, PartialEq, Eq)]
struct Key {
nbytes: u32,
mask: u32,
/// Scaled coordinates, the element-size one (0 or 1) last.
scaled: Vec<u64>,
}
impl Key {
/// `H5D__btree_new_node`'s right key: one chunk past `self` in every
/// dimension, with no storage.
fn right_of(&self) -> Key {
Key {
nbytes: 0,
mask: 0,
scaled: self.scaled.iter().map(|s| s + 1).collect(),
}
}
fn cmp_scaled(&self, other: &Key) -> Ordering {
self.scaled.cmp(&other.scaled)
}
}
#[derive(Debug, Clone)]
struct Node {
level: u8,
left: Option<usize>,
right: Option<usize>,
/// `children.len() + 1` keys once the node holds a child.
keys: Vec<Key>,
/// Chunk addresses in a leaf, node indexes above.
children: Vec<u64>,
}
/// What an insertion below a node did (`H5B__insert_helper`'s outputs).
#[derive(Default)]
struct Ret {
/// The node's new left key (`lt_key_changed`).
lt: Option<Key>,
/// The node's new right key (`rt_key_changed`).
rt: Option<Key>,
/// The node split: the key shared by the halves and the new right node.
split: Option<(Key, usize)>,
}
struct Tree {
nodes: Vec<Node>,
two_k: usize,
}
fn bad(why: &str) -> FormatError {
FormatError::SerializationError(format!("version-1 B-tree chunk index: {why}"))
}
impl Tree {
fn new(k: u16) -> Self {
Self {
nodes: vec![Node {
level: 0,
left: None,
right: None,
keys: Vec::new(),
children: Vec::new(),
}],
two_k: 2 * usize::from(k),
}
}
/// `H5B_insert` of `key` (a chunk after every chunk already inserted).
fn insert(&mut self, key: &Key, addr: u64) -> Result<(), FormatError> {
let r = self.insert_helper(0, key, addr, 64)?;
let Some((md, split)) = r.split else {
return Ok(());
};
// The root split: copy it to a new node and make the root the
// parent of the copy and its new right sibling.
let lt = r.lt.unwrap_or_else(|| self.nodes[0].keys[0].clone());
let rt = match r.rt {
Some(rt) => rt,
None => self.nodes[split]
.keys
.last()
.cloned()
.ok_or_else(|| bad("empty node"))?,
};
let moved = self.nodes[0].clone();
let level = moved.level;
let moved_id = self.nodes.len();
self.nodes.push(moved);
self.nodes[split].left = Some(moved_id);
self.nodes[0] = Node {
level: level + 1,
left: None,
right: None,
keys: vec![lt, md, rt],
children: vec![moved_id as u64, split as u64],
};
Ok(())
}
fn insert_helper(
&mut self,
id: usize,
key: &Key,
addr: u64,
depth: u8,
) -> Result<Ret, FormatError> {
if depth == 0 {
return Err(bad("tree too deep"));
}
let n = self.nodes[id].children.len();
let level = self.nodes[id].level;
let mut ret = Ret::default();
if n == 0 {
// The first chunk (H5B_INS_FIRST): its key and the right key.
let node = &mut self.nodes[id];
node.keys = vec![key.clone(), key.right_of()];
node.children = vec![addr];
return Ok(ret);
}
// Binary search with H5D__btree_cmp3: 1 when the chunk is not below
// the right key, -1 when below the left key, else 0.
let (mut lo, mut hi, mut idx) = (0usize, n, 0usize);
let mut cmp = Ordering::Less;
while lo < hi && cmp != Ordering::Equal {
idx = (lo + hi) / 2;
let node = &self.nodes[id];
cmp = if key.cmp_scaled(&node.keys[idx + 1]) != Ordering::Less {
Ordering::Greater
} else if key.cmp_scaled(&node.keys[idx]) == Ordering::Less {
Ordering::Less
} else {
Ordering::Equal
};
if cmp == Ordering::Less {
hi = idx;
} else {
lo = idx + 1;
}
}
let (mut lt_changed, mut rt_changed) = (false, false);
// The child to add after child `idx`, with its left key.
let mut new_child: Option<(Key, u64)> = None;
match cmp {
Ordering::Less => return Err(bad("chunks out of order")),
Ordering::Greater if idx + 1 < n => {
return Err(bad("cannot place chunk"));
}
Ordering::Greater if level == 0 => {
// Past every chunk of the right-most leaf: a new maximum
// (H5B_INS_RIGHT through `new_node`), which moves the right
// key one chunk past it.
idx = n - 1;
self.nodes[id].keys[idx + 1] = key.right_of();
rt_changed = true;
new_child = Some((key.clone(), addr));
}
Ordering::Equal if level == 0 => {
// Inside the last chunk's range: H5D__btree_insert adds it
// to the right of that chunk; the right key stays.
if key.scaled == self.nodes[id].keys[idx].scaled {
return Err(bad("duplicate chunk"));
}
new_child = Some((key.clone(), addr));
}
_ => {
if cmp == Ordering::Greater {
idx = n - 1;
}
let child = usize::try_from(self.nodes[id].children[idx])
.map_err(|_| bad("bad node index"))?;
let r = self.insert_helper(child, key, addr, depth - 1)?;
if let Some(lt) = r.lt {
self.nodes[id].keys[idx] = lt;
lt_changed = true;
}
if let Some(rt) = r.rt {
self.nodes[id].keys[idx + 1] = rt;
rt_changed = true;
}
if let Some((md, split)) = r.split {
new_child = Some((md, split as u64));
}
}
}
// Pass the node's changed end keys up, as H5B__insert_helper does.
if lt_changed && idx == 0 {
ret.lt = Some(self.nodes[id].keys[0].clone());
}
if rt_changed && idx + 1 >= n {
ret.rt = Some(self.nodes[id].keys[idx + 1].clone());
}
if let Some((md, child)) = new_child {
// A full node splits first; the child goes to the half that
// holds child `idx`.
let (mut target, mut split) = (id, None);
if n == self.two_k {
let s = self.split(id, idx);
let nleft = self.nodes[id].children.len();
if idx >= nleft {
idx -= nleft;
target = s;
}
split = Some(s);
}
// H5B__insert_child (H5B_INS_RIGHT): the new child after child
// `idx`, its left key after that child's.
let node = &mut self.nodes[target];
node.keys.insert(idx + 1, md);
node.children.insert(idx + 1, child);
ret.split = split.map(|s| (self.nodes[s].keys[0].clone(), s));
}
Ok(ret)
}
/// `H5B__split` of the full node `id`, the insertion going after child
/// `idx`; returns the new right node.
fn split(&mut self, id: usize, idx: usize) -> usize {
let node = &self.nodes[id];
let ratio = if node.right.is_none() {
SPLIT_RATIOS[2]
} else if node.left.is_none() {
SPLIT_RATIOS[0]
} else {
SPLIT_RATIOS[1]
};
let mut nleft = (self.two_k as f64 * ratio) as usize;
if idx < nleft && nleft == self.two_k {
nleft -= 1;
} else if idx >= nleft && nleft == 0 {
nleft += 1;
}
let new_id = self.nodes.len();
let right = Node {
level: node.level,
left: Some(id),
right: node.right,
keys: node.keys[nleft..].to_vec(),
children: node.children[nleft..].to_vec(),
};
let old_right = node.right;
self.nodes.push(right);
if let Some(r) = old_right {
self.nodes[r].left = Some(new_id);
}
let node = &mut self.nodes[id];
node.keys.truncate(nleft + 1);
node.children.truncate(nleft);
node.right = Some(new_id);
new_id
}
}
/// Bytes of one node of a chunk B-tree with `ndims` key dimensions (the
/// dataset's rank plus the element-size one).
fn node_size(two_k: usize, ndims: usize, offset_size: usize) -> usize {
let key = 8 + 8 * ndims;
8 + 2 * offset_size + (two_k + 1) * key + two_k * offset_size
}
/// Build the chunk B-tree for `chunks`, given in row-major order of their
/// scaled coordinates, with nodes laid out from `base_address`. `chunk_dims`
/// are the chunk's dimensions (the dataset's rank of them) and `elem_size`
/// the element size, the key's last dimension. Returns the nodes' bytes; the
/// root is at `base_address`. `chunks` must not be empty: an index without
/// chunks has no tree (its address is undefined).
pub(crate) fn build_chunk_btree_v1_at(
chunks: &[ChunkEntry],
chunk_dims: &[u64],
elem_size: u32,
base_address: u64,
offset_size: u8,
) -> Result<Vec<u8>, FormatError> {
if chunks.is_empty() {
return Err(bad("no chunks"));
}
let rank = chunk_dims.len();
let mut tree = Tree::new(CHUNK_BTREE_K);
for c in chunks {
if c.scaled.len() != rank {
return Err(bad("chunk rank differs from the dataset's"));
}
let nbytes = u32::try_from(c.nbytes).map_err(|_| {
FormatError::SerializationError(format!(
"a chunk of {} bytes cannot be indexed by a version-1 B-tree \
(HDF5 1.8 chunks are under 4 GiB)",
c.nbytes
))
})?;
let mut scaled = c.scaled.clone();
scaled.push(0);
let key = Key {
nbytes,
mask: c.filter_mask,
scaled,
};
tree.insert(&key, c.address)?;
}
let os = usize::from(offset_size);
let ndims = rank + 1;
let nsize = node_size(tree.two_k, ndims, os);
let addr_of = |id: usize| base_address + (id * nsize) as u64;
let mut dims: Vec<u64> = chunk_dims.to_vec();
dims.push(u64::from(elem_size));
let mut out = vec![0u8; tree.nodes.len() * nsize];
for (i, node) in tree.nodes.iter().enumerate() {
let d = &mut out[i * nsize..(i + 1) * nsize];
d[0..4].copy_from_slice(b"TREE");
d[4] = 1; // node type: raw data chunks
d[5] = node.level;
let n = u16::try_from(node.children.len()).map_err(|_| bad("node too large"))?;
d[6..8].copy_from_slice(&n.to_le_bytes());
let undef = u64::MAX;
put_addr(&mut d[8..], node.left.map_or(undef, addr_of), os);
put_addr(&mut d[8 + os..], node.right.map_or(undef, addr_of), os);
let mut p = 8 + 2 * os;
for (k, key) in node.keys.iter().enumerate() {
d[p..p + 4].copy_from_slice(&key.nbytes.to_le_bytes());
d[p + 4..p + 8].copy_from_slice(&key.mask.to_le_bytes());
for (j, (&s, &dim)) in key.scaled.iter().zip(&dims).enumerate() {
let off = s
.checked_mul(dim)
.ok_or_else(|| FormatError::Overflow("chunk key offset".into()))?;
d[p + 8 + 8 * j..p + 16 + 8 * j].copy_from_slice(&off.to_le_bytes());
}
p += 8 + 8 * ndims;
if let Some(&child) = node.children.get(k) {
let a = if node.level == 0 {
child
} else {
addr_of(usize::try_from(child).map_err(|_| bad("bad node index"))?)
};
put_addr(&mut d[p..], a, os);
p += os;
}
}
}
Ok(out)
}
fn put_addr(d: &mut [u8], v: u64, os: usize) {
d[..os].copy_from_slice(&v.to_le_bytes()[..os]);
}
#[cfg(test)]
mod tests {
use super::*;
fn build(n: u64) -> Tree {
let mut t = Tree::new(CHUNK_BTREE_K);
for i in 0..n {
let key = Key {
nbytes: 80,
mask: 0,
scaled: vec![i, 0],
};
t.insert(&key, 1000 + i).unwrap();
}
t
}
/// Leaves in order from the root, with their child counts.
fn leaves(t: &Tree, id: usize, out: &mut Vec<usize>) {
let n = &t.nodes[id];
if n.level == 0 {
out.push(n.children.len());
} else {
for &c in &n.children {
leaves(t, c as usize, out);
}
}
}
#[test]
fn sequential_appends_split_as_libhdf5_does() {
// libhdf5 2.0 (h5py, libver=('v108', 'latest')) writes 1000 chunks
// as a root over 17 leaves of 57 chunks and one of 31, with the
// root's right key on the last chunk (9990, 8 for 10-element f8
// chunks).
let t = build(1000);
assert_eq!(t.nodes[0].level, 1);
let mut l = Vec::new();
leaves(&t, 0, &mut l);
let mut want = vec![57; 17];
want.push(31);
assert_eq!(l, want);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![999, 1]);
// 100 000 chunks: three levels, a root of 31 children.
let t = build(100_000);
assert_eq!(t.nodes[0].level, 2);
assert_eq!(t.nodes[0].children.len(), 31);
}
#[test]
fn right_key_moves_every_other_append() {
let t = build(5);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![5, 1]);
let t = build(6);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![5, 1]);
}
#[test]
fn keys_and_siblings_are_consistent() {
let t = build(5000);
for (i, n) in t.nodes.iter().enumerate() {
assert!(n.children.len() <= t.two_k);
assert_eq!(n.keys.len(), n.children.len() + 1);
if let Some(r) = n.right {
assert_eq!(t.nodes[r].left, Some(i));
assert_eq!(n.keys.last(), t.nodes[r].keys.first());
}
}
}
}
+22 -13
View File
@@ -194,10 +194,17 @@ const MAX_DEPTH: u16 = 64;
/// Take `n` records from the traversal's budget, or refuse the tree.
fn spend(budget: &mut usize, n: usize) -> Result<(), FormatError> {
*budget = budget
.checked_sub(n)
.ok_or(FormatError::NestingDepthExceeded)?;
Ok(())
match budget.checked_sub(n) {
Some(left) => {
*budget = left;
Ok(())
}
None => {
// Spent: a walk that goes on after a failure stops here.
*budget = 0;
Err(FormatError::NestingDepthExceeded)
}
}
}
/// Collect all records from a B-tree v2 by traversing from the root.
@@ -504,16 +511,18 @@ fn collect_internal_records<S: Storage + ?Sized>(
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
// We collect child[0] records, then record[0], then child[1], etc.
// After the first child that fails, the others are only touched (see
// `storage::touch`); that error is returned.
// A child that fails does not stop the walk: the others are still
// descended into (their records are dropped with the result), then the
// first error is returned, as when stopping there. A storage that
// records what it lacks (see `storage::touch`) so learns every node the
// walk can reach in one attempt. The record budget is spent as before,
// so the walk is no longer than a successful one.
let mut failed = None;
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
if failed.is_some() {
let len = usize::try_from(node_size)
.unwrap_or(usize::MAX)
.min(1 << 16);
crate::storage::touch(file, child_addr, len);
continue;
if failed.is_some() && *budget == 0 {
// The record budget is spent: the tree is refused, and a walk
// over what is left could be as long as the one it bounds.
break;
}
if let Err(e) = (|| -> Result<(), FormatError> {
if child_depth == 0 {
@@ -553,7 +562,7 @@ fn collect_internal_records<S: Storage + ?Sized>(
}
Ok(())
})() {
failed = Some(e);
failed.get_or_insert(e);
}
}
+1 -1
View File
@@ -899,7 +899,7 @@ mod tests {
fn make_chunk(offsets: Vec<u64>, address: u64, size: u32) -> ChunkInfo {
ChunkInfo {
chunk_size: size,
chunk_size: u64::from(size),
filter_mask: 0,
offsets,
address,
+2 -2
View File
@@ -109,7 +109,7 @@ pub struct ChunkMapping {
/// File byte address of the compressed chunk.
pub file_offset: u64,
/// Size of the compressed chunk in the file.
pub file_size: u32,
pub file_size: u64,
/// Filter mask (0 = all filters applied).
pub filter_mask: u32,
/// Pre-computed row-copy operations for assembling this chunk into output.
@@ -338,7 +338,7 @@ mod tests {
fn make_chunk(offsets: Vec<u64>, address: u64, size: u32) -> ChunkInfo {
ChunkInfo {
chunk_size: size,
chunk_size: u64::from(size),
filter_mask: 0,
offsets,
address,
+26 -23
View File
@@ -314,7 +314,7 @@ pub(crate) fn chunk_req(
chunk_bytes: usize,
wanted: bool,
) -> ExtentReq {
let len = c.chunk_size as usize;
let len = crate::addr::saturating_usize(c.chunk_size);
ExtentReq {
addr: c.address,
len,
@@ -521,7 +521,7 @@ pub fn decompress_all_chunks_with_stats_in<S: Storage + ?Sized>(
#[derive(Debug, Clone)]
pub struct ChunkInfo {
/// Size of chunk data in the file (after compression).
pub chunk_size: u32,
pub chunk_size: u64,
/// Bitmask of filters that were NOT applied (0 = all applied).
pub filter_mask: u32,
/// N-dimensional offset of this chunk in dataset space.
@@ -610,7 +610,7 @@ fn stored_element_size(dt: &Datatype, offset_size: u8) -> u64 {
/// (`H5D__chunk_set_sizes`: "stored datatype size in chunk layout does not
/// match datatype description"). Reading it anyway laid the chunks out with
/// the wrong element size.
pub(crate) fn check_chunk_element_size(
pub fn check_chunk_element_size(
layout: &DataLayout,
datatype: &Datatype,
offset_size: u8,
@@ -1028,7 +1028,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
if node_level == 0 {
chunks.push(stored.len());
stored.push(ChunkInfo {
chunk_size,
chunk_size: u64::from(chunk_size),
filter_mask,
offsets: keys[k..].to_vec(),
address,
@@ -1131,7 +1131,7 @@ pub fn generate_implicit_chunks_in_grid(
}
chunks.push(ChunkInfo {
chunk_size: chunk_byte_size as u32,
chunk_size: chunk_byte_size,
filter_mask: 0,
offsets,
address: base_address.saturating_add(grid_idx.saturating_mul(chunk_byte_size)),
@@ -1190,9 +1190,15 @@ fn read_btree_v2_chunks<S: Storage + ?Sized>(
}
_ => return Err(bad("tree is not a chunk index")),
};
let unfiltered_bytes = checked_chunk_byte_len(chunk_dims, elem_size)?;
let unfiltered_bytes =
u32::try_from(unfiltered_bytes).map_err(|_| bad("chunk larger than 4 GiB"))?;
// A u64 whatever the platform: an unfiltered chunk's size is only
// recorded here; a chunk this platform cannot address fails when read.
let unfiltered_bytes = u64::try_from(
chunk_dims
.iter()
.try_fold(elem_size as u128, |acc, &c| acc.checked_mul(c as u128))
.ok_or_else(|| bad("chunk size overflows"))?,
)
.map_err(|_| bad("chunk size overflows"))?;
let records = collect_btree_v2_records_in(file_data, &header, offset_size, length_size)?;
let mut chunks = Vec::with_capacity(records.len());
@@ -1213,10 +1219,7 @@ fn read_btree_v2_chunks<S: Storage + ?Sized>(
pos += size_len;
let mask = u32::from_le_bytes([data[pos], data[pos + 1], data[pos + 2], data[pos + 3]]);
pos += 4;
(
u32::try_from(size).map_err(|_| bad("stored chunk larger than 4 GiB"))?,
mask,
)
(size, mask)
};
let mut offsets = Vec::with_capacity(rank);
for &dim in chunk_dims {
@@ -1334,9 +1337,9 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
// Single chunk — one chunk covering the entire dataset
let chunk_byte_size = checked_chunk_byte_len(&chunk_dims, elem_size)?;
let (csize, fmask) = if let Some(fs) = single_filtered_size {
(fs as u32, single_filter_mask.unwrap_or(0))
(fs, single_filter_mask.unwrap_or(0))
} else {
(chunk_byte_size as u32, 0)
(chunk_byte_size as u64, 0)
};
vec![ChunkInfo {
chunk_size: csize,
@@ -1488,7 +1491,7 @@ pub fn list_chunks_for_read_in<S: Storage + ?Sized>(
let chunk_bytes = checked_chunk_byte_len(&chunk_dims, elem_size)?;
if let Some(c) = chunks
.iter()
.find(|c| c.address != u64::MAX && c.chunk_size as usize != chunk_bytes)
.find(|c| c.address != u64::MAX && c.chunk_size != chunk_bytes as u64)
{
return Err(FormatError::ChunkedReadError(format!(
"incorrect chunk size returned from index for unfiltered chunk at {:?}: \
@@ -2124,7 +2127,7 @@ pub fn read_chunked_data_indexed_in<S: Storage + ?Sized>(
.iter()
.zip(&hits)
.map(|(m, hit)| {
let len = m.file_size as usize;
let len = crate::addr::saturating_usize(m.file_size);
ExtentReq {
addr: m.file_offset,
len,
@@ -2467,7 +2470,7 @@ mod tests {
// Entries: key[i], child[i] pairs, then final key
for chunk in chunks {
// Key: chunk_size(4) + filter_mask(4) + ndims offsets
buf.extend_from_slice(&chunk.chunk_size.to_le_bytes());
buf.extend_from_slice(&(chunk.chunk_size as u32).to_le_bytes());
buf.extend_from_slice(&chunk.filter_mask.to_le_bytes());
for d in 0..ndims {
let off = if d < chunk.offsets.len() {
@@ -2757,7 +2760,7 @@ mod tests {
}
chunk_infos.push(ChunkInfo {
chunk_size: chunk_bytes as u32,
chunk_size: chunk_bytes as u64,
filter_mask: 0,
offsets: vec![start as u64, 0],
address: data_offset as u64,
@@ -2869,7 +2872,7 @@ mod tests {
.collect();
let stored = crate::filters::compress_chunk(&chunk, &pipeline, 4).unwrap();
chunks.push(ChunkInfo {
chunk_size: stored.len() as u32,
chunk_size: stored.len() as u64,
filter_mask: 0,
offsets: vec![r0 as u64, c0 as u64, 0],
address: file.len() as u64,
@@ -2943,7 +2946,7 @@ mod tests {
let short = crate::filters::compress_chunk(&[1u8; 64], &pipeline, 4).unwrap();
for bad in [5usize, 11, 40] {
chunks[bad].address = file.len() as u64;
chunks[bad].chunk_size = short.len() as u32;
chunks[bad].chunk_size = short.len() as u64;
file.extend_from_slice(&short);
}
for _ in 0..20 {
@@ -3133,7 +3136,7 @@ mod tests {
file_data[data_offset..data_offset + compressed.len()].copy_from_slice(&compressed);
chunk_infos.push(ChunkInfo {
chunk_size: compressed.len() as u32,
chunk_size: compressed.len() as u64,
filter_mask: 0,
offsets: vec![start as u64, 0],
address: data_offset as u64,
@@ -3215,7 +3218,7 @@ mod tests {
file_data[data_offset..data_offset + chunk_size].copy_from_slice(&chunk_bytes);
chunk_infos.push(ChunkInfo {
chunk_size: chunk_size as u32,
chunk_size: chunk_size as u64,
filter_mask: 0,
offsets: vec![row_start as u64, col_start as u64, 0],
address: data_offset as u64,
@@ -3340,7 +3343,7 @@ mod tests {
assert_eq!(c.address, 0x1000 + i as u64 * chunk_byte_size as u64);
assert_eq!(c.offsets, vec![i as u64 * 20]);
assert_eq!(c.filter_mask, 0);
assert_eq!(c.chunk_size, chunk_byte_size as u32);
assert_eq!(c.chunk_size, chunk_byte_size as u64);
}
}
+611 -129
View File
@@ -7,6 +7,7 @@ use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use crate::btree_v1_write;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::checksum::jenkins_lookup3;
use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line};
@@ -19,6 +20,7 @@ use crate::filter_pipeline::{
FilterPipeline,
};
use crate::filters::compress_chunk_masked;
use crate::libver::LibVer;
/// Round a file offset up to the next cache-line boundary.
///
/// This ensures chunk data starts at an address that is a multiple of the
@@ -400,86 +402,83 @@ pub fn split_into_chunks(
chunk_dims: &[u64],
element_size: usize,
) -> Vec<(Vec<u64>, Vec<u8>)> {
let rank = shape.len();
if rank == 0 {
if shape.is_empty() {
return vec![(vec![], raw_data.to_vec())];
}
(0..chunk_count(shape, chunk_dims))
.map(|i| extract_chunk(raw_data, shape, chunk_dims, element_size, i))
.collect()
}
// Compute number of chunks per dimension
let mut num_chunks_per_dim = Vec::with_capacity(rank);
for d in 0..rank {
num_chunks_per_dim.push(shape[d].div_ceil(chunk_dims[d]));
/// Number of chunks of the current extent `shape`.
fn chunk_count(shape: &[u64], chunk_dims: &[u64]) -> u64 {
shape
.iter()
.zip(chunk_dims)
.map(|(&s, &c)| s.div_ceil(c))
.product()
}
/// The `linear_idx`-th chunk (row-major over the chunks of the current
/// extent) of the row-major dataset `raw_data`: its offset in dataset space
/// and its bytes, a whole chunk with the part past the dataset's edge zero.
/// Copied one row (a run along the last dimension) at a time.
fn extract_chunk(
raw_data: &[u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
linear_idx: u64,
) -> (Vec<u64>, Vec<u8>) {
let rank = shape.len();
let mut offsets = vec![0u64; rank];
let mut remaining = linear_idx;
for d in (0..rank).rev() {
let n = shape[d].div_ceil(chunk_dims[d]);
offsets[d] = (remaining % n) * chunk_dims[d];
remaining /= n;
}
let total_chunks: u64 = num_chunks_per_dim.iter().product();
let chunk_elements: usize = chunk_dims.iter().map(|&d| saturating_usize(d)).product();
let mut chunk = vec![0u8; chunk_elements.saturating_mul(element_size)];
// Dataset strides (row-major)
// Elements of the chunk inside the dataset, per dimension.
let valid: Vec<usize> = (0..rank)
.map(|d| saturating_usize(shape[d].saturating_sub(offsets[d]).min(chunk_dims[d])))
.collect();
if valid.contains(&0) {
return (offsets, chunk);
}
let mut ds_strides = vec![1usize; rank];
for i in (0..rank.saturating_sub(1)).rev() {
ds_strides[i] = ds_strides[i + 1] * saturating_usize(shape[i + 1]);
}
// Chunk strides
let mut chunk_strides = vec![1usize; rank];
for i in (0..rank.saturating_sub(1)).rev() {
chunk_strides[i] = chunk_strides[i + 1] * saturating_usize(chunk_dims[i + 1]);
for d in (0..rank - 1).rev() {
ds_strides[d] = ds_strides[d + 1] * saturating_usize(shape[d + 1]);
chunk_strides[d] = chunk_strides[d + 1] * saturating_usize(chunk_dims[d + 1]);
}
let chunk_total_elements: usize = chunk_dims.iter().map(|&d| saturating_usize(d)).product();
let mut result = Vec::with_capacity(saturating_usize(total_chunks));
for linear_idx in 0..total_chunks {
// Convert linear index to chunk grid coordinates
let mut chunk_grid_coords = vec![0u64; rank];
let mut remaining = linear_idx;
for d in (0..rank).rev() {
chunk_grid_coords[d] = remaining % num_chunks_per_dim[d];
remaining /= num_chunks_per_dim[d];
let row = valid[rank - 1] * element_size;
let mut idx = vec![0usize; rank];
loop {
let src: usize = (0..rank)
.map(|d| (saturating_usize(offsets[d]) + idx[d]) * ds_strides[d])
.sum::<usize>()
* element_size;
let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::<usize>() * element_size;
// Whole elements only, as far as `raw_data` reaches.
let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size;
chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]);
// Next row: advance every dimension but the last.
let mut d = rank - 1;
loop {
if d == 0 {
return (offsets, chunk);
}
d -= 1;
idx[d] += 1;
if idx[d] < valid[d] {
break;
}
idx[d] = 0;
}
// Chunk offset in dataset space
let offsets: Vec<u64> = (0..rank)
.map(|d| chunk_grid_coords[d] * chunk_dims[d])
.collect();
// Extract chunk data
let mut chunk_bytes = vec![0u8; chunk_total_elements * element_size];
for flat_idx in 0..chunk_total_elements {
let mut remaining_idx = flat_idx;
let mut ds_flat = 0usize;
let mut out_of_bounds = false;
for d in 0..rank {
let coord_in_chunk = remaining_idx / chunk_strides[d];
remaining_idx %= chunk_strides[d];
let global_coord = saturating_usize(offsets[d]) + coord_in_chunk;
if global_coord >= saturating_usize(shape[d]) {
out_of_bounds = true;
break;
}
ds_flat += global_coord * ds_strides[d];
}
if out_of_bounds {
// Zero-filled (already initialized)
continue;
}
let src_start = ds_flat * element_size;
let dst_start = flat_idx * element_size;
if src_start + element_size <= raw_data.len() {
chunk_bytes[dst_start..dst_start + element_size]
.copy_from_slice(&raw_data[src_start..src_start + element_size]);
}
}
result.push((offsets, chunk_bytes));
}
result
}
/// Parallel compression threshold: use rayon when chunk count exceeds this.
@@ -490,46 +489,68 @@ pub fn split_into_chunks(
#[cfg(feature = "parallel")]
const PARALLEL_COMPRESS_THRESHOLD: usize = 2;
/// Compress all chunks, using parallel compression when beneficial, and
/// return each chunk's stored bytes with its filter mask.
/// Largest chunk compressed in parallel: every thread holds a chunk and its
/// compressed copy at once, so chunks larger than this (up to 4 GiB and
/// more) are compressed one after another.
#[cfg(feature = "parallel")]
const PARALLEL_COMPRESS_MAX_CHUNK_BYTES: u64 = 64 << 20;
/// Extract and compress every chunk of the dataset, returning each chunk's
/// raw size, stored bytes and filter mask, in chunk order.
///
/// Chunks run through the pipeline as libhdf5 runs them
/// ([`compress_chunk_masked`]): an optional filter that fails — LZF or Blosc
/// output no smaller than its input — is skipped and its mask bit set.
///
/// With the `parallel` feature and more than [`PARALLEL_COMPRESS_THRESHOLD`]
/// filtered chunks, compression runs across rayon threads; otherwise it is
/// sequential. Output order matches input order, so per-chunk bytes are
/// identical to the sequential path.
/// Each chunk is extracted just before it is compressed and dropped after,
/// so at most one raw chunk per thread is held. With the `parallel` feature,
/// more than [`PARALLEL_COMPRESS_THRESHOLD`] filtered chunks, and chunks of
/// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across
/// rayon threads; otherwise it is sequential. Output order matches chunk
/// order, so per-chunk bytes are identical to the sequential path.
fn compress_all_chunks(
chunks: &[(Vec<u64>, Vec<u8>)],
raw_data: &[u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
chunk_bytes: u64,
pipeline: &Option<FilterPipeline>,
element_size: u32,
) -> Result<Vec<(Vec<u8>, u32)>, FormatError> {
) -> Result<Vec<(u64, Vec<u8>, u32)>, FormatError> {
let one = |i: u64| -> Result<(u64, Vec<u8>, u32), FormatError> {
let (_, raw) = if shape.is_empty() {
(Vec::new(), raw_data.to_vec())
} else {
extract_chunk(raw_data, shape, chunk_dims, element_size, i)
};
let raw_size = raw.len() as u64;
match pipeline {
Some(pl) => {
let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?;
Ok((raw_size, stored, mask))
}
None => Ok((raw_size, raw, 0)),
}
};
let n = if shape.is_empty() {
1
} else {
chunk_count(shape, chunk_dims)
};
#[cfg(feature = "parallel")]
{
if let Some(pl) = pipeline
&& chunks.len() > PARALLEL_COMPRESS_THRESHOLD
if pipeline.is_some()
&& n > PARALLEL_COMPRESS_THRESHOLD as u64
&& chunk_bytes <= PARALLEL_COMPRESS_MAX_CHUNK_BYTES
{
use rayon::prelude::*;
return chunks
.par_iter()
.map(|(_offsets, chunk_bytes)| compress_chunk_masked(chunk_bytes, pl, element_size))
.collect();
return (0..n).into_par_iter().map(one).collect();
}
}
#[cfg(not(feature = "parallel"))]
let _ = chunk_bytes;
// Sequential fallback
chunks
.iter()
.map(|(_offsets, chunk_bytes)| {
if let Some(pl) = pipeline {
compress_chunk_masked(chunk_bytes, pl, element_size)
} else {
Ok((chunk_bytes.clone(), 0))
}
})
.collect()
(0..n).map(one).collect()
}
/// Build the complete chunked dataset blob (chunk data + index) and return
@@ -550,10 +571,12 @@ pub fn serialize_v4_single_chunk_pub(
filter_mask,
offset_size,
element_size,
4,
)
}
/// Serialize a v4 single chunk layout message.
/// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk
/// layout message.
fn serialize_v4_single_chunk(
chunk_dims: &[u32],
chunk_address: u64,
@@ -561,9 +584,10 @@ fn serialize_v4_single_chunk(
filter_mask: Option<u32>,
offset_size: u8,
element_size: u32,
version: u8,
) -> Vec<u8> {
let mut buf = Vec::new();
buf.push(4); // version
buf.push(version);
buf.push(2); // class = chunked
// flags: bit 0 = unknown meaning in some files, bit 1 = filters for single chunk
@@ -603,8 +627,9 @@ fn serialize_v4_fixed_array(
offset_size: u8,
element_size: u32,
max_bits: u8,
version: u8,
) -> Vec<u8> {
let mut buf = layout_v4_chunked_prefix(chunk_dims, element_size);
let mut buf = layout_v4_chunked_prefix(chunk_dims, element_size, version);
// chunk index type = 3 (Fixed Array)
buf.push(3);
@@ -643,9 +668,9 @@ pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_
}
}
fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32) -> Vec<u8> {
fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32, version: u8) -> Vec<u8> {
let mut buf = Vec::new();
buf.push(4); // version
buf.push(version);
buf.push(2); // class = chunked
let flags: u8 = 0x00;
@@ -671,21 +696,53 @@ pub(crate) fn push_addr(buf: &mut Vec<u8>, addr: u64, offset_size: u8) {
/// Width of the chunk-size field of a filtered chunk index element. Must
/// match the library's `H5D_FARRAY_FILT_COMPUTE_CHUNK_SIZE_LEN` (the EA and
/// B-tree v2 indexes use the same formula):
/// `1 + ((log2(unfiltered chunk bytes) + 8) / 8)`, capped at 8.
pub(crate) fn filtered_chunk_size_len(slots: &[Option<WrittenChunk>]) -> usize {
/// B-tree v2 indexes use the same formula): see [`chunk_size_len`]. Chunks
/// of more than `u32::MAX` bytes are written with layout version 5
/// ([`layout_version_for`]).
pub(crate) fn filtered_chunk_size_len(slots: &[Option<WrittenChunk>], length_size: u8) -> usize {
let max_raw = slots
.iter()
.flatten()
.map(|c| c.raw_size)
.max()
.unwrap_or(1);
let log2_val = if max_raw <= 1 {
chunk_size_len(max_raw, layout_version_for(max_raw), length_size)
}
/// Largest chunk, in bytes, a layout message of version 4 or lower may
/// describe: libhdf5 writes a larger one with version 5
/// (`H5D__chunk_construct`: "chunk size > 4GB requires H5F_LIBVER_V200"),
/// which libhdf5 before 2.0 cannot read.
pub const MAX_V4_CHUNK_BYTES: u64 = u32::MAX as u64;
/// The layout message version clawhdf5 writes for chunks of `chunk_bytes`
/// bytes: 4, or 5 for a chunk larger than [`MAX_V4_CHUNK_BYTES`] (what
/// libhdf5 2.x writes for it; the chunk index is chosen as for version 4).
pub fn layout_version_for(chunk_bytes: u64) -> u8 {
if chunk_bytes > MAX_V4_CHUNK_BYTES {
5
} else {
4
}
}
/// Width libhdf5 gives the stored-size field of a filtered chunk index
/// element (Fixed Array, Extensible Array, v2 B-tree) for chunks of
/// `chunk_bytes` bytes under layout message `layout_version`
/// (`H5D_FARRAY_FILT_COMPUTE_CHUNK_SIZE_LEN` and its EA and B-tree twins):
/// up to version 4, one byte more than the chunk's size needs,
/// `1 + ((log2(chunk_bytes) + 8) / 8)` capped at 8; from version 5 (HDF5
/// 2.0), the file's size of lengths (`length_size`), whatever the chunk.
pub fn chunk_size_len(chunk_bytes: u64, layout_version: u8, length_size: u8) -> usize {
if layout_version >= 5 {
return usize::from(length_size);
}
let log2 = if chunk_bytes <= 1 {
0
} else {
63 - max_raw.leading_zeros()
63 - chunk_bytes.leading_zeros()
};
(1 + ((log2_val + 8) / 8) as usize).min(8)
(1 + ((log2 + 8) / 8) as usize).min(8)
}
/// Append one chunk index element: the chunk's address, plus its stored size
@@ -732,7 +789,7 @@ pub fn build_fixed_array_at(
let os = offset_size as usize;
let num_elements = slots.len();
let chunk_size_bytes = has_filters.then(|| filtered_chunk_size_len(slots));
let chunk_size_bytes = has_filters.then(|| filtered_chunk_size_len(slots, length_size));
let elem_size = os + chunk_size_bytes.map_or(0, |n| n + 4);
let client_id: u8 = if has_filters { 1 } else { 0 };
@@ -827,23 +884,23 @@ pub fn precompress_chunks(
element_size: usize,
options: &ChunkOptions,
) -> Result<PrecompressedChunks, FormatError> {
let chunk_bytes = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
.and_then(|b| u32::try_from(b).ok())
.unwrap_or(0);
let pipeline = options.build_pipeline_for_chunk(element_size as u32, chunk_bytes);
let (_, chunk_bytes) = checked_chunk_dims(chunk_dims, element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
check_huge_chunk_filters(options, chunk_bytes)?;
}
let pipeline = options
.build_pipeline_for_chunk(element_size as u32, u32::try_from(chunk_bytes).unwrap_or(0));
let has_filters = pipeline.is_some();
let pipeline_message = pipeline.as_ref().map(|pl| pl.serialize());
let raw_chunks = split_into_chunks(raw_data, shape, chunk_dims, element_size);
let compressed = compress_all_chunks(&raw_chunks, &pipeline, element_size as u32)?;
let chunks = raw_chunks
.into_iter()
.zip(compressed)
.map(|((_offsets, raw_bytes), (c, mask))| (raw_bytes.len() as u64, c, mask))
.collect();
let chunks = compress_all_chunks(
raw_data,
shape,
chunk_dims,
element_size,
chunk_bytes,
&pipeline,
)?;
Ok(PrecompressedChunks {
chunks,
@@ -855,6 +912,67 @@ pub fn precompress_chunks(
})
}
/// The chunk dimensions as the layout message stores them (each below
/// 2^32), and one chunk's size in bytes. A chunk dimension of 2^32 or more
/// (which HDF5 2.0 can store, in wider fields) is refused, as it is when
/// read: clawhdf5 holds chunk dimensions as `u32`. So is a chunk whose size
/// overflows 64 bits, or that this platform cannot hold in memory (a chunk
/// of 4 GiB or more on a 32-bit target).
fn checked_chunk_dims(
chunk_dims: &[u64],
element_size: usize,
) -> Result<(Vec<u32>, u64), FormatError> {
let dims = chunk_dims
.iter()
.map(|&d| {
u32::try_from(d).map_err(|_| {
FormatError::InvalidChunkDimensions(format!(
"chunk dimension {d} is 2^32 or more, which clawhdf5 does not support"
))
})
})
.collect::<Result<Vec<u32>, _>>()?;
let bytes = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
.filter(|&b| usize::try_from(b).is_ok())
.ok_or_else(|| {
FormatError::Overflow(format!(
"a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space"
))
})?;
Ok((dims, bytes))
}
/// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes:
/// shuffle, deflate, Zstandard, LZ4 (whose HDF5 framing records the size in
/// 64 bits and splits the chunk into blocks) and Fletcher32. The others
/// record the chunk size or their block lengths in 32 bits, or cannot take
/// a buffer that large (h5py's LZF, bitshuffle, bzip2, Blosc), and pcodec is
/// clawhdf5's own; they are refused rather than written into a chunk
/// libhdf5 could not decode.
fn check_huge_chunk_filters(options: &ChunkOptions, chunk_bytes: u64) -> Result<(), FormatError> {
let refused = if let Some(plugin) = &options.plugin {
Some(match plugin {
PluginFilter::Lzf => "LZF",
PluginFilter::Bitshuffle { .. } => "bitshuffle",
PluginFilter::Bzip2 { .. } => "bzip2",
PluginFilter::Blosc { .. } => "Blosc",
})
} else if options.pcodec {
Some("pcodec")
} else {
None
};
match refused {
Some(name) => Err(FormatError::FilterError(format!(
"{name} cannot compress a chunk of {chunk_bytes} bytes (4 GiB or more); \
use smaller chunks, or deflate, Zstandard or LZ4"
))),
None => Ok(()),
}
}
/// Lay out precompressed chunks at `base_address` and build index structures.
///
/// This is the address-dependent half of chunk writing. Call it in Pass 1
@@ -866,6 +984,44 @@ pub fn build_chunked_data_from_precompressed(
base_address: u64,
maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> {
build_chunked_data_from_precompressed_libver(
pre,
base_address,
maxshape,
LibVer::Latest,
LibVer::Latest,
)
}
/// [`build_chunked_data_from_precompressed`] for a file whose low library
/// version bound is `low`: below [`LibVer::V110`] (that is, for HDF5 1.8)
/// every chunked dataset gets a version-3 layout message and a version-1
/// B-tree chunk index, whatever its maximum shape, as libhdf5 writes it;
/// otherwise the version-4 layout and the index libhdf5 picks for it.
///
/// A chunk of 4 GiB or more (over [`MAX_V4_CHUNK_BYTES`]) takes layout
/// message version 5 whatever `low` is, as in libhdf5 (a version-1 B-tree
/// key holds a 32-bit size), and needs a `high` bound of at least
/// [`LibVer::V200`] ([`FormatError::LibverBound`] otherwise).
pub fn build_chunked_data_from_precompressed_libver(
pre: &PrecompressedChunks,
base_address: u64,
maxshape: Option<&[u64]>,
low: LibVer,
high: LibVer,
) -> Result<ChunkedDataResult, FormatError> {
let (_, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
if high < LibVer::V200 {
return Err(FormatError::LibverBound {
what: format!("a chunk of {chunk_bytes} bytes (4 GiB or more)"),
needs: LibVer::V200,
high,
});
}
} else if low < LibVer::V110 {
return build_btree_v1_chunked_data(pre, base_address, maxshape);
}
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
let offset_size: u8 = 8;
let length_size: u8 = 8;
@@ -891,7 +1047,8 @@ pub fn build_chunked_data_from_precompressed(
});
}
let chunk_dims_u32: Vec<u32> = pre.chunk_dims.iter().map(|&d| d as u32).collect();
let (chunk_dims_u32, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, element_size)?;
let version = layout_version_for(chunk_bytes);
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
@@ -915,6 +1072,7 @@ pub fn build_chunked_data_from_precompressed(
ea_address,
offset_size,
element_size as u32,
version,
)
}
ChunkIndexPlan::SingleChunk => {
@@ -932,6 +1090,7 @@ pub fn build_chunked_data_from_precompressed(
filter_mask,
offset_size,
element_size as u32,
version,
)
}
ChunkIndexPlan::FixedArray(grid, nslots) => {
@@ -957,6 +1116,7 @@ pub fn build_chunked_data_from_precompressed(
offset_size,
element_size as u32,
FA_PAGE_BITS,
version,
)
}
ChunkIndexPlan::BTreeV2 => {
@@ -981,6 +1141,7 @@ pub fn build_chunked_data_from_precompressed(
offset_size,
element_size as u32,
node_size,
version,
)
}
};
@@ -992,6 +1153,92 @@ pub fn build_chunked_data_from_precompressed(
})
}
/// Lay out precompressed chunks at `base_address` followed by a version-1
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
/// for a chunked dataset under a low bound of 1.8.
fn build_btree_v1_chunked_data(
pre: &PrecompressedChunks,
base_address: u64,
maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> {
if let Some(ms) = maxshape {
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
if ms.len() != pre.shape.len() {
return Err(bad("rank differs from the shape"));
}
if ms.iter().zip(&pre.shape).any(|(&m, &s)| m < s) {
return Err(bad("smaller than the shape"));
}
}
let offset_size: u8 = 8;
let mut data_buf = Vec::new();
let mut entries = Vec::with_capacity(pre.chunks.len());
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
entries.push(btree_v1_write::ChunkEntry {
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
nbytes: stored.len() as u64,
filter_mask: *filter_mask,
address: base_address + data_buf.len() as u64,
});
data_buf.extend_from_slice(stored);
}
let element_size = u32::try_from(pre.element_size)
.map_err(|_| FormatError::Overflow("element size".into()))?;
// A dataset with no chunks has no tree: its address is undefined, as
// libhdf5 leaves it until the first chunk is written.
let btree_address = if entries.is_empty() {
u64::MAX
} else {
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
let addr = base_address + data_buf.len() as u64;
let tree = btree_v1_write::build_chunk_btree_v1_at(
&entries,
&pre.chunk_dims,
element_size,
addr,
offset_size,
)?;
data_buf.extend_from_slice(&tree);
addr
};
let layout_message =
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
Ok(ChunkedDataResult {
data_bytes: data_buf,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
}
/// A version-3 layout message for a chunked dataset: dimensionality (the
/// rank plus one), the B-tree's address, then each chunk dimension and the
/// element size, four bytes each.
fn serialize_v3_chunked(
chunk_dims: &[u64],
btree_address: u64,
offset_size: u8,
element_size: u32,
) -> Result<Vec<u8>, FormatError> {
let ndims = u8::try_from(chunk_dims.len() + 1)
.map_err(|_| FormatError::Overflow("chunked layout rank".into()))?;
let mut buf = vec![3u8, 2, ndims];
push_addr(&mut buf, btree_address, offset_size);
for &d in chunk_dims {
let d =
u32::try_from(d).map_err(|_| FormatError::Overflow(format!("chunk dimension {d}")))?;
buf.extend_from_slice(&d.to_le_bytes());
}
buf.extend_from_slice(&element_size.to_le_bytes());
Ok(buf)
}
/// Most slots a Fixed Array index may have before we refuse to build it: its
/// data block holds one element per chunk of the *maximum* extent, so a huge
/// finite maxshape with small chunks would otherwise exhaust memory.
@@ -1132,7 +1379,7 @@ fn build_btree_v2_chunk_index_at(
let chunk_size_bytes = has_filters.then(|| {
let slots: Vec<Option<WrittenChunk>> =
records.iter().map(|(_, c)| Some((*c).clone())).collect();
filtered_chunk_size_len(&slots)
filtered_chunk_size_len(&slots, length_size)
});
let record_size = os + chunk_size_bytes.map_or(0, |n| n + 4) + 8 * rank;
let record_size_u16 = u16::try_from(record_size)
@@ -1182,8 +1429,9 @@ fn serialize_v4_btree_v2(
offset_size: u8,
element_size: u32,
node_size: u32,
version: u8,
) -> Vec<u8> {
let mut buf = layout_v4_chunked_prefix(chunk_dims, element_size);
let mut buf = layout_v4_chunked_prefix(chunk_dims, element_size, version);
buf.push(5); // chunk index type = 5 (version-2 B-tree)
buf.extend_from_slice(&node_size.to_le_bytes());
buf.push(BT2_SPLIT_PERCENT);
@@ -1763,7 +2011,7 @@ mod tests {
#[test]
fn serialize_v4_single_chunk_no_filters_roundtrip() {
let msg = serialize_v4_single_chunk(&[20], 0x1000, None, None, 8, 8);
let msg = serialize_v4_single_chunk(&[20], 0x1000, None, None, 8, 8, 4);
let layout = DataLayout::parse(&msg, 8, 8).unwrap();
match layout {
DataLayout::Chunked {
@@ -1788,7 +2036,7 @@ mod tests {
#[test]
fn serialize_v4_single_chunk_with_filters_roundtrip() {
let msg = serialize_v4_single_chunk(&[100], 0x2000, Some(500), Some(0), 8, 8);
let msg = serialize_v4_single_chunk(&[100], 0x2000, Some(500), Some(0), 8, 8, 4);
let layout = DataLayout::parse(&msg, 8, 8).unwrap();
match layout {
DataLayout::Chunked {
@@ -1807,7 +2055,7 @@ mod tests {
#[test]
fn serialize_v4_fixed_array_roundtrip() {
let msg = serialize_v4_fixed_array(&[20], 0x3000, 8, 8, 4);
let msg = serialize_v4_fixed_array(&[20], 0x3000, 8, 8, 4, 4);
let layout = DataLayout::parse(&msg, 8, 8).unwrap();
match layout {
DataLayout::Chunked {
@@ -1851,11 +2099,245 @@ mod tests {
assert_eq!(&fa[28..32], b"FADB");
}
// ---- Chunks of 4 GiB or more (layout message version 5) ----
/// 2^29 + 1 `f64`: 4 GiB + 8 bytes, the smallest `f64` chunk past
/// `u32::MAX`.
const HUGE_DIM: u64 = (1 << 29) + 1;
const HUGE_BYTES: u64 = HUGE_DIM * 8;
fn huge_chunk(address: u64, compressed_size: u64) -> WrittenChunk {
WrittenChunk {
address,
compressed_size,
raw_size: HUGE_BYTES,
filter_mask: 0,
}
}
#[test]
fn chunks_past_u32_max_take_layout_version_5() {
assert_eq!(layout_version_for(0), 4);
assert_eq!(layout_version_for(u64::from(u32::MAX)), 4);
assert_eq!(layout_version_for(u64::from(u32::MAX) + 1), 5);
assert_eq!(layout_version_for(HUGE_BYTES), 5);
}
/// A chunk of 4 GiB or more never goes into a version-1 B-tree (its key
/// holds a 32-bit size): under a 1.8 low bound it still takes layout
/// version 5, and a high bound below 2.0 refuses it, as in libhdf5.
#[test]
fn huge_chunks_ignore_the_v18_low_bound_and_need_v200() {
// One filtered chunk; the stored bytes stand in for its compression.
let pre = PrecompressedChunks {
chunks: vec![(HUGE_BYTES, vec![0u8; 16], 0)],
has_filters: true,
element_size: 8,
shape: vec![HUGE_DIM],
chunk_dims: vec![HUGE_DIM],
pipeline_message: None,
};
let r = build_chunked_data_from_precompressed_libver(
&pre,
4096,
None,
LibVer::V18,
LibVer::Latest,
)
.unwrap();
assert_eq!(r.layout_message[0], 5, "layout message version");
for high in [LibVer::V18, LibVer::V114] {
match build_chunked_data_from_precompressed_libver(&pre, 4096, None, LibVer::V18, high)
{
Err(FormatError::LibverBound { needs, .. }) => assert_eq!(needs, LibVer::V200),
other => panic!("high bound {high}: {:?}", other.map(|r| r.layout_message)),
}
}
}
/// `H5D_FARRAY_FILT_COMPUTE_CHUNK_SIZE_LEN` (and its EA and v2 B-tree
/// twins) in libhdf5 2.2.0: one byte more than the chunk size needs up to
/// layout version 4, the size of lengths from version 5.
#[test]
fn chunk_size_len_follows_libhdf5() {
assert_eq!(chunk_size_len(1, 4, 8), 2);
assert_eq!(chunk_size_len(160, 4, 8), 2);
assert_eq!(chunk_size_len(255, 4, 8), 2);
assert_eq!(chunk_size_len(256, 4, 8), 3);
assert_eq!(chunk_size_len(u64::from(u32::MAX), 4, 8), 5);
assert_eq!(chunk_size_len(1 << 32, 4, 8), 6);
assert_eq!(chunk_size_len(u64::MAX, 4, 8), 8);
assert_eq!(chunk_size_len(HUGE_BYTES, 5, 8), 8);
assert_eq!(chunk_size_len(160, 5, 8), 8);
assert_eq!(chunk_size_len(HUGE_BYTES, 5, 4), 4);
let slots = [Some(huge_chunk(0x1000, 20_000)), None];
assert_eq!(filtered_chunk_size_len(&slots, 8), 8);
let small = [Some(WrittenChunk {
raw_size: 160,
..huge_chunk(0x1000, 100)
})];
assert_eq!(filtered_chunk_size_len(&small, 8), 2);
}
/// A 4 GiB + 8 byte chunk: every layout message is version 5 with the
/// dimensions libhdf5 writes (4 bytes each: 0x20000001 and 8), and a
/// filtered index element stores the chunk's size in 8 bytes.
#[test]
fn huge_chunk_layout_messages_and_index_elements() {
let dims = [HUGE_DIM as u32];
let parsed = |msg: &[u8]| {
assert_eq!(msg[0], 5, "layout message version");
// Class chunked, then (after the flags) 2 dimensions of 4 bytes.
assert_eq!(&msg[1..2], &[2]);
assert_eq!(&msg[3..5], &[2, 4]);
assert_eq!(&msg[5..13], &[1, 0, 0, 0x20, 8, 0, 0, 0]);
match DataLayout::parse(msg, 8, 8).unwrap() {
DataLayout::Chunked {
chunk_dimensions,
chunk_index_type,
..
} => {
assert_eq!(chunk_dimensions, vec![HUGE_DIM as u32, 8]);
chunk_index_type.unwrap()
}
other => panic!("{other:?}"),
}
};
let single = serialize_v4_single_chunk(&dims, 0x800, Some(20_000), Some(0), 8, 8, 5);
assert_eq!(parsed(&single), 1);
// Filtered size (8 bytes), filter mask, address.
assert_eq!(single.len(), 13 + 1 + 8 + 4 + 8);
assert_eq!(
parsed(&serialize_v4_fixed_array(&dims, 0x800, 8, 8, 10, 5)),
3
);
let ea = ea_writer::serialize_v4_extensible_array(&dims, 0x800, 8, 8, 5);
assert_eq!(parsed(&ea), 4);
assert_eq!(
parsed(&serialize_v4_btree_v2(&dims, 0x800, 8, 8, 2048, 5)),
5
);
let slots = [Some(huge_chunk(0x1000, 20_000)), None];
// FAHD: element size (address 8 + size 8 + mask 4) at byte 6.
let fa = build_fixed_array_at(&slots, 8, 8, true, 0x2000);
assert_eq!(fa[6], 20);
// FADB element 0: address, then the stored size in 8 bytes.
let fadb = 28;
let prefix = 4 + 1 + 1 + 8;
assert_eq!(
&fa[fadb + prefix..fadb + prefix + 8],
&0x1000u64.to_le_bytes()
);
assert_eq!(
&fa[fadb + prefix + 8..fadb + prefix + 16],
&20_000u64.to_le_bytes()
);
// AEHD: element size at byte 6 as well.
let ea = ea_writer::build_extensible_array_at(&slots, 8, 8, true, 0x2000);
assert_eq!(&ea[..4], b"EAHD");
assert_eq!(ea[6], 20);
// BTHD: record size (element + 8-byte scaled offset) at bytes 10-11.
let chunk = huge_chunk(0x1000, 20_000);
let (bt, _) =
build_btree_v2_chunk_index_at(1, &[(vec![0], &chunk)], 8, 8, true, 0x2000).unwrap();
assert_eq!(&bt[..4], b"BTHD");
assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28);
}
#[test]
fn huge_chunks_refused_where_unsupported() {
// A chunk dimension of 2^32 or more.
assert!(matches!(
checked_chunk_dims(&[1 << 32], 1),
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
));
// A chunk size that overflows u64.
assert!(matches!(
checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8),
Err(FormatError::Overflow(_))
));
assert_eq!(
checked_chunk_dims(&[HUGE_DIM], 8).unwrap(),
(vec![HUGE_DIM as u32], HUGE_BYTES)
);
// Filters that cannot take a chunk that large.
for plugin in [
PluginFilter::Lzf,
PluginFilter::Bzip2 { level: 9 },
PluginFilter::Blosc {
codec: BloscCodec::Lz4,
level: 5,
shuffle: BloscShuffle::Byte,
},
PluginFilter::Bitshuffle {
block_size: 0,
compression: BitshuffleCompression::Lz4,
},
] {
let options = ChunkOptions {
plugin: Some(plugin),
..Default::default()
};
assert!(matches!(
check_huge_chunk_filters(&options, HUGE_BYTES),
Err(FormatError::FilterError(m)) if m.contains("4 GiB")
));
}
for options in [
ChunkOptions {
deflate_level: Some(6),
shuffle: true,
fletcher32: true,
..Default::default()
},
ChunkOptions {
zstd_level: Some(3),
..Default::default()
},
ChunkOptions {
lz4: true,
..Default::default()
},
] {
check_huge_chunk_filters(&options, HUGE_BYTES).unwrap();
}
}
/// Chunks are extracted row by row: the result is the element-by-element
/// split, edge padding included.
#[test]
fn extract_chunk_matches_elementwise_split() {
let shape = [5u64, 7, 3];
let chunks = [2u64, 3, 2];
let data: Vec<u8> = (0..5 * 7 * 3 * 2).map(|i| i as u8).collect();
let n = chunk_count(&shape, &chunks);
assert_eq!(n, 3 * 3 * 2);
for i in 0..n {
let (offsets, chunk) = extract_chunk(&data, &shape, &chunks, 2, i);
assert_eq!(chunk.len(), 2 * 3 * 2 * 2);
for (e, pair) in chunk.as_chunks::<2>().0.iter().enumerate() {
let c = [e / 6, (e / 2) % 3, e % 2];
let g: Vec<u64> = (0..3).map(|d| offsets[d] + c[d] as u64).collect();
let expect = if (0..3).all(|d| g[d] < shape[d]) {
let flat = ((g[0] * 7 + g[1]) * 3 + g[2]) as usize;
[data[2 * flat], data[2 * flat + 1]]
} else {
[0, 0]
};
assert_eq!(*pair, expect, "chunk {i} element {e}");
}
}
// Data shorter than the shape: the missing elements stay zero.
let (_, chunk) = extract_chunk(&data[..5], &[4], &[4], 2, 0);
assert_eq!(chunk, [0, 1, 2, 3, 0, 0, 0, 0]);
}
// ---- Extensible Array tests ----
#[test]
fn serialize_v4_extensible_array_roundtrip() {
let msg = ea_writer::serialize_v4_extensible_array(&[10], 0x4000, 8, 8);
let msg = ea_writer::serialize_v4_extensible_array(&[10], 0x4000, 8, 8, 4);
let layout = DataLayout::parse(&msg, 8, 8).unwrap();
match layout {
DataLayout::Chunked {
@@ -2020,7 +2502,7 @@ mod tests {
for info in &infos {
// Skipped chunks are stored at the chunk's size (shuffled).
assert_eq!(
info.chunk_size == (c * 8) as u32,
info.chunk_size == (c * 8) as u64,
info.filter_mask != 0,
"{info:?}"
);
+70 -5
View File
@@ -815,6 +815,7 @@ fn datatype_name(dt: &Datatype) -> &'static str {
Datatype::Enumeration { .. } => "Enumeration",
Datatype::VariableLength { .. } => "VariableLength",
Datatype::Array { .. } => "Array",
Datatype::Complex { .. } => "Complex",
}
}
@@ -1419,9 +1420,11 @@ pub fn read_as_f32(raw: &[u8], datatype: &Datatype) -> Result<Vec<f32>, FormatEr
result.push(match format {
FloatFormat::Single => read_f32_bytes(chunk, &order),
FloatFormat::Half => read_f16_bytes(chunk, &order),
// Double rounds; every other supported layout (bfloat16, FP8)
// is exact in f32.
_ => format.decode(chunk, &order) as f32,
// Double rounds.
FloatFormat::Double => format.decode(chunk, &order) as f32,
// Every other supported layout (bfloat16, FP8, FP6, FP4) is
// exact in f32.
FloatFormat::Other(_) => narrow_decoded(format.decode(chunk, &order)),
});
}
return Ok(result);
@@ -1564,6 +1567,9 @@ pub fn read_compound_fields(
datatype: &Datatype,
) -> Result<Vec<CompoundFieldData>, FormatError> {
match datatype {
Datatype::Complex { size, base_type } => {
read_compound_fields(raw, &Datatype::complex_as_compound(*size, base_type))
}
Datatype::Compound { size, members } => {
let elem_size = *size as usize;
if elem_size == 0 {
@@ -1947,7 +1953,12 @@ enum FloatFormat {
Double,
/// Any other IEEE-style layout (implied leading mantissa bit, all-ones
/// exponent for infinity/NaN) whose values are all exact in `f64`:
/// bfloat16, the FP8 formats, and similar.
/// bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1, and similar.
///
/// libhdf5 (checked against 2.2.0) decodes all of them this way, also the
/// OCP MX formats whose specification has no infinity (FP6, FP4) or a
/// single NaN (FP8 E4M3): an all-ones exponent is infinity or NaN, not a
/// finite value. clawhdf5 follows libhdf5 so both read a file alike.
Other(FloatLayout),
}
@@ -2046,7 +2057,9 @@ impl FloatLayout {
if mantissa == 0 {
f64::INFINITY
} else {
f64::NAN
// The NaN libhdf5 converts every NaN to: all mantissa bits
// set (the sign is applied below).
LIBHDF5_NAN
}
} else {
let bias = i64::from(self.exponent_bias);
@@ -2067,6 +2080,27 @@ impl FloatLayout {
}
}
/// The `f64` NaN libhdf5's conversion (`H5T__conv_f_f`) produces from a NaN
/// of a non-native float layout: sign clear, every mantissa bit set.
const LIBHDF5_NAN: f64 = f64::from_bits(0x7FFF_FFFF_FFFF_FFFF);
/// Narrow a value decoded from a non-native float layout to `f32`. Every such
/// value is exact in `f32`; a NaN becomes the NaN libhdf5 gives for
/// `H5T_NATIVE_FLOAT` (sign kept, every mantissa bit set) rather than
/// whatever payload an `as` cast leaves.
fn narrow_decoded(value: f64) -> f32 {
if value.is_nan() {
let sign = if value.is_sign_negative() {
1u32 << 31
} else {
0
};
f32::from_bits(sign | 0x7FFF_FFFF)
} else {
value as f32
}
}
/// `x * 2^power` without `std` (no `powi`/`libm`). `x` is a non-negative
/// integer below 2^53, so it is exact.
fn scale_by_pow2(x: f64, power: i64) -> f64 {
@@ -2425,6 +2459,37 @@ mod tests {
assert!(got[4].is_nan());
}
#[test]
fn fp4_decodes_as_libhdf5_does() {
// H5T_FLOAT_F4E2M1: 4 significant bits in a byte; the high bits are
// padding. libhdf5 2.2.0 reads 0b0110 as +inf and 0b0111 as NaN
// (IEEE-style, although OCP MX FP4 has neither), and returns NaNs
// with every mantissa bit set.
let fp4 = Datatype::FloatingPoint {
size: 1,
byte_order: DatatypeByteOrder::LittleEndian,
bit_offset: 0,
bit_precision: 4,
exponent_location: 1,
exponent_size: 2,
mantissa_location: 0,
mantissa_size: 1,
exponent_bias: 1,
};
let raw = [0x01, 0x05, 0xF5, 0x06, 0x0E, 0x07, 0x0F];
let got = read_as_f64(&raw, &fp4).unwrap();
assert_eq!(
&got[..5],
&[0.5, 3.0, 3.0, f64::INFINITY, f64::NEG_INFINITY]
);
assert_eq!(got[5].to_bits(), 0x7FFF_FFFF_FFFF_FFFF);
assert_eq!(got[6].to_bits(), 0xFFFF_FFFF_FFFF_FFFF);
let got = read_as_f32(&raw, &fp4).unwrap();
assert_eq!(&got[..3], &[0.5, 3.0, 3.0]);
assert_eq!(got[5].to_bits(), 0x7FFF_FFFF);
assert_eq!(got[6].to_bits(), 0xFFFF_FFFF);
}
#[test]
fn full_width_signed_unchanged() {
// Regression: full-width 32-bit signed must be unaffected.
+196 -17
View File
@@ -140,6 +140,19 @@ pub enum Datatype {
base_type: Box<Datatype>,
dimensions: Vec<u32>,
},
/// Class 11: HDF5 2.0 native complex number (`H5T_COMPLEX`, datatype
/// message version 5): two consecutive `base_type` values, real then
/// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float.
///
/// This variant exists for **writing** (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}`
/// stays the default complex encoding. [`Datatype::parse`] still
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
Complex { size: u32, base_type: Box<Datatype> },
}
/// Longest opaque tag that can be stored: its NUL-padded length must fit
@@ -862,19 +875,7 @@ impl Datatype {
actual: size as usize,
});
}
let members = vec![
CompoundMember {
name: String::from("r"),
byte_offset: 0,
datatype: base_type.clone(),
},
CompoundMember {
name: String::from("i"),
byte_offset: base_size as u64,
datatype: base_type,
},
];
Ok((Datatype::Compound { size, members }, pos))
Ok((Self::complex_as_compound(size, &base_type), pos))
}
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
}
@@ -939,7 +940,8 @@ impl Datatype {
.try_for_each(|m| m.datatype.check_unused_bits()),
Datatype::Enumeration { base_type, .. }
| Datatype::VariableLength { base_type, .. }
| Datatype::Array { base_type, .. } => base_type.check_unused_bits(),
| Datatype::Array { base_type, .. }
| Datatype::Complex { base_type, .. } => base_type.check_unused_bits(),
_ => Ok(()),
}
}
@@ -1046,7 +1048,8 @@ impl Datatype {
} else {
0
};
let mut buf = Self::build_header(9, 1, [bf0, bf1, 0], *size);
let version = base_type.min_parent_version().max(1);
let mut buf = Self::build_header(9, version, [bf0, bf1, 0], *size);
buf.extend_from_slice(&base_type.serialize());
buf
}
@@ -1054,7 +1057,11 @@ impl Datatype {
let num = members.len() as u16;
let bf0 = (num & 0xFF) as u8;
let bf1 = ((num >> 8) & 0xFF) as u8;
let mut buf = Self::build_header(6, 3, [bf0, bf1, 0], *size);
let version = members
.iter()
.map(|m| m.datatype.min_parent_version())
.fold(3, u8::max);
let mut buf = Self::build_header(6, version, [bf0, bf1, 0], *size);
let ob = offset_bytes_for_size(*size);
for m in members {
// Null-terminated name
@@ -1097,7 +1104,8 @@ impl Datatype {
base_type,
dimensions,
} => {
let mut buf = Self::build_header(10, 3, [0, 0, 0], self.type_size());
let version = base_type.min_parent_version().max(3);
let mut buf = Self::build_header(10, version, [0, 0, 0], self.type_size());
buf.push(dimensions.len() as u8);
for &d in dimensions {
buf.extend_from_slice(&d.to_le_bytes());
@@ -1152,9 +1160,82 @@ impl Datatype {
};
Self::build_header(7, version, [bf0, 0, 0], *size)
}
Datatype::Complex { size, base_type } => {
// Version 5 (HDF5 2.0), as libhdf5's `H5O__dtype_encode_helper`
// writes it: bit 0 = homogeneous (the only kind libhdf5
// supports), bits 1-2 = form (0, rectangular); the base
// datatype message follows.
let mut buf = Self::build_header(11, 5, [0x01, 0, 0], *size);
buf.extend_from_slice(&base_type.serialize());
buf
}
}
}
/// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`]
/// returns for a class-11 message. Readers that meet a
/// [`Datatype::Complex`] handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size();
Datatype::Compound {
size,
members: vec![
CompoundMember {
name: String::from("r"),
byte_offset: 0,
datatype: base_type.clone(),
},
CompoundMember {
name: String::from("i"),
byte_offset: u64::from(base_size),
datatype: base_type.clone(),
},
],
}
}
/// The lowest datatype message version a type that contains this one
/// may be encoded with. libhdf5 raises a compound, array, variable-length
/// or enum type to the version of its members (`H5O_DTYPE_CHECK_VERSION`
/// in `H5Odtype.c`), so a type holding a native complex (version 5) is
/// itself written as version 5; everything else we write keeps the
/// container's own version.
fn min_parent_version(&self) -> u8 {
match self {
Datatype::Complex { .. } => 5,
Datatype::Compound { members, .. } => members
.iter()
.map(|m| m.datatype.min_parent_version())
.fold(0, u8::max),
Datatype::Enumeration { base_type, .. }
| Datatype::VariableLength { base_type, .. }
| Datatype::Array { base_type, .. } => base_type.min_parent_version(),
_ => 0,
}
}
/// The highest datatype message version in this type's encoding, its
/// members' and base types' included (the version decides which HDF5
/// releases can read it: 1-3 HDF5 1.8, 4 HDF5 1.12, 5 HDF5 2.0).
pub fn max_encoded_version(&self) -> u8 {
let own = self.serialize().first().map_or(0, |b| b >> 4);
let inner = match self {
Datatype::Compound { members, .. } => members
.iter()
.map(|m| m.datatype.max_encoded_version())
.max()
.unwrap_or(0),
Datatype::Enumeration { base_type, .. }
| Datatype::VariableLength { base_type, .. }
| Datatype::Array { base_type, .. }
| Datatype::Complex { base_type, .. } => base_type.max_encoded_version(),
_ => 0,
};
own.max(inner)
}
/// Check that this datatype can be written: every part of it has an
/// on-disk encoding, and the encoding is one the reader (and libhdf5)
/// accepts. [`Self::serialize`] cannot report errors, so the writer calls
@@ -1189,6 +1270,20 @@ impl Datatype {
Datatype::Enumeration { base_type, .. }
| Datatype::VariableLength { base_type, .. }
| Datatype::Array { base_type, .. } => base_type.check_encodable_parts(),
// libhdf5 only builds complex types over IEEE floats
// (`H5Tcomplex_create`), always twice the base size.
Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { size: b, .. } if b.checked_mul(2) == Some(*size) => {
Ok(())
}
Datatype::FloatingPoint { .. } => Err(FormatError::SerializationError(format!(
"complex datatype of size {size} is not twice its base size {}",
base_type.type_size()
))),
_ => Err(FormatError::SerializationError(
"complex datatype base must be a floating-point type".into(),
)),
},
_ => Ok(()),
}
}
@@ -1216,6 +1311,7 @@ impl Datatype {
Datatype::Reference { size, .. } => *size,
Datatype::Enumeration { size, .. } => *size,
Datatype::VariableLength { size, .. } => *size,
Datatype::Complex { size, .. } => *size,
Datatype::Array {
base_type,
dimensions,
@@ -1815,6 +1911,89 @@ mod tests {
));
}
#[test]
fn native_complex_serializes_as_libhdf5_2_0_does() {
use crate::type_builders::{make_native_complex_f32_type, make_native_complex_f64_type};
let dt = make_native_complex_f64_type();
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!(
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
let f32c = make_native_complex_f32_type().serialize();
assert_eq!(
&f32c[..8],
&[0x5b, 0x01, 0x00, 0x00, 0x08, 0x00, 0x00, 0x00]
);
assert_eq!(
&f32c[8..],
&crate::type_builders::make_f32_type().serialize()[..]
);
}
#[test]
fn compound_holding_native_complex_serializes_as_libhdf5_2_0_does() {
// The same type as `test_compound_with_complex_member_from_hdf5_2_0`:
// libhdf5 raises the compound to version 5 for its complex member.
let dt = Datatype::Compound {
size: 24,
members: vec![
CompoundMember {
name: "z".into(),
byte_offset: 0,
datatype: crate::type_builders::make_native_complex_f64_type(),
},
CompoundMember {
name: "k".into(),
byte_offset: 16,
datatype: crate::type_builders::make_i64_type(),
},
],
};
let mut want = vec![
0x56, 0x02, 0x00, 0x00, 0x18, 0x00, 0x00, 0x00, b'z', 0x00, 0x00,
];
want.extend_from_slice(&COMPLEX_F64_HDF5_2_0);
want.extend_from_slice(&[b'k', 0x00, 0x10]);
want.extend_from_slice(&[
0x10, 0x08, 0x00, 0x00, 0x08, 0x00, 0x00, 0x00, 0x00, 0x00, 0x40, 0x00,
]);
assert_eq!(dt.serialize(), want);
// An array of complex is raised to version 5 as well; one without
// stays at version 3.
let arr = Datatype::Array {
base_type: Box::new(crate::type_builders::make_native_complex_f32_type()),
dimensions: vec![2],
};
assert_eq!(arr.serialize()[0], 0x5a);
arr.check_encodable().unwrap();
let plain = Datatype::Array {
base_type: Box::new(crate::type_builders::make_f32_type()),
dimensions: vec![2],
};
assert_eq!(plain.serialize()[0], 0x3a);
}
#[test]
fn native_complex_must_be_twice_an_ieee_float() {
let bad_size = Datatype::Complex {
size: 12,
base_type: Box::new(crate::type_builders::make_f64_type()),
};
assert!(bad_size.check_encodable().is_err());
let int_base = Datatype::Complex {
size: 8,
base_type: Box::new(crate::type_builders::make_i32_type()),
};
assert!(int_base.check_encodable().is_err());
}
#[test]
fn test_reference_object() {
let buf = build_dt_header(7, 1, [0, 0, 0], 8);
+3 -2
View File
@@ -18,9 +18,10 @@ pub(crate) fn serialize_v4_extensible_array(
ea_address: u64,
offset_size: u8,
element_size: u32,
version: u8,
) -> Vec<u8> {
let mut buf = Vec::new();
buf.push(4); // version
buf.push(version);
buf.push(2); // class = chunked
buf.push(0x00); // flags
@@ -87,7 +88,7 @@ pub fn build_extensible_array_at(
ea_base_address: u64,
) -> Vec<u8> {
let os = offset_size as usize;
let chunk_size_bytes = has_filters.then(|| filtered_chunk_size_len(slots));
let chunk_size_bytes = has_filters.then(|| filtered_chunk_size_len(slots, length_size));
let elem_size = os + chunk_size_bytes.map_or(0, |n| n + 4);
let client_id: u8 = if has_filters { 1 } else { 0 };
let arr_off_size = (MAX_NELMTS_BITS as usize).div_ceil(8);
+20
View File
@@ -167,6 +167,19 @@ pub enum FormatError {
VlDataError(String),
/// Serialization error.
SerializationError(String),
/// The file's library version bounds
/// ([`FileWriter::libver_bounds`](crate::file_writer::FileWriter::libver_bounds))
/// do not allow what was asked for: `what` needs the format of HDF5
/// `needs` or later, and the high bound is `high` (or the low bound is
/// above the high one, with `needs` the low bound).
LibverBound {
/// What cannot be written.
what: String,
/// The oldest release whose format holds it.
needs: crate::libver::LibVer,
/// The file's high bound.
high: crate::libver::LibVer,
},
/// Dataset is missing data.
DatasetMissingData,
/// Dataset is missing shape.
@@ -450,6 +463,13 @@ impl fmt::Display for FormatError {
FormatError::SerializationError(msg) => {
write!(f, "serialization error: {msg}")
}
FormatError::LibverBound { what, needs, high } => {
write!(
f,
"{what} needs the HDF5 {needs} file format, above the high \
library version bound ({high})"
)
}
FormatError::DatasetMissingData => {
write!(f, "dataset is missing data")
}
@@ -224,7 +224,7 @@ fn read_element(
};
Ok((
Some(ChunkInfo {
chunk_size: chunk_byte_size as u32,
chunk_size: chunk_byte_size,
filter_mask: 0,
offsets,
address,
@@ -259,7 +259,7 @@ fn read_element(
};
Ok((
Some(ChunkInfo {
chunk_size: chunk_size as u32,
chunk_size,
filter_mask,
offsets,
address,
@@ -946,7 +946,7 @@ mod tests {
assert_eq!(chunks.len(), 2);
assert_eq!(chunks[0].address, base_addr);
assert_eq!(chunks[0].offsets, vec![0]);
assert_eq!(chunks[0].chunk_size, chunk_byte_size as u32);
assert_eq!(chunks[0].chunk_size, chunk_byte_size);
assert_eq!(chunks[1].address, base_addr + chunk_byte_size);
assert_eq!(chunks[1].offsets, vec![20]);
}
+240 -8
View File
@@ -5,12 +5,13 @@
use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use alloc::{format, string::String, vec, vec::Vec};
use crate::attribute::AttributeMessage;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::chunked_write::{
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed, precompress_chunks,
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver,
precompress_chunks,
};
use crate::data_layout::VdsMapping;
use crate::dataspace::{Dataspace, DataspaceType};
@@ -31,6 +32,7 @@ pub use crate::type_builders::ProvenanceConfig;
pub use crate::type_builders::{AttrValue, CompoundTypeBuilder, EnumTypeBuilder};
use crate::datatype::{CharacterSet, Datatype};
use crate::libver::LibVer;
pub(crate) const OFFSET_SIZE: u8 = 8;
pub(crate) const LENGTH_SIZE: u8 = 8;
@@ -168,13 +170,15 @@ pub(crate) fn build_dataset_oh(
attrs: AttrStorage<'_>,
fill_message: &[u8],
refcount: u32,
layout_version: u8,
) -> Result<Vec<u8>, FormatError> {
let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01);
// Versions 3 and 4 encode a contiguous layout the same way.
let mut dl = Vec::new();
dl.push(4); // version
dl.push(layout_version);
dl.push(1); // class = contiguous
// An empty dataset has no storage: its address must be the undefined
// address, as libhdf5 writes it. A real address with size 0 trips
@@ -198,14 +202,16 @@ pub(crate) fn build_compact_dataset_oh(
attrs: AttrStorage<'_>,
fill_message: &[u8],
refcount: u32,
layout_version: u8,
) -> Result<Vec<u8>, FormatError> {
let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01);
// Compact layout message: version=4, class=0, u16 size, inline data
// Compact layout message: version (3 and 4 are the same here), class=0,
// u16 size, inline data
let mut dl = Vec::new();
dl.push(4); // version
dl.push(layout_version);
dl.push(0); // class = compact
dl.extend_from_slice(&(data.len() as u16).to_le_bytes());
dl.extend_from_slice(data);
@@ -1371,6 +1377,10 @@ pub struct FileWriter {
/// file-space strategy (a File Space Info message in the superblock
/// extension).
page_size: Option<u32>,
/// Library version bounds: the low bound picks the format versions
/// written, the high bound limits the features allowed.
low: LibVer,
high: LibVer,
}
impl Default for FileWriter {
@@ -1485,9 +1495,39 @@ impl FileWriter {
alignment_threshold: 0,
alignment_bytes: 0,
page_size: None,
low: LibVer::V110,
high: LibVer::Latest,
}
}
/// Set the library version bounds, as libhdf5's `H5Pset_libver_bounds`
/// (h5py's `libver=(low, high)`): the oldest HDF5 release whose format
/// the file uses (`low`), and the newest whose features it may use
/// (`high`). See [`crate::libver`] for what each bound changes.
///
/// The default, `(LibVer::V110, LibVer::Latest)`, is what clawhdf5 has
/// always written: the HDF5 1.10 format (version-3 superblock, version-4
/// layouts with the 1.10 chunk indexes), readable by HDF5 1.10 and later.
///
/// `(LibVer::V18, LibVer::V18)` writes a file HDF5 1.8 can read — the
/// low bound libhdf5 2.0 uses by default: a version-2 superblock,
/// version-3 layouts, and a version-1 B-tree for every chunked dataset,
/// resizable ones included; [`Self::finish`] then fails with
/// [`FormatError::LibverBound`] for anything HDF5 1.8 cannot read
/// (virtual datasets, a paged file, the 1.12 reference types, native
/// complex numbers). With a low bound of 1.8 and a later high bound
/// such objects are written in the newer format, as libhdf5 writes them;
/// the rest of the file stays readable by 1.8. A chunk of 4 GiB or more
/// always takes HDF5 2.0's version-5 layout (never a version-1 B-tree)
/// and so needs a high bound of at least [`LibVer::V200`].
///
/// A low bound above the high bound makes [`Self::finish`] fail.
pub fn libver_bounds(&mut self, low: LibVer, high: LibVer) -> &mut Self {
self.low = low;
self.high = high;
self
}
/// Set global file alignment: datasets with raw data >= `threshold` bytes
/// will have their data aligned to `bytes` boundary.
///
@@ -1583,6 +1623,33 @@ impl FileWriter {
)));
}
let (low, high) = (self.low, self.high);
let within_bounds = |what: &dyn Fn() -> String, needs: LibVer| {
if needs > high {
Err(FormatError::LibverBound {
what: what(),
needs,
high,
})
} else {
Ok(())
}
};
within_bounds(&|| format!("a low library version bound of {low}"), low)?;
if page_size.is_some() {
within_bounds(&|| "the paged file-space strategy".into(), LibVer::V110)?;
}
// Versions 3 of the layout message and 2 of the superblock are what
// HDF5 1.8 reads; 1.10 added version 4 (with its chunk indexes) and
// version 3. A paged file needs the version-3 superblock whatever
// the low bound (libhdf5 raises it as far as the high bound allows).
let layout_version: u8 = if low < LibVer::V110 { 3 } else { 4 };
let superblock_version: u8 = if low < LibVer::V110 && page_size.is_none() {
2
} else {
3
};
// The group tree, in layout order: groups depth-first from the root,
// then every group's datasets in the same order.
let tree = writer_tree::build(self.root, self.track_order)?;
@@ -1621,9 +1688,20 @@ impl FileWriter {
let ds_attrs = all_ds.iter().flat_map(|d| &d.attrs);
for a in group_attrs.chain(ds_attrs) {
a.datatype.check_encodable()?;
within_bounds(
&|| format!("the datatype of attribute {:?}", a.name),
LibVer::for_datatype_version(a.datatype.max_encoded_version()),
)?;
}
for d in &all_ds {
d.dt.check_encodable()?;
within_bounds(
&|| "a dataset's datatype".into(),
LibVer::for_datatype_version(d.dt.max_encoded_version()),
)?;
if d.virtual_sources.is_some() {
within_bounds(&|| "a virtual dataset".into(), LibVer::V110)?;
}
}
let is_vds: Vec<bool> = all_ds.iter().map(|d| d.virtual_sources.is_some()).collect();
@@ -1749,10 +1827,12 @@ impl FileWriter {
elem_size,
&d.chunk_options,
)?;
let result = build_chunked_data_from_precompressed(
let result = build_chunked_data_from_precompressed_libver(
&pre,
dummy_cursor,
d.maxshape.as_deref(),
low,
high,
)?;
dummy_cursor += result.data_bytes.len() as u64;
let oh = build_chunked_dataset_oh(
@@ -1785,6 +1865,7 @@ impl FileWriter {
},
&d.fill_message,
d.refcount,
layout_version,
)?;
dummy_blobs.push(DataBlob {
data: vec![],
@@ -1804,6 +1885,7 @@ impl FileWriter {
},
&d.fill_message,
d.refcount,
layout_version,
)?;
dummy_blobs.push(DataBlob {
data: vec![],
@@ -1904,13 +1986,15 @@ impl FileWriter {
let base_address = cursor2 as u64;
// Reuse precompressed chunks from Pass 1 — avoids re-compressing
// the same data a second time.
let result = build_chunked_data_from_precompressed(
let result = build_chunked_data_from_precompressed_libver(
dummy_blobs[i]
.precompressed
.as_ref()
.expect("chunked dataset missing precompressed cache"),
base_address,
d.maxshape.as_deref(),
low,
high,
)?;
cursor2 += result.data_bytes.len();
let oh = build_chunked_dataset_oh(
@@ -1944,6 +2028,7 @@ impl FileWriter {
},
&d.fill_message,
d.refcount,
layout_version,
)?;
ds_blobs2.push(DataBlob {
data: vec![],
@@ -1973,6 +2058,7 @@ impl FileWriter {
},
&d.fill_message,
d.refcount,
layout_version,
)?;
let mut data = vec![0u8; padding];
data.extend_from_slice(&d.raw);
@@ -1997,7 +2083,7 @@ impl FileWriter {
let mut buf = Vec::with_capacity(cursor2);
let sb = Superblock {
version: 3,
version: superblock_version,
offset_size: OFFSET_SIZE,
length_size: LENGTH_SIZE,
base_address: 0,
@@ -2783,4 +2869,150 @@ mod tests {
assert_eq!(sb.version, 3);
assert_eq!(sb.page_size, None);
}
fn layout_of(bytes: &[u8], name: &str) -> Vec<u8> {
let sb = Superblock::parse(bytes, 0).unwrap();
let addr = resolve_path_any(bytes, &sb, name).unwrap();
let hdr = ObjectHeader::parse(bytes, addr as usize, 8, 8).unwrap();
hdr.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap()
.data
.clone()
}
#[test]
fn libver_v18_writes_the_1_8_format() {
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("compact").with_f64_data(&[3.0]).compact();
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
fw.create_dataset("none")
.with_f64_data(&[])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 2);
assert_eq!(layout_of(&bytes, "contig")[..2], [3, 1]);
assert_eq!(layout_of(&bytes, "compact")[..2], [3, 0]);
let grow = layout_of(&bytes, "grow");
// Version 3, chunked, 2 dimensions (the element size is the last),
// B-tree address, chunk dims 2 and 8.
assert_eq!(grow[..3], [3, 2, 2]);
assert_eq!(grow[11..], [2, 0, 0, 0, 8, 0, 0, 0]);
let root = u64::from_le_bytes(grow[3..11].try_into().unwrap()) as usize;
assert_eq!(&bytes[root..root + 5], b"TREE\x01");
// No chunks, no tree.
assert_eq!(layout_of(&bytes, "none")[3..11], [0xff; 8]);
assert_eq!(read_dataset_f64(&bytes, "grow"), vec![1.0, 2.0, 3.0]);
assert_eq!(read_dataset_f64(&bytes, "contig"), vec![1.0, 2.0]);
assert_eq!(read_dataset_f64(&bytes, "compact"), vec![3.0]);
}
#[test]
fn default_libver_bounds_keep_the_1_10_format() {
let mut fw = FileWriter::new();
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
let default = fw.finish().unwrap();
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V110, LibVer::Latest);
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
assert_eq!(fw.finish().unwrap(), default);
assert_eq!(layout_of(&default, "contig")[0], 4);
assert_eq!(layout_of(&default, "grow")[..2], [4, 2]);
}
#[test]
fn libver_high_bound_refuses_newer_features() {
let bound = |r: Result<Vec<u8>, FormatError>, needs: LibVer| match r {
Err(FormatError::LibverBound { needs: n, high, .. }) => {
assert_eq!((n, high), (needs, LibVer::V18));
}
other => panic!("expected a bound error, got {other:?}"),
};
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("z")
.with_native_complex_f64_data(&[[1.0, 2.0]]);
bound(fw.finish(), LibVer::V200);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("x").with_f64_data(&[1.0]).set_attr(
"z",
AttrValue::Raw {
datatype: crate::type_builders::make_native_complex_f64_type(),
shape: vec![],
data: vec![0; 16],
},
);
bound(fw.finish(), LibVer::V200);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("r").with_compound_data(
Datatype::Reference {
size: 16,
ref_type: crate::datatype::ReferenceType::Object2,
},
vec![0; 16],
1,
);
bound(fw.finish(), LibVer::V112);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("src").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("vds")
.with_shape(&[2])
.with_f64_data(&[])
.with_virtual_sources(vec![VdsMapping {
source_file: ".".into(),
source_dataset: "src".into(),
source_selection: sel_all(),
virtual_selection: sel_hyper_1d(0, 2),
}]);
bound(fw.finish(), LibVer::V110);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18)
.with_page_size(4096);
bound(fw.finish(), LibVer::V110);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V110, LibVer::V18);
bound(fw.finish(), LibVer::V110);
}
#[test]
fn libver_low_v18_high_latest_allows_newer_objects() {
// As libhdf5 does: the object that needs a newer format gets it,
// the rest of the file keeps the 1.8 format.
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::Latest);
fw.create_dataset("z")
.with_native_complex_f64_data(&[[1.0, 2.0]]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 2);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::Latest)
.with_page_size(4096);
fw.create_dataset("x").with_f64_data(&[1.0]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
assert_eq!(layout_of(&bytes, "x")[0], 3);
}
}
+104 -24
View File
@@ -340,10 +340,23 @@ pub fn decompress_chunk_exact_with<'s>(
} else {
MAX_DECOMPRESS_SIZE
};
let size_hint = if ctx.max_output != 0 {
// The bound is the output's size when only size-preserving
// filters (shuffle, Fletcher32) remain to be undone; before
// any other filter (a second deflate, N-Bit, ...) it is only
// a ceiling, and the output starts smaller and grows.
// Reserving the bound of a 4 GiB chunk for a stage a few MiB
// long doubled the peak memory of reading it (8.5 GiB, now
// 4.0, for `huge_chunks_filtered.h5`'s double deflate).
let exact = pipeline.filters[..i].iter().enumerate().all(|(j, f)| {
filter_skipped(filter_mask, j)
|| matches!(f.filter_id, FILTER_SHUFFLE | FILTER_FLETCHER32)
});
let size_hint = if ctx.max_output == 0 {
input.len().saturating_mul(4).min(1 << 20)
} else if exact {
ctx.max_output
} else {
input.len().saturating_mul(4).min(1 << 20)
input.len().saturating_mul(4).min(ctx.max_output)
};
let inflater = scratch
.inflater
@@ -429,7 +442,10 @@ pub fn compress_chunk_masked(
"more than 32 filters in a pipeline".into(),
));
}
let mut result = data.to_vec();
// The input is not copied: the first filter reads it where it is (a
// chunk may be 4 GiB or more), and each filter's output replaces the
// previous one.
let mut owned: Option<Vec<u8>> = None;
let mut mask = 0u32;
for (i, filter) in pipeline.filters.iter().enumerate() {
let ctx = FilterContext {
@@ -437,7 +453,8 @@ pub fn compress_chunk_masked(
element_size: element_size as usize,
max_output: 0,
};
let out = match filter_registry::encode(&result, &ctx) {
let result: &[u8] = owned.as_deref().unwrap_or(data);
let out = match filter_registry::encode(result, &ctx) {
Ok(out)
if FAIL_UNLESS_SMALLER.contains(&filter.filter_id) && out.len() >= result.len() =>
{
@@ -449,13 +466,13 @@ pub fn compress_chunk_masked(
r => r,
};
match out {
Ok(out) => result = out,
Ok(out) => owned = Some(out),
Err(e @ FormatError::UnsupportedFilter(_)) => return Err(e),
Err(_) if filter.flags & FILTER_FLAG_OPTIONAL != 0 => mask |= 1 << i,
Err(e) => return Err(e),
}
}
Ok((result, mask))
Ok((owned.unwrap_or_else(|| data.to_vec()), mask))
}
/// The filters compiled into this build, sorted by ID (see
@@ -1339,6 +1356,9 @@ fn deflate_compress(data: &[u8], level: u32) -> Result<Vec<u8>, FormatError> {
deflate_bounded(data, level).map_err(FormatError::CompressionError)
}
/// Largest compression output reserved at its worst-case size up front.
const DEFLATE_EXACT_BOUND: usize = 64 << 20;
/// Deflate `data` into a zlib stream in one pass, into a buffer sized for the
/// worst case up front (the same reasoning as [`inflate_bounded`]).
#[cfg(feature = "deflate")]
@@ -1347,24 +1367,44 @@ pub(crate) fn deflate_bounded(data: &[u8], level: u32) -> Result<Vec<u8>, String
// zlib's compressBound, plus the zlib header and trailer.
let bound = data.len() + (data.len() >> 12) + (data.len() >> 14) + (data.len() >> 25) + 13 + 6;
// flate2's Rust backends (zlib-rs, miniz_oxide) zero the whole spare
// capacity on each call, so a large input's worst-case bound would be
// memory held for nothing (4 GiB for a 4 GiB chunk that deflates to a
// few MiB): past 64 MiB the output starts at 1/16 of the bound and
// doubles as needed.
let first = if bound <= DEFLATE_EXACT_BOUND {
bound
} else {
bound / 16
};
let mut out = Vec::new();
out.try_reserve_exact(bound)
out.try_reserve_exact(first)
.map_err(|e| format!("deflate: cannot allocate output: {e}"))?;
let mut deflater = Compress::new(Compression::new(level), true);
loop {
let (in_before, out_before) = (deflater.total_in(), deflater.total_out());
let rest = &data[saturating_usize(in_before)..];
// zlib takes at most u32::MAX input bytes per call, and `Finish`
// ends the stream after the bytes it took: a chunk of 4 GiB or more
// was cut at 4 GiB - 1. Finish only once the rest fits one call.
let flush = if rest.len() > u32::MAX as usize {
FlushCompress::None
} else {
FlushCompress::Finish
};
let status = deflater
.compress_vec(
&data[saturating_usize(in_before)..],
&mut out,
FlushCompress::Finish,
)
.compress_vec(rest, &mut out, flush)
.map_err(|e| format!("deflate: {e}"))?;
match status {
Status::StreamEnd => return Ok(out),
// The bound should make running out of room unreachable; grow
// rather than fail if it happens.
Status::StreamEnd => {
if out.capacity() - out.len() > DEFLATE_EXACT_BOUND {
out.shrink_to_fit();
}
return Ok(out);
}
// Out of room (the bound makes it unreachable below
// `DEFLATE_EXACT_BOUND`): grow rather than fail.
Status::Ok | Status::BufError if out.len() == out.capacity() => out
.try_reserve(out.capacity().max(4096))
.map_err(|e| format!("deflate: cannot allocate output: {e}"))?,
@@ -1395,10 +1435,12 @@ const LZ4_DEFAULT_BLOCK_SIZE: usize = 1 << 30;
/// * The legacy clawhdf5 framing (up to 2.7.0): a 4-byte little-endian size
/// followed by one raw LZ4 block. libhdf5 cannot read it.
///
/// They are told apart unambiguously: an HDF5 chunk is smaller than 4 GiB, so
/// the registered format's big-endian `u64` size always starts with four zero
/// bytes and the whole chunk is at least 12 bytes; a legacy chunk starts with
/// four zero bytes only when it is empty, and is then 5 bytes long.
/// They are told apart unambiguously: for a chunk under 4 GiB the registered
/// format's big-endian `u64` size starts with four zero bytes and the whole
/// chunk is at least 12 bytes; a legacy chunk starts with four zero bytes
/// only when it is empty, and is then 5 bytes long. A chunk of 4 GiB or more
/// (HDF5 2.0) is always the registered format: clawhdf5 never wrote legacy
/// chunks that large.
///
/// Every size read from the payload is bounded against `expected_bytes` (the
/// pipeline's declared chunk size) before it sizes an allocation, so a crafted
@@ -1416,14 +1458,18 @@ fn lz4_decompress(data: &[u8], expected_bytes: usize) -> Result<Vec<u8>, FormatE
"lz4: declared size exceeds chunk size".into(),
));
}
if size > MAX_DECOMPRESS_SIZE {
// Without a chunk size, a ceiling; with one, the chunk size is the
// bound (a chunk may be 4 GiB or more, like a deflated one).
if expected_bytes == 0 && size > MAX_DECOMPRESS_SIZE {
return Err(FormatError::DecompressionError(
"lz4: declared size exceeds limit".into(),
));
}
Ok(())
};
if data.len() >= 12 && data[..4] == [0, 0, 0, 0] {
// A chunk of 4 GiB or more cannot be legacy (its size field is 32
// bits), and its registered-format size does not start with zeros.
if data.len() >= 12 && (data[..4] == [0, 0, 0, 0] || expected_bytes > u32::MAX as usize) {
return lz4_decompress_hdf5(data, check_size);
}
// Legacy clawhdf5 framing: 4-byte LE size + one LZ4 block.
@@ -1441,9 +1487,10 @@ fn lz4_decompress_hdf5(
) -> Result<Vec<u8>, FormatError> {
let err = |m: &str| FormatError::DecompressionError(format!("lz4: {m}"));
let be32 = |b: &[u8]| u32::from_be_bytes([b[0], b[1], b[2], b[3]]) as usize;
// The first four bytes are zero (checked by the caller), so the size is
// the low 32 bits of the big-endian u64.
let orig_size = be32(&data[4..8]);
let orig_size = u64::from_be_bytes([
data[0], data[1], data[2], data[3], data[4], data[5], data[6], data[7],
]);
let orig_size = usize::try_from(orig_size).map_err(|_| err("chunk too large"))?;
check_size(orig_size)?;
let block_size = be32(&data[8..12]).min(orig_size);
if block_size == 0 && orig_size != 0 {
@@ -2717,6 +2764,39 @@ mod tests {
assert_eq!(decompressed, data);
}
/// A chunk's size bounds an LZ4 chunk, not the 256 MiB ceiling for an
/// unknown size: a 300 MiB chunk was refused ("declared size exceeds
/// limit").
#[test]
#[cfg(feature = "lz4")]
fn lz4_chunks_over_256_mib_decode() {
let mut data = vec![0u8; 300 << 20];
data[12345] = 7;
let compressed = lz4_compress(&data, &[]).unwrap();
let decompressed = lz4_decompress(&compressed, data.len()).unwrap();
assert!(decompressed == data);
// Without a chunk size the ceiling still applies.
assert!(lz4_decompress(&compressed, 0).is_err());
}
/// A chunk of 4 GiB or more is always in the registered framing, whose
/// big-endian size then does not start with four zero bytes: its size
/// is read whole (here larger than the chunk, so refused before any
/// allocation), not taken for a legacy 4-byte size.
#[test]
#[cfg(all(feature = "lz4", target_pointer_width = "64"))]
fn lz4_chunks_of_4_gib_use_the_registered_framing() {
let chunk = (1usize << 32) + 8;
let mut data = ((chunk + 8) as u64).to_be_bytes().to_vec();
data.extend_from_slice(&(1u32 << 30).to_be_bytes());
data.extend_from_slice(&[0; 8]);
let err = lz4_decompress(&data, chunk).unwrap_err();
assert!(
matches!(&err, FormatError::DecompressionError(m) if m.contains("exceeds chunk size")),
"{err:?}"
);
}
#[test]
#[cfg(feature = "lz4")]
fn pipeline_lz4_only() {
+4 -4
View File
@@ -383,7 +383,7 @@ fn parse_fa_element(
offset_size: u8,
element_size: u8,
chunk_byte_size: u64,
) -> Result<Option<(u64, u32, u32)>, FormatError> {
) -> Result<Option<(u64, u64, u32)>, FormatError> {
let os = offset_size as usize;
if client_id == 0 {
// Non-filtered: element is just the chunk address.
@@ -393,7 +393,7 @@ fn parse_fa_element(
return Ok(None);
}
let address = read_offset(file_data, abs, offset_size)?;
Ok(Some((address, chunk_byte_size as u32, 0)))
Ok(Some((address, chunk_byte_size, 0)))
} else {
// Filtered: address(offset_size) + chunk_size(variable) + filter_mask(4)
let es = element_size as usize;
@@ -418,7 +418,7 @@ fn parse_fa_element(
file_data[fm_off + 2],
file_data[fm_off + 3],
]);
Ok(Some((address, chunk_size as u32, filter_mask)))
Ok(Some((address, chunk_size, filter_mask)))
}
}
@@ -688,7 +688,7 @@ mod tests {
assert_eq!(c.address, base_addr + i as u64 * chunk_byte_size as u64);
assert_eq!(c.offsets, vec![i as u64 * 20]);
assert_eq!(c.filter_mask, 0);
assert_eq!(c.chunk_size, chunk_byte_size as u32);
assert_eq!(c.chunk_size, chunk_byte_size as u64);
}
}
+166 -32
View File
@@ -10,7 +10,7 @@ use crate::addr::to_usize;
use crate::btree_v2::{BTreeV2Header, find_btree_v2_records_in};
use crate::error::FormatError;
use crate::filter_pipeline::FilterPipeline;
use crate::storage::{Storage, Window, len_usize, read_exact_at};
use crate::storage::{Storage, Window, len_usize, read_exact_at, read_upto};
/// Parsed fractal heap header (signature "FRHP").
#[derive(Debug, Clone)]
@@ -132,6 +132,9 @@ fn heap_id_type(first: u8) -> Result<u8, FormatError> {
const BTREE_HUGE_INDIRECT: u8 = 1;
const BTREE_HUGE_INDIRECT_FILTERED: u8 = 2;
/// Most child entries [`FractalHeapHeader::hint_managed_blocks`] walks.
const MAX_HINTED_BLOCKS: usize = 4096;
impl FractalHeapHeader {
/// Parse a fractal heap header at the given offset.
pub fn parse(
@@ -667,15 +670,8 @@ impl FractalHeapHeader {
"fractal heap: maximum recursion depth exceeded".into(),
));
}
let block_offset_bytes = (self.max_heap_size as usize).div_ceil(8);
let iblock_header = 5 + offset_size as usize + block_offset_bytes;
let nrows_usize = nrows as usize;
// Rows below max_direct_rows hold direct blocks; rows at/above hold
// child indirect blocks. (NOT the FRHP "starting rows" field.)
let start_indirect = self.max_direct_rows();
let max_direct_rows = nrows_usize.min(start_indirect);
// The block up to its last child entry. The walk below reads
// entries in order and stops at the one covering the target, which
// the geometry alone locates, so the first window ends there: a
@@ -685,32 +681,11 @@ impl FractalHeapHeader {
// over the whole block. Either window holds what it was asked for or
// ends at the end of the file, so its bounds checks are the
// whole-file ones.
let direct_entry = usize::from(offset_size)
+ if self.filter_pipeline.is_some() {
usize::from(self.length_size) + 4
} else {
0
};
let direct_entries = max_direct_rows.saturating_mul(usize::from(self.table_width));
let entries_len = |n: usize| {
n.min(direct_entries)
.saturating_mul(direct_entry)
.saturating_add(
n.saturating_sub(direct_entries)
.saturating_mul(usize::from(offset_size)),
)
};
let all_entries = direct_entries.saturating_add(
nrows_usize
.saturating_sub(start_indirect)
.saturating_mul(usize::from(self.table_width)),
);
let block_len = iblock_header.saturating_add(entries_len(all_entries));
let layout = self.indirect_layout(nrows_usize, offset_size);
let block_len = layout.len_upto(layout.all_entries);
let target_entry = self.indirect_entry_for(nrows_usize, iblock_heap_offset, target_offset);
let first_len = target_entry.map_or(block_len, |i| {
iblock_header
.saturating_add(entries_len(i.saturating_add(1)))
.min(block_len)
layout.len_upto(i.saturating_add(1)).min(block_len)
});
let mut next = self.walk_indirect_block(
&Window::read(file, iblock_addr as u64, first_len)?,
@@ -755,6 +730,140 @@ impl FractalHeapHeader {
}
}
/// Where the child entries of an indirect block of `nrows` rows are.
fn indirect_layout(&self, nrows: usize, offset_size: u8) -> IndirectLayout {
let block_offset_bytes = (self.max_heap_size as usize).div_ceil(8);
// Rows below max_direct_rows hold direct blocks; rows at/above hold
// child indirect blocks. (NOT the FRHP "starting rows" field.)
let start_indirect = self.max_direct_rows();
let direct_entries = nrows
.min(start_indirect)
.saturating_mul(usize::from(self.table_width));
IndirectLayout {
header: 5 + usize::from(offset_size) + block_offset_bytes,
direct_entry: usize::from(offset_size)
+ if self.filter_pipeline.is_some() {
usize::from(self.length_size) + 4
} else {
0
},
direct_entries,
indirect_entry: usize::from(offset_size),
all_entries: direct_entries.saturating_add(
nrows
.saturating_sub(start_indirect)
.saturating_mul(usize::from(self.table_width)),
),
}
}
/// Hint the heap's root block (see [`Storage::hint`]): every managed
/// object is read through it, and a storage that fetches between
/// attempts can fetch it along with whatever else the attempt missed
/// (the name index read before any object, say).
pub fn hint_root_block<S: Storage + ?Sized>(&self, file: &S) {
if is_undefined(self.root_block_address, self.offset_size) {
return;
}
let len = if self.current_rows_in_root_indirect_block == 0 {
if self.filter_pipeline.is_some() {
self.root_direct_block_filtered_size
} else {
self.starting_block_size
}
} else {
let layout = self.indirect_layout(
usize::from(self.current_rows_in_root_indirect_block),
self.offset_size,
);
layout.len_upto(layout.all_entries) as u64
};
file.hint(
self.root_block_address,
usize::try_from(len).unwrap_or(usize::MAX),
);
}
/// Hint every managed direct block of the heap (see
/// [`Storage::hint`]), for a caller about to read all of its objects (a
/// dense group's listing). The indirect blocks leading to them are read
/// here, as every object read goes through them, a few levels deep and
/// up to [`MAX_HINTED_BLOCKS`] entries; direct blocks are only hinted.
/// Nothing is returned and no error: a storage that has the file in
/// memory skips it, and the objects are read (and checked) as before.
pub fn hint_managed_blocks<S: Storage + ?Sized>(&self, file: &S) {
if file.as_contiguous().is_some()
|| is_undefined(self.root_block_address, self.offset_size)
|| self.current_rows_in_root_indirect_block == 0
{
// A direct root block is what `hint_root_block` hints.
return;
}
let mut budget = MAX_HINTED_BLOCKS;
self.hint_indirect_block(
file,
self.root_block_address,
usize::from(self.current_rows_in_root_indirect_block),
0,
&mut budget,
);
}
fn hint_indirect_block<S: Storage + ?Sized>(
&self,
file: &S,
addr: u64,
nrows: usize,
depth: usize,
budget: &mut usize,
) {
let os = self.offset_size;
let layout = self.indirect_layout(nrows, os);
let len = layout.len_upto(layout.all_entries);
if depth > 4 || len > 1 << 20 {
return;
}
let Ok(bytes) = read_upto(file, addr, len) else {
return;
};
if bytes.len() < len || bytes.get(..4) != Some(b"FHIB".as_slice()) {
return;
}
let start_indirect = self.max_direct_rows();
let mut pos = layout.header;
for row in 0..nrows {
for _ in 0..self.table_width {
let Some(left) = budget.checked_sub(1) else {
return;
};
*budget = left;
let Ok(child) = read_offset(&bytes, pos, os) else {
return;
};
if row < start_indirect {
let size = if self.filter_pipeline.is_some() {
match read_offset(&bytes, pos + usize::from(os), self.length_size) {
Ok(n) => n,
Err(_) => return,
}
} else {
self.block_size_for_row(row)
};
pos += layout.direct_entry;
if !is_undefined(child, os) {
file.hint(child, usize::try_from(size).unwrap_or(usize::MAX));
}
} else {
pos += layout.indirect_entry;
if !is_undefined(child, os) {
let rows = self.rows_for_size(self.block_size_for_row(row));
self.hint_indirect_block(file, child, usize::from(rows), depth + 1, budget);
}
}
}
}
}
/// Which child entry of an indirect block (numbered in walk order:
/// direct rows, then indirect rows) covers `target_offset`, from the
/// doubling-table geometry alone — the entry
@@ -939,6 +1048,31 @@ impl FractalHeapHeader {
}
}
/// Where an indirect block's child entries are: after its header, the
/// direct blocks' entries (address, and for a filtered heap the stored
/// size and filter mask), then the child indirect blocks' (address).
struct IndirectLayout {
header: usize,
direct_entry: usize,
direct_entries: usize,
indirect_entry: usize,
all_entries: usize,
}
impl IndirectLayout {
/// Bytes from the block's start to the end of its first `n` entries.
fn len_upto(&self, n: usize) -> usize {
let direct = n.min(self.direct_entries);
self.header
.saturating_add(direct.saturating_mul(self.direct_entry))
.saturating_add(
n.min(self.all_entries)
.saturating_sub(direct)
.saturating_mul(self.indirect_entry),
)
}
}
/// A managed direct block's location, extent and (for a filtered heap) its
/// stored size and filter mask.
/// The child of an indirect block that covers a heap offset.
+117 -6
View File
@@ -4,7 +4,7 @@
use alloc::{string::String, vec::Vec};
use crate::addr::checked_addr;
use crate::btree_v1::collect_symbol_table_nodes_in;
use crate::btree_v1::{BTreeV1Node, collect_symbol_table_nodes_in};
use crate::error::FormatError;
use crate::local_heap::LocalHeap;
use crate::message_type::MessageType;
@@ -48,7 +48,7 @@ pub fn resolve_v1_group_entries_in<S: Storage + ?Sized>(
offset_size: u8,
length_size: u8,
) -> Result<Vec<GroupEntry>, FormatError> {
let entries = v1_group_entries(file_data, sym_table_msg, offset_size, length_size)?;
let entries = v1_group_entries(file_data, sym_table_msg, offset_size, length_size, true)?;
if entries.iter().any(|e| e.name.is_empty()) {
return Err(FormatError::InvalidLinkName);
}
@@ -57,11 +57,16 @@ pub fn resolve_v1_group_entries_in<S: Storage + ?Sized>(
/// Every entry of a v1 group, empty names included — for looking a name up,
/// which never matches an empty name.
///
/// With `hint_headers` (a listing, whose children are usually opened
/// next), each entry's object header is hinted (see
/// [`Storage::hint`]) as soon as its symbol table node is read.
pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
file_data: &S,
sym_table_msg: &SymbolTableMessage,
offset_size: u8,
length_size: u8,
hint_headers: bool,
) -> Result<Vec<GroupEntry>, FormatError> {
// Parse local heap
let heap = LocalHeap::parse_in(
@@ -93,12 +98,23 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
// `storage::touch` does); that error is returned.
let mut failed = None;
for snod_addr in snod_addrs {
let snod = checked_addr(snod_addr)
.and_then(|a| SymbolTableNode::parse_in(file_data, a, offset_size));
if hint_headers && let Ok(snod) = &snod {
for entry in &snod.entries {
if entry.object_header_address != u64::MAX {
file_data.hint(
entry.object_header_address,
crate::object_header::OBJECT_HEADER_HINT_LEN,
);
}
}
}
if failed.is_some() {
let _ = SymbolTableNode::parse_in(file_data, snod_addr, offset_size);
continue;
}
let mut node = || -> Result<(), FormatError> {
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
let node = || -> Result<(), FormatError> {
let snod = snod?;
for entry in &snod.entries {
// Like libhdf5, look at the heap's free list only once a name
// is needed: an empty group with a damaged heap still lists.
@@ -126,6 +142,95 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
}
}
/// The entry called `name` in a v1 group, looked up as libhdf5 looks it up
/// (`H5G__stab_lookup`: `H5B_find` down the group's B-tree, then
/// `H5G__node_found` in one symbol table node) instead of by reading every
/// entry: at each node, the child whose key interval holds the name
/// (left key < name <= right key, keys being names in the local heap,
/// compared bytewise as `strcmp` does) is found by binary search. Reads
/// O(depth) nodes and names, where listing reads the whole group.
///
/// `Ok(None)` when the search does not lead to the name. In a group whose
/// B-tree is out of name order (damaged, or made by hand) that does not
/// prove it absent, so callers then fall back to reading every entry;
/// libhdf5 would report it missing.
pub(crate) fn find_v1_entry<S: Storage + ?Sized>(
file_data: &S,
sym_table_msg: &SymbolTableMessage,
name: &str,
offset_size: u8,
length_size: u8,
) -> Result<Option<GroupEntry>, FormatError> {
let heap = LocalHeap::parse_in(
file_data,
checked_addr(sym_table_msg.local_heap_address)?,
offset_size,
length_size,
)?;
// The keys and names are read from the heap's data segment one at a
// time: hint it (see `Storage::hint`).
file_data.hint(
heap.data_segment_address,
usize::try_from(heap.data_segment_size).map_or(1 << 20, |n| n.min(1 << 20)),
);
// As in listing: the heap's free list is checked before a name is used.
let mut heap_checked = false;
let mut name_at = |offset: u64| -> Result<String, FormatError> {
if !heap_checked {
heap.validate_free_list_in(file_data, length_size)?;
heap_checked = true;
}
heap.read_string_in(file_data, offset)
};
let want = name.as_bytes();
let mut address = sym_table_msg.btree_address;
for _ in 0..=crate::btree_v1::MAX_BTREE_DEPTH {
let node =
BTreeV1Node::parse_in(file_data, checked_addr(address)?, offset_size, length_size)?;
if node.node_type != 0 {
return Err(FormatError::InvalidBTreeNodeType(node.node_type));
}
// H5B_find's binary search with H5G__node_cmp3: go left when the
// name sorts at or before the left key, right when after the right
// key; otherwise this child holds it.
let (mut lo, mut hi) = (0usize, node.children.len());
let mut child = None;
while lo < hi {
let i = lo + (hi - lo) / 2;
let (Some(&left), Some(&right)) = (node.keys.get(i), node.keys.get(i + 1)) else {
return Ok(None);
};
if want <= name_at(left)?.as_bytes() {
hi = i;
} else if want > name_at(right)?.as_bytes() {
lo = i + 1;
} else {
child = Some(node.children[i]);
break;
}
}
let Some(child) = child else {
return Ok(None);
};
if node.node_level > 0 {
address = child;
continue;
}
let snod = SymbolTableNode::parse_in(file_data, checked_addr(child)?, offset_size)?;
for entry in &snod.entries {
if name_at(entry.link_name_offset)?.as_bytes() == want {
return Ok(Some(GroupEntry {
name: String::from(name),
object_header_address: entry.object_header_address,
cache_type: entry.cache_type,
}));
}
}
return Ok(None);
}
Err(FormatError::NestingDepthExceeded)
}
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
/// are the local-heap offset of the link's target path, and the entry's object
/// header address is undefined.
@@ -298,7 +403,13 @@ pub fn resolve_path_in<S: Storage + ?Sized>(
let mut current_sym_table = root_sym_table.clone();
for (i, component) in components.iter().enumerate() {
let entries = v1_group_entries(file_data, &current_sym_table, offset_size, length_size)?;
let entries = v1_group_entries(
file_data,
&current_sym_table,
offset_size,
length_size,
false,
)?;
let found = entries.iter().find(|e| e.name == *component);
match found {
+176 -11
View File
@@ -101,18 +101,50 @@ fn resolve_compact_entries(
Ok(entries)
}
/// What a version-2 B-tree header takes with 8-byte offsets and lengths
/// (22 bytes of fields, the root node's address and record count, and the
/// checksum), rounded up: hinted before one is read.
const BTREE_V2_HEADER_HINT_LEN: usize = 64;
/// The fractal heap of a dense group. The name index's header and the
/// heap's root block are read next, whatever the lookup: they are hinted
/// (see [`Storage::hint`]) so that a storage fetching between attempts
/// gets them in the same round trip as the heap's header.
fn dense_heap<S: Storage + ?Sized>(
file_data: &S,
link_info: &LinkInfoMessage,
fh_addr: u64,
offset_size: u8,
length_size: u8,
) -> Result<FractalHeapHeader, FormatError> {
if let Some(btree_addr) = link_info.btree_name_index_address {
file_data.hint(btree_addr, BTREE_V2_HEADER_HINT_LEN);
}
let fh =
FractalHeapHeader::parse_in(file_data, checked_addr(fh_addr)?, offset_size, length_size)?;
fh.hint_root_block(file_data);
Ok(fh)
}
/// Visit every link in dense storage (fractal heap + B-tree v2 name index).
///
/// With `hint_headers` (a listing, whose children are usually opened
/// next), the object header of every hard link is hinted (see
/// [`Storage::hint`]) as soon as the link is read, even after a failure.
fn for_each_dense_link<S: Storage + ?Sized>(
file_data: &S,
link_info: &LinkInfoMessage,
fh_addr: u64,
offset_size: u8,
length_size: u8,
hint_headers: bool,
mut visit: impl FnMut(LinkMessage),
) -> Result<(), FormatError> {
// Parse fractal heap
let fh =
FractalHeapHeader::parse_in(file_data, checked_addr(fh_addr)?, offset_size, length_size)?;
let fh = dense_heap(file_data, link_info, fh_addr, offset_size, length_size)?;
if hint_headers {
// Every link is read: so is every block of the heap.
fh.hint_managed_blocks(file_data);
}
// Parse B-tree v2 for name index
let btree_addr = link_info
@@ -144,11 +176,27 @@ fn for_each_dense_link<S: Storage + ?Sized>(
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
// Read managed object from fractal heap
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size);
let link = fh
.read_managed_object_in(file_data, id_bytes, offset_size)
.and_then(|d| parse_link(&d, offset_size));
if hint_headers
&& let Ok(Some(LinkMessage {
link_target:
LinkTarget::Hard {
object_header_address,
},
..
})) = &link
{
file_data.hint(
*object_header_address,
crate::object_header::OBJECT_HEADER_HINT_LEN,
);
}
if failed.is_some() {
continue;
}
match link_data.and_then(|d| parse_link(&d, offset_size)) {
match link {
Ok(Some(link)) => visit(link),
Ok(None) => {}
Err(e) => failed = Some(e),
@@ -175,6 +223,7 @@ fn resolve_dense_entries<S: Storage + ?Sized>(
fh_addr,
offset_size,
length_size,
true,
|link| {
if let LinkTarget::Hard {
object_header_address,
@@ -247,8 +296,7 @@ fn links_named<S: Storage + ?Sized>(
return Ok(found);
};
let fh =
FractalHeapHeader::parse_in(file_data, checked_addr(fh_addr)?, offset_size, length_size)?;
let fh = dense_heap(file_data, &link_info, fh_addr, offset_size, length_size)?;
let btree_addr = link_info
.btree_name_index_address
.ok_or_else(|| FormatError::PathNotFound(String::from("no B-tree v2 name index")))?;
@@ -265,6 +313,7 @@ fn links_named<S: Storage + ?Sized>(
fh_addr,
offset_size,
length_size,
false,
|link| {
if link.name == name {
found.push(link);
@@ -336,6 +385,27 @@ fn lookup_link<S: Storage + ?Sized>(
length_size: u8,
) -> Result<Option<LinkTarget>, FormatError> {
if is_v1_group(object_header) {
// Down the group's B-tree, as libhdf5 looks a name up; only when
// that does not find a hard link of that name is every entry read
// (a soft link, a group whose B-tree is out of order). A storage
// error (a read a restartable storage has not fetched yet) is
// returned as is: reading every entry would not get further.
if let Some(sym_msg) = object_header
.messages
.iter()
.find(|m| m.msg_type == MessageType::SymbolTable)
{
let stm = SymbolTableMessage::parse(&sym_msg.data, offset_size)?;
match group_v1::find_v1_entry(file_data, &stm, name, offset_size, length_size) {
Ok(Some(e)) if e.object_header_address != u64::MAX => {
return Ok(Some(LinkTarget::Hard {
object_header_address: e.object_header_address,
}));
}
Err(e @ FormatError::Storage(_)) => return Err(e),
_ => {}
}
}
let entries = resolve_group_entries(file_data, object_header, offset_size, length_size)?;
if let Some(e) = entries
.iter()
@@ -409,7 +479,7 @@ fn resolve_child_core<S: Storage + ?Sized>(
let not_found = || FormatError::PathNotFound(String::from(name));
let header = ObjectHeader::parse_in(file_data, checked_addr(group_address)?, os, ls)?;
if !is_v2_group(&header) || is_v1_group(&header) {
return resolve_group_children_in(file_data, superblock, group_address)?
return group_children(file_data, superblock, group_address, false)?
.into_iter()
.find(|e| e.name == name)
.map(|e| e.object_header_address)
@@ -575,6 +645,18 @@ fn resolve_group_children_core<S: Storage + ?Sized>(
file_data: &S,
superblock: &Superblock,
group_address: u64,
) -> Result<Vec<GroupEntry>, FormatError> {
group_children(file_data, superblock, group_address, true)
}
/// [`resolve_group_children`]; with `hint_headers`, every child's object
/// header is hinted (see [`Storage::hint`]) as soon as its address is
/// known, for a listing whose children are opened next.
fn group_children<S: Storage + ?Sized>(
file_data: &S,
superblock: &Superblock,
group_address: u64,
hint_headers: bool,
) -> Result<Vec<GroupEntry>, FormatError> {
let os = superblock.offset_size;
let ls = superblock.length_size;
@@ -589,7 +671,10 @@ fn resolve_group_children_core<S: Storage + ?Sized>(
.find(|m| m.msg_type == MessageType::SymbolTable)
.ok_or_else(|| FormatError::PathNotFound(String::from("no symbol table message")))?;
let stm = SymbolTableMessage::parse(&sym_msg.data, os)?;
let all = group_v1::resolve_v1_group_entries_in(file_data, &stm, os, ls)?;
let all = group_v1::v1_group_entries(file_data, &stm, os, ls, hint_headers)?;
if all.iter().any(|e| e.name.is_empty()) {
return Err(FormatError::InvalidLinkName);
}
if all.iter().any(group_v1::is_v1_soft_link) {
soft = group_v1::v1_soft_links_in(file_data, &stm, os, ls)?;
}
@@ -615,7 +700,7 @@ fn resolve_group_children_core<S: Storage + ?Sized>(
};
let link_info = find_link_info(&header, os)?;
if let Some(fh_addr) = link_info.fractal_heap_address {
for_each_dense_link(file_data, &link_info, fh_addr, os, ls, visit)?;
for_each_dense_link(file_data, &link_info, fh_addr, os, ls, hint_headers, visit)?;
} else {
for msg in &header.messages {
if msg.msg_type == MessageType::Link
@@ -737,7 +822,7 @@ fn resolve_group_entries<S: Storage + ?Sized>(
let stm = SymbolTableMessage::parse(&sym_msg.data, offset_size)?;
// A lookup: an entry with an empty name (which fails a listing) is
// skipped by the name comparison, as in libhdf5.
group_v1::v1_group_entries(file_data, &stm, offset_size, length_size)
group_v1::v1_group_entries(file_data, &stm, offset_size, length_size, false)
} else if is_v2_group(object_header) {
resolve_v2_group_entries_in(file_data, object_header, offset_size, length_size)
} else {
@@ -920,6 +1005,86 @@ mod tests {
assert_eq!(values, vec![22.5, 23.1, 21.8]);
}
/// `v1_groups_400.h5` from its superblock on (it has a user block).
fn v1_groups_400() -> (Vec<u8>, Superblock) {
let all: &[u8] = include_bytes!("../tests/fixtures/v1_groups_400.h5");
let data = all[signature::find_signature(all).unwrap()..].to_vec();
let sb = Superblock::parse(&data, 0).unwrap();
(data, sb)
}
/// Every child of a v1 group resolves by name, down the group's B-tree,
/// to the address the listing gives, reading a small part of what the
/// listing reads; a name the group does not hold is not found.
#[test]
fn v1_lookup_down_the_btree_agrees_with_the_listing() {
let (data, sb) = v1_groups_400();
let children = resolve_group_children(&data, &sb, sb.root_group_address).unwrap();
assert_eq!(children.len(), 401);
for c in &children {
let path = format!("/{}", c.name);
assert_eq!(
resolve_path_any(&data, &sb, &path).unwrap(),
c.object_header_address,
"{path}"
);
}
for missing in ["/g0400", "/a", "/g", "/g00000", "/zz", "/x0"] {
assert!(
matches!(
resolve_path_any(&data, &sb, missing),
Err(FormatError::PathNotFound(_))
),
"{missing}"
);
}
let st = crate::storage::CountingStorage::new(data.clone());
resolve_group_children_in(&st, &sb, sb.root_group_address).unwrap();
let listing = st.bytes_read();
st.reset();
let last = children.last().unwrap();
assert_eq!(
resolve_path_any_in(&st, &sb, &format!("/{}", last.name)).unwrap(),
last.object_header_address
);
assert!(
st.bytes_read() * 8 < listing,
"lookup read {} bytes, listing {listing}",
st.bytes_read()
);
}
/// A v1 group whose B-tree is out of name order (a name changed in the
/// heap so that it sorts past every key) is still looked up by reading
/// every entry, as before the lookup went down the B-tree.
#[test]
fn v1_lookup_falls_back_when_the_btree_is_out_of_order() {
let (mut data, sb) = v1_groups_400();
let at: Vec<usize> = data
.windows(6)
.enumerate()
.filter(|(_, w)| *w == b"g0200\0")
.map(|(i, _)| i)
.collect();
assert_eq!(at.len(), 1, "one heap string");
data[at[0]] = b'~';
let children = resolve_group_children(&data, &sb, sb.root_group_address).unwrap();
let moved = children.iter().find(|c| c.name == "~0200").unwrap();
assert_eq!(
resolve_path_any(&data, &sb, "/~0200").unwrap(),
moved.object_header_address
);
assert!(resolve_path_any(&data, &sb, "/g0200").is_err());
for c in &children {
let path = format!("/{}", c.name);
assert_eq!(
resolve_path_any(&data, &sb, &path).unwrap(),
c.object_header_address,
"{path}"
);
}
}
#[test]
fn path_not_found_v2() {
let file_data: &[u8] = include_bytes!("../tests/fixtures/v2_groups.h5");
+2
View File
@@ -61,6 +61,7 @@ pub mod addr;
pub mod attribute;
pub mod attribute_info;
pub mod btree_v1;
mod btree_v1_write;
pub mod btree_v2;
mod btree_v2_write;
mod bulk_alloc;
@@ -107,6 +108,7 @@ pub mod group_v1;
pub mod group_v2;
#[cfg(feature = "parallel")]
pub mod lane_partition;
pub mod libver;
pub mod link_info;
pub mod link_message;
pub mod local_heap;
+93
View File
@@ -0,0 +1,93 @@
//! Library version bounds for writing: which HDF5 releases can read a file.
//!
//! libhdf5 picks the version of every object it writes from the file's
//! *low* bound (`H5Pset_libver_bounds`; h5py's `libver=`): the oldest
//! format version that holds the object, but never older than the one the
//! low bound names. The *high* bound caps it: a feature that needs a newer
//! format than the high bound is an error. [`LibVer`] names the same
//! releases, and [`crate::file_writer::FileWriter::libver_bounds`] sets them.
//!
//! What the low bound changes in what clawhdf5 writes:
//!
//! | | low [`LibVer::V18`] | low [`LibVer::V110`] or later (the default) |
//! |---|---|---|
//! | superblock | version 2 | version 3 |
//! | data layout message | version 3 | version 4 |
//! | chunk index | version-1 B-tree (every chunked dataset) | single chunk, Fixed Array, Extensible Array or version-2 B-tree, as libhdf5 picks |
//!
//! Everything else (version-2 object headers, link and group-info messages,
//! dense storage in fractal heaps with version-2 B-trees, filter pipeline
//! version 2, fill value version 3, datatype versions up to 3) is the same
//! and already readable by HDF5 1.8.
//!
//! What the high bound refuses: anything that needs 1.10 (virtual datasets,
//! the paged file-space strategy) above [`LibVer::V18`], the 1.12 reference
//! types (datatype version 4) above [`LibVer::V110`], and HDF5 2.0's native
//! complex numbers (datatype version 5) above [`LibVer::V114`].
//! `libver_bounds(LibVer::V18, LibVer::V18)` therefore writes a file HDF5
//! 1.8 can read, or fails.
use core::fmt;
/// An HDF5 library release, as a bound on the file format versions a writer
/// may use (libhdf5's `H5F_libver_t`). Ordered oldest first.
///
/// There is no `Earliest`: clawhdf5 cannot write the pre-1.8 format
/// (symbol-table groups, version-1 object headers).
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash)]
#[non_exhaustive]
pub enum LibVer {
/// HDF5 1.8 (`H5F_LIBVER_V18`, h5py `'v108'`).
V18,
/// HDF5 1.10 (`H5F_LIBVER_V110`, h5py `'v110'`).
V110,
/// HDF5 1.12 (`H5F_LIBVER_V112`, h5py `'v112'`).
V112,
/// HDF5 1.14 (`H5F_LIBVER_V114`, h5py `'v114'`).
V114,
/// HDF5 2.0 (`H5F_LIBVER_V200`).
V200,
/// The newest format this build of clawhdf5 writes
/// (`H5F_LIBVER_LATEST`, h5py `'latest'`).
Latest,
}
impl LibVer {
/// The release a datatype message of this version first appeared in:
/// versions 1-3 are readable by HDF5 1.8, 4 needs 1.12, 5 needs 2.0.
pub(crate) fn for_datatype_version(version: u8) -> Self {
match version {
0..=3 => LibVer::V18,
4 => LibVer::V112,
_ => LibVer::V200,
}
}
}
impl fmt::Display for LibVer {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
f.write_str(match self {
LibVer::V18 => "1.8",
LibVer::V110 => "1.10",
LibVer::V112 => "1.12",
LibVer::V114 => "1.14",
LibVer::V200 => "2.0",
LibVer::Latest => "latest",
})
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn ordered_oldest_first() {
assert!(LibVer::V18 < LibVer::V110);
assert!(LibVer::V114 < LibVer::V200);
assert!(LibVer::V200 < LibVer::Latest);
assert_eq!(LibVer::for_datatype_version(3), LibVer::V18);
assert_eq!(LibVer::for_datatype_version(4), LibVer::V112);
assert_eq!(LibVer::for_datatype_version(5), LibVer::V200);
}
}
+42 -1
View File
@@ -163,6 +163,12 @@ impl ObjectHeader {
offset_size: u8,
length_size: u8,
) -> Result<ObjectHeader, FormatError> {
// The first chunk is read once the prefix says how long it is: say
// so (see `Storage::hint`), for a storage that fetches between
// attempts.
if file.as_contiguous().is_none() {
file.hint(offset, OBJECT_HEADER_HINT_LEN);
}
// The longest prefix of either version, in one read. It holds the
// whole prefix or ends at the end of the file, so its bounds checks
// are the whole-file ones.
@@ -279,10 +285,20 @@ impl ObjectHeader {
let mut spans = ChunkSpans::new(file.len(), offset, length)?;
let mut chunk0_count = 0usize;
let mut next = 0usize;
let hints = file.as_contiguous().is_none();
while let Some((chunk_offset, chunk_length)) = spans.get(next) {
let chunk = read_exact_at(file, chunk_offset, chunk_length)?;
let known = spans.len;
let count =
Self::parse_v1_messages(&chunk, offset_size, length_size, messages, &mut spans)?;
// The continuation chunks this one names are read next.
if hints {
for i in known..spans.len {
if let Some((o, l)) = spans.get(i) {
file.hint(o, l);
}
}
}
// Only the first chunk's messages are held to the prefix count.
if next == 0 {
chunk0_count = count;
@@ -295,7 +311,13 @@ impl ObjectHeader {
/// The messages of one version-1 chunk: each checked and appended to
/// `messages` (NIL ones dropped), each continuation added to `spans`.
/// Returns how many messages (NIL ones included) the chunk holds.
#[inline(never)]
///
/// Inlined into the chunk loop: kept out of line (`#[inline(never)]`,
/// 4313917), the call cost `ObjectHeader::parse` about 2.5 ns per
/// header, 4% on small version-1 headers (A/B builds, 2026-09-27; see
/// `BENCHMARKS.md`). Without an attribute the compiler keeps it out of
/// line too.
#[inline]
fn parse_v1_messages(
data: &[u8],
offset_size: u8,
@@ -477,8 +499,17 @@ impl ObjectHeader {
// whenever a message no longer fits), up to the same bound as a
// version-1 header.
let mut spans = ChunkSpans::new(file.len(), base as u64, chunk0_msg_end.saturating_add(4))?;
// The continuation chunks a chunk names are read next (see
// `Storage::hint`).
let hints = file.as_contiguous().is_none();
if hints {
for &(o, l) in &continuations {
file.hint(o as u64, l);
}
}
while let Some((cont_offset, cont_length)) = continuations.pop() {
spans.add(cont_offset as u64, cont_length)?;
let known = continuations.len();
Self::parse_v2_continuation(
file,
cont_offset as u64,
@@ -489,6 +520,11 @@ impl ObjectHeader {
&mut messages,
&mut continuations,
)?;
if hints {
for &(o, l) in &continuations[known..] {
file.hint(o as u64, l);
}
}
}
Ok(ObjectHeader {
@@ -634,6 +670,11 @@ impl ObjectHeader {
}
}
/// What an object header is hinted to take before its prefix is read (see
/// [`Storage::hint`]): the first chunk of a typical dataset's header. A
/// longer header is read all the same.
pub(crate) const OBJECT_HEADER_HINT_LEN: usize = 512;
/// Longest version-2 object header prefix: signature(4) + version(1) +
/// flags(1) + times(16) + attribute phase change(4) + chunk-0 size(8).
const V2_PREFIX_MAX: usize = 34;
+2 -2
View File
@@ -255,7 +255,7 @@ pub fn decompress_chunks_lane_partitioned_in<S: Storage + ?Sized>(
for &local in &indices {
let index = batch.start + local;
let chunk_info = &chunks[index];
let size = chunk_info.chunk_size as usize;
let size = crate::addr::saturating_usize(chunk_info.chunk_size);
let raw_chunk = raw_bytes.get(index, &reqs[index])?;
let decompressed = decompress_chunk_exact(
@@ -426,7 +426,7 @@ mod tests {
for i in 0..8u64 {
let len = if short && i == 5 { 16 } else { 32 };
infos.push(ChunkInfo {
chunk_size: len as u32,
chunk_size: len as u64,
filter_mask: 0,
offsets: vec![i * 8],
address: file.len() as u64,
+118 -1
View File
@@ -243,6 +243,58 @@ fn copy_overlap(
}
}
/// Copy the part of unfiltered chunk `chunk` (`chunk_bytes` long, shape
/// `chunk_shape`) that overlaps the box into `out`, reading only the runs of
/// the overlap from the file. The whole chunk must still lie inside the
/// file, as it must when it is fetched whole.
#[allow(clippy::too_many_arguments)]
fn read_unfiltered_overlap<S: Storage + ?Sized>(
file_data: &S,
chunk: &crate::chunked_read::ChunkInfo,
chunk_bytes: usize,
chunk_shape: &[u64],
elem_size: usize,
out: &mut [u8],
box_start: &[u64],
box_extent: &[u64],
) -> Result<(), FormatError> {
let rank = chunk_shape.len();
let origin = &chunk.offsets[..rank];
let (lo, extent): (Vec<u64>, Vec<u64>) = (0..rank)
.map(|d| {
let lo = origin[d].max(box_start[d]);
let hi = origin[d]
.saturating_add(chunk_shape[d])
.min(box_start[d] + box_extent[d]);
(lo, hi.saturating_sub(lo))
})
.unzip();
let file_len = crate::storage::len_usize(file_data);
let base = crate::addr::to_usize(chunk.address)?;
if base > file_len || chunk_bytes > file_len - base {
return Err(FormatError::UnexpectedEof {
expected: base.saturating_add(chunk_bytes),
available: file_len,
});
}
let overlap = Selection::Hyperslab {
start: lo.iter().zip(origin).map(|(l, o)| l - o).collect(),
stride: vec![1; rank],
count: extent.clone(),
block: vec![1; rank],
};
let rows = crate::gather::gather_storage(
file_data,
chunk.address,
chunk_bytes,
chunk_shape,
elem_size,
&overlap,
)?;
copy_overlap(&rows, &lo, &extent, out, box_start, box_extent, elem_size);
Ok(())
}
/// Read `selection` without materialising the whole dataset, when that is
/// possible and worthwhile. `Ok(None)` means "use the full-read path": an
/// `All`/`None`/invalid selection, a layout this doesn't handle (compact,
@@ -281,9 +333,45 @@ pub fn read_selection_in<S: Storage + ?Sized>(
offset_size: u8,
length_size: u8,
selection: &Selection,
) -> Result<Option<Vec<u8>>, FormatError> {
read_selection_filled_in(
file_data,
layout,
dataspace,
elem_size,
pipeline,
offset_size,
length_size,
selection,
None,
)
}
/// [`read_selection_in`] for a dataset whose fill value is `fill` (one
/// element's bytes; `None` or all zeros is the default fill): the elements
/// of a chunked dataset's selection that lie in chunks never written read as
/// `fill`, as they do in a full read. Only the chunks the selection's
/// bounding box overlaps are read, so a selection of a few elements of a
/// dataset whose chunks are 4 GiB or more costs one decoded chunk (a
/// filtered chunk has to be decoded whole) or, unfiltered, only the bytes
/// it selects.
#[allow(clippy::too_many_arguments)]
pub fn read_selection_filled_in<S: Storage + ?Sized>(
file_data: &S,
layout: &DataLayout,
dataspace: &Dataspace,
elem_size: usize,
pipeline: Option<&FilterPipeline>,
offset_size: u8,
length_size: u8,
selection: &Selection,
fill: Option<&[u8]>,
) -> Result<Option<Vec<u8>>, FormatError> {
let dims = &dataspace.dimensions;
if dims.is_empty() || elem_size == 0 {
// A fill value that is not one element's bytes is the full path's to
// interpret.
let odd_fill = fill.is_some_and(|f| !f.is_empty() && f.len() != elem_size);
if dims.is_empty() || elem_size == 0 || odd_fill {
return Ok(None);
}
let total = dataspace.checked_num_elements()?;
@@ -341,8 +429,18 @@ pub fn read_selection_in<S: Storage + ?Sized>(
return Ok(None);
}
let mut boxed = alloc_output(checked_byte_len(box_elements, elem_size)?)?;
if let Some(fill) = fill.filter(|f| <[u8]>::len(f) == elem_size && f.iter().any(|&b| b != 0)) {
for element in boxed.chunks_exact_mut(elem_size) {
element.copy_from_slice(fill);
}
}
match layout {
// No chunk was ever written: every element is the fill value.
DataLayout::Chunked {
btree_address: None,
..
} if fill.is_some() => {}
DataLayout::Chunked {
btree_address: Some(_),
..
@@ -373,6 +471,25 @@ pub fn read_selection_in<S: Storage + ?Sized>(
})
})
.collect();
// An unfiltered chunk of a file that is not in memory: fetch
// only the rows the box needs, not the whole chunk (which may be
// 4 GiB or more).
let (direct, wanted): (Vec<_>, Vec<_>) = wanted.into_iter().partition(|c| {
file_data.as_contiguous().is_none()
&& pipeline.is_none_or(|pl| all_filters_skipped(pl, c.filter_mask))
});
for chunk in direct {
read_unfiltered_overlap(
file_data,
chunk,
chunk_bytes,
&chunk_shape,
elem_size,
&mut boxed,
&box_start,
&box_extent,
)?;
}
// Their stored bytes, batch by batch when the file is not in
// memory; each batch's chunks are decoded into this thread's
// reusable buffers before the next batch is fetched.
+30
View File
@@ -89,6 +89,21 @@ pub trait Storage {
fn as_contiguous(&self) -> Option<&[u8]> {
None
}
/// A hint that `[offset, offset + len)` is about to be read by the
/// same operation: a parser that has just learnt where the next
/// structures are (a node's children, a structure's body) says so
/// before it reads them one at a time. Nothing is read and nothing
/// fails. The default ignores it, as does every backend that reads
/// when asked; the browser's restartable reader, which fetches over
/// the network between attempts, fetches hinted bytes along with the
/// bytes an attempt actually missed, so structures a parser only
/// reaches after a miss arrive in the same round trip. Results never
/// depend on hints.
#[inline]
fn hint(&self, offset: u64, len: usize) {
let _ = (offset, len);
}
}
impl Storage for [u8] {
@@ -148,6 +163,11 @@ impl<T: Storage + ?Sized> Storage for &T {
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
#[inline]
fn hint(&self, offset: u64, len: usize) {
(**self).hint(offset, len)
}
}
impl<T: Storage + ?Sized> Storage for Box<T> {
@@ -170,6 +190,11 @@ impl<T: Storage + ?Sized> Storage for Box<T> {
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
#[inline]
fn hint(&self, offset: u64, len: usize) {
(**self).hint(offset, len)
}
}
#[cfg(feature = "std")]
@@ -193,6 +218,11 @@ impl<T: Storage + ?Sized> Storage for std::sync::Arc<T> {
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
#[inline]
fn hint(&self, offset: u64, len: usize) {
(**self).hint(offset, len)
}
}
/// `storage.len()` as the `usize` the parsers' end-of-file errors report
@@ -90,6 +90,10 @@ impl SymbolTableNode {
offset: u64,
offset_size: u8,
) -> Result<SymbolTableNode, FormatError> {
// The entries are read once the header says how many there are:
// say so (see `Storage::hint`), for libhdf5's default node size
// (group leaf K = 4: 8 entries of 40 bytes with 8-byte offsets).
file.hint(offset, 8 + 8 * (2 * usize::from(offset_size) + 24));
// signature(4) + version(1) + reserved(1) + number_of_symbols(2) = 8
let header = read_exact_at(file, offset, 8)?;
@@ -137,6 +137,37 @@ pub fn make_f32_type() -> Datatype {
}
}
/// numpy `complex64` the way h5py stores it: a compound `{r: f32, i: f32}`.
/// Every HDF5 reader opens it; h5py reads it back as `complex64`.
pub fn make_complex_f32_type() -> Datatype {
Datatype::complex_as_compound(8, &make_f32_type())
}
/// numpy `complex128` the way h5py stores it: a compound `{r: f64, i: f64}`.
pub fn make_complex_f64_type() -> Datatype {
Datatype::complex_as_compound(16, &make_f64_type())
}
/// HDF5 2.0's native complex type `H5T_COMPLEX_IEEE_F32LE` (datatype class
/// 11). Only libhdf5 2.0 and newer (h5py built on it) can read a file that
/// uses it; older libhdf5, including h5dump 1.14, refuses the object.
/// Prefer [`make_complex_f32_type`] unless the consumer wants class 11.
pub fn make_native_complex_f32_type() -> Datatype {
Datatype::Complex {
size: 8,
base_type: Box::new(make_f32_type()),
}
}
/// HDF5 2.0's native complex type `H5T_COMPLEX_IEEE_F64LE` (datatype class
/// 11); see [`make_native_complex_f32_type`] for who can read it.
pub fn make_native_complex_f64_type() -> Datatype {
Datatype::Complex {
size: 16,
base_type: Box::new(make_f64_type()),
}
}
pub fn make_i32_type() -> Datatype {
Datatype::FixedPoint {
size: 4,
@@ -640,6 +671,70 @@ impl DatasetBuilder {
self
}
/// Store complex numbers, each `[re, im]`, as h5py does for numpy
/// `complex64`: a compound `{r, i}` of `f32` ([`make_complex_f32_type`]),
/// readable by every HDF5 library. For HDF5 2.0's native complex type use
/// [`Self::with_native_complex_f32_data`].
pub fn with_complex_f32_data(&mut self, data: &[[f32; 2]]) -> &mut Self {
self.set_complex(
make_complex_f32_type(),
data.as_flattened(),
f32::to_le_bytes,
)
}
/// Store complex numbers, each `[re, im]`, as h5py does for numpy
/// `complex128`: a compound `{r, i}` of `f64` ([`make_complex_f64_type`]).
pub fn with_complex_f64_data(&mut self, data: &[[f64; 2]]) -> &mut Self {
self.set_complex(
make_complex_f64_type(),
data.as_flattened(),
f64::to_le_bytes,
)
}
/// Store complex numbers, each `[re, im]`, as HDF5 2.0's native complex
/// type `H5T_COMPLEX_IEEE_F32LE` (datatype class 11). The bytes are the
/// same as [`Self::with_complex_f32_data`]; only the datatype differs.
/// h5py on libhdf5 2.0+ reads it as `complex64`; libhdf5 1.x cannot open
/// the dataset at all, so this is opt-in.
pub fn with_native_complex_f32_data(&mut self, data: &[[f32; 2]]) -> &mut Self {
self.set_complex(
make_native_complex_f32_type(),
data.as_flattened(),
f32::to_le_bytes,
)
}
/// Store complex numbers, each `[re, im]`, as HDF5 2.0's native complex
/// type `H5T_COMPLEX_IEEE_F64LE` (class 11); see
/// [`Self::with_native_complex_f32_data`].
pub fn with_native_complex_f64_data(&mut self, data: &[[f64; 2]]) -> &mut Self {
self.set_complex(
make_native_complex_f64_type(),
data.as_flattened(),
f64::to_le_bytes,
)
}
fn set_complex<T: Copy, const N: usize>(
&mut self,
datatype: Datatype,
parts: &[T],
le: fn(T) -> [u8; N],
) -> &mut Self {
self.datatype = Some(datatype);
let mut b = Vec::with_capacity(parts.len() * N);
for &v in parts {
b.extend_from_slice(&le(v));
}
self.data = Some(b);
if self.shape.is_none() {
self.shape = Some(vec![(parts.len() / 2) as u64]);
}
self
}
/// Write a compound (struct) dataset.
pub fn with_compound_data(
&mut self,
@@ -151,7 +151,7 @@ fn crafted() -> (Vec<u8>, Chunked, Vec<ChunkInfo>) {
// v1 B-tree key (size, filter mask, offsets + 0) then the child
// address.
let mut pat = Vec::new();
pat.extend_from_slice(&c.chunk_size.to_le_bytes());
pat.extend_from_slice(&(c.chunk_size as u32).to_le_bytes());
pat.extend_from_slice(&c.filter_mask.to_le_bytes());
// The key holds one offset per dimension plus the element offset
// (0); `offsets` may or may not list that last one.
@@ -173,7 +173,7 @@ fn crafted() -> (Vec<u8>, Chunked, Vec<ChunkInfo>) {
assert!(
chunks
.iter()
.all(|c| c.chunk_size == HUGE && c.address == blob)
.all(|c| c.chunk_size == u64::from(HUGE) && c.address == blob)
);
(bytes, ds, chunks)
}
@@ -313,7 +313,7 @@ fn large_reads_are_fetched_in_batches() {
let data: Vec<u8> = (0..2 * CHUNK).map(|i| (i % 251) as u8).collect();
let chunks: Vec<ChunkInfo> = (0..40u64)
.map(|i| ChunkInfo {
chunk_size: CHUNK as u32,
chunk_size: CHUNK as u64,
filter_mask: 0,
offsets: vec![i * CHUNK as u64],
address: (i % 2) * CHUNK as u64,
@@ -383,6 +383,169 @@ else:
assert_eq!((fields[1].name.as_str(), im), ("i", vec![2.0, 4.0]));
}
/// Complex data both ways we write it: h5py's compound `{r, i}` (the
/// default, readable everywhere) and HDF5 2.0's native complex type (class
/// 11, opt-in). h5py on libhdf5 2.0+ must read the native datasets and
/// attributes as numpy `complex64`/`complex128`, and a compound holding a
/// native complex member (written as datatype version 5, as libhdf5 does).
/// Skips the native checks when h5py's libhdf5 predates 2.0; also runs
/// h5dump 2.x when `CLAWHDF5_H5DUMP2` names one.
#[test]
#[ignore = "requires Python h5py module"]
fn h5py_reads_our_complex_datasets_and_attributes() {
use clawhdf5_format::type_builders::{
make_complex_f64_type, make_native_complex_f32_type, make_native_complex_f64_type,
};
let path = std::env::temp_dir().join("clawhdf5_test_complex.h5");
let native_ok = h5py_read(
&path,
"import h5py; print(int(getattr(h5py.get_config(), 'has_native_complex', False)))",
) == "1";
let z64 = [[1.5f32, -2.0], [0.0, 3.25], [-7.0, 1.0e-3]];
let z128 = [[1.0f64, 2.0], [-3.5, 4.0e300], [0.25, -0.0]];
let big: Vec<[f64; 2]> = (0..600).map(|k| [k as f64, -(k as f64) / 4.0]).collect();
let c128 = |re: f64, im: f64| [re.to_le_bytes(), im.to_le_bytes()].concat();
let mut fw = FileWriter::new();
fw.create_dataset("compound128")
.with_complex_f64_data(&z128)
.set_attr(
"c",
AttrValue::Raw {
datatype: make_complex_f64_type(),
shape: vec![],
data: c128(1.0, -1.0),
},
);
if native_ok {
fw.create_dataset("native64")
.with_native_complex_f32_data(&z64)
.set_attr(
"c",
AttrValue::Raw {
datatype: make_native_complex_f64_type(),
shape: vec![2],
data: [c128(0.5, -1.5), c128(2.0, 3.0)].concat(),
},
);
fw.create_dataset("native128")
.with_native_complex_f64_data(&z128);
fw.create_dataset("native_chunked")
.with_native_complex_f64_data(&big)
.with_shape(&[20, 30])
.with_chunks(&[7, 16])
.with_deflate(4);
// A compound with a native complex member, and an array of them.
let rec = CompoundTypeBuilder::new()
.field("z", make_native_complex_f64_type())
.i64_field("k")
.build();
let raw = [
c128(1.0, 2.0),
7i64.to_le_bytes().to_vec(),
c128(-1.0, 0.5),
(-8i64).to_le_bytes().to_vec(),
]
.concat();
fw.create_dataset("records").with_compound_data(rec, raw, 2);
let arr: Vec<u8> = [1.0f32, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]
.iter()
.flat_map(|v| v.to_le_bytes())
.collect();
fw.create_dataset("arrays")
.with_array_data(make_native_complex_f32_type(), &[2], arr, 2);
fw.set_root_attr(
"zroot",
AttrValue::Raw {
datatype: make_native_complex_f32_type(),
shape: vec![],
data: [4.0f32.to_le_bytes(), (-4.0f32).to_le_bytes()].concat(),
},
);
}
std::fs::write(&path, fw.finish().unwrap()).unwrap();
let script = format!(
r#"
import h5py, json, numpy as np
f = h5py.File('{}', 'r')
def z(a): return [[float(np.real(v)), float(np.imag(v))] for v in np.asarray(a).ravel()]
out = {{}}
d = f['compound128']
out['compound128'] = [str(d.dtype), z(d[()]), str(d.attrs['c'].dtype), z(d.attrs['c'])]
if {native}:
for n in ['native64', 'native128', 'native_chunked']:
d = f[n]
out[n] = [str(d.dtype), list(d.shape), z(d[()]), d.id.get_type().get_class()]
a = f['native64'].attrs['c']
out['attr'] = [str(a.dtype), z(a)]
r = f['records'][()]
out['records'] = [str(r.dtype['z']), z(r['z']), r['k'].tolist()]
a = f['arrays'][()]
out['arrays'] = [str(a.dtype), list(a.shape), z(a)]
a = f.attrs['zroot']
out['zroot'] = [str(a.dtype), z(a)]
out['CLASS'] = h5py.h5t.COMPLEX
print(json.dumps(out))
"#,
path.display(),
native = if native_ok { "True" } else { "False" },
);
let v: serde_json::Value = serde_json::from_str(&h5py_read(&path, &script)).unwrap();
let pairs = |p: &[[f64; 2]]| serde_json::json!(p);
assert_eq!(v["compound128"][0], "complex128");
assert_eq!(v["compound128"][1], pairs(&z128));
assert_eq!(v["compound128"][2], "complex128");
assert_eq!(v["compound128"][3], pairs(&[[1.0, -1.0]]));
if !native_ok {
eprintln!("HDF5 < 2.0: native complex not checked");
return;
}
let z64_wide: Vec<[f64; 2]> = z64.iter().map(|p| [p[0].into(), p[1].into()]).collect();
let class_complex = v["CLASS"].clone();
for (name, dtype, shape, values) in [
("native64", "complex64", vec![3], z64_wide.clone()),
("native128", "complex128", vec![3], z128.to_vec()),
("native_chunked", "complex128", vec![20, 30], big.clone()),
] {
assert_eq!(v[name][0], dtype, "{name}");
assert_eq!(v[name][1], serde_json::json!(shape), "{name}");
assert_eq!(v[name][2], pairs(&values), "{name}");
// Stored as class 11, not converted from a compound.
assert_eq!(v[name][3], class_complex, "{name}");
}
assert_eq!(v["attr"][0], "complex128");
assert_eq!(v["attr"][1], pairs(&[[0.5, -1.5], [2.0, 3.0]]));
assert_eq!(v["records"][0], "complex128");
assert_eq!(v["records"][1], pairs(&[[1.0, 2.0], [-1.0, 0.5]]));
assert_eq!(v["records"][2], serde_json::json!([7, -8]));
assert_eq!(v["arrays"][0], "complex64");
assert_eq!(v["arrays"][1], serde_json::json!([2, 2]));
assert_eq!(
v["arrays"][2],
pairs(&[[1.0, 2.0], [3.0, 4.0], [5.0, 6.0], [7.0, 8.0]])
);
assert_eq!(v["zroot"][0], "complex64");
assert_eq!(v["zroot"][1], pairs(&[[4.0, -4.0]]));
// h5dump from libhdf5 2.x, when available (Debian's 1.14 cannot read
// class 11 at all).
if let Ok(h5dump) = std::env::var("CLAWHDF5_H5DUMP2") {
let o = std::process::Command::new(&h5dump)
.arg(&path)
.output()
.expect("run h5dump");
let text = String::from_utf8_lossy(&o.stdout);
assert!(
o.status.success(),
"h5dump: {}",
String::from_utf8_lossy(&o.stderr)
);
assert!(text.contains("H5T_COMPLEX"), "{text}");
}
}
#[test]
#[ignore = "requires Python h5py module"]
fn read_h5py_generated_enum() {
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-gpu"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "GPU-accelerated vector operations for rustyhdf5 using wgpu compute shaders"
description = "GPU vector distance computation for clawhdf5 using wgpu compute shaders (not HDF5 I/O)"
license = "MIT"
repository = "https://git.redclaw.dev/quantumclaw/clawhdf5"
readme = "README.md"
+47 -10
View File
@@ -1,25 +1,62 @@
# clawhdf5-gpu
[![crates.io](https://img.shields.io/crates/v/clawhdf5-gpu.svg)](https://crates.io/crates/clawhdf5-gpu)
[![docs.rs](https://docs.rs/clawhdf5-gpu/badge.svg)](https://docs.rs/clawhdf5-gpu)
GPU vector distance computation through [wgpu](https://wgpu.rs) and
hand-written WGSL compute shaders: upload a set of vectors once, then run
cosine or L2 top-k searches, dot products, distance matrices and norms
against them on Vulkan, Metal, DirectX 12 or OpenGL.
GPU-accelerated vector operations for clawhdf5 using wgpu compute shaders.
This crate does **not** read or write HDF5: dataset I/O in clawhdf5 is
CPU-only. It is a vector-search accelerator used optionally by
[`clawhdf5-agent`](../clawhdf5-agent/README.md) (its `gpu` feature exposes
`gpu_search::GpuSearchBackend` and a GPU arm of `strategy::search_with_metrics`;
`HDF5Memory::search` itself uses the HNSW index on the CPU).
## Features
Not on crates.io yet; depend on it from git:
- GPU-accelerated distance computations (L2, cosine)
- wgpu-based compute shaders for cross-platform GPU support
- Float16 support via `half` crate
```toml
[dependencies]
clawhdf5-gpu = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
```rust,no_run
use clawhdf5_gpu::GpuAccelerator;
let accel = GpuAccelerator::new().unwrap();
let distances = accel.l2_distances(&query, &vectors).unwrap();
// Fall back to a CPU path when there is no usable GPU.
let mut gpu = match GpuAccelerator::new() {
Ok(g) => g,
Err(_) => return,
};
let dim = 128;
let vectors = vec![0.5f32; 1000 * dim]; // 1000 vectors, row-major
gpu.upload_vectors(&vectors, dim).unwrap();
let norms = gpu.compute_norms_gpu(&vectors, dim).unwrap();
gpu.upload_norms(&norms).unwrap();
let query = vec![1.0f32; dim];
let top10 = gpu.cosine_search(&query, 10).unwrap(); // (index, similarity), best first
let near10 = gpu.l2_search(&query, 10).unwrap(); // (index, distance), nearest first
```
`GpuAccelerator` also has `is_available`, `device_info`,
`batch_cosine_search`, `batch_dot_product`, `distance_matrix`,
`compute_norms`, and `f16_to_f32_batch`/`f32_to_f16_batch`. Vector sets
larger than the device's largest storage buffer binding are split into
chunks and the results merged. A GPU→CPU readback waits at most 30 s and
then fails with `GpuError::BufferMap` instead of hanging.
## Features
| Feature | Default | What |
|---|---|---|
| `gpu-wgpu` | yes | the wgpu implementation. Without it `GpuAccelerator::new()` returns `GpuError::NotCompiled` and `is_available()` is `false`. |
No C is compiled, but wgpu talks to the system's graphics drivers at run
time; the crate is exempt from CI's "no C in the default build" check for
that reason.
## License
MIT
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-io"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "I/O abstraction layer for rustyhdf5"
description = "I/O adapters for clawhdf5 (buffers, mmap, prefetch)"
license = "MIT"
repository = "https://git.redclaw.dev/quantumclaw/clawhdf5"
readme = "README.md"
+45 -15
View File
@@ -1,24 +1,54 @@
# clawhdf5-io
[![crates.io](https://img.shields.io/crates/v/clawhdf5-io.svg)](https://crates.io/crates/clawhdf5-io)
[![docs.rs](https://docs.rs/clawhdf5-io/badge.svg)](https://docs.rs/clawhdf5-io)
I/O building blocks under [`clawhdf5`](../clawhdf5/README.md): the
`HDF5Read`/`HDF5ReadWrite` traits with in-memory, borrowed, file and
memory-mapped readers, plus several experimental modules (async reads, an
HSDS client, a VOL-style trait, sub-filing, prefetch, and an MPI connector).
The facade uses it for memory-mapped reads (`MmapReader`, and the private
copy-on-write mapping that applies a metadata cache image).
I/O abstraction layer for clawhdf5.
Remote files are **not** read through this crate: HTTP(S) and object
stores go through `clawhdf5_format::storage::Storage` and
[`clawhdf5-remote`](../clawhdf5-remote/README.md).
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-io = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = ["mmap"] }
```
## Main items
| Item | What |
|---|---|
| `HDF5Read`, `HDF5ReadWrite` | byte-level read/write traits; `MemoryReader`, `BorrowedReader`, `FileReader`, `FileWriter` implement them |
| `MmapReader`, `MmapReadWrite` (`mmap`) | memory-mapped files through `memmap2`; `HDF5Read::private_copy` gives a copy-on-write view |
| `prefetch::PrefetchReader`, `prefetch::SweepDetector`, `sweep` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction |
| `ParallelConfig` | lane partitioning for parallel chunk decoding |
| `vol::VirtualObjectLayer`, `vol::NativeVol` | a backend-agnostic object-layer trait (modelled on libhdf5's VOL) |
| `async_read` (`async`) | tokio-based `AsyncHDF5Read` and `AsyncHDF5File` |
| `hsds::HsdsClient` (`hsds`) | a REST client for an HSDS server |
| `subfiling` | splitting one logical file across several physical files |
| `mpi_vol::MpiVol` (`mpi-io`) | an MPI connector: see below |
### MPI (`mpi-io`)
`MpiVol` is **not collective MPI-IO**. Reads are root-read + broadcast
(rank 0 reads the file with `std::fs::read`, parses the dataset and
broadcasts the bytes); writes gather every rank's shard to rank 0, which
writes the merged dataset. It does not call `MPI_File_read_at_all` or any
other MPI-IO routine. Collective I/O is on the [roadmap](../../ROADMAP.md).
`clawhdf5-bench`'s `mpi_io_bench` binary exercises it.
## Features
- Memory-mapped file access (`mmap` feature)
- Async I/O via Tokio (`async` feature)
- HSDS remote access (`hsds` feature)
- Prefetching and sweep optimizations
## Usage
```rust
use clawhdf5_io::MmapReader;
let reader = MmapReader::open("data.h5").unwrap();
```
| Feature | Default | What | Builds C |
|---|---|---|---|
| `mmap` | no (the `clawhdf5` facade turns it on) | `MmapReader`, `MmapReadWrite` | no |
| `async` | no | `async_read` (tokio) | no |
| `hsds` | no | `hsds` (reqwest, and `async`) | yes: reqwest's default TLS is native-tls (OpenSSL on Linux) |
| `mpi-io` | no | a real `MpiVol` (without it `MpiVol::new_world` returns an error) | yes: `mpi-sys` needs an MPI installation and libclang |
## License
+7 -5
View File
@@ -1,11 +1,8 @@
# clawhdf5-migrate
[![crates.io](https://img.shields.io/crates/v/clawhdf5-migrate.svg)](https://crates.io/crates/clawhdf5-migrate)
[![docs.rs](https://img.shields.io/docsrs/clawhdf5-migrate)](https://docs.rs/clawhdf5-migrate)
CLI tool to migrate a SQLite agent-memory database in the `memory_chunks` / `sessions` / `entities` / `relations` layout (table and
column names are configurable) to a
[clawhdf5-agent](https://crates.io/crates/clawhdf5-agent) store. This is **not**
[clawhdf5-agent](../clawhdf5-agent/README.md) store. This is **not**
ZeroClaw's schema — ZeroClaw keeps memories in a single `memories` table and
does not use clawhdf5.
@@ -15,10 +12,15 @@ the knowledge graph (entities and relations) are carried over.
## Installation
Not on crates.io yet; install from a checkout:
```bash
cargo install clawhdf5-migrate
cargo install --path crates/clawhdf5-migrate
```
It builds C: `rusqlite` is built with its `bundled` feature, which compiles
SQLite (so no system libsqlite is needed, but a C compiler is).
## Usage
```bash
+38
View File
@@ -0,0 +1,38 @@
# clawhdf5-napi
> **Status: does not work end to end.** The TypeScript package built on
> this crate (`packages/clawhdf5-node`) has never run successfully, is
> unpublished, and is not built or tested in CI. See "The Node.js package
> does not work" in [`docs/known-issues.md`](../../docs/known-issues.md).
> Fix it and add CI, or remove it, before depending on it.
A Node.js native addon ([napi-rs](https://napi.rs), N-API 9) exposing
[`clawhdf5-agent`](../clawhdf5-agent/README.md) as a `ClawhdfMemory`
class. It wraps `clawhdf5_agent::openclaw::ClawhdfBackend` (the agent's
`search` with re-ranking and confidence on) and the consolidation engine.
It was written for an OpenClaw integration that is not being pursued
([`docs/openclaw.md`](../../docs/openclaw.md)).
## What the addon exposes
`ClawhdfMemory.create(path, dim)`, `.open(path)`, `.openOrCreate(path,
dim)`, and on an instance: `search`, `get`, `write`, `ingestMarkdown`,
`exportMarkdown`, `save`, `saveBatch`, `stats`, `compact`, `tickSession`,
`flushWal`, `walPendingCount`, `runConsolidation`, and the ephemeral tier
(`enableEphemeral`, `ephemeralSet`/`Get`/`Delete`, `ephemeralStats`,
`promoteEphemeral`). napi-rs converts names and `#[napi(object)]` fields to
camelCase.
## Build
```bash
cargo build --release -p clawhdf5-napi # the Rust cdylib
# the .node package: npm install -g @napi-rs/cli; cd packages/clawhdf5-node; napi build --platform --release
```
It links against Node's N-API through `napi-sys` (a `-sys` crate), so it is
exempt from CI's "no C in the default build" check.
## License
MIT
+1 -1
View File
@@ -3,7 +3,7 @@ name = "clawhdf5-netcdf4"
version = "2.7.0"
edition = "2024"
rust-version.workspace = true
description = "NetCDF-4 read support built on rustyhdf5 — pure Rust, no C dependencies"
description = "NetCDF-4 read support built on clawhdf5 — pure Rust, no C dependencies"
license = "MIT"
repository = "https://git.redclaw.dev/quantumclaw/clawhdf5"
readme = "README.md"
+57 -11
View File
@@ -1,25 +1,71 @@
# clawhdf5-netcdf4
[![crates.io](https://img.shields.io/crates/v/clawhdf5-netcdf4.svg)](https://crates.io/crates/clawhdf5-netcdf4)
[![docs.rs](https://docs.rs/clawhdf5-netcdf4/badge.svg)](https://docs.rs/clawhdf5-netcdf4)
Read NetCDF-4 files in pure Rust. NetCDF-4 files are HDF5 files with
conventions for dimensions, coordinate variables and attributes; this crate
reads them through the [`clawhdf5`](../clawhdf5/README.md) facade, with no
libnetcdf or libhdf5. Read-only: NetCDF-3 (classic) files are not HDF5 and
are not supported.
NetCDF-4 read support built on clawhdf5 — pure Rust, no C dependencies.
Not on crates.io yet; depend on it from git:
## Features
- Read NetCDF-4 / HDF5-backed `.nc` files
- Dimension, variable, and CF convention support
- Climate and scientific data access
```toml
[dependencies]
clawhdf5-netcdf4 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
```rust,no_run
use clawhdf5_netcdf4::NetCDF4File;
let nc = NetCDF4File::open("climate.nc").unwrap();
let temp = nc.variable("temperature").unwrap();
let nc = NetCDF4File::open("climate.nc")?;
for dim in nc.dimensions()? {
println!("{}: {} (unlimited = {})", dim.name, dim.size, dim.is_unlimited);
}
let mut temp = nc.variable("temperature")?;
let dims: Vec<&str> = temp.dimensions().iter().map(|d| d.name.as_str()).collect();
println!("{:?} over {:?}", temp.shape()?, dims);
let cf = temp.cf_attributes()?;
println!("units: {:?}", cf.units);
// scale_factor/add_offset applied; _FillValue and missing_value become NaN
let values: Vec<f64> = temp.read_f64()?;
# Ok::<(), clawhdf5_netcdf4::Error>(())
```
## API
| Item | What |
|---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable |
Variables and dimensions follow netCDF-C:
- A variable's dimensions are the ones the file names: the ids in its
`_Netcdf4Coordinates` attribute, else the dimension scales its
`DIMENSION_LIST` references, found in its group or a parent group. Only
an axis the file names no dimension for (an HDF5 file not written by a
netCDF library) gets the first dimension of the group of the same size,
else an anonymous `dim_<size>`.
- Dimension scales that are only dimensions are not variables; a dataset
`_nc4_non_coord_<name>` is the variable `<name>`.
- A variable along an unlimited dimension has the dimension's length:
`shape` is that length, and the reads return that many values, the
records the variable has not written as its fill value (`_FillValue`,
else netCDF's default for the type; NaN from `read_f64`).
`stored_shape` is the HDF5 dataset's extent.
No cargo features. Tests compare against files written by netCDF4-python,
h5py dimension scales, h5netcdf and xarray, variable by variable with what
netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them
with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is
not installed). What the HDF5 reader underneath cannot read is listed in
[`docs/known-issues.md`](../../docs/known-issues.md).
## License
MIT
+230 -95
View File
@@ -2,7 +2,8 @@
//!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited).
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length
//! is the largest extent of the variables attached to them.
use std::collections::HashMap;
@@ -21,73 +22,255 @@ pub struct Dimension {
pub is_unlimited: bool,
}
/// Extract dimensions from an HDF5 group (root or subgroup).
/// A dimension scale of one group: the dataset that defines a dimension.
#[derive(Debug, Clone)]
pub(crate) struct Scale {
/// Object header address of the scale's dataset (what a variable's
/// `DIMENSION_LIST` references).
pub address: u64,
/// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists).
pub dimid: Option<i64>,
/// Index of its dimension in [`GroupDims::dims`].
pub dim: usize,
}
/// The dimensions a group defines, with the scales that define them.
#[derive(Debug, Clone, Default)]
pub(crate) struct GroupDims {
/// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order).
pub dims: Vec<Dimension>,
/// The dimension scales behind `dims`; empty when the group has no
/// dimension scale and `dims` were inferred from 1-D datasets.
pub scales: Vec<Scale>,
}
impl GroupDims {
/// The dimension defined by the scale at `address`.
pub fn by_address(&self, address: u64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.address == address)
.map(|s| &self.dims[s.dim])
}
/// The dimension whose scale has `_Netcdf4Dimid` `id`.
pub fn by_dimid(&self, id: i64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.dimid == Some(id))
.map(|s| &self.dims[s.dim])
}
}
/// The dimensions of an HDF5 group (root or subgroup).
///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. The dimension
/// size is the dataset's first (and typically only) shape extent. Unlimited dimensions
/// have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace.
pub(crate) fn extract_dimensions(
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`. A group with no dimension
/// scale at all (not written by a netCDF library) gets one dimension per
/// 1-D dataset instead.
pub(crate) fn group_dims(
file: &clawhdf5::File,
group: &clawhdf5::Group<'_>,
) -> Result<Vec<Dimension>, Error> {
) -> Result<GroupDims, Error> {
let addresses: HashMap<String, u64> = group.entries()?.into_iter().collect();
let dataset_names = group.datasets()?;
let mut dims = Vec::new();
let mut seen_dimids: HashMap<i64, usize> = HashMap::new();
// (dimid, dimension, scale address), in discovery order.
let mut found: Vec<(Option<i64>, Dimension, u64)> = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let attrs = ds.attrs()?;
// Check if this is a dimension scale
if !is_dimension_scale(&attrs) {
continue;
}
let Some(&address) = addresses.get(ds_name) else {
continue;
};
let shape = ds.shape()?;
let size = shape.first().copied().unwrap_or(0);
let is_unlimited = check_unlimited(file, group, ds_name);
let dimid = get_dimid(&attrs);
let is_unlimited = is_unlimited(&ds);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dim = Dimension {
name: ds_name.clone(),
size,
is_unlimited,
};
if let Some(id) = dimid {
seen_dimids.insert(id, dims.len());
}
dims.push(dim);
found.push((get_dimid(&attrs), dim, address));
}
// Sort by dimid if available, otherwise keep discovery order
if !seen_dimids.is_empty() {
let mut pairs: Vec<(i64, Dimension)> = Vec::new();
let mut unordered = Vec::new();
for (i, dim) in dims.into_iter().enumerate() {
let id = seen_dimids
.iter()
.find(|(_, idx)| **idx == i)
.map(|(k, _)| *k);
if let Some(id) = id {
pairs.push((id, dim));
} else {
unordered.push(dim);
if found.is_empty() {
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let mut dims = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited: is_unlimited(&ds),
});
}
}
pairs.sort_by_key(|(id, _)| *id);
dims = pairs.into_iter().map(|(_, d)| d).collect();
dims.extend(unordered);
return Ok(GroupDims {
dims,
scales: Vec::new(),
});
}
Ok(dims)
// By dimid; scales without one keep their discovery order after those
// with one (the sort is stable).
found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0)));
let mut out = GroupDims::default();
for (i, (dimid, dim, address)) in found.into_iter().enumerate() {
out.dims.push(dim);
out.scales.push(Scale {
address,
dimid,
dim: i,
});
}
Ok(out)
}
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !is_pure_dimension(attrs);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// The dimension scale attached to each axis of a variable, from its
/// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length
/// sequence of object references per axis) — the address of the scale
/// netCDF-C takes for the axis, or `None` for an axis with none. netCDF-C's
/// `dimscale_visitor` lets `H5DSiterate_scales` visit every scale attached
/// to the axis and keeps the last, so with several (h5py's `attach_scale`
/// twice) the last one is the axis's dimension. `None` overall when the
/// attribute is missing or not in that form.
pub(crate) fn dimension_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<Option<u64>>> {
use clawhdf5_format::data_read::read_object_references;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::vl_data::VlResolver;
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("DIMENSION_LIST") else {
return None;
};
let Datatype::VariableLength {
is_string: false,
base_type,
..
} = datatype
else {
return None;
};
let sb = file.superblock();
let base_size = usize::try_from(base_type.type_size()).ok()?;
let sequences = VlResolver::new_in(file.storage(), sb.offset_size, sb.length_size)
.sequences(data, base_size)
.ok()?;
sequences
.iter()
.map(|refs| {
let refs = read_object_references(refs, base_type, sb.offset_size).ok()?;
Some(refs.iter().rev().find(|r| !r.is_null()).map(|r| r.address))
})
.collect()
}
/// Whether a dataset is a dimension scale that is only a dimension, not a
/// netCDF variable: netCDF-C and h5netcdf give it this `NAME`, and netCDF-C
/// does not list it among the variables.
pub(crate) fn is_pure_dimension(attrs: &HashMap<String, AttrValue>) -> bool {
is_dimension_scale(attrs)
&& matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
)
}
/// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
pub(crate) fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") {
return class == "DIMENSION_SCALE";
}
@@ -95,7 +278,7 @@ fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
}
/// Get the _Netcdf4Dimid attribute value if present.
fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
match attrs.get("_Netcdf4Dimid") {
Some(AttrValue::I64(id)) => Some(*id),
Some(AttrValue::U64(id)) => Some(*id as i64),
@@ -103,56 +286,8 @@ fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
}
}
/// Check if a dimension is unlimited by inspecting the HDF5 dataspace max_dimensions.
///
/// A dimension is unlimited when `max_dimensions[0] == u64::MAX` in the HDF5 dataspace.
fn check_unlimited(_file: &clawhdf5::File, group: &clawhdf5::Group<'_>, ds_name: &str) -> bool {
let ds = match group.dataset(ds_name) {
Ok(ds) => ds,
Err(_) => return false,
};
match ds.max_dimensions() {
Ok(Some(max_dims)) => max_dims.first().copied() == Some(u64::MAX),
_ => false,
}
}
/// Extract dimensions from an HDF5 group using both dimension scale attributes
/// and variable DIMENSION_LIST references.
///
/// This is a more robust approach that also discovers dimensions from variables
/// that reference them, even when dimension scales aren't explicitly set.
pub(crate) fn extract_dimensions_from_datasets(
group: &clawhdf5::Group<'_>,
file: &clawhdf5::File,
) -> Result<Vec<Dimension>, Error> {
// First try the standard approach with DIMENSION_SCALE
let mut dims = extract_dimensions(file, group)?;
// If we found dimensions, return them
if !dims.is_empty() {
return Ok(dims);
}
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let dataset_names = group.datasets()?;
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
// This 1-D dataset could be a coordinate variable / dimension
let is_unlimited = check_unlimited(file, group, ds_name);
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited,
});
}
}
Ok(dims)
/// Whether a dataset's first axis is unlimited (`max_dimensions[0] ==
/// u64::MAX` in its dataspace).
fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool {
matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX))
}
+23 -17
View File
@@ -9,12 +9,15 @@ use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension};
use crate::error::Error;
use crate::variable::{self, Variable};
use crate::scope::{self, Scope};
use crate::variable::Variable;
/// A NetCDF-4 group corresponding to an HDF5 group.
pub struct NetCDF4Group<'f> {
/// Group name.
name: String,
/// Path of the group from the root (`/`-separated).
path: String,
/// Underlying HDF5 file.
file: &'f clawhdf5::File,
/// Underlying HDF5 group.
@@ -25,11 +28,13 @@ impl<'f> NetCDF4Group<'f> {
/// Create a new NetCDF4Group from an HDF5 group.
pub(crate) fn new(
name: String,
path: String,
file: &'f clawhdf5::File,
hdf5_group: clawhdf5::Group<'f>,
) -> Self {
Self {
name,
path,
file,
hdf5_group,
}
@@ -40,27 +45,22 @@ impl<'f> NetCDF4Group<'f> {
&self.name
}
/// List dimensions defined in this group.
/// List dimensions defined in this group (not those of its parent
/// groups, which its variables can also use).
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5_group, self.file)
Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims)
}
/// List variables in this group.
/// List variables in this group: its datasets, except the dimension
/// scales that are only dimensions. Their dimensions can be defined in
/// this group or a parent group.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let dims = self.dimensions()?;
variable::build_variables(&self.hdf5_group, &dims)
Scope::new(self.file, &self.path)?.variables()
}
/// Get a specific variable by name.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let dims = self.dimensions()?;
let ds = self
.hdf5_group
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = crate::variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
scope::variable_at(self.file, &self.path, name)
}
/// Read all attributes of this group.
@@ -79,12 +79,18 @@ impl<'f> NetCDF4Group<'f> {
.hdf5_group
.group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), self.file, hdf5_group))
Ok(NetCDF4Group::new(
name.to_string(),
format!("{}/{name}", self.path),
self.file,
hdf5_group,
))
}
/// List dataset (variable) names in this group.
/// The names of this group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5_group.datasets()?)
Scope::new(self.file, &self.path)?.variable_names()
}
}
+19 -13
View File
@@ -27,6 +27,7 @@ pub mod cf;
pub mod dimension;
pub mod error;
pub mod group;
mod scope;
pub mod types;
pub mod variable;
@@ -75,25 +76,25 @@ impl NetCDF4File {
/// List dimensions defined in the root group.
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5.root(), &self.hdf5)
Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims)
}
/// List all variables in the root group.
/// List all variables in the root group: its datasets, except the
/// dimension scales that are only dimensions (netCDF-C does not list
/// them either).
pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> {
let dims = self.dimensions()?;
variable::build_variables(&self.hdf5.root(), &dims)
scope::Scope::new(&self.hdf5, "/")?.variables()
}
/// The names of the root group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variable_names()
}
/// Get a specific variable by name from the root group.
pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> {
let dims = self.dimensions()?;
let ds = self
.hdf5
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
scope::variable_at(&self.hdf5, "", name)
}
/// Read all global (root group) attributes.
@@ -112,7 +113,12 @@ impl NetCDF4File {
.hdf5
.group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), &self.hdf5, hdf5_group))
Ok(NetCDF4Group::new(
name.to_string(),
name.to_string(),
&self.hdf5,
hdf5_group,
))
}
/// Access the underlying HDF5 file for advanced operations.
+231
View File
@@ -0,0 +1,231 @@
//! A group's variables and the dimensions they are defined on.
//!
//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the
//! file, never by size: the dimension ids in its `_Netcdf4Coordinates`
//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension
//! scales its `DIMENSION_LIST` attribute references, looked up in the
//! variable's group and then each parent group up to the root. Only an axis
//! with neither (a file not written by a netCDF library) gets a dimension
//! by size. Dimension scales that are only dimensions are not variables, and
//! a variable stored as `_nc4_non_coord_<name>` (a variable sharing a
//! dimension's name without being its coordinate variable) is `<name>`.
use std::collections::{HashMap, HashSet};
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension, GroupDims};
use crate::error::Error;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// A group, with the dimensions visible from it.
pub(crate) struct Scope<'f> {
file: &'f clawhdf5::File,
group: clawhdf5::Group<'f>,
/// This group's dimensions, then its parent's, and so on to the root's.
levels: Vec<GroupDims>,
}
impl<'f> Scope<'f> {
/// The group at `path` (`/`-separated from the root; `""` or `"/"` is
/// the root).
pub fn new(file: &'f clawhdf5::File, path: &str) -> Result<Self, Error> {
let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect();
let mut levels = Vec::with_capacity(parts.len() + 1);
for n in (0..=parts.len()).rev() {
let group = file.group(&parts[..n].join("/"))?;
levels.push(dimension::group_dims(file, &group)?);
}
let group = file.group(&parts.join("/"))?;
Ok(Self {
file,
group,
levels,
})
}
/// The group's datasets, as `(dataset name, object header address)` in
/// listing order.
fn datasets(&self) -> Result<Vec<(String, u64)>, Error> {
let datasets: HashSet<String> = self.group.datasets()?.into_iter().collect();
Ok(self
.group
.entries()?
.into_iter()
.filter(|(name, _)| datasets.contains(name))
.collect())
}
/// The group's variables: every dataset but the dimension scales that
/// are only dimensions.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let mut variables = Vec::new();
for (ds_name, address) in self.datasets()? {
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
continue;
}
variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?);
}
Ok(variables)
}
/// The names of the group's variables.
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
let mut names = Vec::new();
for (ds_name, address) in self.datasets()? {
let attrs = self.file.dataset_at(address)?.attrs()?;
if !dimension::is_pure_dimension(&attrs) {
names.push(nc_name(&ds_name));
}
}
Ok(names)
}
/// The variable called `name`: the dataset `_nc4_non_coord_<name>` if
/// there is one, else the dataset `<name>` unless it is only a
/// dimension.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let datasets = self.datasets()?;
let prefixed = format!("{NON_COORD_PREFIX}{name}");
let address = datasets
.iter()
.find(|(n, _)| *n == prefixed)
.or_else(|| datasets.iter().find(|(n, _)| n == name))
.map(|&(_, address)| address)
.ok_or_else(not_found)?;
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
return Err(not_found());
}
self.variable_from(nc_name(name), address, ds, attrs)
}
fn variable_from(
&self,
name: String,
address: u64,
ds: clawhdf5::Dataset<'f>,
attrs: HashMap<String, AttrValue>,
) -> Result<Variable<'f>, Error> {
let shape = ds.shape()?;
let dims = self.variable_dims(address, &attrs, &shape);
Ok(Variable::new(name, ds, dims, attrs))
}
/// The first dimension, searching this group and then its ancestors,
/// that `find` picks.
fn find<'a>(
&'a self,
find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>,
) -> Option<Dimension> {
self.levels.iter().find_map(find).cloned()
}
/// The dimensions of the dataset at `address`, one per axis of `shape`,
/// as netCDF-C resolves them (see the module docs).
fn variable_dims(
&self,
address: u64,
attrs: &HashMap<String, AttrValue>,
shape: &[u64],
) -> Vec<Dimension> {
let rank = shape.len();
let mut dims: Vec<Option<Dimension>> = vec![None; rank];
if rank == 0 {
return Vec::new();
}
// A coordinate variable is the scale of its (first) dimension.
dims[0] = self.levels[0].by_address(address).cloned();
if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) {
for (slot, id) in dims.iter_mut().zip(ids) {
if slot.is_none() {
*slot = self.find(|level| level.by_dimid(id));
}
}
}
if dims.iter().any(Option::is_none)
&& let Some(scales) =
dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank)
{
for (slot, scale) in dims.iter_mut().zip(scales) {
if slot.is_none()
&& let Some(scale) = scale
{
*slot = self.find(|level| level.by_address(scale));
}
}
}
// Neither: the first dimension of this group of the same size not
// already taken by another such axis, else an anonymous one.
let own = &self.levels[0].dims;
let mut used = vec![false; own.len()];
dims.into_iter()
.zip(shape)
.map(|(dim, &size)| {
dim.unwrap_or_else(|| {
match own
.iter()
.enumerate()
.find(|&(i, d)| !used[i] && d.size == size)
{
Some((i, d)) => {
used[i] = true;
d.clone()
}
None => Dimension {
name: format!("dim_{size}"),
size,
is_unlimited: false,
},
}
})
})
.collect()
}
}
/// The variable `name` of the group at `group_path`; `name` may itself be
/// a path (`"sub/var"`), relative to that group.
pub(crate) fn variable_at<'f>(
file: &'f clawhdf5::File,
group_path: &str,
name: &str,
) -> Result<Variable<'f>, Error> {
match name.trim_start_matches('/').rsplit_once('/') {
Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}"))
.map_err(|_| Error::VariableNotFound(name.to_string()))?
.variable(leaf),
None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')),
}
}
/// The netCDF name of the dataset `ds_name`.
fn nc_name(ds_name: &str) -> String {
ds_name
.strip_prefix(NON_COORD_PREFIX)
.unwrap_or(ds_name)
.to_string()
}
/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension
/// of each axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
+208 -81
View File
@@ -2,12 +2,18 @@
//!
//! Variables in NetCDF-4 are HDF5 datasets. This module wraps them with
//! dimension associations and CF attribute support.
//!
//! A variable along an unlimited dimension has that dimension's length in
//! netCDF, even when fewer records of it have been written (its HDF5 dataset
//! is shorter): [`Variable::shape`] is the netCDF shape and the reads return
//! that many values, the unwritten ones as the fill value, as netCDF-C does.
//! [`Variable::stored_shape`] is the dataset's extent.
use std::collections::HashMap;
use clawhdf5::AttrValue;
use crate::cf::{self, CfAttributes};
use crate::cf::{self, CfAttributes, FillValue};
use crate::dimension::Dimension;
use crate::error::Error;
use crate::types::{NcType, dtype_to_nctype};
@@ -20,18 +26,23 @@ pub struct Variable<'f> {
dataset: clawhdf5::Dataset<'f>,
/// Dimensions associated with this variable.
dims: Vec<Dimension>,
/// Cached attributes.
attrs_cache: Option<HashMap<String, AttrValue>>,
/// The dataset's attributes.
attrs: HashMap<String, AttrValue>,
}
impl<'f> Variable<'f> {
/// Create a new Variable wrapping an HDF5 dataset.
pub(crate) fn new(name: String, dataset: clawhdf5::Dataset<'f>, dims: Vec<Dimension>) -> Self {
pub(crate) fn new(
name: String,
dataset: clawhdf5::Dataset<'f>,
dims: Vec<Dimension>,
attrs: HashMap<String, AttrValue>,
) -> Self {
Self {
name,
dataset,
dims,
attrs_cache: None,
attrs,
}
}
@@ -40,13 +51,27 @@ impl<'f> Variable<'f> {
&self.name
}
/// The dimensions of this variable.
/// The dimensions of this variable, one per axis: the ones the file
/// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its
/// group or a parent group. An axis the file gives no dimension (a file
/// not written by a netCDF library) gets the first dimension of the
/// variable's group of the same size, else an anonymous `dim_<size>`.
pub fn dimensions(&self) -> &[Dimension] {
&self.dims
}
/// The shape of this variable (dimension sizes).
/// The shape of this variable as netCDF reports it: along an unlimited
/// dimension, the dimension's current length (the longest variable on
/// it), even if fewer records of this variable have been written;
/// otherwise the dataset's extent. The reads return this many values.
pub fn shape(&self) -> Result<Vec<u64>, Error> {
Ok(nc_shape(&self.dataset.shape()?, &self.dims))
}
/// The extent of the HDF5 dataset: what has been written. It differs
/// from [`shape`](Self::shape) only along an unlimited dimension that
/// another variable has more records of.
pub fn stored_shape(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.shape()?)
}
@@ -58,59 +83,88 @@ impl<'f> Variable<'f> {
/// Read all attributes as a HashMap.
pub fn attrs(&mut self) -> Result<&HashMap<String, AttrValue>, Error> {
if self.attrs_cache.is_none() {
self.attrs_cache = Some(self.dataset.attrs()?);
}
Ok(self
.attrs_cache
.as_ref()
.expect("invariant: attrs_cache is Some after initialization"))
Ok(&self.attrs)
}
/// Extract CF convention attributes.
pub fn cf_attributes(&mut self) -> Result<CfAttributes, Error> {
let attrs = self.attrs()?;
Ok(cf::extract_cf_attributes(attrs))
Ok(cf::extract_cf_attributes(&self.attrs))
}
/// Read data as f64 with scale_factor/add_offset applied.
///
/// Missing values (matching `_FillValue` or `missing_value`) become NaN.
/// If no scale_factor or add_offset attributes exist, returns the raw f64 data.
/// Records along an unlimited dimension that this variable has not
/// written (see [`shape`](Self::shape)) are NaN.
pub fn read_f64(&mut self) -> Result<Vec<f64>, Error> {
let raw = self.dataset.read_f64()?;
let cf = self.cf_attributes()?;
Ok(cf::apply_scale_offset(&raw, &cf))
let cf = cf::extract_cf_attributes(&self.attrs);
self.padded(cf::apply_scale_offset(&raw, &cf), || Ok(f64::NAN))
}
/// Read raw data as f64 without any scale/offset transformation.
///
/// Unwritten records along an unlimited dimension read as the fill
/// value (`_FillValue`, else netCDF's default for the type), as in the
/// other `read_raw_*` methods and [`read_string`](Self::read_string).
pub fn read_raw_f64(&self) -> Result<Vec<f64>, Error> {
Ok(self.dataset.read_f64()?)
self.padded_read(self.dataset.read_f64()?)
}
/// Read raw data as f32 without any scale/offset transformation.
pub fn read_raw_f32(&self) -> Result<Vec<f32>, Error> {
Ok(self.dataset.read_f32()?)
self.padded_read(self.dataset.read_f32()?)
}
/// Read raw data as i32 without any scale/offset transformation.
pub fn read_raw_i32(&self) -> Result<Vec<i32>, Error> {
Ok(self.dataset.read_i32()?)
self.padded_read(self.dataset.read_i32()?)
}
/// Read raw data as i64 without any scale/offset transformation.
pub fn read_raw_i64(&self) -> Result<Vec<i64>, Error> {
Ok(self.dataset.read_i64()?)
self.padded_read(self.dataset.read_i64()?)
}
/// Read raw data as u64 without any scale/offset transformation.
pub fn read_raw_u64(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.read_u64()?)
self.padded_read(self.dataset.read_u64()?)
}
/// Read raw data as strings.
pub fn read_string(&self) -> Result<Vec<String>, Error> {
Ok(self.dataset.read_string()?)
self.padded_read(self.dataset.read_string()?)
}
/// The fill value netCDF-C gives the variable's unwritten values: its
/// `_FillValue`, else the default fill value of its type (`NC_FILL_*`).
fn fill_value(&self) -> Result<FillValue, Error> {
if let Some(fill) = cf::extract_cf_attributes(&self.attrs).fill_value {
return Ok(fill);
}
Ok(default_fill(self.nc_type()?))
}
/// `data`, read in the dataset's extent, laid out in the variable's
/// netCDF shape with the fill value in the positions not written.
fn padded_read<T: Clone + FromFill>(&self, data: Vec<T>) -> Result<Vec<T>, Error> {
self.padded(data, || Ok(T::from_fill(&self.fill_value()?)))
}
/// Like [`padded_read`](Self::padded_read), padding with what `fill`
/// returns (called only when there is something to pad).
fn padded<T: Clone>(
&self,
data: Vec<T>,
fill: impl FnOnce() -> Result<T, Error>,
) -> Result<Vec<T>, Error> {
let extent = self.dataset.shape()?;
let shape = nc_shape(&extent, &self.dims);
if shape == extent {
return Ok(data);
}
pad(data, &extent, &shape, fill()?)
}
/// Read raw bytes without any type conversion.
@@ -122,23 +176,23 @@ impl<'f> Variable<'f> {
let dtype = self.dataset.dtype()?;
match dtype {
clawhdf5::DType::F64 => {
let vals = self.dataset.read_f64()?;
let vals = self.read_raw_f64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
}
clawhdf5::DType::F32 => {
let vals = self.dataset.read_f32()?;
let vals = self.read_raw_f32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
}
clawhdf5::DType::I32 => {
let vals = self.dataset.read_i32()?;
let vals = self.read_raw_i32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
}
clawhdf5::DType::I64 => {
let vals = self.dataset.read_i64()?;
let vals = self.read_raw_i64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
}
clawhdf5::DType::U64 => {
let vals = self.dataset.read_u64()?;
let vals = self.read_raw_u64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
}
other => {
@@ -169,64 +223,137 @@ impl std::fmt::Debug for Variable<'_> {
}
}
/// Build variables from a group's datasets and associated dimensions.
pub(crate) fn build_variables<'f>(
group: &clawhdf5::Group<'f>,
available_dims: &[Dimension],
) -> Result<Vec<Variable<'f>>, Error> {
let dataset_names = group.datasets()?;
let mut variables = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
// Associate dimensions with this variable.
// First try DIMENSION_LIST attribute, then fall back to shape matching.
let var_dims = match_dimensions_to_variable(&shape, available_dims);
variables.push(Variable::new(ds_name.clone(), ds, var_dims));
/// The netCDF shape of a variable whose dataset has `extent`: along an
/// unlimited dimension the dimension's length, which is at least the extent.
fn nc_shape(extent: &[u64], dims: &[Dimension]) -> Vec<u64> {
if dims.len() != extent.len() {
return extent.to_vec();
}
Ok(variables)
extent
.iter()
.zip(dims)
.map(|(&e, d)| if d.is_unlimited { e.max(d.size) } else { e })
.collect()
}
/// Match dimensions to a variable based on shape.
///
/// For each axis of the variable, find a dimension with matching size.
/// If multiple dimensions have the same size, prefer exact name matching
/// from the convention order.
pub(crate) fn match_dimensions_to_variable(
shape: &[u64],
available_dims: &[Dimension],
) -> Vec<Dimension> {
let mut result = Vec::with_capacity(shape.len());
// Track which dimensions have been used to avoid duplicates
let mut used = vec![false; available_dims.len()];
for &dim_size in shape {
let mut matched = false;
// Find a dimension with matching size that hasn't been used yet
for (i, dim) in available_dims.iter().enumerate() {
if !used[i] && dim.size == dim_size {
result.push(dim.clone());
used[i] = true;
matched = true;
/// `data`, row-major in `extent`, placed in a row-major array of `shape`
/// (as many axes, each at least as long) filled with `fill`.
fn pad<T: Clone>(data: Vec<T>, extent: &[u64], shape: &[u64], fill: T) -> Result<Vec<T>, Error> {
let too_big = || Error::TypeError(format!("variable of shape {shape:?} is too large"));
let to_usize = |dims: &[u64]| -> Result<Vec<usize>, Error> {
dims.iter()
.map(|&d| usize::try_from(d).map_err(|_| too_big()))
.collect()
};
let (extent, shape) = (to_usize(extent)?, to_usize(shape)?);
let total = shape
.iter()
.try_fold(1usize, |n, &d| n.checked_mul(d))
.ok_or_else(too_big)?;
if extent.len() != shape.len()
|| extent.iter().zip(&shape).any(|(e, s)| e > s)
|| extent.iter().product::<usize>() != data.len()
{
return Err(Error::TypeError(format!(
"{} values of extent {extent:?} do not fit shape {shape:?}",
data.len()
)));
}
let (Some((&row, outer)), Some(&row_stride)) = (extent.split_last(), shape.last()) else {
return Ok(data);
};
let mut out = vec![fill; total];
if row == 0 {
return Ok(out);
}
// The position of the current row along each outer axis.
let mut index = vec![0usize; outer.len()];
for chunk in data.chunks_exact(row) {
let offset = index.iter().zip(&shape).fold(0, |o, (&i, &n)| o * n + i);
out[offset * row_stride..][..row].clone_from_slice(chunk);
for (i, &n) in index.iter_mut().zip(outer).rev() {
*i += 1;
if *i < n {
break;
}
}
if !matched {
// Create an anonymous dimension for unmatched sizes
result.push(Dimension {
name: format!("dim_{dim_size}"),
size: dim_size,
is_unlimited: false,
});
*i = 0;
}
}
Ok(out)
}
result
/// netCDF's default fill value for a type (`NC_FILL_*` in `netcdf.h`).
fn default_fill(nc_type: NcType) -> FillValue {
match nc_type {
NcType::Byte => FillValue::Int(-127),
NcType::UByte => FillValue::UInt(255),
NcType::Short => FillValue::Int(-32767),
NcType::UShort => FillValue::UInt(65535),
NcType::Int => FillValue::Int(-2_147_483_647),
NcType::UInt => FillValue::UInt(4_294_967_295),
NcType::Int64 => FillValue::Int(-9_223_372_036_854_775_806),
NcType::UInt64 => FillValue::UInt(18_446_744_073_709_551_614),
NcType::Float => FillValue::Float(f64::from(9.969_21e36_f32)),
NcType::Double => FillValue::Float(9.969_209_968_386_869e36),
NcType::String => FillValue::String(String::new()),
NcType::Char => FillValue::Int(0),
}
}
/// A fill value converted to the element type of a read, as the read
/// converts the stored values.
trait FromFill {
fn from_fill(fill: &FillValue) -> Self;
}
macro_rules! numeric_from_fill {
($($t:ty),*) => {$(
impl FromFill for $t {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::Float(v) => *v as $t,
FillValue::Int(v) => *v as $t,
FillValue::UInt(v) => *v as $t,
FillValue::String(_) => <$t>::default(),
}
}
}
)*};
}
numeric_from_fill!(f64, f32, i32, i64, u64);
impl FromFill for String {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::String(s) => s.clone(),
_ => String::new(),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn pad_places_rows() {
// (2, 1) written of (2, 4): each row padded, not the tail.
let out = pad(vec![1, 2], &[2, 1], &[2, 4], 0).unwrap();
assert_eq!(out, vec![1, 0, 0, 0, 2, 0, 0, 0]);
// Leading axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 2], &[3, 2], -1).unwrap();
assert_eq!(out, vec![1, 2, 3, 4, -1, -1]);
// Nothing written.
let out = pad(Vec::<i32>::new(), &[0, 3], &[2, 3], 7).unwrap();
assert_eq!(out, vec![7; 6]);
// 3-D, middle axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 1, 2], &[2, 2, 2], 0).unwrap();
assert_eq!(out, vec![1, 2, 0, 0, 3, 4, 0, 0]);
}
#[test]
fn pad_rejects_wrong_length() {
assert!(pad(vec![1, 2, 3], &[2, 2], &[3, 2], 0).is_err());
assert!(pad(vec![1, 2, 3, 4], &[2, 2], &[1, 4], 0).is_err());
}
}
+466 -1
View File
@@ -4,7 +4,7 @@
use std::process::Command;
use clawhdf5_netcdf4::{AttrValue, NetCDF4File};
use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File};
// ---------------------------------------------------------------------------
// Helpers
@@ -377,3 +377,468 @@ ds.close()
let names = file.variable("name").unwrap().read_string().unwrap();
assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]);
}
// ===========================================================================
// Unlimited dimensions: the length netCDF-C reports
// ===========================================================================
/// An unlimited dimension's length is the largest extent of the variables
/// using it, in any group (netCDF-C's `nc4_find_dim_len`), not its dimension
/// scale's extent (which netCDF-C leaves at 0): variables of different
/// lengths, one in a subgroup, a dimension no variable has written, a
/// coordinate variable, a subgroup's own unlimited dimension. Compared with
/// what netCDF4-python reports for the same file.
#[test]
fn unlimited_dimension_lengths_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("unlimited.nc");
let script = format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w", format="NETCDF4") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("rec", None)
f.createDimension("x", 3)
f.createVariable("t", "f8", ("time", "x"))[0:2, :] = np.ones((2, 3))
f.createVariable("a", "i4", ("time",))[0:4] = np.arange(4)
f.createVariable("e", "i4", ("empty",))
f.createVariable("rec", "f4", ("rec",))[0:3] = [1, 2, 3]
f.createVariable("r", "f4", ("x", "rec"))[:, 0:5] = np.ones((3, 5))
g = f.createGroup("sub")
g.createVariable("c", "i4", ("time",))[0:6] = np.arange(6)
g.createDimension("srec", None)
g.createVariable("s", "i4", ("srec", "x"))[0:1, :] = np.ones((1, 3))
with nc.Dataset({path:?}) as f:
for grp in (f, f.groups["sub"]):
for name, d in grp.dimensions.items():
print(grp.path, name, len(d), d.isunlimited())
"#,
path = path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let expected: Vec<String> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(str::to_string)
.collect();
// time: t has 2 records, a 4 and sub/c 6; rec: the coordinate variable
// has 3, r 5; empty: nothing written.
for want in [
"/ time 6 True",
"/ empty 0 True",
"/ rec 5 True",
"/ x 3 False",
"/sub srec 1 True",
] {
assert!(
expected.iter().any(|l| l == want),
"netCDF4 reports {expected:?}"
);
}
let file = NetCDF4File::open(&path).unwrap();
let sub = file.group("sub").unwrap();
let mut got = Vec::new();
for (path, dims) in [
("/", file.dimensions().unwrap()),
("/sub", sub.dimensions().unwrap()),
] {
for d in dims {
let unlimited = if d.is_unlimited { "True" } else { "False" };
got.push(format!("{path} {} {} {unlimited}", d.name, d.size));
}
}
assert_eq!(got, expected);
}
// ===========================================================================
// Variables' dimensions, shapes and values as netCDF4-python reports them
// ===========================================================================
/// Whether python can import `module`.
fn python_has(module: &str) -> bool {
Command::new(python())
.args(["-c", &format!("import {module}")])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
/// h5netcdf is not in every interop environment (CI installs it; a local
/// `.venv` may not have it), so its tests skip without it even under
/// `CLAWHDF5_REQUIRE_INTEROP=1`.
macro_rules! skip_if_no_h5netcdf {
() => {
if !python_has("h5netcdf") {
eprintln!("SKIP: python3 with h5netcdf not available");
return;
}
};
}
/// Every variable of the file at `path`, in every group, as netCDF4-python
/// reports it: `"<group path> <name> (<dims>) (<shape>)"` and its values
/// (numeric variables; element by element with masking off, so unwritten
/// records are the fill value), sorted by the description.
///
/// Values are read one element at a time because netCDF-C 4.9.3 lays out a
/// whole-variable read of a variable shorter than an unlimited dimension
/// that is not its first wrongly (the written values first, then the fill);
/// element reads, and reads of one index of the leading axis, are right.
fn netcdf4_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
let script = r#"
import sys
import numpy as np
import netCDF4 as nc
def walk(g):
for name, v in g.variables.items():
v.set_auto_mask(False)
head = "%s %s (%s) (%s)" % (g.path, name, ",".join(v.dimensions), ",".join(map(str, v.shape)))
vals = []
if v.dtype != str and v.dtype.kind in "iuf":
vals = [repr(float(v[i])) for i in np.ndindex(v.shape)]
print(head + "|" + " ".join(vals))
for sub in g.groups.values():
walk(sub)
with nc.Dataset(sys.argv[1]) as f:
walk(f)
"#;
let out = Command::new(python())
.args(["-c", script, &path.display().to_string()])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let mut view: Vec<(String, Vec<f64>)> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(|line| {
let (head, vals) = line.split_once('|').unwrap();
let vals = vals
.split_whitespace()
.map(|v| v.parse().unwrap())
.collect();
(head.to_string(), vals)
})
.collect();
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// The same view of the file through clawhdf5-netcdf4.
fn clawhdf5_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
fn describe(
group_path: &str,
vars: Vec<clawhdf5_netcdf4::Variable<'_>>,
) -> Vec<(String, Vec<f64>)> {
vars.into_iter()
.map(|v| {
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
let shape: Vec<String> = v.shape().unwrap().iter().map(u64::to_string).collect();
let head = format!(
"{group_path} {} ({}) ({})",
v.name(),
dims.join(","),
shape.join(",")
);
let vals = match v.nc_type().unwrap() {
NcType::String | NcType::Char => Vec::new(),
_ => v.read_raw_f64().unwrap(),
};
(head, vals)
})
.collect()
}
fn walk(
group_path: &str,
group: &clawhdf5_netcdf4::NetCDF4Group<'_>,
out: &mut Vec<(String, Vec<f64>)>,
) {
out.extend(describe(group_path, group.variables().unwrap()));
for name in group.group_names().unwrap() {
walk(
&format!("{group_path}/{name}"),
&group.group(&name).unwrap(),
out,
);
}
}
let file = NetCDF4File::open(path).unwrap();
let mut view = describe("/", file.variables().unwrap());
for name in file.group_names().unwrap() {
walk(&format!("/{name}"), &file.group(&name).unwrap(), &mut view);
}
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// clawhdf5-netcdf4 reports the same variables, dimensions, shapes and
/// values (bit for bit, NaN equal to NaN) as netCDF4-python.
fn assert_same_view(path: &std::path::Path) {
let want = netcdf4_view(path);
let got = clawhdf5_view(path);
let heads = |v: &[(String, Vec<f64>)]| v.iter().map(|(h, _)| h.clone()).collect::<Vec<_>>();
assert_eq!(heads(&got), heads(&want), "variables differ from netCDF4's");
for ((head, got), (_, want)) in got.iter().zip(&want) {
let same = got.len() == want.len()
&& got
.iter()
.zip(want)
.all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan()));
assert!(same, "{head}: got {got:?}, netCDF4 reads {want:?}");
}
}
/// The reproducer of the known-issues entry: `a` is on the unlimited `time`
/// (5 long through `b`) with 2 records, not on an anonymous `dim_2`; the
/// pure dimension scales `time` and `empty` are not variables; `a` has
/// shape (5,) and reads its 3 unwritten records as the fill value.
#[test]
fn variable_dimensions_come_from_the_file() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("repro.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("x", 3)
f.createVariable("a", "i4", ("time",))[0:2] = [1, 2]
f.createVariable("b", "f4", ("time", "x"))[0:5, :] = np.arange(15).reshape(5, 3)
f.createVariable("e", "i4", ("empty",))
f.createVariable("c", "i4", ("x",))[:] = [7, 8, 9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["a", "b", "c", "e"]);
assert!(matches!(
file.variable("time"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
let a = file.variable("a").unwrap();
assert_eq!(a.dimensions()[0].name, "time");
assert_eq!(a.shape().unwrap(), [5]);
assert_eq!(a.stored_shape().unwrap(), [2]);
assert_eq!(
a.read_raw_i32().unwrap(),
[1, 2, -2_147_483_647, -2_147_483_647, -2_147_483_647]
);
}
/// Dimensions of one size are told apart by the file, not by order: `p`
/// and `q` are both 2 long, and `v(q, p)`, `same(p, p)` (one dimension
/// twice), a scalar, `q`'s coordinate variable, a variable called `p` that
/// is not `p`'s coordinate variable (stored as `_nc4_non_coord_p`), and
/// variables in a subgroup and a sub-subgroup on dimensions of their
/// ancestors.
#[test]
fn equal_size_and_inherited_dimensions_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("dims.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("p", 2)
f.createDimension("q", 2)
f.createVariable("v", "i4", ("q", "p"))[:] = np.array([[1, 2], [3, 4]])
f.createVariable("same", "i4", ("p", "p"))[:] = np.array([[5, 6], [7, 8]])
f.createVariable("s", "f8", ())[...] = 3.5
f.createVariable("q", "f4", ("q",))[:] = [0, 1]
f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]])
g = f.createGroup("g")
g.createDimension("r", 2)
g.createVariable("w", "i4", ("r", "q", "p"))[:] = np.arange(8).reshape(2, 2, 2)
h = g.createGroup("h")
h.createVariable("z", "i4", ("p", "r"))[:] = np.array([[1, 2], [3, 4]])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let v = file.variable("v").unwrap();
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["q", "p"]);
let p = file.variable("p").unwrap();
assert!(!p.is_coordinate());
assert!(file.variable("q").unwrap().is_coordinate());
let s = file.variable("s").unwrap();
assert!(s.dimensions().is_empty());
assert_eq!(s.shape().unwrap(), Vec::<u64>::new());
let z = file
.group("g")
.unwrap()
.group("h")
.unwrap()
.variable("z")
.unwrap();
let dims: Vec<&str> = z.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["p", "r"]);
}
/// Variables shorter than their unlimited dimension have its length and
/// read the fill value (`_FillValue`, else netCDF's default for the type)
/// where nothing was written — also when the unlimited dimension is not
/// the first; `read_f64` gives NaN there.
#[test]
fn unwritten_records_read_as_fill_like_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("pad.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("t", None)
f.createDimension("x", 2)
f.createVariable("a", "i4", ("t",))[0:2] = [1, 2]
f.createVariable("f", "f4", ("x", "t"), fill_value=-5.0)[:, 0:1] = np.array([[1], [2]])
f.createVariable("d", "f8", ("t",))[0:4] = [1, 2, 3, 4]
f.createVariable("u", "u8", ("t",))[0:1] = [1]
f.createVariable("b", "i1", ("t", "x"))[0:3, :] = np.ones((3, 2))
f.createVariable("st", str, ("t",))[0] = "hi"
g = f.createGroup("g")
g.createVariable("k", "f4", ("t",))[0:1] = [9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut f = file.variable("f").unwrap();
assert_eq!(f.shape().unwrap(), [2, 4]);
assert_eq!(f.stored_shape().unwrap(), [2, 1]);
assert_eq!(
f.read_raw_f32().unwrap(),
[1.0, -5.0, -5.0, -5.0, 2.0, -5.0, -5.0, -5.0]
);
let read = f.read_f64().unwrap();
assert_eq!(read[0], 1.0);
assert!(read[1].is_nan() && read[7].is_nan());
let st = file.variable("st").unwrap();
assert_eq!(st.read_string().unwrap(), ["hi", "", "", ""]);
assert_eq!(st.shape().unwrap(), [4]);
}
/// A file with HDF5 dimension scales but none of netCDF's own attributes
/// (h5py's `dims` API): the dimensions come from `DIMENSION_LIST`, so
/// `v(q, p)` is not `v(p, q)` although both are 2 long; with two scales
/// attached to one axis (`w`), netCDF-C takes the last.
#[test]
fn h5py_dimension_scales_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("scales.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
with h5py.File({path:?}, "w") as f:
f["p"] = np.arange(2.0)
f["q"] = np.arange(2.0) + 10
f["p"].make_scale("p")
f["q"].make_scale("q")
f["v"] = np.arange(4).reshape(2, 2)
f["v"].dims[0].attach_scale(f["q"])
f["v"].dims[1].attach_scale(f["p"])
f["w"] = np.arange(2)
f["w"].dims[0].attach_scale(f["p"])
f["w"].dims[0].attach_scale(f["q"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files h5netcdf writes (its own implementation of the netCDF-4
/// conventions over h5py): an unlimited dimension, equal sizes, a subgroup
/// on inherited dimensions, a scalar.
#[test]
fn h5netcdf_file_matches_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_h5netcdf!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("h5netcdf.nc");
run_python(&format!(
r#"
import h5netcdf
import numpy as np
with h5netcdf.File({path:?}, "w") as f:
f.dimensions = {{"p": 2, "q": 2, "t": None}}
f.create_variable("v", ("q", "p"), "i4")[...] = np.array([[1, 2], [3, 4]])
f.create_variable("q", ("q",), "f4")[...] = [0, 1]
f.create_variable("same", ("p", "p"), "i4")[...] = np.array([[5, 6], [7, 8]])
a = f.create_variable("a", ("t", "p"), "f8")
f.resize_dimension("t", 3)
a[...] = np.ones((3, 2))
f.create_variable("short", ("t",), "i4")
g = f.create_group("g")
g.dimensions = {{"r": 2}}
g.create_variable("w", ("r", "q", "p"), "i4")[...] = np.arange(8).reshape(2, 2, 2)
g.create_variable("s", (), "f8")[...] = 2.5
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files xarray writes, through netCDF4 and (when installed) h5netcdf:
/// coordinates, two dimensions of one size, an unlimited dimension.
#[test]
fn xarray_files_match_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_xarray!();
let dir = tempfile::tempdir().unwrap();
let mut engines = vec!["netcdf4"];
if python_has("h5netcdf") {
engines.push("h5netcdf");
} else {
eprintln!("SKIP: xarray with engine h5netcdf (h5netcdf not available)");
}
for engine in engines {
let path = dir.path().join(format!("xarray_{engine}.nc"));
run_python(&format!(
r#"
import numpy as np
import xarray as xr
ds = xr.Dataset(
{{
"temp": (("time", "lat", "lon"), np.arange(12.0).reshape(3, 2, 2)),
"grid": (("lon", "lat"), np.array([[1, 2], [3, 4]], dtype="i4")),
"scalar": ((), 1.5),
}},
coords={{"time": [0.0, 6.0, 12.0], "lat": [10.0, 20.0], "lon": [5.0, 6.0]}},
)
ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
}
@@ -754,3 +754,53 @@ fn test_dimension_struct_equality() {
};
assert_ne!(d1, d3);
}
/// A dimension scale that is only a dimension (netCDF-C's `NAME`) is not a
/// variable, and `_nc4_non_coord_<name>` is the variable `<name>`, found in
/// place of the scale of the same name.
#[test]
fn test_pure_dimensions_hidden_and_non_coord_names() {
let pure = "This is a netCDF dimension but not a netCDF variable. 2";
let mut b = FileBuilder::new();
b.create_dataset("x")
.with_f32_data(&[0.0, 0.0])
.with_shape(&[2])
.set_attr("CLASS", AttrValue::String("DIMENSION_SCALE".into()))
.set_attr("NAME", AttrValue::String(pure.into()))
.set_attr("_Netcdf4Dimid", AttrValue::I64(0));
b.create_dataset("_nc4_non_coord_x")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3]);
b.create_dataset("v")
.with_f64_data(&[5.0, 6.0])
.with_shape(&[2]);
let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap();
let dims = file.dimensions().unwrap();
assert_eq!(dims.len(), 1);
assert_eq!(dims[0].name, "x");
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["v", "x"]);
let x = file.variable("x").unwrap();
assert_eq!(x.name(), "x");
assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]);
assert!(!x.is_coordinate());
// No DIMENSION_LIST: `v` gets `x` by size, as before.
assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x");
}
/// `variable` still takes a path relative to the group, as it did when it
/// opened the dataset by path.
#[test]
fn test_variable_by_path() {
let file = NetCDF4File::from_bytes(make_grouped_netcdf4()).unwrap();
let pressure = file.variable("surface/pressure").unwrap();
assert_eq!(pressure.name(), "pressure");
assert_eq!(pressure.read_raw_f64().unwrap(), [1013.25, 1012.0, 1011.5]);
assert!(file.variable("/time").is_ok());
assert!(matches!(
file.variable("nowhere/pressure"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
}
+15 -7
View File
@@ -1,8 +1,5 @@
# clawhdf5-py
[![crates.io](https://img.shields.io/crates/v/clawhdf5-py.svg)](https://crates.io/crates/clawhdf5-py)
[![docs.rs](https://docs.rs/clawhdf5-py/badge.svg)](https://docs.rs/clawhdf5-py)
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
`clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5.
@@ -40,7 +37,12 @@ with clawhdf5.File("data.h5", "r") as f:
`dtype.metadata['enum']`), complex, `S<n>` fixed strings, `object` for
variable-length strings (`bytes` values) and sequences (array values),
`V<n>` opaque, array types, and compounds as structured dtypes.
Other types raise `TypeError`.
Non-IEEE floats (HDF5 2.x's bfloat16, FP8, FP6 and FP4) read as h5py
3.16 reads them: as the narrowest IEEE float that holds them (`float32`
for bfloat16, `float16` for the 1-byte formats) in the file's byte order,
with libhdf5's values (`tests/test_small_floats.py`); inside a compound or
array type they raise `TypeError`, and writing them raises
`NotImplementedError`. Other types raise `TypeError`.
- Keys are h5py's: integers, slices with a positive step, `...`, one
increasing list of integers, compound field names. Each maps onto a
hyperslab selection. `None` and negative steps are refused
@@ -61,6 +63,8 @@ with clawhdf5.File("data.h5", "r") as f:
`clawhdf5.InternalError`, a `RuntimeError`.
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`).
- Also as in h5py: `File.mode` (`'r'`, or `'r+'` for a writable file), `File.flush()` (a no-op:
edits are already synced), `Dataset.chunks`.
## Remote files
@@ -96,8 +100,11 @@ f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
written on `close()`.
writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request).
## Editing a file in place
@@ -130,7 +137,8 @@ with clawhdf5.File("data.h5", "r+") as f:
`str` is stored as a fixed-length UTF-8 string (h5py stores a
variable-length one), so h5py reads it back as `bytes`.
- Not supported (`NotImplementedError`, nothing written): creating or
deleting datasets, groups and attributes, writing compound fields by
deleting datasets and groups, deleting attributes (creating and
replacing them works, compact or dense), writing compound fields by
name, variable-length data, HDF5 array types, and whatever
`FileEditor` refuses (listed in `docs/known-issues.md`).
+120 -4
View File
@@ -8,10 +8,16 @@
//! returns become the array's buffer as they are: the `Vec<u8>` is handed to
//! numpy without a copy and viewed as the dtype.
//!
//! Anything this mapping cannot describe exactly — non-IEEE floats, integers
//! with padding bits, VAX byte order, references, bitfields, time, and
//! variable-length members inside compounds or arrays — is a `TypeError`,
//! never a best-effort guess.
//! The one exception is a dataset or attribute of a non-IEEE float
//! (bfloat16, FP8 E4M3/E5M2, FP6, FP4, ...): like h5py, it reads as the
//! narrowest of `float16`/`float32`/`float64` that holds every value of the
//! format exactly, in the file's byte order, and the values are converted
//! (as libhdf5 converts them, NaN bits included).
//!
//! Anything this mapping cannot describe exactly — non-IEEE floats inside
//! compounds or arrays, integers with padding bits, VAX byte order,
//! references, bitfields, time, and variable-length members inside compounds
//! or arrays — is a `TypeError`, never a best-effort guess.
use std::collections::HashMap;
@@ -35,6 +41,13 @@ pub(crate) enum Layout {
VlString { utf8: bool },
/// Variable-length sequence of a fixed-size base type.
VlSequence,
/// A non-IEEE float (`source`), converted to the IEEE float of `size`
/// bytes numpy reports (see [`widened_float`]).
Float {
source: Datatype,
size: usize,
big_endian: bool,
},
}
/// Everything needed to turn a dataset's or attribute's bytes into numpy.
@@ -139,6 +152,74 @@ fn float_format(dt: &Datatype) -> PyResult<String> {
Ok(format!("{}f{size}", byte_order_char(byte_order, *size)?))
}
/// Whether `dt` is an IEEE 754 binary16/32/64 float numpy reads as it is.
pub(crate) fn is_ieee_float(dt: &Datatype) -> bool {
float_format(dt).is_ok()
}
/// The IEEE float h5py reads a non-IEEE float as (`TypeFloatID.py_dtype` in
/// h5py 3.16): the first of `float16`, `float32`, `float64` at least as wide
/// whose mantissa holds the format's and whose normal exponent range covers
/// it — bfloat16 as `float32`, FP8/FP6/FP4 as `float16` — in the file's byte
/// order. Returns `(size in bytes, big endian)`, or `None` when no IEEE type
/// holds it (or the byte order is VAX).
fn widened_float(dt: &Datatype) -> Option<(usize, bool)> {
let Datatype::FloatingPoint {
size,
byte_order,
exponent_size,
mantissa_size,
exponent_bias,
..
} = dt
else {
return None;
};
let big_endian = match byte_order {
DatatypeByteOrder::LittleEndian => false,
DatatypeByteOrder::BigEndian => true,
DatatypeByteOrder::Vax => return None,
};
if *exponent_size >= 32 {
return None;
}
let max_exp = (1i64 << exponent_size) - i64::from(*exponent_bias) - 1;
let min_exp = 1 - i64::from(*exponent_bias);
// numpy's finfo: (itemsize, nmant, maxexp, minexp).
[(2, 10, 16, -14), (4, 23, 128, -126), (8, 52, 1024, -1022)]
.into_iter()
.find(|&(bytes, nmant, maxexp, minexp)| {
bytes >= *size && *mantissa_size <= nmant && max_exp <= maxexp && min_exp >= minexp
})
.map(|(bytes, ..)| (bytes as usize, big_endian))
}
/// Encode `values` as IEEE floats of `size` bytes. Every value is exact in
/// the target (`widened_float` chose it so); a NaN keeps its sign and gets
/// every mantissa bit set, as libhdf5 converts NaNs.
fn encode_floats(values: &[f64], size: usize, big_endian: bool) -> Vec<u8> {
let mut out = Vec::with_capacity(values.len() * size);
for &v in values {
let bits: u64 = if v.is_nan() {
let sign = u64::from(v.is_sign_negative()) << (size * 8 - 1);
sign | ((1u64 << (size * 8 - 1)) - 1)
} else {
match size {
2 => u64::from(clawhdf5_format::float16::f32_to_f16_bits(v as f32)),
4 => u64::from((v as f32).to_bits()),
_ => v.to_bits(),
}
};
let bytes = bits.to_le_bytes();
if big_endian {
out.extend(bytes[..size].iter().rev());
} else {
out.extend_from_slice(&bytes[..size]);
}
}
out
}
/// `r`/`i` compounds of two identical IEEE floats are complex numbers in h5py.
fn complex_format(
size: u32,
@@ -232,6 +313,9 @@ fn np_dtype_with_metadata<'py>(
/// read as they are.
pub(crate) fn fixed_dtype<'py>(py: Python<'py>, dt: &Datatype) -> PyResult<Bound<'py, PyAny>> {
match dt {
Datatype::Complex { size, base_type } => {
fixed_dtype(py, &Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint { .. } => np_dtype(py, int_format(dt)?),
Datatype::FloatingPoint { .. } => np_dtype(py, float_format(dt)?),
Datatype::String { size, charset, .. } => {
@@ -366,6 +450,25 @@ impl Converter {
vl_unit: 0,
})
}
Datatype::FloatingPoint { .. } if !is_ieee_float(dt) => {
let Some((size, big_endian)) = widened_float(dt) else {
// The error names the layout.
return Err(float_format(dt).expect_err("not IEEE"));
};
let order = if big_endian { '>' } else { '<' };
let dtype = np_dtype(py, format!("{order}f{size}"))?;
Ok(Self {
view: dtype.clone().unbind(),
dtype: dtype.unbind(),
layout: Layout::Float {
source: dt.clone(),
size,
big_endian,
},
elem_size: dt.type_size() as usize,
vl_unit: 0,
})
}
_ => {
let dtype = fixed_dtype(py, dt)?;
Ok(Self {
@@ -416,6 +519,19 @@ impl Converter {
(Elements::Bytes(bytes), Layout::Fixed) => {
bytes_as_array(py, bytes, self.view.bind(py), shape)
}
(
Elements::Bytes(bytes),
Layout::Float {
source,
size,
big_endian,
},
) => {
let values = clawhdf5_format::data_read::read_as_f64(&bytes, source)
.map_err(|e| unsupported(e.to_string()))?;
let converted = encode_floats(&values, *size, *big_endian);
bytes_as_array(py, converted, self.view.bind(py), shape)
}
(Elements::Bytes(bytes), Layout::Subarray(dims)) => {
let mut full = shape.to_vec();
full.extend_from_slice(dims);
+10 -2
View File
@@ -47,8 +47,14 @@ fn not_implemented(what: impl std::fmt::Display) -> PyErr {
/// `edit_helpers._convert_array`), or why they cannot be written.
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
match dt {
Datatype::Complex { .. } => Ok("complex"),
Datatype::FixedPoint { .. } => Ok("int"),
Datatype::FloatingPoint { .. } => Ok("float"),
Datatype::FloatingPoint { .. } if crate::convert::is_ieee_float(dt) => Ok("float"),
// Read as a wider IEEE float; writing would need the reverse
// conversion (rounding into bfloat16, FP8, ...).
Datatype::FloatingPoint { .. } => Err(not_implemented(
"writing non-IEEE floats (bfloat16, FP8, FP6, FP4, ...)",
)),
Datatype::Enumeration {
base_type, members, ..
} => {
@@ -102,7 +108,9 @@ fn check_exact(dt: &Datatype) -> PyResult<()> {
Datatype::Compound { members, .. } => {
members.iter().try_for_each(|m| check_exact(&m.datatype))
}
Datatype::Array { base_type, .. } => check_exact(base_type),
Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
check_exact(base_type)
}
Datatype::String {
padding: StringPadding::NullPad,
..
+23 -1
View File
@@ -161,6 +161,10 @@ pub(crate) enum DatasetData {
I64(Vec<i64>),
I32(Vec<i32>),
U8(Vec<u8>),
/// numpy `complex64`, `[re, im]` pairs, written as h5py does.
C64(Vec<[f32; 2]>),
/// numpy `complex128`, `[re, im]` pairs, written as h5py does.
C128(Vec<[f64; 2]>),
}
/// Specification for a dataset to be written.
@@ -281,6 +285,14 @@ pub(crate) fn apply_dataset_spec(
DatasetData::U8(v) => {
db.with_u8_data(v);
}
// h5py's compound `{r, i}`, not HDF5 2.0's native complex type: it
// is what h5py writes (3.16 included) and every libhdf5 can read it.
DatasetData::C64(v) => {
db.with_complex_f32_data(v);
}
DatasetData::C128(v) => {
db.with_complex_f64_data(v);
}
}
if !spec.shape.is_empty() {
db.with_shape(&spec.shape);
@@ -314,9 +326,19 @@ pub(crate) fn extract_numpy_data(
"int64" => DatasetData::I64(flat.extract::<Vec<i64>>()?),
"int32" => DatasetData::I32(flat.extract::<Vec<i32>>()?),
"uint8" => DatasetData::U8(flat.extract::<Vec<u8>>()?),
// A 1-D complex array viewed as floats is its (re, im) parts in order.
"complex64" => {
let parts: Vec<f32> = flat.call_method1("view", ("<f4",))?.extract()?;
DatasetData::C64(parts.as_chunks::<2>().0.to_vec())
}
"complex128" => {
let parts: Vec<f64> = flat.call_method1("view", ("<f8",))?.extract()?;
DatasetData::C128(parts.as_chunks::<2>().0.to_vec())
}
_ => {
return Err(PyErr::new::<pyo3::exceptions::PyTypeError, _>(format!(
"unsupported numpy dtype: {dtype_str}; expected float64, float32, int64, int32, or uint8"
"unsupported numpy dtype: {dtype_str}; expected float64, float32, int64, int32, \
uint8, complex64 or complex128"
)));
}
};
@@ -0,0 +1,92 @@
"""Non-IEEE floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1) read
as h5py 3.16 reads them: as the narrowest IEEE float that holds every value
(bfloat16 as float32, the 1-byte formats as float16), in the file's byte
order, with the values libhdf5 converts them to.
The fixture was written by libhdf5 2.2.0; the JSON next to it holds what
libhdf5 2.2.0 itself returns for every element (see gen_mx_floats.py)."""
import json
import os
import shutil
import numpy as np
import pytest
import clawhdf5
FIXTURES = os.path.join(os.path.dirname(__file__), "..", "..", "clawhdf5", "tests", "fixtures")
FILE = os.path.join(FIXTURES, "mx_floats_hdf5_2_2.h5")
REFERENCE = json.load(open(os.path.join(FIXTURES, "mx_floats_hdf5_2_2.json")))["objects"]
# What h5py 3.16 reports for each (checked 2026-09-28).
DTYPES = {
"bf16le": "<f4",
"bf16be": ">f4",
"f8e4m3": "<f2",
"f8e5m2": "<f2",
"f6e2m3": "<f2",
"f6e2m3_pad": "<f2",
"f6e3m2": "<f2",
"f6e3m2_pad": "<f2",
"f4e2m1": "<f2",
"f4e2m1_pad": "<f2",
}
def _expected(name, dtype):
"""The libhdf5 2.2.0 values as `dtype`, NaN bits included (sign kept,
every mantissa bit set)."""
out = []
for text in REFERENCE[name]["f64"]:
if text.startswith("nan:"):
negative = int(text[4:], 16) >> 63
bits = dtype.itemsize * 8
word = (negative << (bits - 1)) | ((1 << (bits - 1)) - 1)
out.append(np.frombuffer(word.to_bytes(dtype.itemsize, "little"), dtype.newbyteorder("<"))[0])
else:
out.append(float(text))
return np.array(out, dtype=dtype.newbyteorder("<")).astype(dtype)
def test_values_match_libhdf5_2_2():
assert set(DTYPES) == set(REFERENCE)
with clawhdf5.File(FILE, "r") as f:
for name, dtype in DTYPES.items():
ds = f[name]
assert ds.dtype == np.dtype(dtype), name
assert ds.dtype.str == dtype, name
want = _expected(name, np.dtype(dtype))
got = ds[()]
assert got.dtype.str == dtype
assert got.tobytes() == want.tobytes(), name
# Selections convert the same way.
assert ds[3:9].tobytes() == want[3:9].tobytes(), name
assert ds[[0, 5, 7]].tobytes() == want[[0, 5, 7]].tobytes(), name
assert ds[5].tobytes() == want[5].tobytes(), name
if REFERENCE[name]["attribute"]:
attr = f.attrs[name]
assert attr.dtype.str == dtype
assert attr.tobytes() == want.tobytes(), name
def test_matches_h5py(h5py):
with clawhdf5.File(FILE, "r") as ours, h5py.File(FILE, "r") as theirs:
for name in DTYPES:
assert ours[name].dtype == theirs[name].dtype, name
assert ours[name][()].tobytes() == theirs[name][()].tobytes(), name
if name in theirs.attrs:
assert ours.attrs[name].dtype == theirs.attrs[name].dtype
assert ours.attrs[name].tobytes() == theirs.attrs[name].tobytes()
def test_writing_is_refused(tmp_path):
path = tmp_path / "copy.h5"
shutil.copyfile(FILE, path)
with clawhdf5.File(str(path), "r+") as f:
with pytest.raises(NotImplementedError, match="non-IEEE"):
f["bf16le"][0] = 1.0
with pytest.raises(NotImplementedError, match="non-IEEE"):
f["f4e2m1"][:] = np.zeros(16)
with open(path, "rb") as a, open(FILE, "rb") as b:
assert a.read() == b.read()
@@ -247,6 +247,51 @@ def test_roundtrip_uint8(tmp_h5):
assert result.dtype == np.uint8
@pytest.mark.parametrize("dtype", [np.complex64, np.complex128])
def test_roundtrip_complex(tmp_h5, dtype):
"""Complex arrays are written as h5py writes them (a compound {r, i});
h5py and clawhdf5 both read them back as the same numpy complex dtype."""
import h5py
original = (np.arange(12).reshape(3, 4) * (1.5 - 0.25j)).astype(dtype)
with clawhdf5.File(tmp_h5, "w") as f:
f.create_dataset("z", data=original)
f.create_dataset("zc", data=original, chunks=(2, 2), compression="gzip")
with clawhdf5.File(tmp_h5, "r") as f:
for name in ["z", "zc"]:
result = f[name][:]
assert result.dtype == dtype
np.testing.assert_array_equal(result, original)
with h5py.File(tmp_h5, "r") as f:
for name in ["z", "zc"]:
assert f[name].dtype == dtype
assert f[name].id.get_type().get_class() == h5py.h5t.COMPOUND
np.testing.assert_array_equal(f[name][:], original)
def test_read_native_complex_from_h5py(tmp_h5):
"""HDF5 2.0's native complex type (class 11), written through h5py's
low-level API, reads as numpy complex."""
import h5py
from h5py import h5s, h5t
if not getattr(h5py.get_config(), "has_native_complex", False):
pytest.skip("h5py's libhdf5 predates 2.0")
original = np.array([1 + 2j, -3.5 + 0j, 0 - 1e-3j])
with h5py.File(tmp_h5, "w") as f:
for name, t, dt in [
(b"n64", h5t.COMPLEX_IEEE_F32LE, np.complex64),
(b"n128", h5t.COMPLEX_IEEE_F64LE, np.complex128),
]:
d = h5py.h5d.create(f.id, name, t, h5s.create_simple((3,)))
d.write(h5s.ALL, h5s.ALL, original.astype(dt), mtype=t)
with clawhdf5.File(tmp_h5, "r") as f:
for name, dt in [("n64", np.complex64), ("n128", np.complex128)]:
result = f[name][:]
assert result.dtype == dt
np.testing.assert_array_equal(result, original.astype(dt))
# ---------------------------------------------------------------------------
# Test: chunked + compressed datasets
# ---------------------------------------------------------------------------
+31
View File
@@ -114,3 +114,34 @@ let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes,
directory with range support (the server the tests use), and
`cargo run -p clawhdf5-remote --example read_url -- URL [DATASET]` lists a
file and prints what it cost.
## Other front ends
- `h5rs` (built with `--features remote`, or `remote-https`) takes URLs as
FILE arguments: [`clawhdf5-tools`](../clawhdf5-tools/README.md).
- Python: `clawhdf5.File("http://…")` and `File.open_url(url, ...)` go
through this crate: [`clawhdf5-py`](../clawhdf5-py/README.md).
- The browser does **not** use this crate (its cache fetches by blocking);
`clawhdf5-wasm`'s `openUrl` has its own restartable cache:
[`examples/wasm-viewer`](../../examples/wasm-viewer/README.md).
## Limits
Files a SWMR writer is still appending to cannot be followed remotely
(the file is pinned at open, so growth is `RemoteError::FileChanged`); the
block size is fixed rather than taken from a paged file's page size; the
cloud backends are built and unit-tested but have not been run against a
real bucket. The full list is under "Remote files (`clawhdf5-remote`)
limits" in [`docs/known-issues.md`](../../docs/known-issues.md); the design
is milestone M3 of [`docs/design/range-reads.md`](../../docs/design/range-reads.md).
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## License
MIT
+22 -2
View File
@@ -109,7 +109,16 @@ virtual datasets. Known differences from
h5dump:
- Floats print at their own precision (a `float32` 0.1 prints as `0.1`),
which for some values is more digits than h5dump's `%g`.
which for some values is more digits than h5dump's `%g`; non-finite
values print as `Inf`, `-Inf` and `NaN` (h5dump: `inf`, `-inf`, `nan`,
`-nan`).
- HDF5 2.x's predefined small floats print under h5dump 2.x's names
(`H5T_FLOAT_BFLOAT16LE`, `H5T_FLOAT_F8E4M3`, ..., `H5T_FLOAT_F4E2M1`;
`ls -v` as h5ls 2.x does: `FP4 E2M1 4-bit float`).
`tests/mx_floats_dump.rs` checks this, and the values, against the output
of h5dump 2.2.0 stored with the fixture. h5dump 1.14 describes these types
instead (`8-bit floating-point 4-bit precision`); other non-IEEE floats
print as an `H5T_FLOAT { ... }` block.
- A compound nested in a compound prints inline (`{ 1, 2.5 }`) where
h5dump prints it as an indented block, one member per line; only the
outer compound is a block.
@@ -306,4 +315,15 @@ CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 cargo test -p clawhd
The interop tests write their files with h5py and compare with h5ls, h5stat,
h5dump and h5diff; each skips when what it needs is missing unless
`CLAWHDF5_REQUIRE_INTEROP=1`.
`CLAWHDF5_REQUIRE_INTEROP=1`. `tests/remote.rs` runs every subcommand on
URLs against a local range server.
This crate also holds the interop tests of the library's in-place editor
(`clawhdf5::FileEditor`), since they use `h5rs check` and h5dump on every
edited file and compare index and heap structures with what libhdf5 makes
of the same edits:
```bash
CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 \
cargo test -p clawhdf5-tools --test edit_interop --test edit_coverage_interop
```
+2 -2
View File
@@ -916,7 +916,7 @@ impl Checker<'_> {
bad.push("has size 0".into());
} else if !filtered
&& let Some(cb) = chunk_bytes
&& u64::from(c.chunk_size) != cb
&& c.chunk_size != cb
{
bad.push(format!(
"is {} bytes; an unfiltered chunk is {cb}",
@@ -933,7 +933,7 @@ impl Checker<'_> {
);
}
}
self.extent(c.address, u64::from(c.chunk_size), path);
self.extent(c.address, c.chunk_size, path);
}
if reported > 50 {
self.problem(
+88
View File
@@ -59,9 +59,85 @@ pub fn is_ieee(dt: &Datatype) -> bool {
) == std
}
/// A float datatype predefined by libhdf5 2.x besides the IEEE ones.
struct SmallFloat {
/// h5dump's name (`H5T_FLOAT_F8E4M3`).
ddl: &'static str,
/// h5ls's description (`FP8 E4M3 8-bit float`).
long: &'static str,
/// The short name `ls` lists (`float8-e4m3`).
short: &'static str,
}
/// The libhdf5 2.x predefined float `dt` is equal to (as `H5Tequal` sees it:
/// the same size, byte order, precision, offset, fields and bias), if any.
/// The 1-byte types are predefined little-endian only.
fn small_float(dt: &Datatype) -> Option<SmallFloat> {
let Datatype::FloatingPoint {
size,
byte_order,
bit_offset,
bit_precision,
exponent_location,
exponent_size,
mantissa_location,
mantissa_size,
exponent_bias,
} = dt
else {
return None;
};
if *bit_offset != 0 || *mantissa_location != 0 {
return None;
}
let f = |ddl, long, short| Some(SmallFloat { ddl, long, short });
let fields = (
*size,
*bit_precision,
*exponent_location,
*exponent_size,
*mantissa_size,
*exponent_bias,
);
match (fields, byte_order) {
((2, 16, 7, 8, 7, 127), DatatypeByteOrder::LittleEndian) => f(
"H5T_FLOAT_BFLOAT16LE",
"bfloat16 16-bit little-endian float",
"bfloat16",
),
((2, 16, 7, 8, 7, 127), DatatypeByteOrder::BigEndian) => f(
"H5T_FLOAT_BFLOAT16BE",
"bfloat16 16-bit big-endian float",
"bfloat16-be",
),
((1, 8, 3, 4, 3, 7), DatatypeByteOrder::LittleEndian) => {
f("H5T_FLOAT_F8E4M3", "FP8 E4M3 8-bit float", "float8-e4m3")
}
((1, 8, 2, 5, 2, 15), DatatypeByteOrder::LittleEndian) => {
f("H5T_FLOAT_F8E5M2", "FP8 E5M2 8-bit float", "float8-e5m2")
}
((1, 6, 3, 2, 3, 1), DatatypeByteOrder::LittleEndian) => {
f("H5T_FLOAT_F6E2M3", "FP6 E2M3 6-bit float", "float6-e2m3")
}
((1, 6, 2, 3, 2, 3), DatatypeByteOrder::LittleEndian) => {
f("H5T_FLOAT_F6E3M2", "FP6 E3M2 6-bit float", "float6-e3m2")
}
((1, 4, 1, 2, 1, 1), DatatypeByteOrder::LittleEndian) => {
f("H5T_FLOAT_F4E2M1", "FP4 E2M1 4-bit float", "float4-e2m1")
}
_ => None,
}
}
/// Short name used by `ls`: `int32`, `float64-be`, `string[3]`, ...
pub fn short(dt: &Datatype) -> String {
if let Some(f) = small_float(dt) {
return f.short.into();
}
match dt {
Datatype::Complex { size, base_type } => {
short(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint {
size,
signed,
@@ -136,7 +212,13 @@ fn cset_word(c: &CharacterSet) -> &'static str {
/// h5ls -v style description.
pub fn long(dt: &Datatype) -> String {
if let Some(f) = small_float(dt) {
return f.long.into();
}
match dt {
Datatype::Complex { size, base_type } => {
long(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint {
size,
signed,
@@ -277,6 +359,8 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
u64::from(*size) * 8,
order_suffix(byte_order)
)),
// As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
Datatype::BitField {
size, byte_order, ..
} => Some(format!(
@@ -414,6 +498,9 @@ fn string_ddl(
/// hdf5-json type object.
pub fn json(dt: &Datatype) -> J {
match dt {
Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint { .. } => {
json!({"class": "H5T_INTEGER", "base": atomic_ddl(dt)})
}
@@ -511,6 +598,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str {
match dt {
Datatype::Complex { .. } => "compound",
Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time",
+1 -1
View File
@@ -224,7 +224,7 @@ pub fn allocated_bytes(h5: &H5, info: &DsInfo) -> Result<u64> {
let ds = info.ds.as_ref().map_err(Clone::clone)?;
chunks(h5, layout, ds, dt)?
.iter()
.map(|c| u64::from(c.chunk_size))
.map(|c| c.chunk_size)
.sum()
}
DataLayout::Virtual { .. } => 0,
+5
View File
@@ -183,6 +183,11 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into());
};
match dt {
Datatype::Complex { size, base_type } => self.decode(
&Datatype::complex_as_compound(*size, base_type),
b,
depth + 1,
),
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()),
+736
View File
@@ -0,0 +1,736 @@
//! Files written with `FileBuilder::libver_bounds(LibVer::V18, LibVer::V18)`
//! must be readable by HDF5 1.8. Every writer feature is written under that
//! bound and read back by HDF5 1.8.23's h5dump (values dumped in binary and
//! compared), by h5py (libhdf5 2.x), by h5dump 1.14, by clawhdf5 and by
//! `h5rs check --data`; then `FileEditor` grows, appends to and annotates
//! the file (splitting version-1 B-tree nodes) and every reader checks it
//! again. The version-1 B-trees we write are compared node by node with
//! the ones libhdf5 writes for the same data under h5py's
//! `libver=('v108', 'latest')`.
//!
//! HDF5 1.8 is found through `CLAWHDF5_H5DUMP18` (the path to its h5dump)
//! or at `~/.cache/hdf5-1.8.23/bin/h5dump`, where
//! `scripts/build-hdf5-1.8.sh` builds it; without it the 1.8 checks are
//! skipped (CI has no HDF5 1.8), even with `CLAWHDF5_REQUIRE_INTEROP=1`.
//! h5py/numpy (`CLAWHDF5_PYTHON`) and h5dump are needed otherwise; they
//! skip when missing unless `CLAWHDF5_REQUIRE_INTEROP=1`.
use std::path::{Path, PathBuf};
use std::process::{Command, Output};
use clawhdf5::{AttrValue, File, FileBuilder, FileEditor, LibVer, Selection};
use clawhdf5_format::datatype::{CharacterSet, Datatype, StringPadding};
use clawhdf5_format::file_writer::{CompoundTypeBuilder, EnumTypeBuilder};
use clawhdf5_format::type_builders::make_i32_type;
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
fn available(cmd: &str, args: &[&str]) -> bool {
Command::new(cmd)
.args(args)
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
fn tools_ok() -> bool {
let ok =
available(&python(), &["-c", "import h5py, numpy"]) && available("h5dump", &["--version"]);
if !ok {
assert!(
!interop_required(),
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy or h5dump is not available"
);
eprintln!("SKIP: h5py/numpy or h5dump not available");
}
ok
}
/// HDF5 1.8's h5dump, when there is one.
fn h5dump18() -> Option<PathBuf> {
let p = match std::env::var_os("CLAWHDF5_H5DUMP18") {
Some(p) => PathBuf::from(p),
None => PathBuf::from(std::env::var_os("HOME")?).join(".cache/hdf5-1.8.23/bin/h5dump"),
};
let o = Command::new(&p).arg("--version").output().ok()?;
let v = String::from_utf8_lossy(&o.stdout).to_string();
if !v.contains("1.8.") {
eprintln!("SKIP 1.8 checks: {} is not HDF5 1.8 ({v})", p.display());
return None;
}
Some(p)
}
fn py(script: &str) -> String {
let o = Command::new(python())
.args(["-c", script])
.output()
.expect("run python");
assert!(
o.status.success(),
"python failed:\n{script}\nSTDOUT: {}\nSTDERR: {}",
String::from_utf8_lossy(&o.stdout),
String::from_utf8_lossy(&o.stderr)
);
String::from_utf8_lossy(&o.stdout).trim().to_string()
}
fn text(o: &Output) -> String {
format!(
"{}{}",
String::from_utf8_lossy(&o.stdout),
String::from_utf8_lossy(&o.stderr)
)
}
fn tmpdir() -> tempfile::TempDir {
tempfile::TempDir::new_in(env!("CARGO_TARGET_TMPDIR")).unwrap()
}
/// The hyperslab of `count` elements from `start`.
fn block(start: &[u64], count: &[u64]) -> Selection {
Selection::Hyperslab {
start: start.to_vec(),
stride: vec![1; start.len()],
count: count.to_vec(),
block: vec![1; start.len()],
}
}
fn le<T: Copy, const N: usize>(v: &[T], f: impl Fn(T) -> [u8; N]) -> Vec<u8> {
v.iter().flat_map(|&x| f(x)).collect()
}
/// The datasets of the test file and the bytes each holds (little-endian,
/// row-major, as h5py's `tobytes()` and h5dump's `-b LE` give them).
struct Expect {
datasets: Vec<(String, Vec<u8>)>,
}
impl Expect {
fn set(&mut self, name: &str, bytes: Vec<u8>) {
match self.datasets.iter_mut().find(|(n, _)| n == name) {
Some(e) => e.1 = bytes,
None => self.datasets.push((name.to_string(), bytes)),
}
}
}
const MANY: usize = 100_000;
/// Write every feature under the 1.8 bound.
fn write_file(path: &Path) -> Expect {
let mut e = Expect {
datasets: Vec::new(),
};
let mut b = FileBuilder::new();
b.libver_bounds(LibVer::V18, LibVer::V18);
// Contiguous, with dense attributes (more than 8).
let v: Vec<f64> = (0..1000).map(|i| i as f64 * 0.5).collect();
let d = b.create_dataset("contig").with_f64_data(&v);
for i in 0..12 {
d.set_attr(&format!("a{i:02}"), AttrValue::I64(i));
}
e.set("/contig", le(&v, f64::to_le_bytes));
b.create_dataset("empty").with_f64_data(&[]);
e.set("/empty", vec![]);
b.create_dataset("scalar")
.with_f64_data(&[2.5])
.with_shape(&[]);
e.set("/scalar", 2.5f64.to_le_bytes().to_vec());
let v: Vec<i32> = (0..16).map(|i| i * 3 - 7).collect();
b.create_dataset("compact").with_i32_data(&v).compact();
e.set("/compact", le(&v, i32::to_le_bytes));
let v: Vec<f32> = (0..20).map(|i| i as f32 / 3.0).collect();
b.create_dataset("f32").with_f32_data(&v);
e.set("/f32", le(&v, f32::to_le_bytes));
let v: Vec<f32> = vec![0.5, -2.0, 1024.0, 0.0];
b.create_dataset("f16").with_f16_data(&v);
e.set(
"/f16",
le(&v, |x| {
clawhdf5_format::float16::f32_to_f16_bits(x).to_le_bytes()
}),
);
// Chunked with every built-in filter HDF5 1.8 has, and edge chunks.
let v: Vec<f32> = (0..60 * 70).map(|i| (i % 97) as f32 * 1.25).collect();
b.create_dataset("chunked")
.with_f32_data(&v)
.with_shape(&[60, 70])
.with_chunks(&[16, 16])
.with_deflate(6)
.with_shuffle()
.with_fletcher32();
e.set("/chunked", le(&v, f32::to_le_bytes));
// What the 1.10 indexes would be: Extensible Array, version-2 B-tree,
// Fixed Array, single chunk. All become version-1 B-trees.
let v: Vec<i32> = (0..25).collect();
b.create_dataset("resizable")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[4]);
e.set("/resizable", le(&v, i32::to_le_bytes));
let v: Vec<i64> = (0..30).map(|i| i * 1_000_000_007).collect();
b.create_dataset("resizable2")
.with_i64_data(&v)
.with_shape(&[5, 6])
.with_maxshape(&[u64::MAX, u64::MAX])
.with_chunks(&[2, 4]);
e.set("/resizable2", le(&v, i64::to_le_bytes));
let v: Vec<i32> = (0..10).map(|i| -i).collect();
b.create_dataset("fixedmax")
.with_i32_data(&v)
.with_maxshape(&[100])
.with_chunks(&[3])
.with_deflate(1);
e.set("/fixedmax", le(&v, i32::to_le_bytes));
let v: Vec<i32> = (0..8).map(|i| i * i).collect();
b.create_dataset("single")
.with_i32_data(&v)
.with_chunks(&[8]);
e.set("/single", le(&v, i32::to_le_bytes));
// Enough chunks for a three-level tree, and a 2-D two-level one.
let v: Vec<i32> = (0..MANY as i32).map(|i| i ^ 0x5a5a).collect();
b.create_dataset("many")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[1]);
e.set("/many", le(&v, i32::to_le_bytes));
let v: Vec<f64> = (0..100 * 100).map(|i| i as f64).collect();
b.create_dataset("grid")
.with_f64_data(&v)
.with_shape(&[100, 100])
.with_chunks(&[1, 1]);
e.set("/grid", le(&v, f64::to_le_bytes));
b.create_dataset("empty_chunked")
.with_f64_data(&[])
.with_maxshape(&[u64::MAX])
.with_chunks(&[10]);
e.set("/empty_chunked", vec![]);
let v: Vec<i32> = (0..12).collect();
b.create_dataset("filled")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[5])
.with_fill_value(&(-9i32).to_le_bytes());
e.set("/filled", le(&v, i32::to_le_bytes));
// Datatypes.
let raw = b"abcdehello\0\0\0\0\0".to_vec();
b.create_dataset("strings").with_compound_data(
Datatype::String {
size: 5,
padding: StringPadding::NullPad,
charset: CharacterSet::Ascii,
},
raw.clone(),
3,
);
e.set("/strings", raw);
let ct = CompoundTypeBuilder::new()
.i32_field("a")
.f64_field("b")
.build();
let mut raw = Vec::new();
for i in 0..5i32 {
raw.extend_from_slice(&i.to_le_bytes());
raw.extend_from_slice(&(f64::from(i) * 1.5).to_le_bytes());
}
b.create_dataset("compound")
.with_compound_data(ct, raw.clone(), 5);
e.set("/compound", raw);
let et = EnumTypeBuilder::i32_based()
.value("RED", 0)
.value("GREEN", 1)
.value("BLUE", 7)
.build();
let v = [0, 7, 1, 1, 0];
b.create_dataset("enum").with_enum_i32_data(et, &v);
e.set("/enum", le(&v, i32::to_le_bytes));
let v: Vec<i32> = (0..24).collect();
b.create_dataset("array").with_array_data(
make_i32_type(),
&[2, 3],
le(&v, i32::to_le_bytes),
4,
);
e.set("/array", le(&v, i32::to_le_bytes));
let v: Vec<i64> = vec![-1, 0, i64::MAX];
b.create_dataset("i64").with_i64_data(&v);
e.set("/i64", le(&v, i64::to_le_bytes));
// Groups: compact with links of every kind, dense (links and
// attributes), and tracking creation order.
let mut g = b.create_group("g_compact");
g.create_dataset("x").with_i32_data(&[1, 2, 3]);
e.set("/g_compact/x", le(&[1i32, 2, 3], i32::to_le_bytes));
g.add_soft_link("soft", "/contig");
g.add_external_link("ext", "other.h5", "/data");
g.set_attr("title", AttrValue::String("compact group".into()));
b.add_group(g.finish());
b.add_hard_link("hard", "/g_compact/x");
let mut g = b.create_group("g_dense");
for i in 0..20 {
let v = [i, i + 1];
g.create_dataset(&format!("d{i:02}")).with_i32_data(&v);
e.set(&format!("/g_dense/d{i:02}"), le(&v, i32::to_le_bytes));
}
for i in 0..12 {
g.set_attr(&format!("attr{i:02}"), AttrValue::F64(f64::from(i) / 4.0));
}
b.add_group(g.finish());
let mut g = b.create_group("g_order");
g.track_order(true);
for i in 0..10 {
let n = format!("z{}", 9 - i);
g.create_dataset(&n).with_i32_data(&[i]);
e.set(&format!("/g_order/{n}"), i.to_le_bytes().to_vec());
}
for i in 0..10 {
g.set_attr(&format!("b{}", 9 - i), AttrValue::I64(i));
}
b.add_group(g.finish());
b.create_dataset("deep/er/path").with_i32_data(&[42]);
e.set("/deep/er/path", 42i32.to_le_bytes().to_vec());
b.set_attr("version", AttrValue::F64(1.8));
b.set_attr("ints", AttrValue::I64Array(vec![1, -2, 3]));
b.set_attr("text", AttrValue::String("readable by 1.8".into()));
b.set_attr(
"texts",
AttrValue::StringArray(vec!["a".into(), "bb".into(), "ccc".into()]),
);
b.set_attr("big", AttrValue::U64(u64::MAX));
b.write(path).unwrap();
e
}
/// Structural facts of a file: superblock and layout message versions.
fn check_versions(path: &Path) {
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
let bytes = std::fs::read(path).unwrap();
assert_eq!(bytes[8], 2, "superblock version");
let f = File::open(path).unwrap();
let sb = f.superblock().clone();
for name in [
"contig",
"compact",
"chunked",
"resizable",
"resizable2",
"many",
] {
let addr = clawhdf5_format::group_v2::resolve_path_any(&bytes, &sb, name).unwrap();
let oh = ObjectHeader::parse(&bytes, addr as usize, 8, 8).unwrap();
let layout = oh
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap();
assert_eq!(layout.data[0], 3, "layout version of {name}");
}
}
/// clawhdf5 reads every dataset's bytes back.
fn check_ours(path: &Path, e: &Expect) {
let f = File::open(path).unwrap();
for (name, want) in &e.datasets {
let got = f
.dataset(name)
.unwrap()
.read_selection(&Selection::All)
.unwrap();
assert!(got == *want, "our read of {name} differs");
}
let attrs = f.dataset("contig").unwrap().attrs().unwrap();
assert!((12..=13).contains(&attrs.len()), "{attrs:?}");
}
/// h5py (libhdf5 2.x) reads every dataset's bytes back.
fn check_h5py(path: &Path, e: &Expect, dir: &Path) {
let mut script = format!(
"import h5py, numpy as np\nf = h5py.File({p:?}, 'r')\n",
p = path.to_str().unwrap()
);
for (i, (name, want)) in e.datasets.iter().enumerate() {
let exp = dir.join(format!("expect{i}.bin"));
std::fs::write(&exp, want).unwrap();
script.push_str(&format!(
"a = np.ascontiguousarray(f[{name:?}][()]).tobytes()\n\
assert a == open({x:?}, 'rb').read(), {name:?}\n",
x = exp.to_str().unwrap()
));
}
script.push_str(
"assert f.attrs['text'] in (b'readable by 1.8', 'readable by 1.8')\n\
assert list(f.attrs['ints']) == [1, -2, 3]\n\
assert len(f['g_dense'].attrs) == 12 and len(f['g_dense']) == 20\n\
assert list(f['g_order']) == ['z9', 'z8', 'z7', 'z6', 'z5', 'z4', 'z3', 'z2', 'z1', 'z0']\n\
assert f['g_compact/soft'].shape == (1000,)\n\
assert f['resizable'].maxshape == (None,)\n\
assert f['filled'].fillvalue == -9\n\
print('ok')\n",
);
assert_eq!(py(&script), "ok");
}
/// h5dump (1.14) and `h5rs check --data` accept the file.
fn check_tools(path: &Path) {
let p = path.to_str().unwrap();
let o = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["check", "--data", "-q", p])
.output()
.unwrap();
assert!(o.status.success(), "h5rs check --data {p}:\n{}", text(&o));
let o = Command::new("h5dump")
.args(["-o", "/dev/null", p])
.output()
.unwrap();
assert!(
o.status.success() && o.stderr.is_empty(),
"h5dump {p}:\n{}",
text(&o)
);
}
/// HDF5 1.8's h5dump reads the whole file exactly as h5dump 1.14 does, and
/// dumps each numeric dataset's values as the bytes we wrote.
fn check_18(h5dump: &Path, path: &Path, e: &Expect, dir: &Path) {
let p = path.to_str().unwrap();
let o = Command::new(h5dump).arg(p).output().unwrap();
assert!(
o.status.success() && o.stderr.is_empty(),
"h5dump 1.8 {p}:\n{}",
text(&o)
);
let out18 = String::from_utf8_lossy(&o.stdout).to_string();
assert!(out18.contains("EXTERNAL_LINK \"ext\""), "{out18}");
let o = Command::new("h5dump").arg(p).output().unwrap();
assert!(o.status.success(), "h5dump {p}:\n{}", text(&o));
// HDF5 1.8 has no name for IEEE half floats.
let out = String::from_utf8_lossy(&o.stdout).replace(
"H5T_IEEE_F16LE",
"16-bit little-endian floating-point 16-bit precision",
);
if let Some((n, (a, b))) = out18
.lines()
.zip(out.lines())
.enumerate()
.find(|(_, (a, b))| a != b)
{
panic!("h5dump 1.8 and h5dump differ at line {}:\n{a}\n{b}", n + 1);
}
assert_eq!(out18.lines().count(), out.lines().count());
// `-b` writes nothing for compounds, enums, arrays, strings and half
// floats (HDF5 1.8
// has no native half float), for h5py's files too: those are covered
// by the text above.
let skip = ["/f16", "/strings", "/compound", "/enum", "/array"];
for (i, (name, want)) in e.datasets.iter().enumerate() {
if want.is_empty() || skip.contains(&name.as_str()) {
continue;
}
let bin = dir.join(format!("dump18_{i}.bin"));
let o = Command::new(h5dump)
.args(["-d", name, "-b", "LE", "-o"])
.arg(&bin)
.arg(p)
.output()
.unwrap();
assert!(o.status.success(), "h5dump 1.8 -d {name}:\n{}", text(&o));
let got = std::fs::read(&bin).unwrap();
assert!(got == *want, "HDF5 1.8 read of {name} differs");
}
}
fn check_all(path: &Path, e: &Expect, dir: &Path, h5dump: Option<&Path>) {
check_ours(path, e);
check_h5py(path, e, dir);
check_tools(path);
if let Some(h) = h5dump {
check_18(h, path, e, dir);
}
}
#[test]
fn every_feature_reads_in_hdf5_1_8() {
if !tools_ok() {
return;
}
let h18 = h5dump18();
let dir = tmpdir();
let path = dir.path().join("v18.h5");
let mut e = write_file(&path);
check_versions(&path);
check_all(&path, &e, dir.path(), h18.as_deref());
// FileEditor: grow the unlimited datasets (splitting B-tree nodes),
// overwrite values in filtered and unfiltered chunks, set attributes
// (compact and dense).
let mut ed = FileEditor::open(&path).unwrap();
let mut res: Vec<i32> = (0..25).collect();
for round in 0..300usize {
let n = res.len() as u64;
let add = 1 + (round % 9) as u64;
ed.resize("resizable", &[n + add]).unwrap();
let vals: Vec<i32> = (0..add).map(|k| (n + k) as i32 * 7 - 3).collect();
ed.write_values("resizable", &block(&[n], &[add]), &vals)
.unwrap();
res.extend(&vals);
}
e.set("/resizable", le(&res, i32::to_le_bytes));
let mut many: Vec<i32> = (0..MANY as i32).map(|i| i ^ 0x5a5a).collect();
ed.resize("many", &[MANY as u64 + 500]).unwrap();
let vals: Vec<i32> = (0..500).collect();
ed.write_values("many", &block(&[MANY as u64], &[500]), &vals)
.unwrap();
many.extend(&vals);
many[12_345] = -1;
ed.write_values("many", &block(&[12_345], &[1]), &[-1i32])
.unwrap();
e.set("/many", le(&many, i32::to_le_bytes));
let mut chunked: Vec<f32> = (0..60 * 70).map(|i| (i % 97) as f32 * 1.25).collect();
let row: Vec<f32> = (0..70).map(|i| -(i as f32)).collect();
ed.write_values("chunked", &block(&[31, 0], &[1, 70]), &row)
.unwrap();
chunked[31 * 70..32 * 70].copy_from_slice(&row);
e.set("/chunked", le(&chunked, f32::to_le_bytes));
ed.resize("resizable2", &[9, 6]).unwrap();
let mut r2: Vec<i64> = (0..30).map(|i| i * 1_000_000_007).collect();
r2.extend(std::iter::repeat_n(0, 24));
e.set("/resizable2", le(&r2, i64::to_le_bytes));
ed.set_attr("contig", "added", &AttrValue::F64(3.25))
.unwrap();
ed.set_attr("g_dense", "attr03", &AttrValue::F64(-1.0))
.unwrap();
ed.set_attr("/", "text", &AttrValue::String("edited".into()))
.unwrap();
drop(ed);
check_ours(&path, &e);
let f = File::open(&path).unwrap();
assert!(matches!(
f.dataset("contig").unwrap().attr("added").unwrap(),
Some(AttrValue::F64(v)) if v == 3.25
));
drop(f);
// Everything but the root's "text" attribute is as check_h5py expects.
py(&format!(
"import h5py\nf = h5py.File({p:?}, 'r')\n\
assert f.attrs['text'] in (b'edited', 'edited')\n\
assert f['contig'].attrs['added'] == 3.25\n\
assert f['g_dense'].attrs['attr03'] == -1.0\n",
p = path.to_str().unwrap()
));
let o = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["check", "--data", "-q", path.to_str().unwrap()])
.output()
.unwrap();
assert!(o.status.success(), "h5rs check after edits:\n{}", text(&o));
if let Some(h) = h18.as_deref() {
check_18(h, &path, &e, dir.path());
}
// h5py (libhdf5 2.x) goes on appending to the 1.8-format file, and 1.8
// still reads it.
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'r+') as f:\n\
\x20 d = f['resizable']\n\
\x20 n = d.shape[0]\n\
\x20 d.resize((n + 100,))\n\
\x20 d[n:] = np.arange(100, dtype='<i4') + 5000\n",
p = path.to_str().unwrap()
));
res.extend((0..100).map(|k| 5000 + k));
e.set("/resizable", le(&res, i32::to_le_bytes));
check_ours(&path, &e);
if let Some(h) = h18.as_deref() {
check_18(h, &path, &e, dir.path());
}
}
// ---- The trees against libhdf5's ----
/// One node of a version-1 chunk B-tree: level, and per key its stored
/// size and offsets; children below.
#[derive(Debug, PartialEq)]
struct TreeNode {
level: u8,
keys: Vec<(u32, Vec<u64>)>,
children: Vec<TreeNode>,
}
fn read_tree(bytes: &[u8], addr: u64, ndims: usize, sizes: bool) -> TreeNode {
let a = addr as usize;
assert_eq!(&bytes[a..a + 4], b"TREE");
let level = bytes[a + 5];
let n = u16::from_le_bytes([bytes[a + 6], bytes[a + 7]]) as usize;
let mut p = a + 24;
let mut keys = Vec::new();
let mut kids = Vec::new();
for i in 0..=n {
let size = u32::from_le_bytes(bytes[p..p + 4].try_into().unwrap());
let offs = (0..ndims)
.map(|d| u64::from_le_bytes(bytes[p + 8 + 8 * d..p + 16 + 8 * d].try_into().unwrap()))
.collect();
keys.push((if sizes { size } else { 0 }, offs));
p += 8 + 8 * ndims;
if i < n {
kids.push(u64::from_le_bytes(bytes[p..p + 8].try_into().unwrap()));
p += 8;
}
}
let children = if level > 0 {
kids.iter()
.map(|&c| read_tree(bytes, c, ndims, sizes))
.collect()
} else {
Vec::new()
};
TreeNode {
level,
keys,
children,
}
}
/// Where two trees first differ (libhdf5's first).
fn first_difference(a: &TreeNode, b: &TreeNode, at: &str) -> Option<String> {
if a.level != b.level || a.keys.len() != b.keys.len() {
return Some(format!(
"{at}: level {} with {} keys vs level {} with {} keys",
a.level,
a.keys.len(),
b.level,
b.keys.len()
));
}
if let Some(i) = (0..a.keys.len()).find(|&i| a.keys[i] != b.keys[i]) {
return Some(format!("{at} key {i}: {:?} vs {:?}", a.keys[i], b.keys[i]));
}
a.children
.iter()
.zip(&b.children)
.enumerate()
.find_map(|(i, (x, y))| first_difference(x, y, &format!("{at}/{i}")))
}
/// The chunk B-tree of dataset `name`: (tree, ndims) from its layout.
fn tree_of(path: &Path, name: &str, sizes: bool) -> TreeNode {
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
let bytes = std::fs::read(path).unwrap();
let f = File::open(path).unwrap();
let sb = f.superblock().clone();
let addr = clawhdf5_format::group_v2::resolve_path_any(&bytes, &sb, name).unwrap();
let oh = ObjectHeader::parse(&bytes, addr as usize, 8, 8).unwrap();
let l = &oh
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap()
.data;
assert_eq!((l[0], l[1]), (3, 2), "{name}: layout v3, chunked");
let ndims = l[2] as usize;
let root = u64::from_le_bytes(l[3..11].try_into().unwrap());
read_tree(&bytes, root, ndims, sizes)
}
/// Our version-1 chunk B-trees are libhdf5's, node for node (levels, child
/// counts, every key's offsets, and its chunk size where the chunks are
/// the same bytes), for 1-D, 2-D and 3-D datasets with two- and three-level
/// trees, filtered or not.
///
/// libhdf5 inserts each chunk into the tree when it leaves its chunk cache.
/// A whole-dataset write with no cache (`rdcc_nbytes=0`, or chunks larger
/// than the cache) inserts them in row-major order, as we build the tree;
/// with the default cache small chunks of a 1-D dataset still arrive in
/// order, but those of a multi-dimensional one arrive in the order the
/// cache's hash evicts them, which gives the same keys in differently
/// filled nodes. Both trees index the same chunks; we do not model the
/// cache.
#[test]
fn chunk_btrees_match_libhdf5() {
if !tools_ok() {
return;
}
let dir = tmpdir();
let theirs = dir.path().join("libhdf5.h5");
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'w', libver=('v108', 'latest'), rdcc_nbytes=0) as f:\n\
\x20 f.create_dataset('d1000', data=np.arange(10000.0), chunks=(10,))\n\
\x20 f.create_dataset('d999', data=np.arange(9990.0), chunks=(10,))\n\
\x20 f.create_dataset('big', data=np.arange(100000, dtype='<i4'), chunks=(1,), maxshape=(None,))\n\
\x20 f.create_dataset('grid', data=np.arange(10000.0).reshape(100, 100), chunks=(1, 1))\n\
\x20 f.create_dataset('cube', data=np.arange(27000, dtype='<i2').reshape(30, 30, 30), chunks=(2, 3, 5))\n\
\x20 f.create_dataset('gz', data=np.arange(20000, dtype='<i8') % 13, chunks=(7,), compression='gzip')\n",
p = theirs.to_str().unwrap()
));
let ours = dir.path().join("ours.h5");
let mut b = FileBuilder::new();
b.libver_bounds(LibVer::V18, LibVer::V18);
let v: Vec<f64> = (0..10000).map(f64::from).collect();
b.create_dataset("d1000")
.with_f64_data(&v)
.with_chunks(&[10]);
b.create_dataset("d999")
.with_f64_data(&v[..9990])
.with_chunks(&[10]);
let v: Vec<i32> = (0..100_000).collect();
b.create_dataset("big")
.with_i32_data(&v)
.with_chunks(&[1])
.with_maxshape(&[u64::MAX]);
let v: Vec<f64> = (0..10000).map(f64::from).collect();
b.create_dataset("grid")
.with_f64_data(&v)
.with_shape(&[100, 100])
.with_chunks(&[1, 1]);
let raw: Vec<u8> = (0..27000i16).flat_map(|x| x.to_le_bytes()).collect();
b.create_dataset("cube")
.with_compound_data(
clawhdf5_format::datatype::Datatype::FixedPoint {
size: 2,
byte_order: clawhdf5_format::datatype::DatatypeByteOrder::LittleEndian,
signed: true,
bit_offset: 0,
bit_precision: 16,
},
raw,
27000,
)
.with_shape(&[30, 30, 30])
.with_chunks(&[2, 3, 5]);
let v: Vec<i64> = (0..20000).map(|i| i % 13).collect();
b.create_dataset("gz")
.with_i64_data(&v)
.with_chunks(&[7])
.with_deflate(4);
b.write(&ours).unwrap();
for (name, sizes) in [
("d1000", true),
("d999", true),
("big", true),
("grid", true),
("cube", true),
// Compressed sizes differ between zlib-rs and zlib.
("gz", false),
] {
let a = tree_of(&theirs, name, sizes);
let b = tree_of(&ours, name, sizes);
if let Some(d) = first_difference(&a, &b, "root") {
panic!("{name}: our chunk B-tree differs from libhdf5's: {d}");
}
}
}
@@ -0,0 +1,111 @@
//! `h5rs dump` and `h5rs ls` of the small floats libhdf5 2.x predefines
//! (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1), against the output
//! of h5dump 2.2.0 and h5ls 2.2.0 stored next to the fixture (the Debian
//! h5dump CI installs, 1.14.x, predates these types). See
//! `crates/clawhdf5/tests/fixtures/gen_mx_floats.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into the lines outside `DATA { ... }` blocks and the values
/// inside them, per block.
fn split(ddl: &str) -> (Vec<&str>, Vec<Vec<String>>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
let mut values: Option<Vec<String>> = None;
for line in ddl.lines() {
match &mut values {
None if line.trim() == "DATA {" => values = Some(Vec::new()),
None => frame.push(line),
Some(v) if line.trim() == "}" => {
data.push(std::mem::take(v));
values = None;
}
Some(v) => {
let (_, body) = line.split_once("):").expect("an indexed data line");
v.extend(
body.split(',')
.map(str::trim)
.filter(|s| !s.is_empty())
.map(String::from),
);
}
}
}
(frame, data)
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("mx_floats_hdf5_2_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("mx_floats_hdf5_2_2.ddl")).unwrap();
let (our_frame, our_data) = split(&ours);
let (ref_frame, ref_data) = split(&reference);
// Everything but the values is h5dump's byte for byte: the datatypes
// print as H5T_FLOAT_BFLOAT16LE, H5T_FLOAT_F4E2M1, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), 17);
assert_eq!(our_data.len(), ref_data.len());
// Values: h5dump prints `%g` (6 significant digits) and inf/-inf/nan/
// -nan; h5rs prints the shortest string that round-trips and Inf/-Inf/NaN.
for (block, (ours, theirs)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(ours.len(), theirs.len(), "block {block}");
for (i, (o, r)) in ours.iter().zip(theirs).enumerate() {
let o: f64 = o.parse().unwrap();
let r: f64 = r.parse().unwrap();
let same = if r.is_nan() || r.is_infinite() {
o.is_nan() == r.is_nan() && (r.is_nan() || o == r)
} else {
// Both round to the same f32 (a bfloat16 subnormal prints
// as its shortest f32 string, `9.1835e-41`, fewer digits
// than h5dump's), or agree to h5dump's 6 digits.
o as f32 == r as f32 || (o - r).abs() <= 5e-6 * r.abs()
};
assert!(same, "block {block}[{i}]: h5rs {o}, h5dump {r}");
}
}
}
#[test]
fn ls_names_small_floats_like_h5ls_2_2() {
let path = fixture("mx_floats_hdf5_2_2.h5");
// h5ls 2.2.0 -v, `Type:` line of each dataset.
for (name, long, short) in [
("bf16le", "bfloat16 16-bit little-endian float", "bfloat16"),
("bf16be", "bfloat16 16-bit big-endian float", "bfloat16-be"),
("f8e4m3", "FP8 E4M3 8-bit float", "float8-e4m3"),
("f8e5m2", "FP8 E5M2 8-bit float", "float8-e5m2"),
("f6e2m3", "FP6 E2M3 6-bit float", "float6-e2m3"),
("f6e3m2", "FP6 E3M2 6-bit float", "float6-e3m2"),
("f4e2m1", "FP4 E2M1 4-bit float", "float4-e2m1"),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
let listing = h5rs(&["ls", path.to_str().unwrap()]);
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
+51
View File
@@ -0,0 +1,51 @@
# clawhdf5-wasm
clawhdf5's HDF5 and NetCDF-4 reader compiled to WebAssembly with
wasm-bindgen, for the browser (and Node). Read-only. Two ways in:
- `open(bytes)` — a file already in memory (a dropped file, a fetched
blob);
- `openUrl(url, opts)` — a file on a web server, read by HTTP range
requests as each call needs its bytes, without downloading it
(range-read milestone M4, [`docs/design/range-reads.md`](../../docs/design/range-reads.md)).
Both give `list`, `info`, `attrs`, `read` and `readHyperslab`; the remote
file's methods return promises, and `stats()` counts requests and bytes.
The JavaScript API, options, limits, package size and tests are documented
with the demo page, [`examples/wasm-viewer/README.md`](../../examples/wasm-viewer/README.md).
## Layout
- `src/core.rs` — the reader over any `clawhdf5_format::storage::Storage`
(`Reader::open_storage`), plain Rust and tested natively.
- `src/lazy.rs` — the restartable "NeedBytes" cache behind `openUrl`: a
call runs as a pass over the blocks fetched so far; a pass that misses is
abandoned, the missing (and hinted) blocks are fetched, and the pass is
run again. No block is evicted while a call runs.
- `js/remote.js` — the HTTP side: `fetch` with `Range`, checking every
answer (a `206` with exactly the bytes asked for, same ETag/Last-Modified
and length) so a call fails rather than return another file's bytes.
- `src/lib.rs` — the wasm-bindgen exports.
## Build and test
```bash
rustup target add wasm32-unknown-unknown
cargo install wasm-bindgen-cli --version 0.2.129 # must equal the crate's wasm-bindgen
bash examples/wasm-viewer/build.sh # -> examples/wasm-viewer/pkg/
cargo test -p clawhdf5-wasm # native: h5py_interop, lazy, vl_strings
bash examples/wasm-viewer/test/run.sh # Node + headless Chromium (not in CI)
```
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus cargo test -p
clawhdf5-wasm --test lazy` compares every corpus file read lazily with the
same file read from bytes.
Built without `mmap` and `parallel` and without the Zstd and SZIP filters
(they link C): such datasets fail with `unsupported filter`. No C is
compiled; `publish = false` (it is distributed as the package
`build.sh` makes).
## License
MIT
+3
View File
@@ -487,6 +487,9 @@ pub fn describe(dt: &Datatype) -> String {
}
}
match dt {
Datatype::Complex { size, base_type } => {
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint {
size,
signed,
+193 -6
View File
@@ -25,6 +25,14 @@
//! pass per block it needs. In practice it is one pass per *wave* of
//! misses: a chunked read asks for all the chunks of a batch at once.
//!
//! Parsers also say what they are about to read ([`Storage::hint`]: a
//! node's children, a structure's body, a listed group's child headers).
//! A pass that misses fetches the hinted blocks it lacks too, as far as the
//! operation's fetch budget allows, so what the parser would only have
//! reached on the next pass arrives in the same round trip; a pass that
//! misses nothing ignores its hints, so they never add a round trip, and
//! results never depend on them.
//!
//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no
//! operation is in flight. Blocks fetched for bulk reads (raw data: a
//! `read_ranges` call, or a read longer than a block) go first, so reading
@@ -50,6 +58,14 @@ pub const DEFAULT_MAX_FETCH: u64 = 512 << 20;
/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run.
pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)";
/// Most blocks one pass records as hinted (see [`Storage::hint`]).
const MAX_HINTED: usize = 1 << 16;
/// Total length of `ranges`.
fn ranges_len(ranges: &[Range<u64>]) -> u64 {
ranges.iter().map(|r| r.end - r.start).sum()
}
/// Settings of a [`LazyStorage`].
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct LazyConfig {
@@ -94,6 +110,9 @@ pub struct LazyStats {
pub bytes_fetched: u64,
/// Blocks evicted to stay within the budget.
pub evictions: u64,
/// Blocks asked for because a parser hinted it would read them
/// ([`Storage::hint`]), not because a pass missed them.
pub hinted_blocks: u64,
/// Bytes cached now.
pub cached_bytes: u64,
}
@@ -125,6 +144,9 @@ struct State {
/// Blocks the current pass missed, and whether a small read wanted
/// them (metadata).
missing: HashMap<u64, bool>,
/// Blocks the current pass was told it is about to read and that are
/// not cached ([`Storage::hint`]), in file order.
hinted: BTreeSet<u64>,
/// Blocks a bulk read missed that have not been supplied yet: kept
/// as bulk when they arrive.
bulk_pending: BTreeSet<u64>,
@@ -155,6 +177,19 @@ pub struct Operation<'a> {
}
impl Operation<'_> {
/// One pass of `f`, as [`LazyStorage::attempt`], but a pass that misses
/// also asks for the blocks it was hinted it would read, as many as fit
/// in what is left of the operation's budget ([`LazyConfig::max_fetch`])
/// after the blocks it missed.
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
let left = self
.storage
.config
.max_fetch
.saturating_sub(self.fetched.get());
self.storage.attempt_within(f, left)
}
/// Count `ranges` against the operation's budget
/// ([`LazyConfig::max_fetch`]) before they are fetched: an error, and
/// nothing counted, if they would take it past the budget.
@@ -225,20 +260,70 @@ impl LazyStorage {
/// Run one pass of `f` over this storage. `Done` when `f` read nothing
/// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s
/// result is dropped (it may be an error caused by the miss, or a
/// result built around one).
/// result built around one). Hints are not followed; see
/// [`Operation::attempt`].
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
self.attempt_within(f, 0)
}
/// [`attempt`](Self::attempt), adding to a pass that misses the hinted
/// blocks it lacks while everything asked for stays within `budget`
/// bytes.
fn attempt_within<T>(&self, f: impl FnOnce() -> T, budget: u64) -> Step<T> {
{
let mut st = lock(&self.state);
st.missing.clear();
st.hinted.clear();
st.stats.passes += 1;
}
let out = f();
let missing = std::mem::take(&mut lock(&self.state).missing);
let (missing, hinted) = {
let mut st = lock(&self.state);
(
std::mem::take(&mut st.missing),
std::mem::take(&mut st.hinted),
)
};
if missing.is_empty() {
return Step::Done(out);
}
drop(out);
Step::Need(self.runs(missing))
let need = self.runs(&missing);
if hinted.is_empty() || budget == 0 {
return Step::Need(need);
}
// The hinted blocks still missing, in file order, while they fit.
let bs = self.config.block_size;
let mut total = ranges_len(&need);
let mut wanted = missing.clone();
{
let st = lock(&self.state);
for i in hinted {
if wanted.contains_key(&i) || st.blocks.contains_key(&i) {
continue;
}
let n = (self.len - i * bs).min(bs);
if total + n > budget {
break;
}
total += n;
wanted.insert(i, true);
}
}
if wanted.len() == missing.len() {
return Step::Need(need);
}
let with_hints = self.runs(&wanted);
// Filling holes between runs can add blocks: never let the hints
// take the pass past the budget.
if ranges_len(&with_hints) > budget {
return Step::Need(need);
}
{
let mut st = lock(&self.state);
st.stats.hinted_blocks += (wanted.len() - missing.len()) as u64;
}
Step::Need(with_hints)
}
/// The bytes of the file at `offset`, fetched for a range a pass asked
@@ -309,7 +394,7 @@ impl LazyStorage {
) -> Result<T, String> {
let op = self.operation();
loop {
match self.attempt(&mut f) {
match op.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges)?;
@@ -350,14 +435,14 @@ impl LazyStorage {
/// blocks, a one-block hole between two runs filled so they merge
/// (unless the hole is cached: it would be fetched again), each at
/// most `max_request` long.
fn runs(&self, missing: HashMap<u64, bool>) -> Vec<Range<u64>> {
fn runs(&self, missing: &HashMap<u64, bool>) -> Vec<Range<u64>> {
let bs = self.config.block_size;
let mut wanted: Vec<u64> = missing.keys().copied().collect();
wanted.sort_unstable();
let mut st = lock(&self.state);
// Remember which blocks only bulk reads asked for: they are kept
// as bulk once supplied.
for (&i, &metadata) in &missing {
for (&i, &metadata) in missing {
if metadata {
st.bulk_pending.remove(&i);
} else {
@@ -499,6 +584,25 @@ impl Storage for LazyStorage {
self.len
}
fn hint(&self, offset: u64, len: usize) {
// A hint is about a structure, not bulk data: at most a block or
// 1 MiB of it is followed, and at most `MAX_HINTED` blocks a pass,
// whatever a hostile file makes a parser hint.
let len = (len as u64).min(self.config.block_size.max(1 << 20));
let Some(span) = self.span(offset, len) else {
return;
};
let mut st = lock(&self.state);
for i in span {
if st.hinted.len() >= MAX_HINTED {
return;
}
if !st.blocks.contains_key(&i) {
st.hinted.insert(i);
}
}
}
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
let mut spans = Vec::with_capacity(ranges.len());
let mut total = 0u64;
@@ -797,6 +901,89 @@ mod tests {
}
}
#[test]
fn hinted_blocks_come_with_a_miss_and_never_alone() {
let data = file(16 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
serve(&s, &data, &[3072..4096]);
let op = s.operation();
// Hints alone: the pass is done, nothing is fetched.
let step = op.attempt(|| {
s.hint(8 * 1024, 100);
owned(s.read_at(3072, 8))
});
assert!(matches!(step, Step::Done(Ok(_))), "{step:?}");
// With a miss, the hinted blocks not cached come too (block 3 is
// cached; 10..=11 is one run, 14 another).
let pass = || {
s.hint(3072, 10);
s.hint(10 * 1024 + 1000, 100);
s.hint(14 * 1024, 1);
owned(s.read_at(0, 8))
};
let Step::Need(need) = op.attempt(pass) else {
panic!("block 0 is missing");
};
assert_eq!(
need,
vec![0..1024, 10 * 1024..12 * 1024, 14 * 1024..15 * 1024]
);
// A plain attempt does not follow hints.
let Step::Need(plain) = s.attempt(pass) else {
panic!("block 0 is missing");
};
assert_eq!(plain, vec![0..1024]);
serve(&s, &data, &need);
let Step::Done(got) = op.attempt(pass) else {
panic!("everything was supplied");
};
assert_eq!(got.unwrap(), &data[..8]);
assert_eq!(s.stats().hinted_blocks, 3);
}
#[test]
fn hints_stay_within_the_fetch_budget() {
// A budget of 3 blocks: the missed block and the first two hinted
// ones fit, the rest are left out; a hint never makes a call fail.
// (Hinted blocks two apart would be merged with the hole between
// them, which would not fit: then no hint is followed.)
let data = file(64 * 1024);
let mut c = config(1024, 1 << 20);
c.max_fetch = 3 * 1024;
let s = LazyStorage::new(data.len() as u64, c);
let got = s
.run_blocking(
|| {
for i in 0..20 {
s.hint(20 * 1024 + i * 3072, 1);
}
owned(s.read_at(0, 8))
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
assert_eq!(got, &data[..8]);
let st = s.stats();
assert_eq!(
(st.bytes_fetched, st.hinted_blocks),
(3 * 1024, 2),
"{st:?}"
);
// A hint longer than the file, or past its end, is harmless.
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
s.run_blocking(
|| {
s.hint(0, usize::MAX);
s.hint(u64::MAX, 10);
owned(s.read_at(0, 8))
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
}
#[test]
fn a_failed_fetch_is_an_error_not_data() {
let data = file(4096);
+1 -1
View File
@@ -391,7 +391,7 @@ impl Http {
) -> Result<T, JsError> {
let op = storage.operation();
loop {
match storage.attempt(&mut f) {
match op.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges).map_err(js_err)?;
+158 -47
View File
@@ -146,8 +146,8 @@ fn transcript(api: &impl Api) -> Vec<String> {
/// takes), and agrees with the in-memory one: the same values, and an error
/// wherever it has one (a malformed file can fail at a different check,
/// with a different message, when read by ranges). Returns what the lazy
/// reader fetched and its transcript.
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
/// reader fetched (requests, bytes, passes) and its transcript.
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, u64, Vec<String>) {
let ctx = format!("{name} (blocks of {block} B)");
let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec())));
let local = Reader::open(data.to_vec());
@@ -156,7 +156,7 @@ fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
(Ok(r), Ok(l), Ok(z)) => (r, l, z),
(Err(r), Err(_), Err(z)) => {
assert_eq!(z, r, "{ctx}: open error");
return (0, 0, Vec::new());
return (0, 0, 0, Vec::new());
}
(r, l, z) => panic!(
"{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}",
@@ -184,7 +184,7 @@ fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
}
assert_eq!(got.len(), local.len(), "{ctx}: transcript length");
let st = lazy.storage.stats();
(st.requests, st.bytes_fetched, got)
(st.requests, st.bytes_fetched, st.passes, got)
}
fn config(block: u64) -> LazyConfig {
@@ -223,7 +223,7 @@ fn builder_file() -> Vec<u8> {
fn builder_files_read_the_same_at_every_block_size() {
let data = builder_file();
for block in [512, 4096, 1 << 20] {
let (requests, _, lines) = check_equal("builder", &data, block);
let (requests, _, _, lines) = check_equal("builder", &data, block);
assert!(requests > 0);
// The transcript covers every object, values included.
assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok")));
@@ -315,7 +315,7 @@ fn h5py_and_netcdf4_files_read_the_same_lazily() {
for name in ["fixture.h5", "fixture.nc"] {
let data = std::fs::read(dir.path().join(name)).unwrap();
for block in [512, 64 * 1024] {
let (_, _, lines) = check_equal(name, &data, block);
let (_, _, _, lines) = check_equal(name, &data, block);
assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2);
}
}
@@ -460,25 +460,32 @@ fn corpus_files_read_the_same_lazily() {
}
files.sort();
assert!(!files.is_empty(), "no HDF5 files under {dirs}");
let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64);
let (mut requests, mut bytes, mut passes, mut total) = (0u64, 0u64, 0u64, 0u64);
for f in &files {
let data = std::fs::read(f).unwrap();
total += data.len() as u64;
let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
let (r, b, p, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
requests += r;
bytes += b;
passes += p;
}
eprintln!(
"{} files ({total} bytes): {requests} requests, {bytes} bytes fetched",
"{} files ({total} bytes): {passes} passes, {requests} requests, {bytes} bytes fetched",
files.len()
);
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
/// What one call cost on a file opened lazily (the open not counted).
#[derive(Debug, Clone, Copy)]
struct Cost {
passes: u64,
requests: u64,
bytes: u64,
}
/// Open `data` lazily at `block`-byte blocks, run `op`, and return its
/// result and what it cost (the open not counted).
fn cost_of<T>(data: &[u8], block: u64, op: impl Fn(&Reader) -> T) -> (T, Cost) {
let lazy = Lazy::open(
data.to_vec(),
LazyConfig {
@@ -488,21 +495,73 @@ fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
)
.unwrap();
let before = lazy.storage.stats();
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
let out = lazy.call(op);
let after = lazy.storage.stats();
(
after.passes - before.passes,
after.requests - before.requests,
out,
Cost {
passes: after.passes - before.passes,
requests: after.requests - before.requests,
bytes: after.bytes_fetched - before.bytes_fetched,
},
)
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
let (got, cost) = cost_of(data, block, |r| r.list(path));
assert_eq!(got.unwrap(), want);
(cost.passes, cost.requests)
}
/// What reading the dataset at `path` whole takes on a file opened lazily
/// at `block`-byte blocks (the open not counted), checking the values
/// against the in-memory read.
fn read_cost(data: &[u8], path: &str, block: u64) -> Cost {
let want = Reader::open(data.to_vec())
.unwrap()
.read(path, None)
.unwrap();
let (got, cost) = cost_of(data, block, |r| r.read(path, None));
assert_eq!(got.unwrap(), want, "{path}");
cost
}
/// An h5py file of `n` datasets of 256 `f32` each (`d0` ... ) in the root
/// group, written with `libver`.
fn h5py_many(dir: &Path, libver: &str, n: usize) -> Vec<u8> {
let path = dir.join(format!("{libver}_{n}.h5"));
let script = format!(
"import h5py, numpy as np\n\
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
\x20 for i in range({n}):\n\
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
std::fs::read(&path).unwrap()
}
/// Listing a group reads every child's object header, and its index (B-tree
/// and symbol table nodes, or B-tree v2 and heap blocks) before that. Each
/// pass asks for every node of a level it is missing, not the first one
/// only, so the passes (network round trips) grow with the depth of the
/// index, not with the number of children: 2000 children with headers
/// scattered over 512-byte blocks list in a handful of passes, where each
/// header block used to cost its own.
/// pass asks for every node of the index it can reach (a failed node does
/// not stop the walk), the blocks it has been told it reads next (a node's
/// body, a symbol table node's entries, the heap's blocks, each child's
/// header: `Storage::hint`), so the passes (network round trips) follow the
/// depth of the index, not the number of children: 2000 children with
/// headers scattered over 512-byte blocks list in a handful of passes,
/// where each header block used to cost its own.
#[test]
fn listing_a_large_group_takes_a_few_passes() {
let mut b = FileBuilder::new();
@@ -514,39 +573,63 @@ fn listing_a_large_group_takes_a_few_passes() {
let data = b.finish().unwrap();
let (passes, requests) = listing_cost(&data, "/many", 512);
eprintln!("FileBuilder, 600 children: {passes} passes, {requests} requests");
assert!(passes <= 6, "{passes} passes");
// 5 before hints (2026-09-27), 102 before the walks went on past a miss.
assert!(passes <= 4, "{passes} passes");
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
return;
}
let dir = tempfile::tempdir().unwrap();
for libver in ["earliest", "latest"] {
let path = dir.path().join(format!("{libver}.h5"));
let script = format!(
"import h5py, numpy as np\n\
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
\x20 for i in range(2000):\n\
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let data = std::fs::read(&path).unwrap();
// Most passes each may take: 8 and 11 before hints (2026-09-27).
for (libver, most) in [("earliest", 5), ("latest", 6)] {
let data = h5py_many(dir.path(), libver, 2000);
let (passes, requests) = listing_cost(&data, "/", 512);
eprintln!("h5py libver={libver}, 2000 children: {passes} passes, {requests} requests");
assert!(passes <= 12, "{libver}: {passes} passes");
assert!(passes <= most, "{libver}: {passes} passes");
}
}
/// Opening one dataset of a large group looks its name up, not the whole
/// group: in a v1 (symbol table) group down its B-tree as libhdf5 does, in
/// a dense group down its name index. Reading one small dataset of 2000
/// fetches a few blocks, where it used to read every symbol table node and
/// name of a v1 group (529 requests, 333 kB at 512-byte blocks, before
/// 2026-09-27).
#[test]
fn reading_one_dataset_of_a_large_group_fetches_a_few_blocks() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = tempfile::tempdir().unwrap();
for (libver, most_passes, most_requests) in [("earliest", 7, 6), ("latest", 8, 9)] {
let data = h5py_many(dir.path(), libver, 2000);
for path in ["/d0", "/d1234", "/d1999"] {
let cost = read_cost(&data, path, 512);
eprintln!("h5py libver={libver}, 2000 children, read {path}: {cost:?}");
assert!(
cost.passes <= most_passes && cost.requests <= most_requests,
"{libver} {path}: {cost:?}"
);
assert!(cost.bytes <= 64 * 1024, "{libver} {path}: {cost:?}");
}
}
}
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
/// file costs lazily, at 1 MiB and 64 KiB blocks (a measurement, printed).
/// file costs lazily, and opening it and reading one dataset whole
/// (`CLAWHDF5_WASM_READ`, by default the middle dataset of the listing),
/// at 1 MiB and 64 KiB blocks (a measurement, printed).
#[test]
fn listing_cost_of_a_given_file() {
let Ok(path) = std::env::var("CLAWHDF5_WASM_LIST_FILE") else {
@@ -563,16 +646,44 @@ fn listing_cost_of_a_given_file() {
)
.unwrap();
let open = lazy.storage.stats();
let n = lazy.call(|r| r.list("/")).unwrap().len();
let list = lazy.call(|r| r.list("/")).unwrap();
let st = lazy.storage.stats();
eprintln!(
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {n}: {} passes, {} requests, {} bytes",
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {}: {} passes, {} requests, {} bytes ({} blocks hinted)",
data.len(),
open.requests,
open.passes,
list.len(),
st.passes - open.passes,
st.requests - open.requests,
st.bytes_fetched - open.bytes_fetched
st.bytes_fetched - open.bytes_fetched,
st.hinted_blocks - open.hinted_blocks,
);
let name = std::env::var("CLAWHDF5_WASM_READ").unwrap_or_else(|_| {
let datasets: Vec<_> = list.iter().filter(|c| c.kind == Kind::Dataset).collect();
format!("/{}", datasets[datasets.len() / 2].name)
});
// Open and read on a fresh cache: the open's own cost (the probe
// block and its passes) and then the read's.
let fresh = Lazy::open(
data.clone(),
LazyConfig {
block_size: block,
..LazyConfig::default()
},
)
.unwrap();
let open = fresh.storage.stats();
let values = fresh.call(|r| r.read(&name, None)).unwrap().data.len();
let st = fresh.storage.stats();
eprintln!(
" open + read('{name}') ({values} values): {} passes, {} requests, {} bytes (the open: {} passes, {} requests, {} bytes)",
st.passes,
st.requests,
st.bytes_fetched,
open.passes,
open.requests,
open.bytes_fetched,
);
}
}
+89 -15
View File
@@ -1,27 +1,101 @@
# clawhdf5
[![crates.io](https://img.shields.io/crates/v/clawhdf5.svg)](https://crates.io/crates/clawhdf5)
[![docs.rs](https://docs.rs/clawhdf5/badge.svg)](https://docs.rs/clawhdf5)
The main crate: a pure-Rust HDF5 reader, writer and in-place editor, with no
libhdf5 and, by default, no C code. It wraps
[`clawhdf5-format`](../clawhdf5-format/README.md) (the binary format) and
[`clawhdf5-io`](../clawhdf5-io/README.md) (memory-mapped reads) in an
h5py-like API.
Pure-Rust HDF5 reader/writer — no C dependencies.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Main types
| Type | What it does |
|---|---|
| `File` | Opens a file (`open`, `open_buffered`, `from_bytes`, `open_storage` for any `Storage`), walks groups (`root`, `group`, `dataset`), lists `datasets`/`groups`/`attrs`. `File` is `Send + Sync`: several threads can read one open file. |
| `Dataset` | `shape`, `dtype`, `max_dimensions`, `attrs`; reads `read_f64`/`read_f32`/`read_i32`/`read_i64`/`read_u64`, strings (`read_string`, `read_string_bytes`), variable-length data (`read_vlen`), hyperslabs and point selections (`read_selection`, `read_f64_selection`, ...), zero-copy views of contiguous data (`read_f64_zerocopy`, ...), `verify_provenance`. |
| `FileBuilder` | Writes a new file: datasets of every numeric type, strings, compounds (`CompoundTypeBuilder`), enums, chunked and compressed layouts (deflate, shuffle, Fletcher-32, LZF, and with features LZ4, Zstd, bitshuffle, bzip2, Blosc, pcodec), nested groups, soft/hard/external links, virtual datasets, attribute creation order. Files open in h5py and h5dump. |
| `FileEditor` | Changes an existing file in place without rewriting it: `write_values`/`write_selection`/`write_all`, `resize` of chunked datasets (every chunk index), `set_attr` (compact and dense storage). Anything it cannot do safely is `Error::Unsupported` before any write. |
| `MmapFile`, `LazyFile` | Alternative readers: memory-mapped, and one that reads lazily and caches. |
| `File::open_swmr` | Reads a file a libhdf5 SWMR writer is still appending to (`Dataset::refresh`, bounded retries), as h5py's `swmr=True` reader does. |
## Examples
```rust,no_run
use clawhdf5::{AttrValue, File, FileBuilder, FileEditor, Selection};
// Write
let mut b = FileBuilder::new();
b.create_dataset("sensors/temperature")
.with_f64_data(&[20.5, 21.0, 21.5, 22.0])
.with_shape(&[4])
.with_maxshape(&[u64::MAX]) // unlimited, so it can grow
.with_chunks(&[2])
.with_deflate(4);
b.set_attr("version", AttrValue::I64(1));
b.write("data.h5")?;
// Read
let file = File::open("data.h5")?;
let ds = file.dataset("sensors/temperature")?;
assert_eq!(ds.shape()?, vec![4]);
let values = ds.read_f64()?;
// Edit in place: grow the dataset and fill the new tail
let mut ed = FileEditor::open("data.h5")?;
ed.resize("sensors/temperature", &[6])?;
let tail = Selection::Hyperslab {
start: vec![4],
stride: vec![1],
count: vec![2],
block: vec![1],
};
ed.write_values("sensors/temperature", &tail, &[22.5f64, 23.0])?;
# Ok::<(), clawhdf5::Error>(())
```
Remote files (HTTP range requests, S3/GCS/Azure) are read through
`File::open_storage`; [`clawhdf5-remote`](../clawhdf5-remote/README.md)
provides the storage and its block cache.
## Features
- Read and write HDF5 files entirely in Rust
- Memory-mapped I/O for large files (`mmap` feature, enabled by default)
- Parallel chunk reads via Rayon (`parallel` feature)
- Lazy dataset access for minimal memory usage
- h5py-compatible file output
| Feature | Default | What | Builds C |
|---|---|---|---|
| `mmap` | yes | memory-mapped reads (`File::open` maps the file; `MmapFile`) | no |
| `provenance` | yes | SHA-256 `_provenance_sha256` attributes (`DatasetBuilder::with_provenance`, `Dataset::verify_provenance`) | no |
| `lzf` | yes | LZF filter (32000), h5py's `compression="lzf"` | no |
| `parallel` | no | chunk decoding on a rayon pool | no |
| `lz4` | no | LZ4 filter (32004) | no |
| `pcodec` | no | pcodec filter | no |
| `bitshuffle`, `bzip2`, `blosc` | no | plugin filters 32008, 307, 32001 (read and write) | no (bzip2 uses the pure-Rust `libbz2-rs-sys`) |
| `blosc2`, `zfp` | no | plugin filters 32026 and 32013, **read only** | no |
| `plugin-filters` | no | `lzf`, `bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp` | no |
| `zstd` | no | Zstandard filter (32015) | yes (libzstd) |
| `fast-deflate` | no | zlib-ng instead of the pure-Rust zlib-rs | yes (cmake) |
| `blake3_hash` | no | `provenance::blake3_hash` helpers | yes (`cc`, for blake3's SIMD code) |
| `apple-compression` | no | currently has no effect in this crate (it is not forwarded) | — |
## Usage
SZIP decoding is a `clawhdf5-format` feature (`szip`, links the system
libaec); the facade does not forward it.
```rust
use clawhdf5::File;
## Limits and further reading
let file = File::open("data.h5").unwrap();
let dataset = file.dataset("/group/data").unwrap();
let values: Vec<f64> = dataset.read_1d().unwrap();
```
- What is known not to work, and what was wrong in earlier releases:
[`docs/known-issues.md`](../../docs/known-issues.md) (editor limits, range
reads, external links and external raw data, which are explicit errors).
- Read coverage against libhdf5/h5py on eight public corpora:
[`CONFORMANCE.md`](../../CONFORMANCE.md).
- Read and write speed against libhdf5 and h5py:
[`BENCHMARKS.md`](../../BENCHMARKS.md).
- Range reads and SWMR design: [`docs/design/range-reads.md`](../../docs/design/range-reads.md),
[`docs/design/swmr.md`](../../docs/design/swmr.md).
- Changes: [`CHANGELOG.md`](../../CHANGELOG.md).
## License
+4 -12
View File
@@ -49,17 +49,9 @@ pub(crate) fn encode_elem(
/// element for chunks of `chunk_bytes` bytes (`H5D__earray_idx_create`,
/// `H5D__farray_idx_create`): one byte more than the nominal size needs —
/// except under layout message version 5 (HDF5 2.0's own format), which
/// always uses 8 bytes.
pub(crate) fn chunk_size_len(chunk_bytes: u64, layout_version: u8) -> usize {
if layout_version >= 5 {
return 8;
}
let log2 = if chunk_bytes <= 1 {
0
} else {
63 - chunk_bytes.leading_zeros()
};
(1 + ((log2 + 8) / 8) as usize).min(8)
/// uses the file's size of lengths (`length_size`).
pub(crate) fn chunk_size_len(chunk_bytes: u64, layout_version: u8, length_size: u8) -> usize {
clawhdf5_format::chunked_write::chunk_size_len(chunk_bytes, layout_version, length_size)
}
/// Creation parameters, in the layout message's order.
@@ -226,7 +218,7 @@ impl Ea {
) -> Result<Self, Error> {
let os = img.os;
let elem_size = if filtered {
os as usize + chunk_size_len(chunk_bytes, layout_version) + 4
os as usize + chunk_size_len(chunk_bytes, layout_version, img.ls) + 4
} else {
os as usize
};
+1 -1
View File
@@ -83,7 +83,7 @@ impl Fa {
let os = img.os;
let osz = os as usize;
let elem_size = if filtered {
osz + chunk_size_len(chunk_bytes, layout_version) + 4
osz + chunk_size_len(chunk_bytes, layout_version, img.ls) + 4
} else {
osz
};
+46 -16
View File
@@ -61,6 +61,9 @@ const MSG_FLAG_DONTSHARE: u8 = 0x04;
/// Opening takes an exclusive advisory lock on the file (`flock`, the lock
/// libhdf5 itself takes when file locking is on), so a second editor, or
/// h5py opening the file for writing, fails until the editor is dropped.
/// The drop unlocks the file before closing it, so the lock is gone as soon
/// as the drop returns, even in a program whose other threads spawn
/// processes (a forked child shares the locked descriptor until it execs).
/// Readers that do not lock ([`File`]) can still open it, but see a file
/// that may be mid-update.
///
@@ -114,6 +117,19 @@ pub struct FileEditor {
free: FreeList,
}
impl Drop for FileEditor {
/// Releases the lock explicitly before the file is closed. A `flock`
/// belongs to the open file description and lasts until every
/// descriptor of it is closed; a process another thread forks (any
/// `std::process::Command`) inherits the descriptor and keeps it until
/// it execs, so closing alone could leave the file locked for a moment
/// after the drop. Unlocking through our descriptor releases the lock
/// for all of them.
fn drop(&mut self) {
let _ = self.file.unlock();
}
}
/// Where a layout message keeps the fields an edit may change (offsets in
/// the message body).
#[derive(Debug, Default, Clone, Copy)]
@@ -396,12 +412,24 @@ impl<'t> ChunkedEdit<'t> {
.iter()
.map(|&c| u64::from(c))
.collect();
// Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message
// version 5) are read, but not rewritten: the editor holds a chunk
// it rewrites in memory, compressed and not. Writing values, and a
// resize that prunes or allocates chunks, are refused here, before
// anything is written; growing the extent (late allocation) and
// setting attributes do not come here.
let chunk_bytes = cd
.iter()
.try_fold(t.es as u64, |a, &c| a.checked_mul(c))
.filter(|&b| b <= u64::from(u32::MAX))
.and_then(|b| usize::try_from(b).ok())
.filter(|&b| b <= u32::MAX as usize)
.ok_or_else(|| Error::Unsupported("chunk larger than 4 GiB".into()))?;
.ok_or_else(|| {
Error::Unsupported(
"rewriting chunks of 4 GiB or more (writing values, or a resize that \
prunes or allocates chunks)"
.into(),
)
})?;
let hdr = Header::load(img, t.addr)?;
let layout_msg = hdr.find(MSG_LAYOUT).ok_or(Error::MissingMessage(
clawhdf5_format::message_type::MessageType::DataLayout,
@@ -584,7 +612,8 @@ impl<'t> ChunkedEdit<'t> {
let d = self.hdr.data(img, self.layout_msg)?;
let node_size = u32::from_le_bytes([d[at], d[at + 1], d[at + 2], d[at + 3]]);
let size_len = if filtered {
earray::chunk_size_len(self.chunk_bytes as u64, self.lpos.version) + 4
earray::chunk_size_len(self.chunk_bytes as u64, self.lpos.version, img.ls)
+ 4
} else {
0
};
@@ -655,7 +684,7 @@ impl<'t> ChunkedEdit<'t> {
}
_ => return Err(Error::Unsupported("chunk index missing".into())),
}
img.free(info.address, u64::from(info.chunk_size));
img.free(info.address, info.chunk_size);
Ok(())
}
@@ -1609,7 +1638,7 @@ fn store_chunk(
let len = bytes.len() as u64;
let placed = match existing {
Some(info) if t.pipeline.is_none() => {
if u64::from(info.chunk_size) != len {
if info.chunk_size != len {
return Err(Error::Unsupported(
"unfiltered chunk stored at an unexpected size".into(),
));
@@ -1621,11 +1650,10 @@ fn store_chunk(
// thing in the file (the chunk an append keeps rewriting usually is)
// and can grow there.
Some(info)
if len <= u64::from(info.chunk_size)
|| img.grow_tail(info.address, u64::from(info.chunk_size), len)? =>
if len <= info.chunk_size || img.grow_tail(info.address, info.chunk_size, len)? =>
{
img.write(info.address, &bytes)?;
(len != u64::from(info.chunk_size) || info.filter_mask != mask).then_some(Elem {
(len != info.chunk_size || info.filter_mask != mask).then_some(Elem {
addr: info.address,
size: len,
mask,
@@ -1635,7 +1663,7 @@ fn store_chunk(
let a = img.alloc(len)?;
img.write(a, &bytes)?;
if let Some(info) = existing {
img.free(info.address, u64::from(info.chunk_size));
img.free(info.address, info.chunk_size);
}
Some(Elem {
addr: a,
@@ -1653,14 +1681,16 @@ fn store_chunk(
fn img_read<'a>(f: &'a File, info: &ChunkInfo) -> Result<&'a [u8], Error> {
let start = usize::try_from(info.address)
.map_err(|_| Error::Unsupported("chunk address out of range".into()))?;
f.as_bytes()
.get(start..start + info.chunk_size as usize)
.ok_or_else(|| {
Error::Format(clawhdf5_format::error::FormatError::UnexpectedEof {
expected: start + info.chunk_size as usize,
available: f.as_bytes().len(),
})
let end = usize::try_from(info.chunk_size)
.ok()
.and_then(|n| start.checked_add(n))
.unwrap_or(usize::MAX);
f.as_bytes().get(start..end).ok_or_else(|| {
Error::Format(clawhdf5_format::error::FormatError::UnexpectedEof {
expected: end,
available: f.as_bytes().len(),
})
})
}
fn decode_chunk(

Some files were not shown because too many files have changed in this diff Show More