Files
clawhdf5/crates/clawhdf5-format/README.md
T
osobhandClaude Opus 5.5 01d2a5dc5d
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s
docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00

110 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# clawhdf5-format
The HDF5 file format in pure Rust: parsers and writers for every on-disk
structure, the filter pipeline and its codecs, and the shared type
definitions the other crates use. Most users want the
[`clawhdf5`](../clawhdf5/README.md) facade, which wraps this crate in an
h5py-like API; use this one directly for low-level access or in `no_std`
code.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## What is in it
- **Parsing:** superblock v0–v3 (`superblock`, with the superblock
extension and metadata cache images, `superblock_ext`), object headers v1
and v2 (`object_header`), every header message the readers use
(`datatype`, `dataspace`, `data_layout` v1–v4 including virtual datasets,
`fill_value`, `attribute`, `link_message`, `shared_message`, ...), groups
old and new (`group_v1` symbol tables with local heaps, `group_v2` with
fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single
chunk, implicit, fixed array, extensible array, v2 B-tree).
- **Reading data:** `data_read` (contiguous, compact, chunked),
`partial_read` and `selection` (hyperslabs and points), `vl_data`
(variable-length strings and sequences through the global heap),
`chunk_cache`.
- **Storage:** the `storage::Storage` trait (`read_at`, `read_ranges`,
`len`, `hint`) that every read path goes through, so a file can be read
from memory, a file handle or a remote backend
([`clawhdf5-remote`](../clawhdf5-remote/README.md)).
- **Writing:** `file_writer::FileWriter` and the builders in
`type_builders` (datasets, groups, attributes, compound and enum types,
links, virtual datasets, creation-order tracking); chunk indexes and
dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`,
`ea_writer`, and the version-1 chunk B-tree of `btree_v1_write`). Output
is read by h5py and h5dump; `FileWriter::libver_bounds` (`libver`) picks
the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other
IDs can be registered at run time with `register_filter`). Built in:
deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4,
Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle,
bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
- **Shared pieces:** `float16` (the one IEEE half-precision conversion the
workspace uses), `provenance` (SHA-256 dataset hashes), `checksum`
(Jenkins lookup3 for v2+ structures).
## Example
```rust
use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};
// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3])
.set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();
// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
.unwrap();
assert!(!hdr.messages.is_empty());
```
## Features
| Feature | Default | What | Builds C |
|---|---|---|---|
| `std` | yes | standard library; without it the crate is `no_std` + `alloc` (CI builds it for `thumbv7em-none-eabihf`) | no |
| `checksum` | yes | verify Jenkins lookup3 checksums | no |
| `deflate` | yes | deflate through flate2 | no |
| `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no |
| `system-zlib-decompress` | yes | macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere | no (links the system libz on macOS) |
| `provenance` | yes | SHA-256 provenance hashes | no |
| `lzf` | yes | LZF (32000) | no |
| `parallel` | no | rayon-parallel chunk decoding | no |
| `fast-checksum` | no | hardware CRC32 through `crc32fast` | no |
| `lz4` | no | LZ4 (32004) | no |
| `pcodec` | no | pcodec | no |
| `bitshuffle`, `bzip2`, `blosc` | no | 32008, 307, 32001, read and write | no |
| `blosc2`, `zfp` | no | 32026, 32013, read only | no |
| `plugin-filters` | no | all six plugin filters above | no |
| `lookup-stats` | no | counters for name-lookup benchmarks | no |
| `zstd` | no | Zstandard (32015) | yes (libzstd) |
| `szip` | no | SZIP (4) decoding | links the system libaec (`libaec-dev`) |
| `fast-deflate` | no | zlib-ng | yes (cmake) |
| `system-zlib` | no | the system zlib | yes (`libz-sys`) |
| `blake3_hash` | no | `provenance::blake3_hash` | yes (`cc`) |
## Robustness
Every parser is meant to return an error, never panic, on hostile input:
nine cargo-fuzz targets live in [`fuzz/`](fuzz/README.md), the conformance
sweep includes the HDF Group's CVE corpus
([`CONFORMANCE.md`](../../CONFORMANCE.md)), and header checks follow
libhdf5's. Open gaps are in [`docs/known-issues.md`](../../docs/known-issues.md).
## License
MIT