feat: check HDF5 2.x small floats against libhdf5 2.2.0
Fixture written by libhdf5 2.2.0 (built from tag 2.2.0) through ctypes: every bit pattern of FP4 E2M1, FP6 E2M3/E3M2, FP8 E4M3/E5M2 and a bfloat16 LE/BE set, as datasets and attributes, with what H5Dread/H5Aread return into double and float and the conversion exceptions libhdf5 raises. clawhdf5 already decoded every value as libhdf5 does, including an all-ones exponent as inf/NaN in the OCP formats that have none (documented as a deliberate match in known-issues). - data_read: NaNs of non-native float layouts get libhdf5's bits (sign kept, every mantissa bit set) in f64 and f32. - h5rs dump/ls name these types as h5dump/h5ls 2.x do (H5T_FLOAT_F4E2M1, "FP4 E2M1 4-bit float", float4-e2m1 ...), checked against h5dump 2.2.0's output of the fixture. - Python bindings read them as h5py 3.16 does (float32 for bfloat16, float16 for the 1-byte formats, file byte order, same bytes as h5py); writing them in 'r+' is refused. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -30,6 +30,35 @@
|
||||
Affected v2.1.0 to v2.7.0. `docs/known-issues.md` also gains an open
|
||||
entry found meanwhile: variables' dimensions are matched by size.
|
||||
|
||||
### HDF5 2.x small floats checked against libhdf5 2.2.0 (2026-09-28)
|
||||
- New fixture written by libhdf5 2.2.0 itself (built from tag `2.2.0`,
|
||||
`crates/clawhdf5/tests/fixtures/gen_mx_floats.py`): a dataset and an
|
||||
attribute of each of `H5T_FLOAT_F4E2M1`, `H5T_FLOAT_F6E2M3`,
|
||||
`H5T_FLOAT_F6E3M2` (every bit pattern, also with the unused high bits set),
|
||||
`H5T_FLOAT_F8E4M3`, `H5T_FLOAT_F8E5M2` (all 256) and
|
||||
`H5T_FLOAT_BFLOAT16LE`/`BE` (zeros, subnormals, max, ±inf, NaNs), with
|
||||
what libhdf5 returns for each element into `double` and `float`.
|
||||
`tests/mx_floats_interop.rs` compares `read_f64`, `read_f32` and the
|
||||
attributes bit for bit. No value was wrong: clawhdf5 already decoded every
|
||||
pattern as libhdf5 does, including an all-ones exponent as ±inf/NaN in the
|
||||
formats the OCP MX specification makes finite (FP4, FP6, FP8 E4M3 — a
|
||||
deliberate match, recorded in `docs/known-issues.md`).
|
||||
- NaNs of these formats now have libhdf5's bits: sign kept, every mantissa
|
||||
bit set (`0x7FFF_FFFF_FFFF_FFFF` as `f64`, `0x7FFF_FFFF` as `f32`); they
|
||||
were `f64::NAN`'s bits and an `as f32` cast's.
|
||||
- `h5rs dump` names these types as h5dump 2.x does (`H5T_FLOAT_F4E2M1`,
|
||||
`H5T_FLOAT_BFLOAT16LE`, ...) instead of printing an `H5T_FLOAT { ... }`
|
||||
block; `h5rs ls -v` describes them as h5ls 2.x does (`FP4 E2M1 4-bit
|
||||
float`) and `h5rs ls` lists them as `bfloat16`, `float8-e4m3`,
|
||||
`float6-e2m3`, `float4-e2m1`, ... instead of `float16`/`float8`. Checked
|
||||
against h5dump 2.2.0's output of the fixture (`tests/mx_floats_dump.rs`).
|
||||
- Python bindings: datasets and attributes of these types used to raise
|
||||
`TypeError`; they now read as h5py 3.16 reads them — `float32` for
|
||||
bfloat16, `float16` for the 1-byte formats, in the file's byte order — with
|
||||
the same bytes h5py returns (`tests/test_small_floats.py`). Writing them in
|
||||
`'r+'` raises `NotImplementedError` before anything is written; inside a
|
||||
compound or array type they still raise `TypeError`.
|
||||
|
||||
### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27)
|
||||
- Parsing a version-1 object header was 4% slower than before range-read
|
||||
M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk
|
||||
|
||||
Reference in New Issue
Block a user