docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -42,7 +42,8 @@ AVX2 and on NEON — with the `SDOT` instruction (through inline assembly,
|
||||
since the intrinsic is unstable) on cores that have dotprod, such as the
|
||||
Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8
|
||||
index answers 1.63x the queries per second of the f32 one on x86-64
|
||||
(AVX2) and 1.18x on a Raspberry Pi 5 ([`BENCHMARKS.md`](../../BENCHMARKS.md)).
|
||||
(AVX2; 2026-09-20, machine not recorded, not re-run) and 1.18x on a
|
||||
Raspberry Pi 5 (2026-09-21) ([`BENCHMARKS.md` § Quantising the index copy](../../BENCHMARKS.md#quantising-the-index-copy-quantized_index)).
|
||||
|
||||
The aarch64 code is compiled out on x86, so only the `test-arm64` CI job
|
||||
builds and tests it.
|
||||
|
||||
@@ -43,7 +43,7 @@ let hits = mem.search(&query, "deploy key", &SearchOptions::new(5).with_sources(
|
||||
for h in &hits {
|
||||
println!("{:.3} {}", h.score, h.chunk);
|
||||
}
|
||||
mem.flush_wal()?; // checkpoint now; otherwise one is made after 500 WAL entries (wal_max_entries)
|
||||
mem.flush_wal()?; // checkpoint now; otherwise one is made once the WAL holds more than 500 entries (wal_max_entries)
|
||||
# Ok::<(), clawhdf5_agent::MemoryError>(())
|
||||
```
|
||||
|
||||
|
||||
@@ -77,7 +77,7 @@ assert!(!hdr.messages.is_empty());
|
||||
| `checksum` | yes | verify Jenkins lookup3 checksums | no |
|
||||
| `deflate` | yes | deflate through flate2 | no |
|
||||
| `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no |
|
||||
| `system-zlib-decompress` | yes | no effect (nothing reads it; kept so existing feature lists still build) | no |
|
||||
| `system-zlib-decompress` | yes | macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere | no (links the system libz on macOS) |
|
||||
| `provenance` | yes | SHA-256 provenance hashes | no |
|
||||
| `lzf` | yes | LZF (32000) | no |
|
||||
| `parallel` | no | rayon-parallel chunk decoding | no |
|
||||
|
||||
@@ -24,7 +24,7 @@ clawhdf5-io = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features =
|
||||
|---|---|
|
||||
| `HDF5Read`, `HDF5ReadWrite` | byte-level read/write traits; `MemoryReader`, `BorrowedReader`, `FileReader`, `FileWriter` implement them |
|
||||
| `MmapReader`, `MmapReadWrite` (`mmap`) | memory-mapped files through `memmap2`; `HDF5Read::private_copy` gives a copy-on-write view |
|
||||
| `prefetch::PrefetchReader`, `sweep::SweepDetector` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction |
|
||||
| `prefetch::PrefetchReader`, `prefetch::SweepDetector`, `sweep` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction |
|
||||
| `ParallelConfig` | lane partitioning for parallel chunk decoding |
|
||||
| `vol::VirtualObjectLayer`, `vol::NativeVol` | a backend-agnostic object-layer trait (modelled on libhdf5's VOL) |
|
||||
| `async_read` (`async`) | tokio-based `AsyncHDF5Read` and `AsyncHDF5File` |
|
||||
|
||||
@@ -39,7 +39,7 @@ let values: Vec<f64> = temp.read_f64()?;
|
||||
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
|
||||
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) |
|
||||
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
|
||||
| `Dimension` | `name`, `size`, `is_unlimited` |
|
||||
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is wrongly 0 when it holds records; use the variables' shapes — [known issue](../../docs/known-issues.md#netcdf-4-an-unlimited-dimension-reports-size-0)) |
|
||||
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
|
||||
| `NcType` | the NetCDF type of a variable |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user