feat(agent): new stores use the int8 vector index by default
`MemoryConfig::quantized_index` now defaults to `true`. It holds a quarter of the index memory and, with the exact re-score, is faster at equal recall on every configuration measured: 1.63x the queries per second on x86-64 (AVX2) and 1.18x on a Raspberry Pi 5 (NEON SDOT), with builds 1.8x and 2.3x faster. The one argument for keeping it off — that int8 search was slower on ARM — did not survive being measured. Existing stores do not change. A store written with v2.6.0 or later keeps its persisted setting. One written before the setting existed has no stored value, and it loads as `false` rather than as the new default, so reopening it never changes how its index is held. That case is guarded by a real store written with the v2.5.0 CLI, committed as `tests/fixtures/store_v2_5_0.h5` (6.8 KB): the test asserts it reopens with an f32 index and still searches, and it fails if the load default is changed to `true`. The CLI needed more than a new default. `create --quantized-index` assigned its value straight into the config, so under the new default every CLI-created store would have been forced back to f32 unless the caller knew to ask for int8. It is replaced by `--f32-index`, which only ever switches the default off; `--quantized-index` is still accepted, hidden, as a no-op, and the two conflict. The whole agent suite passes under the new default, including the brute-force recall oracle, now running on int8 plus re-score without being asked to. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
@@ -2,6 +2,29 @@
|
|||||||
|
|
||||||
## Unreleased
|
## Unreleased
|
||||||
|
|
||||||
|
### Upgrade Notes
|
||||||
|
- **New stores use the int8 vector index by default.**
|
||||||
|
`MemoryConfig::quantized_index` now defaults to `true`: a quarter of the
|
||||||
|
index memory, builds 1.8x (x86-64) and 2.3x (Raspberry Pi 5) faster, and
|
||||||
|
searches 1.63x and 1.18x faster at equal recall, measured on every
|
||||||
|
configuration tested. **Existing stores are unaffected** — a store written
|
||||||
|
with v2.6.0 or later keeps its persisted setting, and one written before the
|
||||||
|
setting existed opens as `false` and keeps its f32 index. Set
|
||||||
|
`quantized_index = false`, or pass `create --f32-index` to the CLI, to opt
|
||||||
|
out. The CLI's `--quantized-index` is still accepted but is now a no-op.
|
||||||
|
|
||||||
|
### Defaults
|
||||||
|
- `clawhdf5-agent`: `MemoryConfig::quantized_index` defaults to `true` for new
|
||||||
|
stores. The reason it had been off — that int8 search was slower on ARM —
|
||||||
|
did not survive measurement (see Corrections). Stores that predate the
|
||||||
|
setting still load it as `false`, so reopening one never changes how its
|
||||||
|
index is held; a store written by the v2.5.0 CLI is now a test fixture that
|
||||||
|
guards exactly that, and the test fails if the load default is changed.
|
||||||
|
- `clawhdf5-cli`: `create --f32-index` opts out. `create` used to assign
|
||||||
|
`--quantized-index` straight into the config, which under the new default
|
||||||
|
would have forced every CLI-created store back to f32 unless the caller
|
||||||
|
knew to ask; it now only ever switches the default off.
|
||||||
|
|
||||||
### Performance
|
### Performance
|
||||||
- `clawhdf5-accel`: **`dot_i8` has aarch64 kernels** — `SDOT` for CPUs with
|
- `clawhdf5-accel`: **`dot_i8` has aarch64 kernels** — `SDOT` for CPUs with
|
||||||
the ARMv8.2 dot-product extension (Cortex-A76 and later, Neoverse-N1, every
|
the ARMv8.2 dot-product extension (Cortex-A76 and later, Neoverse-N1, every
|
||||||
|
|||||||
@@ -39,8 +39,11 @@ Cargo workspace with 16 crates under `crates/` (plus `libaec-sys`, an internal F
|
|||||||
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
|
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
|
||||||
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
|
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
|
||||||
(tied to the checkpoint by a generation id; stale/damaged sidecars are
|
(tied to the checkpoint by a generation id; stale/damaged sidecars are
|
||||||
ignored and the index rebuilt). `MemoryConfig::quantized_index` (off by
|
ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by
|
||||||
default, persisted) stores the index's own copy of the embeddings as `i8`,
|
default** for new stores, persisted; stores predating the setting load as
|
||||||
|
`false` and keep their f32 index — guarded by
|
||||||
|
`tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`)
|
||||||
|
stores the index's own copy of the embeddings as `i8`,
|
||||||
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
|
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
|
||||||
at 100K); because quantised distances are approximate and `ef` cannot
|
at 100K); because quantised distances are approximate and `ef` cannot
|
||||||
compensate, the query path then re-scores the candidate pool against the
|
compensate, the query path then re-scores the candidate pool against the
|
||||||
|
|||||||
@@ -437,7 +437,8 @@ ClawhDF5's agent memory design draws from 15+ recent papers:
|
|||||||
vector index (16 / 64 / scale-with-`k` by default) and are stored with the
|
vector index (16 / 64 / scale-with-`k` by default) and are stored with the
|
||||||
file.
|
file.
|
||||||
|
|
||||||
`MemoryConfig::quantized_index` (off by default) stores the HNSW index's own
|
`MemoryConfig::quantized_index` (**on by default** for new stores) holds the
|
||||||
|
HNSW index's own
|
||||||
copy of the embeddings as `i8`, roughly halving a loaded store's memory
|
copy of the embeddings as `i8`, roughly halving a loaded store's memory
|
||||||
(2.72x -> 1.74x the raw vectors at 100k x 384). Quantised distances are
|
(2.72x -> 1.74x the raw vectors at 100k x 384). Quantised distances are
|
||||||
approximate, so the query path re-scores the candidate pool against the exact
|
approximate, so the query path re-scores the candidate pool against the exact
|
||||||
|
|||||||
@@ -134,13 +134,19 @@ pub struct MemoryConfig {
|
|||||||
pub wal_enabled: bool,
|
pub wal_enabled: bool,
|
||||||
pub wal_max_entries: usize,
|
pub wal_max_entries: usize,
|
||||||
/// Store the vector index's own copy of the embeddings as int8 rather than
|
/// Store the vector index's own copy of the embeddings as int8 rather than
|
||||||
/// f32, a quarter of the memory.
|
/// f32, a quarter of the memory. **On by default** for new stores.
|
||||||
///
|
///
|
||||||
/// The index's copy is the single largest part of a loaded store's
|
/// The index's copy is the single largest part of a loaded store's
|
||||||
/// footprint. Quantised distances are approximate, so the candidate pool
|
/// footprint. Quantised distances are approximate, so the candidate pool
|
||||||
/// is re-scored against the cache's exact embeddings before fusion, which
|
/// is re-scored against the cache's exact embeddings before fusion, which
|
||||||
/// restores recall; what it costs is throughput — roughly 13% of queries
|
/// holds recall at the f32 index's level. It is also faster, not slower:
|
||||||
/// per second and 16% of build time at 100K x 384. See `BENCHMARKS.md`.
|
/// at equal recall, 1.63x the queries per second on x86-64 (AVX2) and
|
||||||
|
/// 1.18x on a Raspberry Pi 5 (NEON `SDOT`), with builds 1.8x and 2.3x
|
||||||
|
/// faster. See `BENCHMARKS.md`.
|
||||||
|
///
|
||||||
|
/// Persisted with the store. Stores written before this setting existed
|
||||||
|
/// have no stored value and open as `false`, so reopening an old store
|
||||||
|
/// never changes how its index is held.
|
||||||
///
|
///
|
||||||
/// Has no effect without the `hnsw` feature.
|
/// Has no effect without the `hnsw` feature.
|
||||||
pub quantized_index: bool,
|
pub quantized_index: bool,
|
||||||
@@ -182,7 +188,7 @@ impl MemoryConfig {
|
|||||||
created_at,
|
created_at,
|
||||||
wal_enabled: true,
|
wal_enabled: true,
|
||||||
wal_max_entries: 500,
|
wal_max_entries: 500,
|
||||||
quantized_index: false,
|
quantized_index: true,
|
||||||
hnsw_m: 16,
|
hnsw_m: 16,
|
||||||
hnsw_ef_construction: 64,
|
hnsw_ef_construction: 64,
|
||||||
hnsw_ef_search: 0,
|
hnsw_ef_search: 0,
|
||||||
|
|||||||
@@ -497,6 +497,9 @@ pub fn validate_and_load(
|
|||||||
wal_max_entries: optional_i64_attr(&attrs, "wal_max_entries")
|
wal_max_entries: optional_i64_attr(&attrs, "wal_max_entries")
|
||||||
.and_then(|v| usize::try_from(v).ok())
|
.and_then(|v| usize::try_from(v).ok())
|
||||||
.unwrap_or(500),
|
.unwrap_or(500),
|
||||||
|
// `false`, not the new-store default: a store written before this
|
||||||
|
// setting existed was built with an f32 index, and reopening it must
|
||||||
|
// not silently change that.
|
||||||
quantized_index: optional_bool_attr(&attrs, "quantized_index", false),
|
quantized_index: optional_bool_attr(&attrs, "quantized_index", false),
|
||||||
hnsw_m: optional_i64_attr(&attrs, "hnsw_m")
|
hnsw_m: optional_i64_attr(&attrs, "hnsw_m")
|
||||||
.and_then(|v| usize::try_from(v).ok())
|
.and_then(|v| usize::try_from(v).ok())
|
||||||
|
|||||||
Binary file not shown.
@@ -284,3 +284,63 @@ fn degenerate_hnsw_parameters_do_not_panic() {
|
|||||||
let results = mem.hybrid_search(&vectors[7], "", 1.0, 0.0, 5);
|
let results = mem.hybrid_search(&vectors[7], "", 1.0, 0.0, 5);
|
||||||
assert_eq!(results[0].index, 7, "exact match should still rank first");
|
assert_eq!(results[0].index, 7, "exact match should still rank first");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn new_stores_default_to_the_quantized_index() {
|
||||||
|
// int8 is the default because it is smaller and, with an exact re-score,
|
||||||
|
// faster at equal recall on every platform measured (see BENCHMARKS.md).
|
||||||
|
let dir = TempDir::new().unwrap();
|
||||||
|
let config = MemoryConfig::new(dir.path().join("mem.h5"), "agent", 8);
|
||||||
|
assert!(config.quantized_index);
|
||||||
|
|
||||||
|
let path = config.path.clone();
|
||||||
|
let mut mem = HDF5Memory::create(config).unwrap();
|
||||||
|
let mut seed = 3;
|
||||||
|
let vectors: Vec<Vec<f32>> = (0..40).map(|_| make_vector(&mut seed, 8)).collect();
|
||||||
|
for (i, v) in vectors.iter().enumerate() {
|
||||||
|
mem.save(entry(&format!("c{i}"), v.clone(), "t")).unwrap();
|
||||||
|
}
|
||||||
|
assert_eq!(
|
||||||
|
mem.hybrid_search(&vectors[11], "", 1.0, 0.0, 1)[0].index,
|
||||||
|
11
|
||||||
|
);
|
||||||
|
mem.flush_wal().unwrap();
|
||||||
|
drop(mem);
|
||||||
|
assert!(HDF5Memory::open(&path).unwrap().config().quantized_index);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_store_written_before_the_setting_existed_stays_f32() {
|
||||||
|
// `store_v2_5_0.h5` was written by the v2.5.0 CLI, before
|
||||||
|
// `quantized_index` or the HNSW parameters were persisted, so it carries
|
||||||
|
// none of them. Flipping the default for new stores must not reach back
|
||||||
|
// and change how an existing store's index is held.
|
||||||
|
let dir = TempDir::new().unwrap();
|
||||||
|
let path = dir.path().join("legacy.h5");
|
||||||
|
std::fs::copy(
|
||||||
|
concat!(
|
||||||
|
env!("CARGO_MANIFEST_DIR"),
|
||||||
|
"/tests/fixtures/store_v2_5_0.h5"
|
||||||
|
),
|
||||||
|
&path,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let bytes = std::fs::read(&path).unwrap();
|
||||||
|
assert!(
|
||||||
|
!bytes.windows(15).any(|w| w == b"quantized_index"),
|
||||||
|
"the fixture must predate the setting, or it tests nothing"
|
||||||
|
);
|
||||||
|
|
||||||
|
let mut mem = HDF5Memory::open(&path).unwrap();
|
||||||
|
assert!(
|
||||||
|
!mem.config().quantized_index,
|
||||||
|
"an old store must reopen with an f32 index"
|
||||||
|
);
|
||||||
|
assert_eq!(mem.config().hnsw_m, 16);
|
||||||
|
assert_eq!(mem.config().hnsw_ef_construction, 64);
|
||||||
|
assert_eq!(mem.count(), 6);
|
||||||
|
// And it still searches: entry 3's own embedding finds it first.
|
||||||
|
let hit = mem.hybrid_search(&[3.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "", 1.0, 0.0, 1);
|
||||||
|
assert_eq!(hit[0].index, 3);
|
||||||
|
}
|
||||||
|
|||||||
@@ -28,9 +28,13 @@ enum Commands {
|
|||||||
/// Enable write-ahead log
|
/// Enable write-ahead log
|
||||||
#[arg(long)]
|
#[arg(long)]
|
||||||
wal: bool,
|
wal: bool,
|
||||||
/// Store the vector index's copy of the embeddings as int8, roughly
|
/// Hold the vector index's copy of the embeddings as f32 instead of
|
||||||
/// halving a loaded store's memory at about 13% fewer queries/second
|
/// the default int8 (which uses a quarter of the memory and is faster
|
||||||
|
/// at equal recall)
|
||||||
#[arg(long)]
|
#[arg(long)]
|
||||||
|
f32_index: bool,
|
||||||
|
/// Accepted for compatibility; int8 is now the default
|
||||||
|
#[arg(long, hide = true, conflicts_with = "f32_index")]
|
||||||
quantized_index: bool,
|
quantized_index: bool,
|
||||||
},
|
},
|
||||||
/// Save a memory entry (reads JSON from stdin or --json)
|
/// Save a memory entry (reads JSON from stdin or --json)
|
||||||
@@ -96,11 +100,18 @@ fn run(cli: Cli) -> Result<(), Box<dyn std::error::Error>> {
|
|||||||
agent_id,
|
agent_id,
|
||||||
dim,
|
dim,
|
||||||
wal,
|
wal,
|
||||||
quantized_index,
|
f32_index,
|
||||||
|
quantized_index: _,
|
||||||
} => {
|
} => {
|
||||||
let mut config = MemoryConfig::new(cli.path.clone(), &agent_id, dim);
|
let mut config = MemoryConfig::new(cli.path.clone(), &agent_id, dim);
|
||||||
config.wal_enabled = wal;
|
config.wal_enabled = wal;
|
||||||
config.quantized_index = quantized_index;
|
// Only ever switch *off* the library default: assigning the flag
|
||||||
|
// outright would force every CLI-created store back to f32 unless
|
||||||
|
// the caller knew to ask for int8.
|
||||||
|
if f32_index {
|
||||||
|
config.quantized_index = false;
|
||||||
|
}
|
||||||
|
let config_quantized = config.quantized_index;
|
||||||
let mem = HDF5Memory::create(config)?;
|
let mem = HDF5Memory::create(config)?;
|
||||||
let j = serde_json::json!({
|
let j = serde_json::json!({
|
||||||
"status": "created",
|
"status": "created",
|
||||||
@@ -108,7 +119,7 @@ fn run(cli: Cli) -> Result<(), Box<dyn std::error::Error>> {
|
|||||||
"agent_id": agent_id,
|
"agent_id": agent_id,
|
||||||
"embedding_dim": dim,
|
"embedding_dim": dim,
|
||||||
"wal_enabled": wal,
|
"wal_enabled": wal,
|
||||||
"quantized_index": quantized_index,
|
"quantized_index": config_quantized,
|
||||||
"count": mem.count(),
|
"count": mem.count(),
|
||||||
});
|
});
|
||||||
println!("{}", serde_json::to_string_pretty(&j)?);
|
println!("{}", serde_json::to_string_pretty(&j)?);
|
||||||
|
|||||||
+5
-4
@@ -364,10 +364,11 @@ cargo install --path crates/clawhdf5-cli
|
|||||||
clawhdf5 --path agent.h5 create --agent-id my-agent --dim 384 --wal
|
clawhdf5 --path agent.h5 create --agent-id my-agent --dim 384 --wal
|
||||||
```
|
```
|
||||||
|
|
||||||
Add `--quantized-index` to store the vector index's copy of the embeddings as
|
New stores hold the vector index's copy of the embeddings as int8, which
|
||||||
int8. That roughly halves a loaded store's memory at about 13% fewer queries
|
roughly halves a loaded store's memory and is faster at equal recall — the
|
||||||
per second, with recall unchanged — the query path re-scores candidates
|
query path re-scores candidates against the exact embeddings. Pass
|
||||||
against the exact embeddings. The setting is recorded in the file.
|
`--f32-index` to keep an f32 index instead. The setting is recorded in the
|
||||||
|
file, and stores created before it existed keep their f32 index.
|
||||||
|
|
||||||
Output:
|
Output:
|
||||||
```json
|
```json
|
||||||
|
|||||||
Reference in New Issue
Block a user