From 5b45c60c9c440c2b92b2cd873f646c7623a4e35b Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 23:15:09 -0500 Subject: [PATCH] docs: ObjectHeader::parse A/B after inlining the v1 message loop (at or below 8f59b2e) Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 40 ++++++++++++++++++++++++++++++++++++++++ CHANGELOG.md | 10 ++++++++++ docs/known-issues.md | 9 ++++++++- 3 files changed, 58 insertions(+), 1 deletion(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index f6e8a3c..621f912 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -500,6 +500,46 @@ explain the slower windows. ## Local file speed after range reads +### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank) + +The remaining 4% (below) was the call to the version-1 message loop, which +`4313917` kept out of line with `#[inline(never)]`. Found with A/B builds +changing one piece at a time (perf is not available: `perf_event_paranoid` +4): `#[inline]` on `parse_v1_messages` alone brought +`object_header_parse_x401` from about 24.5–24.9 µs to 23.6–24.0 µs against +8f59b2e's 23.6–24.1 µs (short 4-second rounds); no attribute measured like +`#[inline(never)]`; +creating the chunk list only when a continuation is found measured no +faster on top and was not kept. + +Same method as below: `8f59b2e` built in its own worktree and target +directory, separate binaries alternating, `taskset -c 5 +local_metadata_bench --bench --warm-up-time 3 --measurement-time 10`, every +binary started with the 1-minute load average below 2 (0.19–1.86) and no +`rustc` running. Candidate: `96086ad` (this change). Median (range) of 3 +rounds; run 2 also alternated `main` `425585e`. + +| function | 8f59b2e | 425585e (main) | 96086ad | vs 8f59b2e | +|---|---:|---:|---:|---:| +| run 1: `object_header_parse_x401` | 23.81 µs (23.76–24.00) | | 23.57 µs (23.23–23.89) | **−1.0%** | +| run 1: `snod_parse_all` | 1.840 µs (1.837–1.854) | | 1.839 µs (1.837–1.873) | 0.0% | +| run 1: `btree_v1_walk` | 343 ns (338–349) | | 352 ns (344–363) | +2.7% | +| run 1: `facade_list_400_groups` | 8.03 ms (8.01–8.10) | | 8.05 ms (8.01–8.15) | +0.2% | +| run 2: `object_header_parse_x401` | 24.15 µs (23.74–24.41) | 24.93 µs (24.72–25.05) | 23.52 µs (23.34–23.68) | **−2.6%** | +| run 2: `snod_parse_all` | 1.843 µs (1.830–1.847) | 1.856 µs (1.850–1.865) | 1.861 µs (1.858–1.869) | +1.0% | +| run 2: `btree_v1_walk` | 346 ns (339–366) | 356 ns (347–356) | 357 ns (350–359) | +3.2% | +| run 2: `facade_list_400_groups` | 8.09 ms (8.02–8.10) | 8.12 ms (7.97–8.19) | 8.08 ms (7.98–8.12) | −0.1% | + +- `ObjectHeader::parse` is at or below 8f59b2e (−1.0%, −2.6%) and 5.6% + faster than `main` in the same run. +- `btree_v1_walk` (one walk of a 350 ns B-tree) is 3% above 8f59b2e in + both runs, with overlapping ranges, and is the same on `main` (+0.2% + between `main` and this change): not from this change. The walk's code + changed in `e553153` (after a failed child the siblings are only read, + so the error returns after them; the fixture never takes that path, but + the loop carries the extra state); left as is. +- `snod_parse_all` and the facade listing are within noise. + ### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) `main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` diff --git a/CHANGELOG.md b/CHANGELOG.md index 1518e19..acd24d6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,16 @@ ## Unreleased +### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27) +- Parsing a version-1 object header was 4% slower than before range-read + M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk + message loop, kept out of line since the allocation fix; it is inlined + again. `object_header_parse_x401` is now 1.0% and 2.6% below `8f59b2e` + and 5.6% below `main` (tank, idle, separate binaries alternating; + `BENCHMARKS.md`). The chunk queue's checks are unchanged (65,536 chunks, + cycle refusal, file-size budget, one chunk buffer at a time, libhdf5's + order, overlapping chunks allowed). + ### Conformance: no our-errors or mismatches left (2026-09-27) - The last 3 our-errors and 2 mismatches were documented as not ours but still counted against us. Re-checked with evidence diff --git a/docs/known-issues.md b/docs/known-issues.md index d87d371..b8e47d6 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -9,7 +9,14 @@ deleting it. ## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) -**Status:** open (speed only; values are correct). Measured on an idle tank +**Status:** fixed 2026-09-27 (`96086ad`, branch +`perf/header-parse-and-last-files`): the remaining cost was the call to the +version-1 message loop, kept out of line by `#[inline(never)]`; inlined, +`object_header_parse_x401` is 1.0% and 2.6% below `8f59b2e` in two idle +A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's +speed"), with every chunk-queue check unchanged. + +Was: open (speed only; values are correct). Measured on an idle tank (load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`, separate binaries alternating, 3 rounds (`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):