PiPNN 1/6: add numerical kernels - #1287
Conversation
There was a problem hiding this comment.
Pull request overview
This PR adds the first set of PiPNN “kernel” building blocks to the DiskANN Rust workspace: SIMD-accelerated top‑k selection for partition assignment and leaf neighbor selection, along with supporting SIMD division and a new lower-triangular A·Aᵀ helper in diskann-linalg.
Changes:
- Add a new
diskann-pipnncrate withpartition_kernelandleaf_kernelimplementations plus extensive correctness tests and Criterion benchmarks. - Extend
diskann-wideto supportDivon relevant f32 SIMD types (native, doubled, and scalar/emulated) and add a corresponding division test macro. - Add
diskann_linalg::sgemm_aat_lower(lower-triangle-only AAT) and wire new crate/tests/CI/mutants exclusions into the workspace.
Reviewed changes
Copilot reviewed 26 out of 27 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| diskann-wide/src/test_utils/ops.rs | Adds test_div! macro to validate lane-wise SIMD division correctness. |
| diskann-wide/src/emulated.rs | Adds Div for scalar/emulated Emulated<f32, N, A> to support division in scalar dispatch. |
| diskann-wide/src/doubled.rs | Adds Div for Doubled<T> to support composite SIMD widths. |
| diskann-wide/src/arch/x86_64/v4/f32x8_.rs | Adds AVX Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v4/f32x4_.rs | Adds SSE Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v4/f32x16_.rs | Adds AVX-512 Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v3/f32x8_.rs | Adds AVX Div op mapping + division tests for V3. |
| diskann-wide/src/arch/x86_64/v3/f32x4_.rs | Adds SSE Div op mapping + division tests for V3. |
| diskann-wide/src/arch/x86_64/v3/f32x16_.rs | Adds division tests for the f32x16 V3 path (likely via doubled composition). |
| diskann-wide/src/arch/aarch64/f32x4_.rs | Adds Neon Div op mapping + division tests. |
| diskann-wide/src/arch/aarch64/f32x2_.rs | Adds Neon Div op mapping + division tests. |
| diskann-pipnn/tests/partition_kernel.rs | New integration tests for partition top‑k dispatch correctness and edge cases. |
| diskann-pipnn/tests/leaf_kernel.rs | New integration tests for leaf neighbor top‑k dispatch correctness and edge cases. |
| diskann-pipnn/src/partition_kernel/tests.rs | New unit tests comparing scalar reference vs runtime dispatch and metric contracts. |
| diskann-pipnn/src/partition_kernel.rs | New partition-assignment distance + top‑k kernel with validation and SIMD dispatch. |
| diskann-pipnn/src/lib.rs | New crate root exporting PiPNN kernel modules. |
| diskann-pipnn/src/leaf_kernel/tests.rs | New unit tests for scalar reference parity and workspace behavior. |
| diskann-pipnn/src/leaf_kernel.rs | New fused lower-triangle leaf neighbor kernel with SIMD dispatch and workspace support. |
| diskann-pipnn/Cargo.toml | Defines new diskann-pipnn crate, dev-deps, and benches. |
| diskann-pipnn/benches/kernels.rs | Adds benchmarks for partition top‑k, lower AAT, leaf top‑k, and full leaf workflow. |
| diskann-linalg/tests/sgemm_aat_lower.rs | New tests for lower-triangle AAT behavior and validation errors. |
| diskann-linalg/src/lib.rs | Adds public sgemm_aat_lower API with dimension checks. |
| diskann-linalg/src/faer.rs | Implements sgemm_aat_lower_impl using Faer triangular matmul. |
| Cargo.toml | Adds diskann-pipnn to workspace members and workspace dependencies. |
| Cargo.lock | Records the new diskann-pipnn package entry. |
| .github/workflows/ci.yml | Adds diskann-pipnn to CI test package lists. |
| .cargo/mutants.toml | Adds mutation-test exclusions for kernel code paths and equivalent transformations. |
Comments suppressed due to low confidence (2)
diskann-pipnn/src/leaf_kernel.rs:651
- Same issue as the L2 arm: using
max_simdfor lower clamping can erase NaNs on the Scalar/Emulated backend, making NaN distances rankable. Clamp withlt_simd+selectto preserve NaNs consistently.
Metric::CosineNormalized => {
let distance = F::splat(arch, 1.0) - dot;
zero.max_simd(distance)
}
diskann-pipnn/src/leaf_kernel.rs:664
- The cosine path also uses
zero.max_simd(distance)for clamping, which can collapse NaNs to zero on the Scalar/Emulated backend (viaf32::max). That contradicts the comment about preserving non-rankable NaNs and can change output ordering. Prefer anlt_simd+selectclamp here as well.
let distance = one - cosine;
// Comparisons with NaN are false, so this explicit lower clamp
// preserves non-rankable NaNs while matching the existing PiPNN
// distance formulas for finite values.
zero.max_simd(distance)
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
e204cb9 to
b046174
Compare
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1287 +/- ##
==========================================
- Coverage 90.59% 90.46% -0.14%
==========================================
Files 513 547 +34
Lines 99091 106211 +7120
==========================================
+ Hits 89775 96086 +6311
- Misses 9316 10125 +809
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (2)
diskann-pipnn/src/partition_kernel/tests.rs:20
- The
PartitionTopKcontract forMetric::L2expectsleader_scalesto contain squared leader norms (see docs anddistance(Metric::L2, ..)test). This helper currently populates unsquared norms, which makes the test data inconsistent with the public API contract and could hide contract-related bugs.
let leader_scales = match metric {
Metric::L2 => (0..leaders).map(|leader| (leader + 1) as f32).collect(),
Metric::Cosine => (0..leaders)
.map(|leader| {
diskann-pipnn/src/partition_kernel.rs:61
InvalidFanout’s error message says the maximum is{maximum}, but validation also rejectsfanout > leaders. Whenleaders < maximumthis message is misleading (it implies the only limit is{maximum}). Consider spelling out both constraints in the message so callers immediately see why it failed.
#[error("invalid fanout {fanout} for {leaders} leaders; maximum is {maximum}")]
8fb4e92 to
20ab8a0
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 25 out of 26 changed files in this pull request and generated no new comments.
Suppressed comments (1)
diskann-pipnn/src/partition_kernel.rs:294
- For
Metric::Cosine, NaN norms currently produce a finite distance (1.0) becausedenominator.gt_simd(0)is false for NaN, so the lane falls back tocosine = 0. That makes NaN-derived pairs/leaders “rankable”, which contradicts the module’s stated NaN-rejection behavior and differs fromdiskann-vectorcosine semantics (NaN norms propagate to a NaN similarity/distance). Consider explicitly preserving NaN denominators so the resulting distance stays NaN and is ignored byinsert_topk.
let denominator = row_norm * leader_norm;
let valid = denominator.gt_simd(zero);
let safe_denominator = valid.select(denominator, one);
let cosine = valid.select(dot / safe_denominator, zero);
one - cosine
| check_length("leader scales", input.leader_scales.len(), leader_scales) | ||
| } | ||
|
|
||
| fn checked_area( |
There was a problem hiding this comment.
checked_area, check_length, ShapeOverflow and InvalidBufferLength are duplicated character-for-character with leaf_kernel. Small enough to shrug at now, but with four more PRs coming it's probably worth a src/shape.rs with a shared ShapeError that each kernel error wraps via #[from].
There was a problem hiding this comment.
I kept the two tiny checked-area/length adapters local because they construct different public kernel error types and sit immediately before each module's unsafe accesses. MatrixView adoption removed the other duplicated shape state; introducing a shared wrapped error would enlarge the public error interface for two call sites.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 25 out of 26 changed files in this pull request and generated no new comments.
Suppressed comments (1)
.github/workflows/nightly.yml:23
DISKANN_FEATURESis defined as a folded scalar with commas at line ends. YAML folding inserts spaces at line breaks, producing a value liketracing, experimental_diversity_search,...which can be parsed as having empty/whitespace-prefixed feature names depending on Cargo’s splitting rules. This is brittle and can break thecargo ... --features "${{ env.DISKANN_FEATURES }}"steps.
DISKANN_FEATURES: >-
virtual_storage,spherical-quantization,product-quantization,tracing,
experimental_diversity_search,disk-index,flatbuffers,linalg,codegen,
multi-vector,bftree,inmem2,integration-test
Use output columns as the sole leaf-specific neighbor count and reserve row/column terminology for matrix shapes. BREAKING CHANGE: LeafKernel::new no longer takes k, nearest_neighbors returns (), and kernel input/neighbor/error fields use source-target and point-leader names.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 25 out of 26 changed files in this pull request and generated no new comments.
Suppressed comments (1)
diskann-pipnn/src/leaf_kernel.rs:467
- The comment claims no output or scratch mutation occurs on error, but after
validate(...)the call toprepare_workspace(...)can returnLeafKernelError::Allocationafter partially resizing/fillingworkspace.norms(beforeworkspace.worstis reserved). This makes the comment/documentation inaccurate and could mislead callers relying on workspace immutability on error.
// Validation establishes every shape and active-prefix invariant used by
// unchecked loads below. No output or scratch mutation occurs on error.
validate(call.input, &call.output)?;
|
I went through the full PiPANN implementation, and it appears to be entirely in-memory. How does it handle datasets that cannot fit into memory? One possible out-of-core approach would be to stream the input, load vectors only when processing each leaf partition, compute distances locally, and then stitch the partial graphs into the final graph. However, this would introduce additional I/O and graph-merging overhead. Do we have experimental results for datasets large enough that loading all vectors into memory is infeasible? It would be helpful to understand the memory–build-time tradeoff and whether out-of-core construction has been evaluated. |
Keep PiPNN beside graph policy so later layers can reuse private RobustPrune state without publishing it across a crate boundary. Preserve independent kernel oracles while removing duplicate formula-sharing differential wrappers.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 28 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (2)
diskann/benches/benchmarks_iai/pipnn_kernels.rs:34
PartitionScales::L2requires squared leader norms, but this benchmark fixture fillsleader_squared_normswith unsquared values (1.0 + leader/LEADERS). That makes the benchmark exercise a different scoring formula than the real kernel contract (and can skew perf/regression tracking if the score distribution changes).
Compute actual squared norms here (or rename+document if you intentionally want non-norm values, but that would violate PartitionScales::L2’s stated units).
let leader_squared_norms = (0..LEADERS)
.map(|leader| 1.0 + leader as f32 / LEADERS as f32)
.collect();
diskann/benches/benchmarks_iai/pipnn_kernels.rs:91
setup_leafsizesoutputusingneighbors = leaf_neighbor_count(LEAF_POINTS, LEAF_K), butselect_leaf_neighborscreates the output view withLEAF_Kcolumns instead of the computedneighbors. This works only as long asLEAF_K <= LEAF_POINTS - 1; if the constants change (or if you copy this pattern elsewhere with small leaves), it will panic at runtime due to a shape/len mismatch.
Use the effective neighbor count when constructing the output view to keep the fixture consistent with the kernel API.
LeafInput {
dots: MatrixView::try_from(dots.as_slice(), LEAF_POINTS, LEAF_POINTS).unwrap(),
},
MutMatrixView::try_from(output.as_mut_slice(), LEAF_POINTS, LEAF_K).unwrap(),
&mut workspace,
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 28 out of 29 changed files in this pull request and generated 1 comment.
Suppressed comments (2)
.github/workflows/ci.yml:422
- Same issue as above:
--features diskann/pipnnis not a valid Cargo feature flag and will cause the SDE AVX-512 test job to error before running tests. Switch to--features pipnn.
cargo test --locked \
--package diskann-wide \
--package diskann-vector \
--package diskann-quantization \
--package diskann \
--features diskann/pipnn \
-- --skip compile_tests
diskann/benches/benchmarks_iai/pipnn_kernels.rs:88
- The output buffer is allocated using
neighbors = leaf_neighbor_count(LEAF_POINTS, LEAF_K), but the view passed tonearest_neighborsusesLEAF_Kdirectly. This is only correct whileLEAF_K <= LEAF_POINTS - 1; if the constants change (or this benchmark is copied with a larger K), the view shape will no longer match the allocation and will panic/fail validation. Useneighbors(or recompute it) consistently for the output-column count.
.nearest_neighbors(
LeafInput {
dots: MatrixView::try_from(dots.as_slice(), LEAF_POINTS, LEAF_POINTS).unwrap(),
},
MutMatrixView::try_from(output.as_mut_slice(), LEAF_POINTS, LEAF_K).unwrap(),
&mut workspace,
| cargo test --locked \ | ||
| --package diskann-wide \ | ||
| --package diskann-vector \ | ||
| --package diskann-quantization \ | ||
| --package diskann \ | ||
| --features diskann/pipnn \ | ||
| -- --skip compile_tests \ |
PiPNN (Pick-in-Partitions Nearest Neighbors) builds ANN graph candidates with overlapping partitions and dense matrix work instead of running beam search against a partially built graph for every inserted point. This first layer adds the numerical kernels used by later PiPNN stages. It does not yet build or persist a graph.
PiPNN now lives under
diskann::graph::pipnn; there is no separate implementation crate. This keeps graph construction beside DiskANN graph policy and lets later layers reuse crate-private graph internals without publishing them.Concepts
fanoutis the number of nearest leaders retained for each point, creating overlapping child partitions.kis the number of local companions retained per point; it is construction policy, not final graph degreeR.Code map
diskann-linalg::sgemm_aat_lowercomputesA · Aᵀand writes only the lower triangle. Callers may leave the upper triangle uninitialized.diskann/src/graph/pipnn/kernel_metric.rsowns metric formulas, scale units, zero/NaN behavior, and runtime metric selection shared by both kernels.partition_kernel.rsconverts point-by-leader dot products into sorted nearest leader IDs. Metric-specific scale handling happens before fixed-size top-k insertion.leaf_kernel.rsscans each strict-lower-triangle pair once and updates both endpoint top-k trackers.k <= 3uses fixed-size insertion; largerkuses the dynamic fallback.diskann/tests/pipnn_{partition,leaf}_kernel.rsare public-interface differential tests with independent formulas. Private tests stay beside branch-heavy tracker and workspace seams.diskann/benches/bench_main_iai.rsis the single DiskANN IAI-Callgrind target. This layer registers partition and leaf kernel groups with instruction/cache regression limits.End-to-end flow
Caller computes dense dot products → typed kernel input validates matrix/scales/output →
diskann-wideselects the runtime architecture once → scalar/SIMD chunks convert dots to metric distances → stable top-k insertion writes caller-owned IDs/neighbors.The kernels do not own providers, graph IDs, recursion, edge merging, thread pools, persistence, or search.
Invariants and boundaries
usizearea overflow are validated before dispatch.u32positions; leaf output contains leaf-local target positions.f32::MAXremains rankable.diskann-wide.Review path
kernel_metric.rs: metric formulas, scale kinds, zero thresholds, NaN handling, and scalar equivalents.PartitionKernelvalidation and tracker insertion, then compare scalar tails with SIMD chunks.LeafKernellower-triangle traversal, dual-endpoint updates, fixed/dynamic top-k paths, and workspace reuse.sgemm_aat_lowernever touches the upper triangle.Validation
k/fanout paths, dimensions around lane boundaries, tails, ties, NaN, infinities, signed zero, zero/singleton/capacity inputs, and validation failures.diskann/pipnn; nightly feature/coverage jobs include the moved module.cargo bench -p diskann --bench bench_main_iai --features pipnn,testing.Stack relation
Stack 1/6. This layer supplies numerical selection. #1288 extracts crate-private shared RobustPrune; #1290 adds partition/leaf orchestration.
Stack 1/6 → #1288