Skip to content

PiPNN 2/6: add numerical kernels - #1287

Open
weiyaoluo (SeliMeli) wants to merge 30 commits into
pipnn-stack/02-final-prunefrom
pipnn-stack/01-kernels
Open

PiPNN 2/6: add numerical kernels#1287
weiyaoluo (SeliMeli) wants to merge 30 commits into
pipnn-stack/02-final-prunefrom
pipnn-stack/01-kernels

Conversation

@SeliMeli

@SeliMeli weiyaoluo (SeliMeli) commented Jul 29, 2026

Copy link
Copy Markdown

PiPNN (Pick-in-Partitions Nearest Neighbors) builds ANN graph candidates with overlapping partitions and dense matrix work instead of running beam search against a partially built graph for every inserted point. This numerical layer adds the kernels used by later PiPNN stages. It does not yet build or persist a graph.

PiPNN now lives under diskann::graph::pipnn; there is no separate implementation crate. This keeps graph construction beside DiskANN graph policy and lets later layers reuse crate-private graph internals without publishing them.

Concepts

  • A leader is a sampled point that names a child partition.
  • fanout is the number of nearest leaders retained for each point, creating overlapping child partitions.
  • A leaf is a bounded partition processed with one lower-triangular all-pairs dot-product matrix.
  • Leaf k is the number of local companions retained per point; it is construction policy, not final graph degree R.

Code map

  1. diskann-linalg::sgemm_aat_lower computes A · Aᵀ and writes only the lower triangle. Callers may leave the upper triangle uninitialized.
  2. diskann/src/graph/pipnn/kernel_metric.rs owns metric formulas, scale units, zero/NaN behavior, and runtime metric selection shared by both kernels.
  3. partition_kernel.rs converts point-by-leader dot products into sorted nearest leader IDs. Metric-specific scale handling happens before fixed-size top-k insertion.
  4. leaf_kernel.rs scans each strict-lower-triangle pair once and updates both endpoint top-k trackers. k <= 3 uses fixed-size insertion; larger k uses the dynamic fallback.
  5. partition_kernel.rs and leaf_kernel.rs co-locate private seam tests with independent formula differentials; diskann-linalg/src/lib.rs similarly owns the lower-triangle GEMM tests.

End-to-end flow

Caller computes dense dot products → typed kernel input validates matrix/scales/output → diskann-wide selects the runtime architecture once → scalar/SIMD chunks convert dots to metric distances → stable top-k insertion writes caller-owned IDs/neighbors.

The kernels do not own providers, graph IDs, recursion, edge merging, thread pools, persistence, or search.

Invariants and boundaries

  • Matrix shapes, scale lengths, fanout, output widths, and usize area overflow are validated before dispatch.
  • Partition output contains leader-local u32 positions; leaf output contains leaf-local target positions.
  • Leaf traversal reads only the diagonal and strict lower triangle, updating each unordered pair exactly once.
  • Ties preserve encounter order. NaN candidates are non-rankable; finite f32::MAX remains rankable.
  • Cosine zero/subnormal norms produce zero similarity without erasing unrelated NaN behavior.
  • Scalar tails and dispatched chunks preserve the documented rounding/order contract.
  • PiPNN names no ISA, target feature, or raw architecture intrinsic; dispatch remains owned by diskann-wide.

Review path

  1. Start with kernel_metric.rs: metric formulas, scale kinds, zero thresholds, NaN handling, and scalar equivalents.
  2. Review PartitionKernel validation and tracker insertion, then compare scalar tails with SIMD chunks.
  3. Review LeafKernel lower-triangle traversal, dual-endpoint updates, fixed/dynamic top-k paths, and workspace reuse.
  4. Verify sgemm_aat_lower never touches the upper triangle.
  5. Finish with the co-located independent differentials around 4/8/16-lane and second-chunk boundaries.

Validation

  • 33 co-located PiPNN kernel tests at this layer: 11 private seam tests plus 22 interface differentials; seven lower-triangle GEMM tests live beside sgemm_aat_lower.
  • Differential coverage spans all four metrics, k/fanout paths, dimensions around lane boundaries, tails, ties, NaN, infinities, signed zero, zero/singleton/capacity inputs, and validation failures.
  • x86-64 baseline and AVX-512 SDE jobs explicitly enable the workspace feature pipnn; nightly feature/coverage jobs include the moved module.
  • AArch64 and Windows cross-target checks compile the feature.

Stack relation

Stack 2/6. Depends on #1315, which extracts crate-private shared RobustPrune. This layer supplies numerical selection; #1290 adds partition/leaf orchestration using both lower layers.

Stack 2/6: #1315#1290

@SeliMeli
weiyaoluo (SeliMeli) requested review from a team and a lite review from Copilot July 29, 2026 11:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds the first set of PiPNN “kernel” building blocks to the DiskANN Rust workspace: SIMD-accelerated top‑k selection for partition assignment and leaf neighbor selection, along with supporting SIMD division and a new lower-triangular A·Aᵀ helper in diskann-linalg.

Changes:

  • Add a new diskann-pipnn crate with partition_kernel and leaf_kernel implementations plus extensive correctness tests and Criterion benchmarks.
  • Extend diskann-wide to support Div on relevant f32 SIMD types (native, doubled, and scalar/emulated) and add a corresponding division test macro.
  • Add diskann_linalg::sgemm_aat_lower (lower-triangle-only AAT) and wire new crate/tests/CI/mutants exclusions into the workspace.

Reviewed changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
diskann-wide/src/test_utils/ops.rs Adds test_div! macro to validate lane-wise SIMD division correctness.
diskann-wide/src/emulated.rs Adds Div for scalar/emulated Emulated<f32, N, A> to support division in scalar dispatch.
diskann-wide/src/doubled.rs Adds Div for Doubled<T> to support composite SIMD widths.
diskann-wide/src/arch/x86_64/v4/f32x8_.rs Adds AVX Div op mapping + division tests.
diskann-wide/src/arch/x86_64/v4/f32x4_.rs Adds SSE Div op mapping + division tests.
diskann-wide/src/arch/x86_64/v4/f32x16_.rs Adds AVX-512 Div op mapping + division tests.
diskann-wide/src/arch/x86_64/v3/f32x8_.rs Adds AVX Div op mapping + division tests for V3.
diskann-wide/src/arch/x86_64/v3/f32x4_.rs Adds SSE Div op mapping + division tests for V3.
diskann-wide/src/arch/x86_64/v3/f32x16_.rs Adds division tests for the f32x16 V3 path (likely via doubled composition).
diskann-wide/src/arch/aarch64/f32x4_.rs Adds Neon Div op mapping + division tests.
diskann-wide/src/arch/aarch64/f32x2_.rs Adds Neon Div op mapping + division tests.
diskann-pipnn/tests/partition_kernel.rs New integration tests for partition top‑k dispatch correctness and edge cases.
diskann-pipnn/tests/leaf_kernel.rs New integration tests for leaf neighbor top‑k dispatch correctness and edge cases.
diskann-pipnn/src/partition_kernel/tests.rs New unit tests comparing scalar reference vs runtime dispatch and metric contracts.
diskann-pipnn/src/partition_kernel.rs New partition-assignment distance + top‑k kernel with validation and SIMD dispatch.
diskann-pipnn/src/lib.rs New crate root exporting PiPNN kernel modules.
diskann-pipnn/src/leaf_kernel/tests.rs New unit tests for scalar reference parity and workspace behavior.
diskann-pipnn/src/leaf_kernel.rs New fused lower-triangle leaf neighbor kernel with SIMD dispatch and workspace support.
diskann-pipnn/Cargo.toml Defines new diskann-pipnn crate, dev-deps, and benches.
diskann-pipnn/benches/kernels.rs Adds benchmarks for partition top‑k, lower AAT, leaf top‑k, and full leaf workflow.
diskann-linalg/tests/sgemm_aat_lower.rs New tests for lower-triangle AAT behavior and validation errors.
diskann-linalg/src/lib.rs Adds public sgemm_aat_lower API with dimension checks.
diskann-linalg/src/faer.rs Implements sgemm_aat_lower_impl using Faer triangular matmul.
Cargo.toml Adds diskann-pipnn to workspace members and workspace dependencies.
Cargo.lock Records the new diskann-pipnn package entry.
.github/workflows/ci.yml Adds diskann-pipnn to CI test package lists.
.cargo/mutants.toml Adds mutation-test exclusions for kernel code paths and equivalent transformations.
Comments suppressed due to low confidence (2)

diskann-pipnn/src/leaf_kernel.rs:651

  • Same issue as the L2 arm: using max_simd for lower clamping can erase NaNs on the Scalar/Emulated backend, making NaN distances rankable. Clamp with lt_simd + select to preserve NaNs consistently.
        Metric::CosineNormalized => {
            let distance = F::splat(arch, 1.0) - dot;
            zero.max_simd(distance)
        }

diskann-pipnn/src/leaf_kernel.rs:664

  • The cosine path also uses zero.max_simd(distance) for clamping, which can collapse NaNs to zero on the Scalar/Emulated backend (via f32::max). That contradicts the comment about preserving non-rankable NaNs and can change output ordering. Prefer an lt_simd + select clamp here as well.
            let distance = one - cosine;
            // Comparisons with NaN are false, so this explicit lower clamp
            // preserves non-rankable NaNs while matching the existing PiPNN
            // distance formulas for finite values.
            zero.max_simd(distance)

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel/tests.rs Outdated
@SeliMeli weiyaoluo (SeliMeli) changed the title Pipnn stack/01 kernels PiPNN 1/6: add numerical kernels Jul 29, 2026
Copilot AI review requested due to automatic review settings July 30, 2026 08:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.

@codecov-commenter

Codecov Comments Bot (codecov-commenter) commented Jul 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.97938% with 39 lines in your changes missing coverage. Please review.
✅ Project coverage is 90.27%. Comparing base (59dd048) to head (ba837a8).
⚠️ Report is 12 commits behind head on main.

Files with missing lines Patch % Lines
diskann/src/graph/pipnn/partition_kernel.rs 91.13% 29 Missing ⚠️
diskann/src/graph/pipnn/leaf_kernel.rs 97.90% 9 Missing ⚠️
diskann/src/graph/pipnn/kernel_metric.rs 99.30% 1 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1287      +/-   ##
==========================================
- Coverage   90.59%   90.27%   -0.33%     
==========================================
  Files         513      548      +35     
  Lines       99091   106446    +7355     
==========================================
+ Hits        89775    96095    +6320     
- Misses       9316    10351    +1035     
Flag Coverage Δ
miri 90.27% <95.97%> (-0.33%) ⬇️
unittests 89.97% <95.97%> (-0.31%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
diskann-linalg/src/faer.rs 100.00% <100.00%> (ø)
diskann-linalg/src/lib.rs 99.68% <100.00%> (+1.18%) ⬆️
diskann-wide/src/arch/x86_64/v3/f32x16_.rs 100.00% <ø> (ø)
diskann-wide/src/arch/x86_64/v3/f32x4_.rs 100.00% <ø> (ø)
diskann-wide/src/arch/x86_64/v3/f32x8_.rs 100.00% <ø> (ø)
diskann-wide/src/arch/x86_64/v4/f32x16_.rs 14.11% <ø> (ø)
diskann-wide/src/arch/x86_64/v4/f32x4_.rs 16.90% <ø> (ø)
diskann-wide/src/arch/x86_64/v4/f32x8_.rs 16.90% <ø> (ø)
diskann-wide/src/doubled.rs 86.89% <100.00%> (+0.17%) ⬆️
diskann-wide/src/emulated.rs 98.31% <100.00%> (+0.01%) ⬆️
... and 5 more

... and 311 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Copilot AI review requested due to automatic review settings July 30, 2026 08:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.

Comments suppressed due to low confidence (2)

diskann-pipnn/src/partition_kernel/tests.rs:20

  • The PartitionTopK contract for Metric::L2 expects leader_scales to contain squared leader norms (see docs and distance(Metric::L2, ..) test). This helper currently populates unsquared norms, which makes the test data inconsistent with the public API contract and could hide contract-related bugs.
    let leader_scales = match metric {
        Metric::L2 => (0..leaders).map(|leader| (leader + 1) as f32).collect(),
        Metric::Cosine => (0..leaders)
            .map(|leader| {

diskann-pipnn/src/partition_kernel.rs:61

  • InvalidFanout’s error message says the maximum is {maximum}, but validation also rejects fanout > leaders. When leaders < maximum this message is misleading (it implies the only limit is {maximum}). Consider spelling out both constraints in the message so callers immediately see why it failed.
    #[error("invalid fanout {fanout} for {leaders} leaders; maximum is {maximum}")]

Copilot AI review requested due to automatic review settings July 31, 2026 04:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 25 out of 26 changed files in this pull request and generated no new comments.

Suppressed comments (1)

diskann-pipnn/src/partition_kernel.rs:294

  • For Metric::Cosine, NaN norms currently produce a finite distance (1.0) because denominator.gt_simd(0) is false for NaN, so the lane falls back to cosine = 0. That makes NaN-derived pairs/leaders “rankable”, which contradicts the module’s stated NaN-rejection behavior and differs from diskann-vector cosine semantics (NaN norms propagate to a NaN similarity/distance). Consider explicitly preserving NaN denominators so the resulting distance stays NaN and is ignored by insert_topk.
        let denominator = row_norm * leader_norm;
        let valid = denominator.gt_simd(zero);
        let safe_denominator = valid.select(denominator, one);
        let cosine = valid.select(dot / safe_denominator, zero);
        one - cosine

Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann/src/graph/pipnn/partition_kernel.rs
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated

@partychen juchen-ms (partychen) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice work overall. I found one correctness issue in the cosine handling that should be resolved before merge. The remaining comments are mostly about reducing duplicated or unsafe code and tightening the API contracts.

Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
check_length("leader scales", input.leader_scales.len(), leader_scales)
}

fn checked_area(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

checked_area, check_length, ShapeOverflow and InvalidBufferLength are duplicated character-for-character with leaf_kernel. Small enough to shrug at now, but with four more PRs coming it's probably worth a src/shape.rs with a shared ShapeError that each kernel error wraps via #[from].

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I kept the two tiny checked-area/length adapters local because they construct different public kernel error types and sit immediately before each module's unsafe accesses. MatrixView adoption removed the other duplicated shape state; introducing a shared wrapped error would enlarge the public error interface for two call sites.

Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-linalg/src/lib.rs Outdated
Comment thread diskann-wide/src/emulated.rs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Weiyao, this is progress from the previous mega-PR. I still have some big-picture comments (we covered most of these offline) -

  • Documentation: As I mentioned, we need thorough documentation in the diskann-pipnn crate. The main modules, partition_kernel and leaf_kernel need documentation up top, highlighting the main structures and how they are used - e.g. process_rows_binary/unary and nearest_leaders. Similarly with process_pairs_simd_* and nearest_leaf_neighbors
  • Testing: I am concerned about the lack of testing for partition_kernel.rs and leaf_kernel.rs.
    • I notice some e2e integration tests but these kernels should be thoroughly tested, sweeping different input parameters, architectures and edge cases. This is especially needed given the amount of unsafe code.
    • That brings me to miri - there should be miri tests too.
    • I'm curious why are the tests in a separate submodule to the main files (for partition_kernel.rs and leaf_kernel.rs)? Let's try to keep tests along with the code being tested.
  • Criterion: Since criterion is not a standard part of our library for benchmarking, let us not introduce it for this crate.
  • Kernel dispatch: I left comments about you're disptaching the kernels, please take a look.

Comment thread .cargo/mutants.toml Outdated
Comment thread diskann-pipnn/src/lib.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/src/partition_kernel.rs Outdated
Comment thread diskann-pipnn/tests/leaf_kernel_api.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Comment thread diskann-pipnn/src/leaf_kernel.rs Outdated
Copilot AI review requested due to automatic review settings August 3, 2026 02:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.

Suppressed comments (2)

diskann-pipnn/src/partition_kernel/tests.rs:19

  • PartitionTopK::leader_scales is documented as "squared leader norms for L2" (and cosine uses unsquared norms), but this test helper feeds unsquared values for the L2 case. That makes the test data inconsistent with the public contract and can mask mistakes in distance computation. Consider squaring the L2 norms here so the tests exercise the intended inputs.
    let leader_scales = match metric {
        Metric::L2 => (0..leaders).map(|leader| (leader + 1) as f32).collect(),
        Metric::Cosine => (0..leaders)

diskann-pipnn/src/partition_kernel.rs:252

  • For the L2 path, the SIMD chunk uses mul_add_simd (fused multiply-add) but the scalar tail uses norm - 2.0 * dot (non-fused). This can introduce small rounding differences between SIMD and tail elements, which can change ordering/tie behavior right at SIMD-width boundaries. Use f32::mul_add for the scalar tail so both paths compute the same value shape.
                |dot, norm| F::splat(arch, -2.0).mul_add_simd(dot, norm),
                |dot, norm| norm - 2.0 * dot,

@SeliMeli

Copy link
Copy Markdown
Author

weiyaoluo (@SeliMeli) please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.

@microsoft-github-policy-service agree [company="{your company}"]

Options:

  • (default - no company specified) I have sole ownership of intellectual property rights to my Submissions and I am not making Submissions in the course of work for my employer.
@microsoft-github-policy-service agree
  • (when company given) I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.
@microsoft-github-policy-service agree company="Microsoft"

Contributor License Agreement

@microsoft-github-policy-service agree company="Microsoft"

Copilot AI review requested due to automatic review settings August 3, 2026 10:49
Select architecture, metric, and leaf-width implementations once, then reuse direct diskann-wide function pointers across stripes and leaves.

BREAKING CHANGE: callers construct LeafKernel or PartitionKernel and pass MatrixView-backed inputs and outputs.
Use output columns as the sole leaf-specific neighbor count and reserve row/column terminology for matrix shapes.

BREAKING CHANGE: LeafKernel::new no longer takes k, nearest_neighbors returns (), and kernel input/neighbor/error fields use source-target and point-leader names.
Keep PiPNN beside graph policy so later layers can reuse private RobustPrune state without publishing it across a crate boundary. Preserve independent kernel oracles while removing duplicate formula-sharing differential wrappers.
Remove the submitted DiskANN microbenchmark target and co-locate numerical tests with their implementation files.
@SeliMeli
weiyaoluo (SeliMeli) changed the base branch from main to pipnn-stack/02-final-prune August 6, 2026 09:36
@SeliMeli weiyaoluo (SeliMeli) changed the title PiPNN 1/6: add numerical kernels PiPNN 2/6: add numerical kernels Aug 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants