Fix normalized cosine distance consistency for PQ graphs - #1298
Fix normalized cosine distance consistency for PQ graphs#1298partychen wants to merge 8 commits into
Conversation
Use half squared L2 consistently for every PQ-involved CosineNormalized comparison so in-memory graph search and pruning operate on compatible values. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
This PR makes Metric::CosineNormalized behavior internally consistent across all PQ-involved distance paths by using 0.5 * squared_l2 for query↔PQ, full↔PQ, and PQ↔PQ comparisons. This prevents pruning/search from comparing incompatible distance scales when PQ vectors are involved.
Changes:
- Introduces a shared scaling constant and applies it to
FixedChunkPQTableCosineNormalized distances (query↔PQ and PQ↔PQ). - Adds a scaled lookup-table construction path for
CosineNormalizedso query-time evaluation remains a table lookup. - Adds regression tests to ensure cross-computer consistency and correct hybrid (full/quant) dispatch behavior.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
diskann-providers/src/model/pq/fixed_chunk_pq_table.rs |
Defines the scale constant and applies scaled squared-L2 for CosineNormalized in direct PQ distance paths. |
diskann-providers/src/model/pq/distance/l2.rs |
Adds new_scaled to scale precomputed L2 lookup tables for CosineNormalized preprocessing. |
diskann-providers/src/model/pq/distance/dynamic.rs |
Wires CosineNormalized to the scaled L2 preprocessing and updates QQ dispatch + tests. |
diskann-providers/src/model/graph/provider/async_/distances.rs |
Adds a regression test ensuring hybrid full/quant and quant/quant CosineNormalized paths use the scaled L2 definition. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1298 +/- ##
==========================================
+ Coverage 90.59% 91.49% +0.89%
==========================================
Files 513 516 +3
Lines 99091 98330 -761
==========================================
+ Hits 89775 89963 +188
+ Misses 9316 8367 -949
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
Avoid scaling the standard L2 lookup table, clarify PQ code documentation, and reuse the shared normalized-cosine scale in hybrid tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (4)
diskann-providers/src/model/pq/distance/dynamic.rs:505
- The relative tolerance here (6.3e-7) is tighter than other SIMD-vs-scalar distance tests in this crate (commonly 1e-6). Relaxing to 1e-6 would make this regression test less likely to be flaky across platforms.
assert_relative_eq!(
cosine_normalized.evaluate_similarity(&*code0, &*code1),
expected,
max_relative = 6.3e-7,
);
diskann-providers/src/model/graph/provider/async_/distances.rs:212
- This assertion uses a very tight relative tolerance (1e-7). Using 1e-6 would better match other SIMD-vs-scalar comparisons in diskann-providers and reduce the risk of cross-platform FP flakiness.
assert_relative_eq!(quant_quant, expected_quant_quant, max_relative = 1.0e-7);
diskann-providers/src/model/graph/provider/async_/distances.rs:203
- This assertion uses a very tight relative tolerance (1e-7). Using 1e-6 would better match other SIMD-vs-scalar comparisons in diskann-providers and reduce the risk of cross-platform FP flakiness.
This issue also appears on line 212 of the same file.
assert_relative_eq!(full_quant, expected_full_quant, max_relative = 1.0e-7);
diskann-providers/src/model/pq/distance/dynamic.rs:439
- These new floating-point assertions use a tighter relative tolerance (5e-7) than similar SIMD-vs-scalar comparisons elsewhere in this crate (often 1e-6). Consider relaxing to 1e-6 to reduce cross-arch / compiler flakiness.
This issue also appears on line 501 of the same file.
assert_relative_eq!(query_distance, expected, max_relative = 5.0e-7);
assert_relative_eq!(random_access_distance, expected, max_relative = 5.0e-7);
assert_relative_eq!(
query_distance,
random_access_distance,
Replace the `TableL2::new_scaled` constructor with a `scaled` adapter so the existing `new` stays untouched and construction is separate from the scale. Name the test tolerance, document the scale constant, and note in the docs that ordering is preserved and that full/full pairs use the exact metric. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (1)
diskann-providers/src/model/pq/distance/dynamic.rs:59
- The doc comment claims that for non-normalized operands the
0.5 * squared_l2approximation differs from normalized cosine distance only by a positive factor and therefore preserves candidate ordering. That relationship is not generally true when norms vary; the difference is not just a constant scale, and ordering can change. Please adjust the comment to avoid stating an incorrect guarantee.
/// In other words, half the squared L2 distance equals normalized cosine distance when
/// both operands are normalized, and the two differ by a positive factor otherwise, so
/// candidate ordering is preserved either way.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 8 out of 8 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (2)
diskann-providers/src/model/graph/provider/async_/distances.rs:187
- The hybrid regression test uses non-unit full vectors (
[1, 0, 0, 2],[2, 0, 0, 1]), butMetric::CosineNormalizedis documented/implemented under a unit-norm assumption. Using unit vectors here would make the test’s intent clearer (CosineNormalized == 0.5*squared-L2) and avoid locking in behavior that only matches scaled-L2 for non-normalized inputs.
let table = FixedChunkPQTable::new(
4,
vec![1.0, 0.0, 0.0, 1.0, 2.0, 0.0, 0.0, 2.0].into(),
vec![0, 2, 4].into(),
)
diskann-providers/src/model/graph/provider/async_/distances.rs:141
HybridComputer::newmapsMetric::CosineNormalizedfull/full comparisons ontoMetric::L2with a 0.5 scale. This contradicts the PR description’s stated scope that full/full comparisons remain on the nativeCosineNormalizedimplementation, and it can also change behavior when full vectors are not perfectly unit-norm. Consider keeping the full-precision path onMetric::CosineNormalized(scale 1.0) and relying on the PQ side’s scaled-L2 approximation for compatibility.
This issue also appears on line 183 of the same file.
pub fn new(quant: pq::distance::DistanceComputer<'a>, dim: Option<usize>) -> Self {
let (full_metric, full_scale) = match quant.metric() {
Metric::CosineNormalized => (Metric::L2, pq::COSINE_NORMALIZED_L2_SCALE),
metric => (metric, 1.0),
};
Add unit-norm coverage while preserving the non-unit integer regression that verifies every Hybrid pruning path uses the same scaled-L2 definition. Update the stale HybridComputer docs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Kept full/full scaled-L2 intentionally because Hybrid pruning mixes full/full, full/PQ, and PQ/PQ comparisons. Restoring native CosineNormalized only for full/full would reintroduce incompatible distance definitions, especially for u8/i8. The tests now separate and document unit-vector semantics and integer consistency. |
Keep the existing Product-PQ query approximation unchanged and select the same squared-L2 metric when constructing Hybrid and Quantized pruning distance computers. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Exercise the Quantized PruneStrategy wiring directly and distinguish its regression coverage from the Hybrid operand-combination test. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Problem
For Product-PQ with
CosineNormalized, graph-search queries used raw squared L2 while graph pruning used cosine distance.Vamana pruning compares
distance_ik / distance_jk. For normalized vectors, cosine distance is half squared L2, so mixing the two approximately doubled this ratio. This madealpha = 1.2behave likealpha = 0.6, causing over-pruning and poor recall.Fix
Use raw squared L2 for Product-PQ Hybrid and Quantized pruning, matching the existing PQ graph-search query approximation.
The change is scoped to Product-PQ pruning. Search behavior, public PQ distance APIs, other metrics, and other quantization strategies are unchanged.
Validation
Regression tests cover all Hybrid operand combinations and the Quantized PQ/PQ pruning path.
The final commit was benchmarked on the first 100,000 BigANN SIFT base vectors and first 1,000 queries, converted to unit-normalized
f32as required byCosineNormalized. Exact ground truth was recomputed for this subset. Configuration: 50 PQ chunks,max_fp_vecs_per_prune = 48,max_degree = 64,l_build = 100,alpha = 1.2, andsearch_l = 100.Search results
Graph structure
Checks:
cargo test -p diskann-providers --libcargo clippy --workspace --all-targets -- -D warnings