fix knn_merge_parts for inner product - #2426
Open
qwertyforce wants to merge 2 commits into
Open
Conversation
divyegala
reviewed
Aug 11, 2026
divyegala
left a comment
Contributor
There was a problem hiding this comment.
Hi @qwertyforce, thank you for the PR! I think this change will severely affect the binary size of cuVS, can you please provide a before-after measurement? Alternatively, we could run an element-wise operation to negate the inner product distances before running the merge kernel.
Contributor
Author
|
Hi! i ran ./build.sh libcuvs --allgpuarch --no-nvtx -n (GCC 13, CUDA 12.8) |
Contributor
|
Thanks @qwertyforce , in that case can we please pursue the alternative of running a negation for inner product? |
Contributor
Author
|
Yes, will try to implement and benchmark it |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
By default, when merging the results of sharded multigpu index, knn_merge_parts keeps K smallest values. But inner product is not a distance, it is a measure of similarity. Therefore results are wrong for IP.
In this PR we are adding an additional overload, that receives an argument select_min, we keep backward compatibility and add a new test