Compaction: eliminating disk scan and updating PQ retrain#691
Open
dian-lun-lin wants to merge 5 commits into
Open
Compaction: eliminating disk scan and updating PQ retrain#691dian-lun-lin wants to merge 5 commits into
dian-lun-lin wants to merge 5 commits into
Conversation
…uteLayerInfoFromSources computeLayerInfoFromSources called getNodes(0) on each source graph to count live nodes at level 0. getNodes(0) sequentially seeks through every node record on disk to filter out deleted entries. On a cold page cache this touches large amounts of source data before compaction even begins, significantly delaying the start of actual graph merging. Since every live node is present at level 0 by the HNSW invariant, the count is simply liveNodes.get(s).cardinality() — an in-memory popcount requiring no I/O. Also switch PQ retraining from ProductQuantization.compute() (full k-means++ init) to basePQ.refine() (Lloyd's iterations only, warm-started from the existing codebook). The source codebooks are already trained on the same distribution, so warm-starting converges in far fewer passes with no recall loss.
dian-lun-lin
requested review from
MarkWolters,
ashkrisk,
jshook and
tlwillke
as code owners
July 2, 2026 22:19
dian-lun-lin
marked this pull request as draft
July 2, 2026 22:19
Contributor
|
Before you submit for review:
If you did not complete any of these, then please explain below. |
dian-lun-lin
marked this pull request as ready for review
July 4, 2026 08:05
MarkWolters
requested changes
Jul 9, 2026
The timer was initialized after the extraction call, so the logged duration was always ~0. Addresses PR review feedback.
Source graphs are mapped with MADV_RANDOM (correct for search-time access), which disables kernel readahead and makes compaction's bulk phases fault one page at a time on a cold cache (measured: retrain ~37s, pre-encode ~42s on a disk-cold 10M-node compaction that takes ~3s/~8s warm). Adds ReaderSupplier.prefetch(offset, length) — streams a byte range into the page cache through a separate readahead-enabled descriptor — and OnDiskGraphIndex.prefetchL0Records(minNode, maxNode) on top of it. Each bulk phase warms exactly the records it is about to read, from the worker that will read them: - PQ retrain prefetches its (source, node)-sorted sample ranges before extraction (bounded by training-set size, not file size), - code pre-encode prefetches each chunk's records at task start, - L0 batch processing prefetches each batch's own records at task start (cross-source search reads are data-dependent and stay demand-faulted). The per-thread prefetch buffer is 64KB: streamed bytes are discarded, so the buffer is sized to stay L2-resident. Measured on 56 threads: disk-cold throughput is device-bound (~3.3 GB/s) at every size from 32KB to 4MB, and in the warm case (range already cached, pass is pure overhead) 64KB runs within noise of larger buffers while 4KB is 2-3x slower in both regimes. Transient cache demand is proportional to the in-flight windows, so there is no up-front whole-file streaming pass and no memory-availability gate to mistune: on a box that cannot hold the sources, pages are simply evicted and reads degrade per-page to the old fault-on-demand behavior.
The L0 cross-source candidate search passed beamWidth (= 2x searchTopK) as rerankK, doubling the approximate-phase beam over what the candidate budget needs. A seeding-vs-beam decomposition study (7 paired disk-cold arms, cohere-10M, median-of-3) showed the narrower beam is where the time goes: S=2: L0 165.2s -> 133.4s (-19%), recall 0.5718 -> 0.5660 S=4: L0 260.3s -> 224.6s (-14%), recall 0.5733 -> 0.5659 Query latency on the merged index is unaffected (0.55ms avg both ways). The wider beam's extra candidates were largely pruned by diversity selection, which keeps at most degree edges per node. The same study found warm-start seeding of these searches (from finished neighbors' merged adjacency, reverse candidates, or upper-layer descent) is net-negative: at matched beam width, seeded searches run 8-17% slower with equal recall — the per-search seeding overhead exceeds the few cheap descent hops it saves. Beam width is the whole lever.
Refinement ablations (three datasets, disk-cold, paired same-window runs) show its recall contribution on the merged index is ~0: same-beam arms with and without refinement land within 0.001. Skipping it saves 45-58s, ~20-25% of total compaction time at 10M nodes. What it buys is navigability — query latency on the merged index rises from ~0.55ms to ~0.95ms avg (p99 2.0ms to 3.9ms, cohere-10M) without it. That is a workload tradeoff, not a correctness call: default to compaction throughput, and let latency-sensitive pipelines opt back in with setRefineAfterCompaction(true).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does
Speeds up graph compaction (merging pre-built partition indexes into one) — 5.6–11.5× faster
cold, at roughly the same recall. Five changes:
computeLayerInfoFromSourcescalledgetNodes(0), whichseeks + reads every node record on disk just to count live nodes. Replaced with an in-memory
liveNodes.cardinality()popcount (no I/O). The big one on large datasets.codebook instead of full k-means++ from scratch (far fewer passes, no recall loss).
MADV_RANDOM, so bulk phases faulted onepage at a time. Now each phase (retrain sampling, code pre-encode, L0 batches) streams the exact
records it is about to read into the page cache first.
rerankK = searchTopK— the L0 cross-source search used a beam 2× the candidate budget; theextra candidates are pruned by diversity selection anyway. Halving it cuts ~19% off L0 for
~0.006 recall.
contributes ~0 recall (a query-latency/navigability pass), so it is now opt-in.
The two structural fixes (1. scan elimination, 3. prefetch) remove the disk-bound overhead that
dominated cold compaction; 4./ 5. trim the remaining compute. Recall stays within ~0.006 of main.
main vs PR (perf-fix + prefetch): compaction time & recall
220g heap,
-wi 1 -i 1single fork ("cold" = warmup iteration),drop_cachesbefore each arm so both start disk-cold. Graph degree 32, beam width 100.Recall is search@10 on the compacted graph.
b86a94c8(stock)compaction-perf-fix@8324dd6d= disk-scan fix + warm-start retrain + streamingprefetch + rerankK + refinement-off
Takeaways