feat(pt_expt): auto-select O(N) NeighborGraph builder by default - #5903
feat(pt_expt): auto-select O(N) NeighborGraph builder by default#5903Shaurya2k06 wants to merge 3 commits into
Conversation
📝 WalkthroughWalkthroughChangesThe inference path now uses a shared resolver to select NeighborGraph auto selection
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant DeepEval
participant resolve_auto_graph_builder
participant nv
participant vesin
participant dense
DeepEval->>resolve_auto_graph_builder: Resolve auto builder for DEVICE
resolve_auto_graph_builder->>nv: Check CUDA availability
resolve_auto_graph_builder->>vesin: Check Vesin availability
resolve_auto_graph_builder-->>DeepEval: Return nv, vesin, or dense
DeepEval->>dense: Use dense when no preferred backend is available
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
source/tests/pt_expt/model/test_graph_builder_dispatch.py (1)
146-158: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAssert dispatch, not only parity.
All backends are intentionally value-equivalent, so this passes if
None/"auto"incorrectly falls back todense. Mock or spy on the resolver/concrete builder and assert that the resolved backend is invoked.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@source/tests/pt_expt/model/test_graph_builder_dispatch.py` around lines 146 - 158, Update test_none_and_auto_match_resolved_builder to spy on or mock the resolver/concrete graph builder, then assert that the backend resolved for None and "auto" is actually invoked. Retain the existing output-parity assertions, but ensure the test fails if either input silently falls back to dense.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@deepmd/pt_expt/utils/vesin_graph_builder.py`:
- Around line 11-15: Update the documentation describing the shared
resolve_auto_graph_builder default ladder to state that vesin is selected for
CPU or CUDA fallback only when vesin.torch is importable; otherwise document
that dense is selected.
---
Nitpick comments:
In `@source/tests/pt_expt/model/test_graph_builder_dispatch.py`:
- Around line 146-158: Update test_none_and_auto_match_resolved_builder to spy
on or mock the resolver/concrete graph builder, then assert that the backend
resolved for None and "auto" is actually invoked. Retain the existing
output-parity assertions, but ensure the test fails if either input silently
falls back to dense.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: d02887c0-8595-4f4e-a31f-c9bc7c3749a5
📒 Files selected for processing (9)
deepmd/pt_expt/infer/deep_eval.pydeepmd/pt_expt/model/make_model.pydeepmd/pt_expt/train/training.pydeepmd/pt_expt/utils/neighbor_graph_method.pydeepmd/pt_expt/utils/nv_graph_builder.pydeepmd/pt_expt/utils/vesin_graph_builder.pysource/tests/pt_expt/infer/test_graph_deepeval.pysource/tests/pt_expt/model/test_graph_builder_dispatch.pysource/tests/pt_expt/utils/test_neighbor_graph_method.py
987d303 to
409bb04
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #5903 +/- ##
==========================================
- Coverage 79.47% 79.22% -0.26%
==========================================
Files 1072 1072
Lines 125055 125063 +8
Branches 4541 4541
==========================================
- Hits 99388 99081 -307
- Misses 24043 24357 +314
- Partials 1624 1625 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
409bb04 to
451f50d
Compare
|
Rebased onto current master after #5912 / #5913 landed the training/eager auto path. What changed in this update
CodeQL empty- |
Training keeps CPU on dense (vesin loops frames). Inference auto now shares resolve_auto_graph_builder: CUDA nv→vesin→dense, CPU vesin→dense. Signed-off-by: shaurya2k06 <shaurya2k06@gmail.com>
451f50d to
01253d9
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.
Suppressed comments (3)
deepmd/pt_expt/utils/graph_builder.py:33
- Docstring names the CUDA dependency as
nvalchemiops, but the rest of the module (warnings/errors) refers to the pip packagenvalchemi-toolkit-ops. Using the installable package name here avoids confusion.
* CUDA: ``nv`` if ``nvalchemiops`` is importable, else ``vesin`` if
``vesin.torch`` is importable, else ``dense``.
deepmd/pt_expt/utils/graph_builder.py:56
- When
neighbor_graph_method='auto'falls back todenseon CUDA due to missing optional deps, the warning doesn’t tell users how to enable the faster backends (unlike the training-path warning below). Adding install hints makes the message actionable.
log.warning(
"nvalchemi-toolkit-ops and vesin[torch] are unavailable; falling "
"back from neighbor_graph_method='auto' to the dense graph builder."
)
deepmd/pt_expt/utils/graph_builder.py:31
- PR description/issue state that the model-level/training default should follow the same CPU/CUDA ladder as inference (CPU preferring vesin when available), but the code here explicitly documents (and
resolve_neighbor_graph_methodenforces) a separate training policy that keeps CPU ondense. This means the implementation doesn’t match the stated acceptance criteria unless the PR description/issue closure is adjusted.
This issue also appears in the following locations of the same file:
- line 32
- line 53
Single owner of the inference / DeepEval auto ladder (training uses
:func:`resolve_neighbor_graph_method`, which keeps CPU on ``dense`` because
vesin loops frames in Python and is not safe as a multi-frame training
default):
There was a problem hiding this comment.
🧹 Nitpick comments (1)
deepmd/pt_expt/utils/graph_builder.py (1)
58-60: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick winConsider logging the CPU dense fallback for parity with the CUDA branch.
The CUDA branch logs a warning when it falls back to
dense(Lines 53-56). The CPU branch falls back todensesilently at Line 60. Dense is the O(N²) carry-all builder; silently downgrading to it on CPU (e.g.vesin[torch]not installed) can cause an unexplained performance regression on large systems, with no diagnostic for the user to act on.Add a similar
log.warning(orlog.info) call before returning"dense"on the CPU path, mentioning how to installvesin[torch]. If you make this change, updatetest_auto_resolution's("cpu", False, False, "dense", False)case insource/tests/pt_expt/infer/test_deep_eval_pt_checkpoint.pytowarns=True.♻️ Proposed fix
if is_vesin_torch_available(): return "vesin" + log.warning( + "vesin[torch] is unavailable; falling back from " + "neighbor_graph_method='auto' to the dense graph builder on CPU. " + "Install it with `pip install vesin[torch]` to enable the O(N) " + "vesin graph builder." + ) return "dense"🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@deepmd/pt_expt/utils/graph_builder.py` around lines 58 - 60, Update the CPU fallback in the graph-builder backend resolution function to log a warning or info message before returning "dense", explicitly mentioning installation of vesin[torch]. Also update the test_auto_resolution case for ("cpu", False, False, "dense") so it expects a warning.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@deepmd/pt_expt/utils/graph_builder.py`:
- Around line 58-60: Update the CPU fallback in the graph-builder backend
resolution function to log a warning or info message before returning "dense",
explicitly mentioning installation of vesin[torch]. Also update the
test_auto_resolution case for ("cpu", False, False, "dense") so it expects a
warning.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 46027a52-d036-4bf6-9812-cd1fa7c5ac10
📒 Files selected for processing (4)
deepmd/pt_expt/infer/deep_eval.pydeepmd/pt_expt/utils/graph_builder.pydeepmd/pt_expt/utils/vesin_graph_builder.pysource/tests/pt_expt/infer/test_deep_eval_pt_checkpoint.py
🚧 Files skipped from review as they are similar to previous changes (1)
- deepmd/pt_expt/utils/vesin_graph_builder.py
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated no new comments.
Suppressed comments (2)
deepmd/pt_expt/utils/graph_builder.py:31
- The new helper documents/supports
neighbor_graph_method="auto", but the core builder dispatch (build_neighbor_graph_for_method) still only accepts concrete methods. Outside ofDeepEval._resolve_neighbor_graph_method, passingneighbor_graph_method="auto"into a pt_expt model graph path would still raise aValueErrorfrom the builder dispatcher. Either wire auto-resolution into the model/dispatcher, or clarify here that callers must resolve "auto" before dispatching.
"""Resolve ``neighbor_graph_method="auto"`` to a concrete inference builder.
Single owner of the inference / DeepEval auto ladder (training uses
:func:`resolve_neighbor_graph_method`, which keeps CPU on ``dense`` because
vesin loops frames in Python and is not safe as a multi-frame training
deepmd/pt_expt/utils/vesin_graph_builder.py:16
- This module docstring reads as if
neighbor_graph_method="auto"is a general pt_expt model option, but currently the only in-tree resolver for "auto" is DeepEval (and the graph builder dispatcher itself rejects "auto"). Consider clarifying that "auto" here refers to DeepEval/inference resolution so users don’t try passing "auto" directly into modelneighbor_graph_methodand hit a runtimeValueError.
for ``nf == 1`` inference and CPU use. Inference ``neighbor_graph_method="auto"``
(:func:`~deepmd.pt_expt.utils.graph_builder.resolve_auto_graph_builder`) selects
vesin only when ``vesin.torch`` is importable (CPU always; CUDA only when ``nv``
is unavailable); otherwise it falls back to ``dense``. Training auto keeps CPU
on ``dense`` and never selects vesin. Prefer ``nv`` (:mod:`.nv_graph_builder`)
Closes #5902
Summary
resolve_auto_graph_builder(device)policy for the carry-all NeighborGraph builders: CUDA prefersnvthenvesinthendense; CPU prefersvesinthendense(asestays explicit-only).None/"auto") from hard-coded"dense"to that ladder, and use the same helper in DeepEval (default"auto") and compiled training's eager_forward_graph..pt2artifacts are unchanged.Why existing tests missed this
Builder dispatch already had vesin/nv vs dense energy/force parity, but nothing asserted that the model/DeepEval default would pick an O(N) builder when available, and compiled training still hardcoded dense while claiming to match the eager default-flip. The new resolver unit tests pin the availability ladder; the extended dispatch and DeepEval tests pin value-transparency of
None/"auto"against the resolved concrete builder.Validation
ruff check/ruff format --checkon touched filespytestfocused suite: 20 passed, 1 skipped (nvCUDA-only)source/tests/pt_expt/utils/test_neighbor_graph_method.pysource/tests/pt_expt/model/test_graph_builder_dispatch.py.pt2parityneighbor_graph_methodSummary by CodeRabbit
New Features
Bug Fixes
Tests