[NV] Add H200 DeepSeek-V4-Pro AgentX recipes - #2364
Conversation
Add one 8xH200 aggregated TP8 (EP1/DP1) DeepSeek-V4-Pro FP8 Dynamo-SGLang AgentX recipe with EAGLE MTP and HiCache, sweeping concurrency [1,2,4,8,16] over a fixed serving topology. Uses the Marlin MoE backend and Dynamo header-based session affinity (X-Dynamo-Session-ID). - New recipe agg-h200-tp8-mtp-kvoffload.yaml + master key dsv4-fp8-h200-dynamo-sglang-agentic-agg (decode num-worker 0 so per-GPU accounting reflects the single aggregated TP8 worker). - benchmark_lib.sh: skip the legacy nvext conv-aware CLI routing when a recipe opts into the header path (AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID=true); the existing AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING opt-out and default are unchanged. - launch_h200-dgxc-slurm.sh: dsv4 fp8 model-path routing, srt-slurm v1.0.10 overlay for the agentic recipe, on-demand SGLang/nginx squash imports, and AgentX dataset / HF caches mounted into the multi-node agentic path.
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
2 similar comments
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30310586014 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30310920956 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
|
/stage-results |
|
@cquil11 staged run 30311003032: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-27~r30311003032 This shared staging slot remains available until the next |
The SGLang image is already staged on the H200 cluster at /data/containers/*.sqsh, like the dsr1 multinode path which imports nothing. Drop the import_squash helper and its now-unused NGINX_IMAGE; the multinode path maps SQUASH_FILE into srtslurm.yaml directly.
中文:合并 origin/main 并解决 perf-changelog 冲突
|
Revoking the standing `/reuse-sweep-run` authorization on this PR (removing the bare command comment from 2026-07-28). A bare `/reuse-sweep-run` is standing rather than one-shot, and it has been silently swallowing sweeps here. The run at the current head, 30505397990, shows the gate emitting That matters because a real code change landed after the authorization: commit The last real evidence, 30311003032 (5/5 multi-node agentic), is at Re-authorize with an explicit run ID once a fresh sweep lands. |
|
@csahithi the AgentX/AIPerf harness has been updated, please merge origin/main into your branch and refresh your submission. Additional tuning may be necessary depending on the config. I apologize for any inconvenience. This is an automated message. |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30506436017 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30586537651 |
Add one 8xH200 aggregated TP8 (EP1/DP1) DeepSeek-V4-Pro FP8 Dynamo-SGLang AgentX recipe with EAGLE MTP and HiCache, sweeping concurrency [1,2,4,8,16] over a fixed serving topology. Uses the Marlin MoE backend and Dynamo header-based session affinity (X-Dynamo-Session-ID).