Conversation
GLM-5.3-Flash introduced three copies of "unwrap a parameter before handing it to a kernel that does not understand DTensor", in two different flavours: mhc.unshard_hc_params used full_tensor(), while kda._to_local and nope_dsa_mla._to_local used to_local(). The two differ only when the placement is Shard -- to_local() silently returns this rank's slice, full_tensor() inserts an all-gather -- and neither is safe there, because with grad_placements=None both label the (per-rank different) gradient with the parameter's own placement and communicate nothing on the way back. xtuner/v1/utils/dtensor.py::materialize_full replaces all three: it accepts a plain tensor or a Replicate DTensor and raises on Shard, naming the parameter. Also folds the DTensor-aware LayerNorm that glm52 and glm53 had each defined into xtuner/v1/module/rms_norm/layer_norm.py; it is a generic norm, not a per-model one. Test Plan: tests/utils/test_dtensor.py::TestMaterializeFull covers pass-through, Replicate unwrap, and the Shard rejection. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Includes the GLM-5.3-Flash design assessment (doc/xtuner_glm5p3flash_design.{md,py}, doc/glm5p3flash_vs_glm5p2.md, demo/glm5p3_flash_demo.py) that this and every later F* milestone implements against.
Crops the published FP8 62-shard GLM-5.3-Flash checkpoint to a 5-main-layer
+ original-MTP-layer BF16 checkpoint (~24.9B params) for local validation,
following doc/xtuner_glm5p3flash_design.md F0. Dequantizes 128x128
block-scaled FP8 weights to BF16 and rewrites the nested text_config
schedule fields (layer_types/mlp_layer_types/indexer_types/kda_layers/
full_attn_layers).
Verified against the real checkpoint: 24.95B total params (matches the
design doc's 24.9B target), config loads under pinned transformers==5.17.0,
and Glm5NextForConditionalGeneration.from_pretrained runs a real forward
pass with finite loss. Findings recorded in doc/progress.md, including that
this transformers build's Glm5NextForConditionalGeneration has no MTP
forward path, so the renamed layers.5.* (MTP) weights load as harmless
UNEXPECTED keys.
The official GLM-5.3-Flash release also ships a native BF16 checkpoint
(120 shards, 45 layers, ~313B params, no quantization_config), separate
from the FP8 62-shard release this script was originally built against.
Investigated what would need to change to make the crop script and model
loading compatible with it.
Findings, none of which required a code change:
- The crop script's FP8 dequantization is already conditional per tensor
(triggers only when a `*_scale_inv` companion key exists), so it
transparently passes a native-BF16 source through untouched -- confirmed
via --dry-run (same 3100 tensors / 14 source shards selected) and a real
crop run.
- MoE expert weights are stored per-expert (`mlp.experts.{i}.
{gate,up,down}_proj.weight`, i=0..287) in *both* the FP8 and native BF16
releases -- this is the official on-disk layout, not something
introduced by dequantization. XTuner's own loader already expects
exactly this layout: Glm53TextMoE.to_hf_key_list() expands the internal
fused fused_w1w3.weight/fused_w2.weight into 288 per-expert HF keys for
both load and save, and hf_tensor_to_canonical() reassembles the
per-expert tensors read back into the fused shape. No per-expert-to-fused
conversion step was missing; loading only ever supports this one format,
matching both releases.
Used the native BF16 release to cross-validate F0's FP8 dequant math:
cropped 25B from both sources and diffed tensor-by-tensor. Tensors in
modules_to_not_convert (embed_tokens, layernorm, hc_* -- never quantized)
match bit-for-bit; FP8-quantized tensors (gate_proj, expert weights,
q_a_proj) match bit-for-bit on 92-96% of elements, with the remainder
differing by max_abs_diff ~2.3e-5 / mean_abs_diff ~1e-7-1e-6 -- consistent
with FP8 e4m3's quantization granularity, not a dequant bug.
Since the native-BF16-sourced crop is strictly better (no FP8 round-trip)
and both sources produce an identical on-disk layout, switched the
canonical `~/model/GLM-5.3-Flash-25B` (the path every GLM_5_3_FLASH_PATH-
gated test and sft_glm53_tiny.sh default to) to the native-BF16-sourced
crop. The original FP8-dequantized crop is kept at
`~/model/GLM-5.3-Flash-25B-fp8dequant` as the cross-validation baseline,
no longer the default.
tests/model/test_glm53_25b_crop.py: added
test_glm53_25b_crop_passes_through_native_bf16_weights_unchanged, covering
the no-quantization_config / no-*_scale_inv input path that was previously
untested (the existing fixture always simulated an FP8-quantized source).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
jayhenry
added this pull request to stack #2112
September 23, 2026 06:54
This was referenced Sep 23, 2026
longboat2010
added a commit
to Ascend-SHL-SACT/xtuner
that referenced
this pull request
Sep 26, 2026
Merge InternLM GLM-5.3-Flash PRs InternLM#2105-InternLM#2111 into glm5_2
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack (bottom to top):
feat/glm53flash-materialize-full-f0→main← you are herefeat/glm53flash-f3-kda→feat/glm53flash-materialize-full-f0feat/glm53flash-f4-mhc→feat/glm53flash-f3-kdafeat/glm53flash-f5-nope-dsa→feat/glm53flash-f4-mhcfeat/glm53flash-f1-vl-data→feat/glm53flash-f5-nope-dsafeat/glm53flash-f2-vision-tower→feat/glm53flash-f1-vl-datafeat/glm53flash-f6-text-moe→feat/glm53flash-f2-vision-towerSummary
Stack layer 1/7 of GLM-5.3-Flash support. Two commits:
DTensor. This existed in two unsafe flavours already (mhc.unshard_hc_paramsusedfull_tensor(),kda._to_local/nope_dsa_mla._to_localusedto_local()) — both are wrong underShardplacement (silent wrong-rank slice, or an all-gather with no matching backward communication).xtuner/v1/utils/dtensor.py::materialize_fullreplaces all three call sites project-wide: accepts a plain tensor orReplicateDTensor, raises onShard. Also folds the DTensor-awareLayerNormthat glm52/glm53 had each defined separately into the sharedxtuner/v1/module/rms_norm/layer_norm.py.doc/xtuner_glm5p3flash_design.mdF0. Includes the design assessment doc (doc/xtuner_glm5p3flash_design.{md,py},doc/glm5p3flash_vs_glm5p2.md,demo/glm5p3_flash_demo.py) that this and every later F* milestone in the stack implements against.Verified against the real checkpoint: 24.95B total params (matches the design doc's 24.9B target), config loads under pinned
transformers==5.17.0, andGlm5NextForConditionalGeneration.from_pretrainedruns a real forward pass with finite loss. Also cross-validated the FP8 dequant math against the official native-BF16 checkpoint release (bit-for-bit on unquantized tensors, FP8-consistent tolerance elsewhere) and switched the canonical crop path to the native-BF16 source.Test Plan
tests/utils/test_dtensor.py::TestMaterializeFull: pass-through for a plain tensor, unwrap forReplicate,pytest.raisesforShard.tests/model/test_glm53_25b_crop.py, includingtest_glm53_25b_crop_passes_through_native_bf16_weights_unchangedfor the no-quantization_config input path.