Skip to content

[Feature] Add GLM-5.3-Flash F0: 25B cropped reference checkpoint builder - #2105

Open
jayhenry wants to merge 2 commits into
mainfrom
feat/glm53flash-materialize-full-f0
Open

jayhenry wants to merge 2 commits into
mainfrom
feat/glm53flash-materialize-full-f0

Conversation

@jayhenry

@jayhenry jayhenry commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Stack (bottom to top):

  1. [Feature] Add GLM-5.3-Flash F0: 25B cropped reference checkpoint builder #2105 feat/glm53flash-materialize-full-f0 → main ← you are here
  2. [Feature] Add GLM-5.3-Flash F3: Kimi Delta Attention (KDA) #2106 feat/glm53flash-f3-kda → feat/glm53flash-materialize-full-f0
  3. [Feature] Add GLM-5.3-Flash F4: mHC four-stream residual #2107 feat/glm53flash-f4-mhc → feat/glm53flash-f3-kda
  4. [Feature] Add GLM-5.3-Flash F5: NoPE DSA + KPool indexer + clamped SwiGLU #2108 feat/glm53flash-f5-nope-dsa → feat/glm53flash-f4-mhc
  5. [Feature] Add GLM-5.3-Flash F1: VL data preprocessing pipeline #2109 feat/glm53flash-f1-vl-data → feat/glm53flash-f5-nope-dsa
  6. [Feature] Add GLM-5.3-Flash F2: vision tower + projector (eager) #2110 feat/glm53flash-f2-vision-tower → feat/glm53flash-f1-vl-data
  7. [Feature] Add GLM-5.3-Flash F6 core: text model + MTP + compose model #2111 feat/glm53flash-f6-text-moe → feat/glm53flash-f2-vision-tower

Base is main. Review only this PR's own diff.


Summary

Stack layer 1/7 of GLM-5.3-Flash support. Two commits:

  1. [Refactor] Unify the DTensor unshard contract behind materialize_full — GLM-5.3-Flash's later layers (mHC, KDA, NoPE-DSA) each need to unwrap a parameter before handing it to a kernel that doesn't understand DTensor. This existed in two unsafe flavours already (mhc.unshard_hc_params used full_tensor(), kda._to_local/nope_dsa_mla._to_local used to_local()) — both are wrong under Shard placement (silent wrong-rank slice, or an all-gather with no matching backward communication). xtuner/v1/utils/dtensor.py::materialize_full replaces all three call sites project-wide: accepts a plain tensor or Replicate DTensor, raises on Shard. Also folds the DTensor-aware LayerNorm that glm52/glm53 had each defined separately into the shared xtuner/v1/module/rms_norm/layer_norm.py.
  2. [Feature] Add GLM-5.3-Flash F0: 25B cropped reference checkpoint builder — crops the published FP8 62-shard GLM-5.3-Flash checkpoint to a 5-main-layer + original-MTP-layer BF16 checkpoint (~24.9B params) for local validation, per doc/xtuner_glm5p3flash_design.md F0. Includes the design assessment doc (doc/xtuner_glm5p3flash_design.{md,py}, doc/glm5p3flash_vs_glm5p2.md, demo/glm5p3_flash_demo.py) that this and every later F* milestone in the stack implements against.

Verified against the real checkpoint: 24.95B total params (matches the design doc's 24.9B target), config loads under pinned transformers==5.17.0, and Glm5NextForConditionalGeneration.from_pretrained runs a real forward pass with finite loss. Also cross-validated the FP8 dequant math against the official native-BF16 checkpoint release (bit-for-bit on unquantized tensors, FP8-consistent tolerance elsewhere) and switched the canonical crop path to the native-BF16 source.

Test Plan

  • tests/utils/test_dtensor.py::TestMaterializeFull: pass-through for a plain tensor, unwrap for Replicate, pytest.raises for Shard.
  • tests/model/test_glm53_25b_crop.py, including test_glm53_25b_crop_passes_through_native_bf16_weights_unchanged for the no-quantization_config input path.

jayhenry and others added 2 commits September 23, 2026 06:43
GLM-5.3-Flash introduced three copies of "unwrap a parameter before handing it to a kernel that
does not understand DTensor", in two different flavours: mhc.unshard_hc_params used
full_tensor(), while kda._to_local and nope_dsa_mla._to_local used to_local(). The two differ
only when the placement is Shard -- to_local() silently returns this rank's slice, full_tensor()
inserts an all-gather -- and neither is safe there, because with grad_placements=None both label
the (per-rank different) gradient with the parameter's own placement and communicate nothing on
the way back.

xtuner/v1/utils/dtensor.py::materialize_full replaces all three: it accepts a plain tensor or a
Replicate DTensor and raises on Shard, naming the parameter.

Also folds the DTensor-aware LayerNorm that glm52 and glm53 had each defined into
xtuner/v1/module/rms_norm/layer_norm.py; it is a generic norm, not a per-model one.

Test Plan: tests/utils/test_dtensor.py::TestMaterializeFull covers pass-through, Replicate
unwrap, and the Shard rejection.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Includes the GLM-5.3-Flash design assessment (doc/xtuner_glm5p3flash_design.{md,py}, doc/glm5p3flash_vs_glm5p2.md, demo/glm5p3_flash_demo.py) that this and every later F* milestone implements against.

Crops the published FP8 62-shard GLM-5.3-Flash checkpoint to a 5-main-layer
+ original-MTP-layer BF16 checkpoint (~24.9B params) for local validation,
following doc/xtuner_glm5p3flash_design.md F0. Dequantizes 128x128
block-scaled FP8 weights to BF16 and rewrites the nested text_config
schedule fields (layer_types/mlp_layer_types/indexer_types/kda_layers/
full_attn_layers).

Verified against the real checkpoint: 24.95B total params (matches the
design doc's 24.9B target), config loads under pinned transformers==5.17.0,
and Glm5NextForConditionalGeneration.from_pretrained runs a real forward
pass with finite loss. Findings recorded in doc/progress.md, including that
this transformers build's Glm5NextForConditionalGeneration has no MTP
forward path, so the renamed layers.5.* (MTP) weights load as harmless
UNEXPECTED keys.

The official GLM-5.3-Flash release also ships a native BF16 checkpoint
(120 shards, 45 layers, ~313B params, no quantization_config), separate
from the FP8 62-shard release this script was originally built against.
Investigated what would need to change to make the crop script and model
loading compatible with it.

Findings, none of which required a code change:
- The crop script's FP8 dequantization is already conditional per tensor
  (triggers only when a `*_scale_inv` companion key exists), so it
  transparently passes a native-BF16 source through untouched -- confirmed
  via --dry-run (same 3100 tensors / 14 source shards selected) and a real
  crop run.
- MoE expert weights are stored per-expert (`mlp.experts.{i}.
  {gate,up,down}_proj.weight`, i=0..287) in *both* the FP8 and native BF16
  releases -- this is the official on-disk layout, not something
  introduced by dequantization. XTuner's own loader already expects
  exactly this layout: Glm53TextMoE.to_hf_key_list() expands the internal
  fused fused_w1w3.weight/fused_w2.weight into 288 per-expert HF keys for
  both load and save, and hf_tensor_to_canonical() reassembles the
  per-expert tensors read back into the fused shape. No per-expert-to-fused
  conversion step was missing; loading only ever supports this one format,
  matching both releases.

Used the native BF16 release to cross-validate F0's FP8 dequant math:
cropped 25B from both sources and diffed tensor-by-tensor. Tensors in
modules_to_not_convert (embed_tokens, layernorm, hc_* -- never quantized)
match bit-for-bit; FP8-quantized tensors (gate_proj, expert weights,
q_a_proj) match bit-for-bit on 92-96% of elements, with the remainder
differing by max_abs_diff ~2.3e-5 / mean_abs_diff ~1e-7-1e-6 -- consistent
with FP8 e4m3's quantization granularity, not a dequant bug.

Since the native-BF16-sourced crop is strictly better (no FP8 round-trip)
and both sources produce an identical on-disk layout, switched the
canonical `~/model/GLM-5.3-Flash-25B` (the path every GLM_5_3_FLASH_PATH-
gated test and sft_glm53_tiny.sh default to) to the native-BF16-sourced
crop. The original FP8-dequantized crop is kept at
`~/model/GLM-5.3-Flash-25B-fp8dequant` as the cross-validation baseline,
no longer the default.

tests/model/test_glm53_25b_crop.py: added
test_glm53_25b_crop_passes_through_native_bf16_weights_unchanged, covering
the no-quantization_config / no-*_scale_inv input path that was previously
untested (the existing fixture always simulated an FP8-quantized source).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant