Skip to content

[GLM-5.2 MXFP4][Tuning] Retune MXFP4 fused-MoE for gfx950 (TP4 + TP8) - #4629

Open
nholmber wants to merge 1 commit into
ROCm:mainfrom
nholmber:nholmber/glm52-mxfp4-fmoe-retune-gfx950
Open

[GLM-5.2 MXFP4][Tuning] Retune MXFP4 fused-MoE for gfx950 (TP4 + TP8)#4629
nholmber wants to merge 1 commit into
ROCm:mainfrom
nholmber:nholmber/glm52-mxfp4-fmoe-retune-gfx950

Conversation

@nholmber

@nholmber nholmber commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Motivation

Retune to use mxmoe a4w4 kernels for decode (#4179) giving +5.1% end-to-end serving throughput at TP4 and +2.4%
at TP8 on MI355X, with peaks of +22.4%.

Technical Details

Retunes aiter/configs/model_configs/glm5_fp4_tuned_fmoe.csv for the two GLM-5.2 fused-shared-expert shapes:

model_dim 6144, experts 257, topk 9
  inter_dim 512 (TP4)  12 rows changed
  inter_dim 256 (TP8)  16 rows changed

Test Plan

Validated improvement from updated tuning end-to-end on vLLM nightly, AITER aded0f84, flydsl==0.3.0:, using isl/osl/conc (1024/1024, 8192/1024, 60000/600 x concurrency 4..256).

Coherence and output sanity were checked before benchmarking.

Test Result

TP4 (inter_dim=512)

ISL OSL conc prompts current tok/s retuned tok/s Δ tput current TTFT ms retuned TTFT ms current TPOT ms retuned TPOT ms Δ TPOT
1024 1024 4 32 346.4 348.3 +0.6% 126 121 11.13 11.07 -0.5%
1024 1024 8 64 558.8 608.9 +9.0% 141 136 13.81 12.64 -8.5%
1024 1024 16 96 880.9 983.1 +11.6% 172 172 17.45 15.59 -10.7%
1024 1024 32 128 1289.0 1577.6 +22.4% 286 287 23.65 19.19 -18.8%
1024 1024 64 256 2107.1 2293.6 +8.9% 434 416 28.77 26.43 -8.1%
1024 1024 128 384 3328.2 3432.3 +3.1% 879 876 35.61 34.63 -2.8%
1024 1024 256 768 5018.7 4948.2 -1.4% 1610 1579 47.21 48.08 +1.8%
8192 1024 4 32 316.3 316.3 -0.0% 397 400 12.06 12.06 -0.0%
8192 1024 8 64 474.4 511.0 +7.7% 491 489 15.83 14.62 -7.6%
8192 1024 16 96 703.8 767.1 +9.0% 765 767 21.47 19.62 -8.6%
8192 1024 32 128 937.5 1085.1 +15.7% 1563 1546 31.46 26.92 -14.4%
8192 1024 64 256 1314.3 1389.1 +5.7% 2907 2896 44.20 41.73 -5.6%
8192 1024 128 384 1731.6 1764.4 +1.9% 7062 6979 64.27 63.14 -1.8%
8192 1024 256 768 2100.9 2105.8 +0.2% 14087 13976 103.77 103.77 -0.0%
60000 600 4 32 137.2 137.3 +0.0% 3058 3054 23.37 23.37 +0.0%
60000 600 8 64 161.8 166.1 +2.6% 4198 4064 41.12 40.14 -2.4%
60000 600 16 96 183.1 188.4 +2.9% 6531 6355 74.23 72.08 -2.9%
60000 600 32 128 193.9 200.0 +3.1% 14386 14302 137.09 132.39 -3.4%
60000 600 64 256 201.8 208.2 +3.2% 48447 46595 222.49 216.63 -2.6%
60000 600 128 384 205.6 210.8 +2.5% 188477 184231 231.96 227.22 -2.0%
60000 600 256 768 204.6 208.3 +1.8% 468647 459197 241.62 239.66 -0.8%
block tput TTFT TPOT
1024 / 1024 +7.5% -2.0% -7.0%
8192 / 1024 +5.6% -0.4% -5.6%
60000 / 600 +2.3% -2.1% -2.0%
overall +5.1%

TP8 (inter_dim=256)

ISL OSL conc prompts current tok/s retuned tok/s Δ tput current TTFT ms retuned TTFT ms current TPOT ms retuned TPOT ms Δ TPOT
1024 1024 4 32 347.2 357.7 +3.0% 131 124 11.10 10.77 -3.0%
1024 1024 8 64 634.0 659.8 +4.1% 142 140 12.11 11.63 -4.0%
1024 1024 16 96 1059.7 1156.5 +9.1% 176 174 14.42 13.18 -8.6%
1024 1024 32 128 1744.4 1830.4 +4.9% 255 254 17.33 16.50 -4.8%
1024 1024 64 256 2639.7 2880.5 +9.1% 357 380 22.91 20.91 -8.7%
1024 1024 128 384 4199.8 4347.5 +3.5% 759 754 28.12 27.19 -3.3%
1024 1024 256 768 6356.8 6444.5 +1.4% 1324 1314 37.22 36.76 -1.2%
8192 1024 4 32 323.4 331.8 +2.6% 336 336 11.85 11.54 -2.6%
8192 1024 8 64 547.2 563.7 +3.0% 427 421 13.67 13.25 -3.1%
8192 1024 16 96 860.2 918.2 +6.7% 656 645 17.50 16.36 -6.5%
8192 1024 32 128 1243.9 1279.4 +2.9% 1301 1327 23.48 22.77 -3.0%
8192 1024 64 256 1648.8 1730.5 +5.0% 2452 2461 35.08 33.32 -5.0%
8192 1024 128 384 2156.9 2167.6 +0.5% 5851 5964 51.36 51.05 -0.6%
8192 1024 256 768 2578.5 2545.4 -1.3% 11772 11989 84.30 85.46 +1.4%
60000 600 4 32 151.2 151.7 +0.4% 2571 2686 21.60 21.26 -1.6%
60000 600 8 64 188.0 188.3 +0.1% 3567 3710 35.46 35.13 -0.9%
60000 600 16 96 215.5 216.1 +0.3% 6242 5683 61.62 62.52 +1.5%
60000 600 32 128 229.0 227.3 -0.8% 12447 12788 115.62 116.07 +0.4%
60000 600 64 256 242.8 240.4 -1.0% 23158 24412 219.05 219.49 +0.2%
60000 600 128 384 248.1 245.8 -0.9% 128901 129396 257.03 261.18 +1.6%
60000 600 256 768 246.4 243.9 -1.0% 360731 365000 273.21 274.66 +0.5%
block tput TTFT TPOT
1024 / 1024 +5.0% -0.5% -4.8%
8192 / 1024 +2.7% +0.5% -2.8%
60000 / 600 -0.4% +1.2% +0.2%
overall +2.4%

Submission Checklist

Retunes glm5_fp4_tuned_fmoe.csv for the two GLM-5.2 fused-shared-expert
shapes on MI355X, moving the decode tiers onto the FlyDSL mxmoe a4w4 port.

  model_dim 6144, experts 257, topk 9
  inter_dim 512 (TP4): 12 rows
  inter_dim 256 (TP8): 16 rows

Tuned on gfx950 with the mxmoe path:

  python3 csrc/ck_gemm_moe_2stages_codegen/gemm_moe_tune.py \
    -i <shape>_untuned.csv -o <shape>_tuned.csv \
    --all --mp 8 --batch 100 --warmup 5 --iters 101 \
    --shape_grouped --mxfp4-flydsl --timeout 7200

The stage-2 _f4out variant is kept at inter_dim=512 tok=8192/16384 rather
than taking the tuner's pick: dropping it costs ~1.4% end to end at long
context, where 8192-token prefill chunks land on those tiers.
@nholmber
nholmber requested a review from a team August 7, 2026 15:36
@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 4629 --add-label <label>

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant