Skip to content

MoE-quantized models crash on default grouped_mm inference path (QuantTensor transpose not supported) #2619

Description

Summary

Fused-expert MoE weights quantized by Rtn/Gptq/KQuant (moe=True) are stored as 3D QuantTensors that are storage-only: any shape-movement op (transpose, reshape, view, permute, expand) unconditionally raises RuntimeError in QuantTensor.__torch_function__ (_unsupported_movement, olive/common/quant/tensor.py).

This is a deliberate design choice — see the docstring on _unsupported_movement and _MOE_ONNX_EXPORT_MSG — because pack_to_uint8/unpack_from_uint8 bit-pack multiple logical values from the tensor's last dimension into shared uint8 bytes, and scales/qzeros are grouped along that same last dimension by group_size. A transpose can't be represented as a stride/metadata-only view (the way it is for a normal dense tensor) because the physical bytes interleave values that, after a transpose, would need to belong to different rows/groups. Supporting a real transpose would require unpack → dequantize → dense transpose → re-quantize, which is lossy (new quantization error) and defeats the point of the storage-only representation.

Where this bites in practice

transformers's MoE fused-experts classes (Qwen3MoeExperts, MixtralExperts, GraniteMoeExperts, DeepSeek-V3, etc., via @use_experts_implementation) support multiple runtime forward strategies selectable via model.set_experts_implementation(...) / config._experts_implementation, independent of the checkpoint's storage layout (is_transposed):

  • "eager" — batched per-expert loop, no transpose of the weight tensor.
  • "grouped_mm" — uses torch._grouped_mm, and its _grouped_linear helper calls weight.transpose(-2, -1) before the matmul (see transformers/integrations/moe.py).
  • "batched_mm" — similar considerations may apply (not yet verified).

"grouped_mm" is auto-selected by transformers in this environment even on CPU (confirmed via Qwen3MoeForCausalLM(config).config._experts_implementation == "grouped_mm" on a freshly constructed model, no GPU involved). This means: after quantizing a K-last (is_transposed=False, already fully supported/allow-listed) MoE model with moe=True via any of Olive's three passes, simply calling model(input_ids=...) on the reloaded checkpoint crashes with:

RuntimeError: Shape-movement / view ops (transpose, reshape, view, permute, flatten, expand, …)
are not supported on an Olive QuantTensor: quantization is storage-only and these ops would
require fully dequantizing the packed weight. Access the dense weight explicitly via
``.to_dense()`` if you really need to reshape it (e.g. outside a memory-sensitive path).

Reproduced independently for both Rtn (merged, #2616) and KQuant (#2618) on a locally-constructed tiny Qwen3MoeForCausalLM (2 layers, 4 experts, no download required) — the crash and stack trace are identical for both passes, confirming this is a shared QuantTensor limitation, not a pass-specific bug.

Workaround (already works today, not documented): call model.set_experts_implementation("eager") on the loaded quantized model before running inference. Verified this produces a correct forward pass (no NaNs) for RTN's tiny-model output.

Relationship to the deferred "transposed layout" (is_transposed=True) design

This is a distinct but related gap from the already-known, deliberately-deferred scope of supporting is_transposed=True architectures (gpt-oss, llama4, aria) for quantization itself (tracked informally in prior design discussion, "Option C" — a QuantTensor metadata flag that stores canonically K-last-packed data but declares an externally-transposed shape, applying an extra transpose during _dequantize).

  • is_transposed is about the checkpoint's storage layout at rest (does Olive dare to quantize it at all).
  • This issue is about the inference-time forward-path choice (experts_implementation) crashing regardless of storage layout, because it exercises a dynamic .transpose() call against QuantTensor, not a load-time layout question.

They likely share the same underlying fix, though: if QuantTensor.__torch_function__ is generalized to treat a 2-axis transpose as a supported, metadata-tracked op (toggling an internal "transposed" flag rather than raising) as part of implementing Option C, grouped_mm's runtime transpose call would very plausibly be satisfied by the same code path. This hasn't been designed or verified — in particular, grouped_mm's _grouped_linear immediately feeds the transposed weight into torch._grouped_mm, a real matmul kernel, so at some point the data must be materialized dense regardless of whether the transpose itself is represented lazily; the design would need to account for where/how that materialization happens and confirm it doesn't reintroduce the "storage-only" guarantee's cost concerns.

Suggested follow-up scope (not committed, for discussion)

  1. Immediate (cheap): document the set_experts_implementation("eager") requirement for native PyTorch inference on any moe=True-quantized checkpoint, in docs/source/features/quantization.md, next to the existing "moe and ONNX export" section.
  2. Future (design work): when/if "Option C" (transposed-layout quantization support) is designed and implemented, explicitly evaluate whether it can also make QuantTensor.transpose(-2, -1) a supported, non-raising op — and if so, whether that's sufficient to make grouped_mm/batched_mm inference paths work without requiring users to force "eager" mode.
  3. Verify whether "batched_mm" (the third known experts_implementation) has the same issue, is unaffected, or is unsupported for another reason.

Repro

import torch
from transformers import Qwen3MoeConfig, Qwen3MoeForCausalLM

config = Qwen3MoeConfig(
    vocab_size=32, hidden_size=16, intermediate_size=16, moe_intermediate_size=8,
    num_hidden_layers=1, num_attention_heads=2, num_key_value_heads=2,
    num_experts=2, num_experts_per_tok=1, decoder_sparse_step=1, head_dim=8,
)
model = Qwen3MoeForCausalLM(config)
print(model.config._experts_implementation)  # "grouped_mm", even on CPU

# ... quantize with Rtn/KQuant/Gptq(moe=True), save, reload, then:
# model(input_ids=...)  # crashes with the QuantTensor transpose RuntimeError above
# model.set_experts_implementation("eager"); model(input_ids=...)  # works

Related: #2584, #2610, #2616, #2618.

Activity

  1. titaiwangms commented on Aug 12, 2026

    @titaiwangms
    ContributorAuthor

    Filed the related "Option C" transposed-layout quantization design as a separate issue, since it's a distinct scope (extending what can be quantized) from this issue (fixing an inference-time crash on already-supported architectures) — see #2621 for the full design write-up and discussion of where the two might share a fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions