Summary
Fused-expert MoE weights quantized by Rtn/Gptq/KQuant (moe=True) are stored as 3D QuantTensors that are storage-only: any shape-movement op (transpose, reshape, view, permute, expand) unconditionally raises RuntimeError in QuantTensor.__torch_function__ (_unsupported_movement, olive/common/quant/tensor.py).
This is a deliberate design choice — see the docstring on _unsupported_movement and _MOE_ONNX_EXPORT_MSG — because pack_to_uint8/unpack_from_uint8 bit-pack multiple logical values from the tensor's last dimension into shared uint8 bytes, and scales/qzeros are grouped along that same last dimension by group_size. A transpose can't be represented as a stride/metadata-only view (the way it is for a normal dense tensor) because the physical bytes interleave values that, after a transpose, would need to belong to different rows/groups. Supporting a real transpose would require unpack → dequantize → dense transpose → re-quantize, which is lossy (new quantization error) and defeats the point of the storage-only representation.
Where this bites in practice
transformers's MoE fused-experts classes (Qwen3MoeExperts, MixtralExperts, GraniteMoeExperts, DeepSeek-V3, etc., via @use_experts_implementation) support multiple runtime forward strategies selectable via model.set_experts_implementation(...) / config._experts_implementation, independent of the checkpoint's storage layout (is_transposed):
"eager" — batched per-expert loop, no transpose of the weight tensor.
"grouped_mm" — uses torch._grouped_mm, and its _grouped_linear helper calls weight.transpose(-2, -1) before the matmul (see transformers/integrations/moe.py).
"batched_mm" — similar considerations may apply (not yet verified).
"grouped_mm" is auto-selected by transformers in this environment even on CPU (confirmed via Qwen3MoeForCausalLM(config).config._experts_implementation == "grouped_mm" on a freshly constructed model, no GPU involved). This means: after quantizing a K-last (is_transposed=False, already fully supported/allow-listed) MoE model with moe=True via any of Olive's three passes, simply calling model(input_ids=...) on the reloaded checkpoint crashes with:
RuntimeError: Shape-movement / view ops (transpose, reshape, view, permute, flatten, expand, …)
are not supported on an Olive QuantTensor: quantization is storage-only and these ops would
require fully dequantizing the packed weight. Access the dense weight explicitly via
``.to_dense()`` if you really need to reshape it (e.g. outside a memory-sensitive path).
Reproduced independently for both Rtn (merged, #2616) and KQuant (#2618) on a locally-constructed tiny Qwen3MoeForCausalLM (2 layers, 4 experts, no download required) — the crash and stack trace are identical for both passes, confirming this is a shared QuantTensor limitation, not a pass-specific bug.
Workaround (already works today, not documented): call model.set_experts_implementation("eager") on the loaded quantized model before running inference. Verified this produces a correct forward pass (no NaNs) for RTN's tiny-model output.
Relationship to the deferred "transposed layout" (is_transposed=True) design
This is a distinct but related gap from the already-known, deliberately-deferred scope of supporting is_transposed=True architectures (gpt-oss, llama4, aria) for quantization itself (tracked informally in prior design discussion, "Option C" — a QuantTensor metadata flag that stores canonically K-last-packed data but declares an externally-transposed shape, applying an extra transpose during _dequantize).
is_transposed is about the checkpoint's storage layout at rest (does Olive dare to quantize it at all).
- This issue is about the inference-time forward-path choice (
experts_implementation) crashing regardless of storage layout, because it exercises a dynamic .transpose() call against QuantTensor, not a load-time layout question.
They likely share the same underlying fix, though: if QuantTensor.__torch_function__ is generalized to treat a 2-axis transpose as a supported, metadata-tracked op (toggling an internal "transposed" flag rather than raising) as part of implementing Option C, grouped_mm's runtime transpose call would very plausibly be satisfied by the same code path. This hasn't been designed or verified — in particular, grouped_mm's _grouped_linear immediately feeds the transposed weight into torch._grouped_mm, a real matmul kernel, so at some point the data must be materialized dense regardless of whether the transpose itself is represented lazily; the design would need to account for where/how that materialization happens and confirm it doesn't reintroduce the "storage-only" guarantee's cost concerns.
Suggested follow-up scope (not committed, for discussion)
- Immediate (cheap): document the
set_experts_implementation("eager") requirement for native PyTorch inference on any moe=True-quantized checkpoint, in docs/source/features/quantization.md, next to the existing "moe and ONNX export" section.
- Future (design work): when/if "Option C" (transposed-layout quantization support) is designed and implemented, explicitly evaluate whether it can also make
QuantTensor.transpose(-2, -1) a supported, non-raising op — and if so, whether that's sufficient to make grouped_mm/batched_mm inference paths work without requiring users to force "eager" mode.
- Verify whether
"batched_mm" (the third known experts_implementation) has the same issue, is unaffected, or is unsupported for another reason.
Repro
import torch
from transformers import Qwen3MoeConfig, Qwen3MoeForCausalLM
config = Qwen3MoeConfig(
vocab_size=32, hidden_size=16, intermediate_size=16, moe_intermediate_size=8,
num_hidden_layers=1, num_attention_heads=2, num_key_value_heads=2,
num_experts=2, num_experts_per_tok=1, decoder_sparse_step=1, head_dim=8,
)
model = Qwen3MoeForCausalLM(config)
print(model.config._experts_implementation) # "grouped_mm", even on CPU
# ... quantize with Rtn/KQuant/Gptq(moe=True), save, reload, then:
# model(input_ids=...) # crashes with the QuantTensor transpose RuntimeError above
# model.set_experts_implementation("eager"); model(input_ids=...) # works
Related: #2584, #2610, #2616, #2618.
Summary
Fused-expert MoE weights quantized by
Rtn/Gptq/KQuant(moe=True) are stored as 3DQuantTensors that are storage-only: any shape-movement op (transpose,reshape,view,permute,expand) unconditionally raisesRuntimeErrorinQuantTensor.__torch_function__(_unsupported_movement,olive/common/quant/tensor.py).This is a deliberate design choice — see the docstring on
_unsupported_movementand_MOE_ONNX_EXPORT_MSG— becausepack_to_uint8/unpack_from_uint8bit-pack multiple logical values from the tensor's last dimension into shareduint8bytes, andscales/qzerosare grouped along that same last dimension bygroup_size. A transpose can't be represented as a stride/metadata-only view (the way it is for a normal dense tensor) because the physical bytes interleave values that, after a transpose, would need to belong to different rows/groups. Supporting a real transpose would require unpack → dequantize → dense transpose → re-quantize, which is lossy (new quantization error) and defeats the point of the storage-only representation.Where this bites in practice
transformers's MoE fused-experts classes (Qwen3MoeExperts,MixtralExperts,GraniteMoeExperts, DeepSeek-V3, etc., via@use_experts_implementation) support multiple runtime forward strategies selectable viamodel.set_experts_implementation(...)/config._experts_implementation, independent of the checkpoint's storage layout (is_transposed):"eager"— batched per-expert loop, no transpose of the weight tensor."grouped_mm"— usestorch._grouped_mm, and its_grouped_linearhelper callsweight.transpose(-2, -1)before the matmul (seetransformers/integrations/moe.py)."batched_mm"— similar considerations may apply (not yet verified)."grouped_mm"is auto-selected bytransformersin this environment even on CPU (confirmed viaQwen3MoeForCausalLM(config).config._experts_implementation == "grouped_mm"on a freshly constructed model, no GPU involved). This means: after quantizing a K-last (is_transposed=False, already fully supported/allow-listed) MoE model withmoe=Truevia any of Olive's three passes, simply callingmodel(input_ids=...)on the reloaded checkpoint crashes with:Reproduced independently for both
Rtn(merged, #2616) andKQuant(#2618) on a locally-constructed tinyQwen3MoeForCausalLM(2 layers, 4 experts, no download required) — the crash and stack trace are identical for both passes, confirming this is a sharedQuantTensorlimitation, not a pass-specific bug.Workaround (already works today, not documented): call
model.set_experts_implementation("eager")on the loaded quantized model before running inference. Verified this produces a correct forward pass (no NaNs) for RTN's tiny-model output.Relationship to the deferred "transposed layout" (
is_transposed=True) designThis is a distinct but related gap from the already-known, deliberately-deferred scope of supporting
is_transposed=Truearchitectures (gpt-oss, llama4, aria) for quantization itself (tracked informally in prior design discussion, "Option C" — aQuantTensormetadata flag that stores canonically K-last-packed data but declares an externally-transposed shape, applying an extra transpose during_dequantize).is_transposedis about the checkpoint's storage layout at rest (does Olive dare to quantize it at all).experts_implementation) crashing regardless of storage layout, because it exercises a dynamic.transpose()call againstQuantTensor, not a load-time layout question.They likely share the same underlying fix, though: if
QuantTensor.__torch_function__is generalized to treat a 2-axistransposeas a supported, metadata-tracked op (toggling an internal "transposed" flag rather than raising) as part of implementing Option C,grouped_mm's runtime transpose call would very plausibly be satisfied by the same code path. This hasn't been designed or verified — in particular,grouped_mm's_grouped_linearimmediately feeds the transposed weight intotorch._grouped_mm, a real matmul kernel, so at some point the data must be materialized dense regardless of whether the transpose itself is represented lazily; the design would need to account for where/how that materialization happens and confirm it doesn't reintroduce the "storage-only" guarantee's cost concerns.Suggested follow-up scope (not committed, for discussion)
set_experts_implementation("eager")requirement for native PyTorch inference on anymoe=True-quantized checkpoint, indocs/source/features/quantization.md, next to the existing "moeand ONNX export" section.QuantTensor.transpose(-2, -1)a supported, non-raising op — and if so, whether that's sufficient to makegrouped_mm/batched_mminference paths work without requiring users to force"eager"mode."batched_mm"(the third knownexperts_implementation) has the same issue, is unaffected, or is unsupported for another reason.Repro
Related: #2584, #2610, #2616, #2618.