Skip to main content

QMoEWoQ

Weight-only-quantized mixture-of-experts: the stacked expert weights bind as :class:QTensor and rest quantized; the activation stays floating point.

__init__​

__init__(self, num_experts: 'int', hidden_size: 'int', intermediate_size: 'int', top_k: 'int', *, gate_up_bias: 'bool' = False, down_bias: 'bool' = False, dtype: '_dtype.dtype | DataType | None' = None, device: 'Device | str | None' = None, **options: 'Any') -> 'None'

QMoEWoQ(num_experts, hidden_size, intermediate_size, top_k, *, gate_up_bias=False, down_bias=False, dtype=None, device=None, **options)

Declare the expert stacks; the same counts, flags, and routing kwargs as :class:MoE (dtype declares the LOGICAL element type; None declares the default dtype).

extra_repr​

extra_repr(self) -> 'str'

extra_repr() -> str

One line of per-class detail for :meth:__repr__; a layer prints its geometry here (in_features=64, out_features=256).

forward​

forward(self, x: 'Tensor', router_logits: 'Tensor', shared_output: 'Tensor | None' = None) -> 'Tensor'

forward(*args, **kwargs) -> Any

The computation. Subclasses define it; callers invoke the module itself (m(x)) so the forward hooks run around it.

set_weights​

set_weights(self, gate_up_experts: 'QTensor', down_experts: 'QTensor', *, gate_up_bias: 'Tensor | None' = None, down_bias: 'Tensor | None' = None, gate_experts: 'QTensor | None' = None, gate_bias: 'Tensor | None' = None, e_score_correction_bias: 'Tensor | None' = None) -> 'None'

set_weights(gate_up_experts, down_experts, *, gate_up_bias=None, down_bias=None, gate_experts=None, gate_bias=None, e_score_correction_bias=None) -> None

Bind the quantized expert stacks (schemes travel with them) and the declared floating-point biases.