MoE
Dense mixture-of-experts: a router picks top-k experts per token, each runs its gated FFN, and the outputs combine under the router weights.
__init__
__init__(self, num_experts: 'int', hidden_size: 'int', intermediate_size: 'int', top_k: 'int', *, gate_up_bias: 'bool' = False, down_bias: 'bool' = False, dtype: '_dtype.dtype | DataType | None' = None, device: 'Device | str | None' = None, **options: 'Any') -> 'None'
MoE(num_experts, hidden_size, intermediate_size, top_k, *, gate_up_bias=False, down_bias=False, dtype=None, device=None, **options)
Declare the expert stacks. options forwards the routing kwargs
(routing_mode, renormalize, activation,
swiglu_fusion, gelu_mode, n_group, topk_group,
routed_scaling_factor, sparse_mixer_eps,
apply_router_weight_on_input, swiglu_alpha,
swiglu_beta, swiglu_limit, expert_output_scale); a
mode takes a name ("silu", "softmax_topk") or the bound
enum value; unset kwargs keep the runtime defaults. dtype is
the dtype the expert stacks declare; None declares the default
dtype (:func:~clika_runtime.get_default_dtype). A
load_state_dict casts the checkpoint to it; assign=True
adopts the checkpoint's own dtype instead.
extra_repr
extra_repr(self) -> 'str'
extra_repr() -> str
One line of per-class detail for :meth:__repr__; a layer prints
its geometry here (in_features=64, out_features=256).
forward
forward(self, x: 'Tensor', router_logits: 'Tensor', shared_output: 'Tensor | None' = None) -> 'Tensor'
forward(*args, **kwargs) -> Any
The computation. Subclasses define it; callers invoke the module
itself (m(x)) so the forward hooks run around it.
set_weights
set_weights(self, gate_up_experts: 'Tensor', down_experts: 'Tensor', *, gate_up_bias: 'Tensor | None' = None, down_bias: 'Tensor | None' = None, gate_experts: 'Tensor | None' = None, gate_bias: 'Tensor | None' = None, e_score_correction_bias: 'Tensor | None' = None) -> 'None'
set_weights(gate_up_experts, down_experts, *, gate_up_bias=None, down_bias=None, gate_experts=None, gate_bias=None, e_score_correction_bias=None) -> None
Bind the declared expert stacks (and the declared biases; the
[E] selection bias binds even when not declared).