Skip to main content

//clika-runtime/io.clika.runtime/Ops/moe

moe

[common]
fun moe(input: Tensor, routerLogits: Tensor, fc1Experts: Tensor, fc2Experts: Tensor, topK: Long, fc1Bias: Tensor? = null, fc2Bias: Tensor? = null, fc3Experts: Tensor? = null, fc3Bias: Tensor? = null, eScoreCorrectionBias: Tensor? = null, routerWeights: Tensor? = null, routingMode: MoeRouting? = null, renormalize: Boolean? = null, nGroup: Long? = null, topkGroup: Long? = null, routedScalingFactor: Double? = null, sparseMixerEps: Double? = null, applyRouterWeightOnInput: Boolean? = null, activation: Activation? = null, swigluFusion: SwigluFusion? = null, swigluAlpha: Double? = null, swigluBeta: Double? = null, swigluLimit: Double? = null, geluMode: GeluMode? = null, sharedOutput: Tensor? = null): Tensor

moe(input: Tensor, routerLogits: Tensor, fc1Experts: Tensor, fc2Experts: Tensor, topK: Long, fc1Bias: Tensor? = null, fc2Bias: Tensor? = null, fc3Experts: Tensor? = null, fc3Bias: Tensor? = null, eScoreCorrectionBias: Tensor? = null, routerWeights: Tensor? = null, routingMode: MoeRouting? = null, renormalize: Boolean? = null, nGroup: Long? = null, topkGroup: Long? = null, routedScalingFactor: Double? = null, sparseMixerEps: Double? = null, applyRouterWeightOnInput: Boolean? = null, activation: Activation? = null, swigluFusion: SwigluFusion? = null, swigluAlpha: Double? = null, swigluBeta: Double? = null, swigluLimit: Double? = null, geluMode: GeluMode? = null, sharedOutput: Tensor? = null): the moe operator. Fused Mixture-of-Experts layer: route, run the top-k experts, combine in one call, with no per-expert dispatch from the caller. Per token, router_logits [T, E] select top_k experts under routing_mode; each selected expert applies its own MLP (fc1 [E, F·I, H] → activation → fc2 [E, H, I], with F = 2 for a gated activation, else 1, and an optional multiplicative fc3 [E, I, H] branch); the expert outputs combine under the routing weights. input is [T, H]; the result is [T, H] at input's dtype. Per-expert biases ride fc1_bias / fc2_bias / fc3_bias; e_score_correction_bias and n_group / topk_group / routed_scaling_factor serve the group-limited routing families; router_weights feeds MoeRouting::PreComputed (caller-supplied combine weights); sparse_mixer_eps tunes MoeRouting::SparseMixer. The gated-activation scalars (swiglu_*, gelu_mode) carry their swiglu / geglu meanings; shared_output [T, H] folds a shared-expert branch into the final combine. A config argument left std::nullopt takes the runtime default (SoftmaxTopK routing, renormalized top-k weights).