---
title: "moe"
sidebar_label: "moe"
description: "Kotlin binding reference: moe."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Ops](index.md)/[moe](moe.md)

# moe

[common]\
fun [moe](moe.md)(input: [Tensor](../Tensor/index.md), routerLogits: [Tensor](../Tensor/index.md), fc1Experts: [Tensor](../Tensor/index.md), fc2Experts: [Tensor](../Tensor/index.md), topK: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html), fc1Bias: [Tensor](../Tensor/index.md)? = null, fc2Bias: [Tensor](../Tensor/index.md)? = null, fc3Experts: [Tensor](../Tensor/index.md)? = null, fc3Bias: [Tensor](../Tensor/index.md)? = null, eScoreCorrectionBias: [Tensor](../Tensor/index.md)? = null, routerWeights: [Tensor](../Tensor/index.md)? = null, routingMode: [MoeRouting](../MoeRouting/index.md)? = null, renormalize: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html)? = null, nGroup: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)? = null, topkGroup: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)? = null, routedScalingFactor: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, sparseMixerEps: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, applyRouterWeightOnInput: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html)? = null, activation: [Activation](../Activation/index.md)? = null, swigluFusion: [SwigluFusion](../SwigluFusion/index.md)? = null, swigluAlpha: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, swigluBeta: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, swigluLimit: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, geluMode: [GeluMode](../GeluMode/index.md)? = null, sharedOutput: [Tensor](../Tensor/index.md)? = null): [Tensor](../Tensor/index.md)

`moe(input: Tensor, routerLogits: Tensor, fc1Experts: Tensor, fc2Experts: Tensor, topK: Long, fc1Bias: Tensor? = null, fc2Bias: Tensor? = null, fc3Experts: Tensor? = null, fc3Bias: Tensor? = null, eScoreCorrectionBias: Tensor? = null, routerWeights: Tensor? = null, routingMode: MoeRouting? = null, renormalize: Boolean? = null, nGroup: Long? = null, topkGroup: Long? = null, routedScalingFactor: Double? = null, sparseMixerEps: Double? = null, applyRouterWeightOnInput: Boolean? = null, activation: Activation? = null, swigluFusion: SwigluFusion? = null, swigluAlpha: Double? = null, swigluBeta: Double? = null, swigluLimit: Double? = null, geluMode: GeluMode? = null, sharedOutput: Tensor? = null)`: the `moe` operator. Fused Mixture-of-Experts layer: route, run the top-k experts, combine in one call, with no per-expert dispatch from the caller. Per token, `router_logits [T, E]` select `top_k` experts under `routing_mode`; each selected expert applies its own MLP (`fc1 [E, F·I, H]` → activation → `fc2 [E, H, I]`, with `F` = 2 for a gated activation, else 1, and an optional multiplicative `fc3 [E, I, H]` branch); the expert outputs combine under the routing weights. `input` is `[T, H]`; the result is `[T, H]` at `input`'s dtype. Per-expert biases ride `fc1_bias` / `fc2_bias` / `fc3_bias`; `e_score_correction_bias` and `n_group` / `topk_group` / `routed_scaling_factor` serve the group-limited routing families; `router_weights` feeds `MoeRouting::PreComputed` (caller-supplied combine weights); `sparse_mixer_eps` tunes `MoeRouting::SparseMixer`. The gated-activation scalars (`swiglu_*`, `gelu_mode`) carry their `swiglu` / `geglu` meanings; `shared_output [T, H]` folds a shared-expert branch into the final combine. A config argument left `std::nullopt` takes the runtime default (SoftmaxTopK routing, renormalized top-k weights).