Skip to main content

//clika-runtime/io.clika.runtime/Ops/addRmsNorm

addRmsNorm

[common]
fun addRmsNorm(input: Tensor, residual: Tensor? = null, residual2: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): List<Tensor>

addRmsNorm(input: Tensor, residual: Tensor? = null, residual2: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): the add_rms_norm operator. Fused residual-add + normalization, on either side of the norm. Which data flow runs is inferred from which addends you supply; there is no mode flag: * residual only → {ACT(norm(x + residual)·w + b), x + residual} * post_residual only → {ACT(norm(x)·w + b) + post_residual, UNDEFINED} * both → {ACT(norm(x + residual)·w + b) + post_residual, x + residual} * neither → raises (that is a plain rms_norm/layer_norm) The second result is the PRE-norm sum, the residual stream the next sub-layer reads. Supplying only post_residual forms no such sum, so that element comes back UNDEFINED: read only the first one in that flow. The activation applies to the norm result, beforepost_residual is added; it is the norm's epilogue, not the sum's. Gated activations (SwiGlu / GeGlu / ReGlu) narrow their input and are rejected.