Skip to main content

RMSNorm

RMSNorm(normalized_shape, eps=None, elementwise_affine=True, device=None, dtype=None, *, bias=False, activation=None)

y = x / sqrt(mean(x * x) + eps) * weight over the last dimension, of size normalized_shape.

Args: normalized_shape: the trailing size (an int, or a one-element sequence). eps: added to the mean square for numerical stability; None means 1e-6. elementwise_affine: declare the weight slot; False is not served by this layer (use nn.functional.rms_norm). device: where the layer's tensors live; None declares on the cpu. dtype: the weight dtype the layer declares; None declares the default dtype (:func:~clika_runtime.get_default_dtype). A load_state_dict casts the checkpoint to the declared dtype; assign=True adopts the checkpoint's own dtype instead. bias: also declare a bias slot added after the scale. activation: a fused epilogue applied to every output (a name such as "gelu" or an :class:~clika_runtime.Activation value); None for none.

__init__​

__init__(self, normalized_shape: 'int | Sequence[int]', eps: 'float | None' = None, elementwise_affine: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, bias: 'bool' = False, activation: 'ActivationLike' = None) -> 'None'

Initialize self. See help(type(self)) for accurate signature.

extra_repr​

extra_repr(self) -> 'str'

extra_repr() -> str

One line of per-class detail for :meth:__repr__; a layer prints its geometry here (in_features=64, out_features=256).

forward​

forward(self, input: 'Tensor') -> 'Tensor'

forward(input) -> Tensor

[*, normalized_shape] to the same shape; the output dtype follows the input.

set_weights​

set_weights(self, weight: 'Tensor', bias: 'Tensor | None' = None) -> 'None'

set_weights(weight, bias=None) -> None

Bind the declared slots directly; raises RuntimeError on a shape mismatch.