RMSNorm
RMSNorm(normalized_shape, eps=None, elementwise_affine=True, device=None, dtype=None, *, bias=False, activation=None)
y = x / sqrt(mean(x * x) + eps) * weight over the last dimension, of
size normalized_shape.
Args:
normalized_shape: the trailing size (an int, or a one-element
sequence).
eps: added to the mean square for numerical stability; None
means 1e-6.
elementwise_affine: declare the weight slot; False is not
served by this layer (use nn.functional.rms_norm).
device: where the layer's tensors live; None declares on the
cpu.
dtype: the weight dtype the layer declares; None declares the
default dtype (:func:~clika_runtime.get_default_dtype). A
load_state_dict casts the checkpoint to the declared dtype;
assign=True adopts the checkpoint's own dtype instead.
bias: also declare a bias slot added after the scale.
activation: a fused epilogue applied to every output (a name such
as "gelu" or an :class:~clika_runtime.Activation
value); None for none.
__init__
__init__(self, normalized_shape: 'int | Sequence[int]', eps: 'float | None' = None, elementwise_affine: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, bias: 'bool' = False, activation: 'ActivationLike' = None) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
extra_repr
extra_repr(self) -> 'str'
extra_repr() -> str
One line of per-class detail for :meth:__repr__; a layer prints
its geometry here (in_features=64, out_features=256).
forward
forward(self, input: 'Tensor') -> 'Tensor'
forward(input) -> Tensor
[*, normalized_shape] to the same shape; the output dtype follows
the input.
set_weights
set_weights(self, weight: 'Tensor', bias: 'Tensor | None' = None) -> 'None'
set_weights(weight, bias=None) -> None
Bind the declared slots directly; raises RuntimeError on a shape mismatch.