Skip to main content

LayerNorm

LayerNorm(normalized_shape, eps=1e-05, elementwise_affine=True, bias=True, device=None, dtype=None, *, activation=None)

y = (x - mean) / sqrt(var + eps) * weight + bias with the statistics over the last dimension, of size normalized_shape.

Args: normalized_shape: the trailing size (an int, or a one-element sequence). eps: added to the variance for numerical stability. elementwise_affine: declare the weight (and bias) slots; False is not served by this layer (use nn.functional.layer_norm). bias: declare the bias slot. device: where the layer's tensors live; None declares on the cpu. dtype: the weight dtype the layer declares; None declares the default dtype (:func:~clika_runtime.get_default_dtype). A load_state_dict casts the checkpoint to the declared dtype; assign=True adopts the checkpoint's own dtype instead. activation: a fused epilogue applied to every output (a name such as "gelu" or an :class:~clika_runtime.Activation value); None for none.

__init__​

__init__(self, normalized_shape: 'int | Sequence[int]', eps: 'float' = 1e-05, elementwise_affine: 'bool' = True, bias: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, activation: 'ActivationLike' = None) -> 'None'

Initialize self. See help(type(self)) for accurate signature.

extra_repr​

extra_repr(self) -> 'str'

extra_repr() -> str

One line of per-class detail for :meth:__repr__; a layer prints its geometry here (in_features=64, out_features=256).

forward​

forward(self, input: 'Tensor') -> 'Tensor'

forward(input) -> Tensor

[*, normalized_shape] to the same shape; the output dtype follows the input.

set_weights​

set_weights(self, weight: 'Tensor', bias: 'Tensor | None' = None) -> 'None'

set_weights(weight, bias=None) -> None

Bind the declared slots directly; raises RuntimeError on a shape mismatch.