LayerNorm
LayerNorm(normalized_shape, eps=1e-05, elementwise_affine=True, bias=True, device=None, dtype=None, *, activation=None)
y = (x - mean) / sqrt(var + eps) * weight + bias with the statistics
over the last dimension, of size normalized_shape.
Args:
normalized_shape: the trailing size (an int, or a one-element
sequence).
eps: added to the variance for numerical stability.
elementwise_affine: declare the weight (and bias) slots;
False is not served by this layer (use
nn.functional.layer_norm).
bias: declare the bias slot.
device: where the layer's tensors live; None declares on the
cpu.
dtype: the weight dtype the layer declares; None declares the
default dtype (:func:~clika_runtime.get_default_dtype). A
load_state_dict casts the checkpoint to the declared dtype;
assign=True adopts the checkpoint's own dtype instead.
activation: a fused epilogue applied to every output (a name such
as "gelu" or an :class:~clika_runtime.Activation
value); None for none.
__init__
__init__(self, normalized_shape: 'int | Sequence[int]', eps: 'float' = 1e-05, elementwise_affine: 'bool' = True, bias: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, activation: 'ActivationLike' = None) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
extra_repr
extra_repr(self) -> 'str'
extra_repr() -> str
One line of per-class detail for :meth:__repr__; a layer prints
its geometry here (in_features=64, out_features=256).
forward
forward(self, input: 'Tensor') -> 'Tensor'
forward(input) -> Tensor
[*, normalized_shape] to the same shape; the output dtype follows
the input.
set_weights
set_weights(self, weight: 'Tensor', bias: 'Tensor | None' = None) -> 'None'
set_weights(weight, bias=None) -> None
Bind the declared slots directly; raises RuntimeError on a shape mismatch.