Skip to main content

Module

Base class for neural-network modules.

Subclass it, register children, parameters, and buffers by attribute assignment in __init__ (after super().__init__()), and define forward. Call the module itself (m(x)), never forward directly: the call runs the registered forward hooks around it.

Every registry walk, state-dict operation, placement change, and the repr visits the tree in registration order; a module or tensor reachable under two names appears once, under the first.

__init__​

__init__(self) -> 'None'

Initialize self. See help(type(self)) for accurate signature.

add_module​

add_module(self, name: 'str', module: "'Module | None'") -> 'None'

add_module(name, module) -> None

File module as a child under name (what assigning a :class:Module attribute does); its tensors then enumerate under the dotted prefix name..

apply​

apply(self: '_M', fn: "Callable[['Module'], None]") -> '_M'

apply(fn) -> Module

Call fn(module) on every module in the tree, children before their parent (an initializer sweep is the typical use). Returns the module.

bfloat16​

bfloat16(self: '_M') -> '_M'

bfloat16() -> Module

Cast the floating-point tensors to bfloat16.

buffers​

buffers(self, recurse: 'bool' = True) -> 'Iterator[Tensor]'

buffers(recurse=True) -> Iterator[Tensor]

The buffers of this module and, with recurse=True, of every child.

children​

children(self) -> "Iterator['Module']"

children() -> Iterator[Module]

The direct child modules, in registration order.

cpu​

cpu(self: '_M') -> '_M'

cpu() -> Module

Move the tree to the cpu; the same as to("cpu").

cuda​

cuda(self: '_M', device: 'int | DeviceLike | None' = None) -> '_M'

cuda(device=None) -> Module

Move the tree to a CUDA device: the first one by default, or the given index / device.

double​

double(self: '_M') -> '_M'

double() -> Module

Cast the floating-point tensors to float64.

eval​

eval(self: '_M') -> '_M'

eval() -> Module

Set the tree to inference mode; the same as train(False).

extra_repr​

extra_repr(self) -> 'str'

extra_repr() -> str

One line of per-class detail for :meth:__repr__; a layer prints its geometry here (in_features=64, out_features=256).

float​

float(self: '_M') -> '_M'

float() -> Module

Cast the floating-point tensors to float32.

forward​

forward(self, *args: 'Any', **kwargs: 'Any') -> 'Any'

forward(*args, **kwargs) -> Any

The computation. Subclasses define it; callers invoke the module itself (m(x)) so the forward hooks run around it.

get_buffer​

get_buffer(self, target: 'str') -> 'Tensor'

get_buffer(target) -> Tensor

The buffer at the dotted path target. Raises AttributeError when the path names nothing, or names something that is not a buffer.

get_extra_state​

get_extra_state(self) -> 'Any'

get_extra_state() -> Any

Extra state a subclass wants in its :meth:state_dict beside the tensors (a picklable value). Override together with :meth:set_extra_state; the base class has none.

get_parameter​

get_parameter(self, target: 'str') -> 'Parameter'

get_parameter(target) -> Parameter

The parameter at the dotted path target ("decoder.scale"). Raises AttributeError when the path names nothing, or names something that is not a parameter.

get_submodule​

get_submodule(self, target: 'str') -> "'Module'"

get_submodule(target) -> Module

The module at the dotted path target ("encoder.layers.3"); "" is the module itself. Raises AttributeError when a path segment names nothing, or names something that is not a module.

half​

half(self: '_M') -> '_M'

half() -> Module

Cast the floating-point tensors to float16.

load_state_dict​

load_state_dict(self, state_dict: 'Mapping[str, Tensor]', strict: 'bool' = True, assign: 'bool' = False, *, pack: 'bool' = False, weight_residency: 'WeightResidencyName | None' = None) -> '_IncompatibleKeys'

load_state_dict(state_dict, strict=True, assign=False, *, pack=False, weight_residency=None) -> _IncompatibleKeys

Bind a flat dict[str, Tensor] checkpoint onto the tree by dotted names. By default each value is copied into the existing tensor (its dtype converts on the way; shapes must match); assign=True binds the incoming tensors themselves without copying, which is how a module declared under meta init (no storage) receives its weights; the same tensor object under two keys then becomes one shared parameter object (a tied weight), exactly as the two keys share storage. That sharing is read through the checkpoint's own handle until the layer packs its weight (the first forward, or pack=True); a pack may embed the values it read, so a later in-place edit goes through the module's parameter (model.up.bias.add_(1.0)), which rebinds the slot and drops the stale pack; a write through the checkpoint handle after the pack is silent, neither refused nor seen, and the forward keeps the packed values. Under weight_residency="borrowed" the bound tensor stays the live copy, so the checkpoint handle stays a live road. pack=True packs every bound weight of the tree's layers for serving at the end of this load instead of on the first forward (a keyword this runtime adds; a checkpoint read for its footprint or a preload wants it). weight_residency sets how the layers bound by this load keep their weights at rest, one mode for the whole load: "packed" keeps one packed copy of each weight for serving and drops the raw one; "borrowed" serves from the bound tensor itself and keeps no packed copy, so the checkpoint's storage stays the one copy; "auto" packs a weight when the pack fits the free memory and otherwise serves from the bound tensor; "per_call" keeps no serving form and rebuilds what a forward needs on every call, the memory-pressure fallback; None keeps each layer's own default. Plain parameters have no residency and ignore it. strict=True raises RuntimeError listing the missing and the unexpected keys; with strict=False they are only reported in the returned _IncompatibleKeys (missing_keys, unexpected_keys). A copy that fails (a shape mismatch, a slot without storage) raises regardless of strict.

modules​

modules(self) -> "Iterator['Module']"

modules() -> Iterator[Module]

Every module in the tree, this one first, each once.

named_buffers​

named_buffers(self, prefix: 'str' = '', recurse: 'bool' = True, remove_duplicate: 'bool' = True) -> 'Iterator[tuple[str, Tensor]]'

named_buffers(prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Tensor]]

(dotted name, tensor) buffer pairs in registration order, persistent and non-persistent alike; the arguments read as in :meth:named_parameters.

named_children​

named_children(self) -> "Iterator[tuple[str, 'Module']]"

named_children() -> Iterator[tuple[str, Module]]

(name, module) pairs of the direct children, in registration order; a child registered under two names appears once.

named_modules​

named_modules(self, memo: "set['Module'] | None" = None, prefix: 'str' = '', remove_duplicate: 'bool' = True) -> "Iterator[tuple[str, 'Module']]"

named_modules(memo=None, prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Module]]

(dotted name, module) pairs over the whole tree, this module first (under prefix), children in registration order. A module reachable twice is yielded once unless remove_duplicate=False.

named_parameters​

named_parameters(self, prefix: 'str' = '', recurse: 'bool' = True, remove_duplicate: 'bool' = True) -> 'Iterator[tuple[str, Tensor]]'

named_parameters(prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Tensor]]

(dotted name, tensor) pairs in registration order, prefix prepended to every name. recurse=False stops at this module's own parameters; remove_duplicate=False also yields a tensor reachable under several names once per name.

parameters​

parameters(self, recurse: 'bool' = True) -> 'Iterator[Tensor]'

parameters(recurse=True) -> Iterator[Tensor]

The parameters of this module and, with recurse=True, of every child; a tensor shared under several names appears once.

register_buffer​

register_buffer(self, name: 'str', tensor: 'Tensor | None', persistent: 'bool' = True) -> 'None'

register_buffer(name, tensor, persistent=True) -> None

File tensor as a buffer: module state that is not a learned parameter (a mask, a precomputed table). The buffer reads back as self.<name> and moves with the module. Persistent buffers appear in :meth:state_dict; persistent=False keeps one out. A :class:~clika_runtime.nn.Buffer carries its own persistent flag, which wins. None files an empty slot.

register_forward_hook​

register_forward_hook(self, hook: 'ForwardHook', *, prepend: 'bool' = False, with_kwargs: 'bool' = False, always_call: 'bool' = False) -> 'RemovableHandle'

register_forward_hook(hook, *, prepend=False, with_kwargs=False, always_call=False) -> RemovableHandle

Run hook(module, args, output) after every forward. Returning None keeps the output; returning a value replaces it. With with_kwargs=True the hook is hook(module, args, kwargs, output). Hooks run in registration order; prepend=True puts this one first. always_call=True also runs the hook (with output=None) when the forward raises. handle.remove() unregisters it.

register_forward_pre_hook​

register_forward_pre_hook(self, hook: 'ForwardPreHook', *, prepend: 'bool' = False, with_kwargs: 'bool' = False) -> 'RemovableHandle'

register_forward_pre_hook(hook, *, prepend=False, with_kwargs=False) -> RemovableHandle

Run hook(module, args) before every forward. Returning None leaves the arguments alone; returning a value replaces them (a single value is wrapped into a one-tuple). With with_kwargs=True the hook is hook(module, args, kwargs) and a replacement is the pair (new_args, new_kwargs). Hooks run in registration order; prepend=True puts this one first. handle.remove() unregisters it.

register_module​

register_module(self, name: 'str', module: "'Module | None'") -> 'None'

register_module(name, module) -> None

The same as :meth:add_module.

register_parameter​

register_parameter(self, name: 'str', param: 'Parameter | None') -> 'None'

register_parameter(name, param) -> None

File param under name (what assigning a :class:~clika_runtime.nn.Parameter attribute does). None files an empty slot that reads back as None and takes no part in the parameter walks.

set_extra_state​

set_extra_state(self, state: 'Any') -> 'None'

set_extra_state(state) -> None

Restore what :meth:get_extra_state produced; the counterpart :meth:load_state_dict calls.

set_submodule​

set_submodule(self, target: 'str', module: "'Module'", strict: 'bool' = False) -> 'None'

set_submodule(target, module, strict=False) -> None

Replace the module at the dotted path target with module. strict=True requires a module to exist there already; with the default, a missing last segment is created.

state_dict​

state_dict(self, *, destination: 'dict[str, Any] | None' = None, prefix: 'str' = '', keep_vars: 'bool' = False) -> 'dict[str, Any]'

state_dict(*, destination=None, prefix="", keep_vars=False) -> dict[str, Tensor]

The module tree's tensors as one flat dict with dotted keys, in registration order: every parameter and every persistent buffer, prefix prepended to each key. The values are plain tensor handles sharing the parameters' storage; keep_vars=True hands out the :class:~clika_runtime.nn.Parameter / :class:~clika_runtime.nn.Buffer objects themselves. A destination dict receives the entries in place.

A layer that holds its weights in a bound layer object contributes every weight it declares, packed for serving or not, so a dict taken after a forward is complete. Loading is the round-trip contract: build the module, then :meth:load_state_dict the dict into it.

to​

to(self: '_M', *args: 'Any', **kwargs: 'Any') -> '_M'

to(device=None, dtype=None, non_blocking=False) -> Module

Move the tree to a device, cast it to a dtype, or both, in place; returns the module. Accepts to(device) (a :class:Device or a string such as "cuda:0"), to(dtype), to(device, dtype), to(tensor) (that tensor's device and dtype), or the device= / dtype= keywords; a dtype is a dtype object (clika_runtime.float16), its name ("half") or a DataType value. Only floating-point tensors cast; integer and boolean buffers keep their dtype. Parameter and Buffer objects keep their identity, so references held elsewhere stay valid. A layer that cannot serve the request (a quantized weight asked to change dtype) raises; refusals are never swallowed.

to_empty​

to_empty(self: '_M', *, device: 'DeviceLike | None', recurse: 'bool' = True) -> '_M'

to_empty(*, device, recurse=True) -> Module

Give every parameter and buffer fresh, uninitialized storage of its own shape and dtype on device (None keeps each tensor's device), without copying any values. A storage-free tensor from a meta-init region re-declares on the target and stays storage-free (zero bytes) until load_state_dict(..., assign=True) binds its value; every other tensor gets uninitialized storage there. recurse=False stops at this module's own tensors.

train​

train(self: '_M', mode: 'bool' = True) -> '_M'

train(mode=True) -> Module

Set the tree's training flag (a module may branch on it, as dropout does). Returns the module.