Module
Base class for neural-network modules.
Subclass it, register children, parameters, and buffers by attribute
assignment in __init__ (after super().__init__()), and define
forward. Call the module itself (m(x)), never forward
directly: the call runs the registered forward hooks around it.
Every registry walk, state-dict operation, placement change, and the
repr visits the tree in registration order; a module or tensor
reachable under two names appears once, under the first.
__init__
__init__(self) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
add_module
add_module(self, name: 'str', module: "'Module | None'") -> 'None'
add_module(name, module) -> None
File module as a child under name (what assigning a
:class:Module attribute does); its tensors then enumerate under
the dotted prefix name..
apply
apply(self: '_M', fn: "Callable[['Module'], None]") -> '_M'
apply(fn) -> Module
Call fn(module) on every module in the tree, children before
their parent (an initializer sweep is the typical use). Returns
the module.
bfloat16
bfloat16(self: '_M') -> '_M'
bfloat16() -> Module
Cast the floating-point tensors to bfloat16.
buffers
buffers(self, recurse: 'bool' = True) -> 'Iterator[Tensor]'
buffers(recurse=True) -> Iterator[Tensor]
The buffers of this module and, with recurse=True, of every
child.
children
children(self) -> "Iterator['Module']"
children() -> Iterator[Module]
The direct child modules, in registration order.
cpu
cpu(self: '_M') -> '_M'
cpu() -> Module
Move the tree to the cpu; the same as to("cpu").
cuda
cuda(self: '_M', device: 'int | DeviceLike | None' = None) -> '_M'
cuda(device=None) -> Module
Move the tree to a CUDA device: the first one by default, or the given index / device.
double
double(self: '_M') -> '_M'
double() -> Module
Cast the floating-point tensors to float64.
eval
eval(self: '_M') -> '_M'
eval() -> Module
Set the tree to inference mode; the same as train(False).
extra_repr
extra_repr(self) -> 'str'
extra_repr() -> str
One line of per-class detail for :meth:__repr__; a layer prints
its geometry here (in_features=64, out_features=256).
float
float(self: '_M') -> '_M'
float() -> Module
Cast the floating-point tensors to float32.
forward
forward(self, *args: 'Any', **kwargs: 'Any') -> 'Any'
forward(*args, **kwargs) -> Any
The computation. Subclasses define it; callers invoke the module
itself (m(x)) so the forward hooks run around it.
get_buffer
get_buffer(self, target: 'str') -> 'Tensor'
get_buffer(target) -> Tensor
The buffer at the dotted path target. Raises AttributeError
when the path names nothing, or names something that is not a
buffer.
get_extra_state
get_extra_state(self) -> 'Any'
get_extra_state() -> Any
Extra state a subclass wants in its :meth:state_dict beside the
tensors (a picklable value). Override together with
:meth:set_extra_state; the base class has none.
get_parameter
get_parameter(self, target: 'str') -> 'Parameter'
get_parameter(target) -> Parameter
The parameter at the dotted path target ("decoder.scale").
Raises AttributeError when the path names nothing, or names
something that is not a parameter.
get_submodule
get_submodule(self, target: 'str') -> "'Module'"
get_submodule(target) -> Module
The module at the dotted path target ("encoder.layers.3");
"" is the module itself. Raises AttributeError when a path
segment names nothing, or names something that is not a module.
half
half(self: '_M') -> '_M'
half() -> Module
Cast the floating-point tensors to float16.
load_state_dict
load_state_dict(self, state_dict: 'Mapping[str, Tensor]', strict: 'bool' = True, assign: 'bool' = False, *, pack: 'bool' = False, weight_residency: 'WeightResidencyName | None' = None) -> '_IncompatibleKeys'
load_state_dict(state_dict, strict=True, assign=False, *, pack=False, weight_residency=None) -> _IncompatibleKeys
Bind a flat dict[str, Tensor] checkpoint onto the tree by
dotted names. By default each value is copied into the existing
tensor (its dtype converts on the way; shapes must match);
assign=True binds the incoming tensors themselves without
copying, which is how a module declared under meta init (no
storage) receives its weights; the same tensor object under two
keys then becomes one shared parameter object (a tied weight),
exactly as the two keys share storage. That sharing is read through
the checkpoint's own handle until the layer packs its weight (the
first forward, or pack=True); a pack may embed the values it
read, so a later in-place edit goes through the module's parameter
(model.up.bias.add_(1.0)), which rebinds the slot and drops the
stale pack; a write through the checkpoint handle after the pack is
silent, neither refused nor seen, and the forward keeps the packed
values. Under weight_residency="borrowed" the bound tensor stays
the live copy, so the checkpoint handle stays a live road.
pack=True packs every
bound weight of the tree's layers for serving at the end of this
load instead of on the first forward (a keyword this runtime adds;
a checkpoint read for its footprint or a preload wants it).
weight_residency sets how the layers bound by this load keep
their weights at rest, one mode for the whole load: "packed"
keeps one packed copy of each weight for serving and drops the raw
one; "borrowed" serves from the bound tensor itself and keeps no
packed copy, so the checkpoint's storage stays the one copy;
"auto" packs a weight when the pack fits the free memory and
otherwise serves from the bound tensor; "per_call" keeps no
serving form and rebuilds what a forward needs on every call, the
memory-pressure fallback; None keeps each layer's own default.
Plain parameters have no residency and ignore it.
strict=True raises RuntimeError
listing the missing and the unexpected keys; with strict=False
they are only reported in the returned _IncompatibleKeys
(missing_keys, unexpected_keys). A copy that fails (a shape
mismatch, a slot without storage) raises regardless of strict.
modules
modules(self) -> "Iterator['Module']"
modules() -> Iterator[Module]
Every module in the tree, this one first, each once.
named_buffers
named_buffers(self, prefix: 'str' = '', recurse: 'bool' = True, remove_duplicate: 'bool' = True) -> 'Iterator[tuple[str, Tensor]]'
named_buffers(prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Tensor]]
(dotted name, tensor) buffer pairs in registration order,
persistent and non-persistent alike; the arguments read as in
:meth:named_parameters.
named_children
named_children(self) -> "Iterator[tuple[str, 'Module']]"
named_children() -> Iterator[tuple[str, Module]]
(name, module) pairs of the direct children, in registration
order; a child registered under two names appears once.
named_modules
named_modules(self, memo: "set['Module'] | None" = None, prefix: 'str' = '', remove_duplicate: 'bool' = True) -> "Iterator[tuple[str, 'Module']]"
named_modules(memo=None, prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Module]]
(dotted name, module) pairs over the whole tree, this module
first (under prefix), children in registration order. A module
reachable twice is yielded once unless remove_duplicate=False.
named_parameters
named_parameters(self, prefix: 'str' = '', recurse: 'bool' = True, remove_duplicate: 'bool' = True) -> 'Iterator[tuple[str, Tensor]]'
named_parameters(prefix="", recurse=True, remove_duplicate=True) -> Iterator[tuple[str, Tensor]]
(dotted name, tensor) pairs in registration order, prefix
prepended to every name. recurse=False stops at this module's
own parameters; remove_duplicate=False also yields a tensor
reachable under several names once per name.
parameters
parameters(self, recurse: 'bool' = True) -> 'Iterator[Tensor]'
parameters(recurse=True) -> Iterator[Tensor]
The parameters of this module and, with recurse=True, of every
child; a tensor shared under several names appears once.
register_buffer
register_buffer(self, name: 'str', tensor: 'Tensor | None', persistent: 'bool' = True) -> 'None'
register_buffer(name, tensor, persistent=True) -> None
File tensor as a buffer: module state that is not a learned
parameter (a mask, a precomputed table). The buffer reads back as
self.<name> and moves with the module. Persistent buffers
appear in :meth:state_dict; persistent=False keeps one out.
A :class:~clika_runtime.nn.Buffer carries its own persistent
flag, which wins. None files an empty slot.
register_forward_hook
register_forward_hook(self, hook: 'ForwardHook', *, prepend: 'bool' = False, with_kwargs: 'bool' = False, always_call: 'bool' = False) -> 'RemovableHandle'
register_forward_hook(hook, *, prepend=False, with_kwargs=False, always_call=False) -> RemovableHandle
Run hook(module, args, output) after every forward. Returning
None keeps the output; returning a value replaces it. With
with_kwargs=True the hook is hook(module, args, kwargs, output). Hooks run in registration order; prepend=True puts
this one first. always_call=True also runs the hook (with
output=None) when the forward raises. handle.remove()
unregisters it.
register_forward_pre_hook
register_forward_pre_hook(self, hook: 'ForwardPreHook', *, prepend: 'bool' = False, with_kwargs: 'bool' = False) -> 'RemovableHandle'
register_forward_pre_hook(hook, *, prepend=False, with_kwargs=False) -> RemovableHandle
Run hook(module, args) before every forward. Returning None
leaves the arguments alone; returning a value replaces them (a
single value is wrapped into a one-tuple). With with_kwargs=True
the hook is hook(module, args, kwargs) and a replacement is the
pair (new_args, new_kwargs). Hooks run in registration order;
prepend=True puts this one first. handle.remove()
unregisters it.
register_module
register_module(self, name: 'str', module: "'Module | None'") -> 'None'
register_module(name, module) -> None
The same as :meth:add_module.
register_parameter
register_parameter(self, name: 'str', param: 'Parameter | None') -> 'None'
register_parameter(name, param) -> None
File param under name (what assigning a
:class:~clika_runtime.nn.Parameter attribute does). None
files an empty slot that reads back as None and takes no part
in the parameter walks.
set_extra_state
set_extra_state(self, state: 'Any') -> 'None'
set_extra_state(state) -> None
Restore what :meth:get_extra_state produced; the counterpart
:meth:load_state_dict calls.
set_submodule
set_submodule(self, target: 'str', module: "'Module'", strict: 'bool' = False) -> 'None'
set_submodule(target, module, strict=False) -> None
Replace the module at the dotted path target with module.
strict=True requires a module to exist there already; with the
default, a missing last segment is created.
state_dict
state_dict(self, *, destination: 'dict[str, Any] | None' = None, prefix: 'str' = '', keep_vars: 'bool' = False) -> 'dict[str, Any]'
state_dict(*, destination=None, prefix="", keep_vars=False) -> dict[str, Tensor]
The module tree's tensors as one flat dict with dotted keys, in
registration order: every parameter and every persistent buffer,
prefix prepended to each key. The values are plain tensor
handles sharing the parameters' storage; keep_vars=True hands
out the :class:~clika_runtime.nn.Parameter /
:class:~clika_runtime.nn.Buffer objects themselves. A
destination dict receives the entries in place.
A layer that holds its weights in a bound layer object contributes
every weight it declares, packed for serving or not, so a dict taken
after a forward is complete. Loading is the round-trip contract:
build the module, then :meth:load_state_dict the dict into it.
to
to(self: '_M', *args: 'Any', **kwargs: 'Any') -> '_M'
to(device=None, dtype=None, non_blocking=False) -> Module
Move the tree to a device, cast it to a dtype, or both, in place;
returns the module. Accepts to(device) (a :class:Device or a
string such as "cuda:0"), to(dtype), to(device, dtype),
to(tensor) (that tensor's device and dtype), or the device=
/ dtype= keywords; a dtype is a dtype object
(clika_runtime.float16), its name ("half") or a DataType
value. Only floating-point tensors cast; integer
and boolean buffers keep their dtype. Parameter and Buffer objects
keep their identity, so references held elsewhere stay valid. A
layer that cannot serve the request (a quantized weight asked to
change dtype) raises; refusals are never swallowed.
to_empty
to_empty(self: '_M', *, device: 'DeviceLike | None', recurse: 'bool' = True) -> '_M'
to_empty(*, device, recurse=True) -> Module
Give every parameter and buffer fresh, uninitialized storage of its
own shape and dtype on device (None keeps each tensor's
device), without copying any values. A storage-free tensor from a
meta-init region re-declares on the target and stays storage-free
(zero bytes) until load_state_dict(..., assign=True) binds its
value; every other tensor gets uninitialized storage there.
recurse=False stops at this module's own tensors.
train
train(self: '_M', mode: 'bool' = True) -> '_M'
train(mode=True) -> Module
Set the tree's training flag (a module may branch on it, as
dropout does). Returns the module.