ClikaRT::nn::Linear
class
Header: ClikaRT/nn/linear.h
Inherits: ClikaRT::nn::Module
Inherited by: ClikaRT::nn::QLinear, ClikaRT::nn::QLinearWoQ
Affine transformation of the trailing dimension: .
The module owns its weights: construct with make(...), or declare shapes and bind a checkpoint via load_state_dict. The first forward packs the weight into the backend's kernel layout (or LoadOptions::pack_on_load packs at the end of the load); the pack is the one resident copy, and named_parameters() / state_dict() read the weight back from it. Whether a dense weight packs at all follows the device's free-memory budget (kAutoPackMaxFreeFraction): a weight whose copy would not fit is served off its bound storage instead, logged either way; LoadOptions::weight_residency pins the posture for a load. to(dtype) restores, casts, and re-packs on the next forward.
Shape: input [*, in_features] (any leading batch dims); weight [out_features, in_features]; bias [out_features], optional; output [*, out_features].
auto lin = ClikaRT::nn::Linear::make(W, b);
auto y = lin->forward(x); // [batch, in] -> [batch, out]
Static member functions
make(Tensor, OptionalTensor, LinearOptions)
static std::shared_ptr<Linear> make(
Tensor weight,
OptionalTensor bias = {},
LinearOptions options = {}
)
Construct FROM tensors (a static factory, so a bad config is a clean error), THE posture dispatch: geometry, dtype and device are all read off weight, the slots are declared AND bound in this one call (an optional bias [out] binds beside it; a bias slot exists exactly when bias is passed; options.bias is not consulted here), and the returned object's dynamic type follows the file doc's dispatch rule: options.output_quant ⇒ QLinear; weight.is_quantized() ⇒ QLinearWoQ (its geometry read off the payload by the quantized orientation law in the file doc); else Linear (the HF [out, in] dense layout). options.device overrides the placement (the payloads move there); options.dtype pins a dtype (dense payloads cast to it; a quantized payload's codes are never cast). The first forward packs. Raises ClikaRT::Error on an undefined / non-2-D weight.
Declared in ClikaRT/nn/linear.h, line 195
make(int64_t, int64_t, LinearOptions)
static std::shared_ptr<Linear> make(
std::int64_t in_features,
std::int64_t out_features,
LinearOptions options = {}
)
Construct FROM tensors (a static factory, so a bad config is a clean error), THE posture dispatch: geometry, dtype and device are all read off weight, the slots are declared AND bound in this one call (an optional bias [out] binds beside it; a bias slot exists exactly when bias is passed; options.bias is not consulted here), and the returned object's dynamic type follows the file doc's dispatch rule: options.output_quant ⇒ QLinear; weight.is_quantized() ⇒ QLinearWoQ (its geometry read off the payload by the quantized orientation law in the file doc); else Linear (the HF [out, in] dense layout). options.device overrides the placement (the payloads move there); options.dtype pins a dtype (dense payloads cast to it; a quantized payload's codes are never cast). The first forward packs. Raises ClikaRT::Error on an undefined / non-2-D weight.
Declared in ClikaRT/nn/linear.h, line 215
Member functions
~Linear()
~Linear() override
Releases the module's packed weights.
Declared in ClikaRT/nn/linear.h, line 222
kind()
virtual LinearKind kind() const noexcept
The module's SERVING POSTURE (see LinearKind; posture, not class identity). A plain Linear reports Dense until a quantized payload binds, then WeightQuantized; QLinearWoQ / QLinear report their pinned posture constantly.
Declared in ClikaRT/nn/linear.h, line 228
share_weight()
void share_weight(Tensor weight)
Adopt weight as this module's weight WITHOUT a copy: two handles, one storage (the tie: e.g. an lm head serving off the embedding table). Total and idempotent: a module whose weight already landed (a checkpoint bind, a prior share, a live pack) ignores the call and returns ok. A shared weight packs only when its copy fits the device's free-memory budget; otherwise the shared storage stays the ONE resident copy and forward serves off it (LoadOptions::weight_residency pins the posture for a load); a shared parameter enumerates once, under its first registered name. Raises ClikaRT::Error on a cross-device share or a pinned dtype the table does not already carry (either would force the copy the tie exists to avoid). NOTE: a placement move (to(device)) re-homes each registry slot individually, so a tied pair moved across devices unties into two copies; re-tie with share_weight after a cross-device move. A dtype move on a shared weight is refused for the same reason (cast the SOURCE instead).
Declared in ClikaRT/nn/linear.h, line 246
set_weights()
void set_weights(Tensor weight, OptionalTensor bias = {})
Bind the declared slots positionally: weight, dense in EITHER orientation (the HF [out, in] or the transposed [in, out], validated against the declared feature counts; the ambiguous square case is read as [out, in]), or a QUANTIZED payload (its scheme rides the tensor and states its orientation, the file doc's law, validated against the declared counts; the serving posture follows it), and, when the module was declared with one, bias ([out]). The declared placement wins (the tensors move to it); an unpinned slot adopts the payload's dtype, a pinned one casts a DENSE payload to the pin (quantized codes never cast). Re-binding after a pack drops the pack; the next forward re-packs from the new weights. Raises ClikaRT::Error on a geometry mismatch or a bias without a declared bias slot.
Declared in ClikaRT/nn/linear.h, line 264
to_impl(StreamOrDevice)
virtual Result<void> to_impl(StreamOrDevice where) override
Move to a placement (a Device, or a Stream, the lane this Linear computes on and lands its outputs on) or cast to a dtype. Before the pack this moves the registry slots (a still-fake slot re-declares its metadata there). Once packed, a placement move REBUILDS the pack on the target; a dtype cast restores the raw weight from the pack, casts, and re-packs on the next forward, never a stale packed form computing on the old placement. A dtype cast PINS that dtype for every later re-bind, whether it runs before or after the first bind (an explicit cast is a dtype choice; a subsequent payload must not silently revert it). On a quantized posture (kind() != Dense) a dtype cast is Unsupported: the weight stays quantized at rest; dequantize explicitly if a dense copy is wanted. can_cast_to_impl gives the same verdict without moving a byte: the tree-wide check behind Module::to(dtype) asks it before any slot is cast.
Declared in ClikaRT/nn/linear.h, line 282
to_impl(DataType)
Declared in ClikaRT/nn/linear.h, line 283
can_cast_to_impl()
Whether THIS module's own slots can be cast to dtype (children are asked by the tree walk, never here). A failure is the refusal to(dtype) would raise; Module::to(dtype) asks every module of the tree BEFORE it casts any, so a refused cast leaves the tree untouched and names the leaf. Default: a dense module casts. Override in a leaf that pins its dtype.
Declared in ClikaRT/nn/linear.h, line 284
forward()
Apply the bound weight: x [*, in] -> [*, out] (then bias + activation if configured). The first call packs the weight into the backend kernel layout (once, thread-safe); every later call reuses the pack: the dense matmul or the weight-only-quantized matmul, per the bound payload.
Throws
ClikaRT::Error: on a shape/dtype mismatch or a still-fake slot.
Declared in ClikaRT/nn/linear.h, line 293
Protected member functions
apply_weight_residency()
virtual void apply_weight_residency(
WeightResidency residency,
std::uint64_t& stamped,
std::uint64_t& class_pinned
) override
The load-scope residency stamp (nn::LoadOptions::weight_residency). Adopts the override as this module's requested residency, exactly as if it had been constructed with it, unless the class pins its residency (the quantized postures serve Packed alone), in which case the pin holds and the module counts in class_pinned. Applies to packs that have not yet run.
Declared in ClikaRT/nn/linear.h, line 302
initialize_impl()
virtual Result<void> initialize_impl() override
The pack hook (Module::initialize_impl): packs the bound weight into the backend kernel layout now (idempotent, thread-safe). Every declared slot must hold a real (loaded) tensor; a still-fake slot is a clean error. After the pack the weight slot reads back from the pack.
Declared in ClikaRT/nn/linear.h, line 309
Linear()
CLIKART_LOCAL Linear()
Declared in ClikaRT/nn/linear.h, line 311
Linear(void)
explicit CLIKART_LOCAL Linear(void* impl) noexcept
Subclass construction: adopts a derived detail::LinearImpl as the one pImpl. The parameter is void*, NOT the typed impl pointer, on purpose: a protected member of an exported class IS exported on PE (dllexport exports every member), and an exported signature must not name a detail:: type; do not "fix" this to the typed pointer.
Declared in ClikaRT/nn/linear.h, line 317
Protected data members
impl_
std::unique_ptr<detail::LinearImpl> impl_
Declared in ClikaRT/nn/linear.h, line 318