Skip to main content

ClikaRT::nn

namespace

Namespaces​

NameDescription
ClikaRT::nn::fused

Classes​

NameDescription
AdaptiveAvgPool1dAverage pooling to a fixed 1-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_avg_pool1d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveAvgPool1dOptionsSettings of AdaptiveAvgPool1d: chain the setters, or assign the fields.
AdaptiveAvgPool2dAverage pooling to a fixed 2-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_avg_pool2d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveAvgPool2dOptionsSettings of AdaptiveAvgPool2d: chain the setters, or assign the fields.
AdaptiveAvgPool3dAverage pooling to a fixed 3-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_avg_pool3d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveAvgPool3dOptionsSettings of AdaptiveAvgPool3d: chain the setters, or assign the fields.
AdaptiveMaxPool1dMax pooling to a fixed 1-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_max_pool1d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveMaxPool1dOptionsSettings of AdaptiveMaxPool1d: chain the setters, or assign the fields.
AdaptiveMaxPool2dMax pooling to a fixed 2-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_max_pool2d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveMaxPool2dOptionsSettings of AdaptiveMaxPool2d: chain the setters, or assign the fields.
AdaptiveMaxPool3dMax pooling to a fixed 3-D output size, the window derived per output position; the input is channels-last. The module form of ClikaRT::ops::adaptive_max_pool3d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AdaptiveMaxPool3dOptionsSettings of AdaptiveMaxPool3d: chain the setters, or assign the fields.
AddAdds other, scaled by alpha, to input elementwise with broadcasting: input + alpha * other, then the optional fused activation. The module form of ClikaRT::ops::add, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AddOptionsSettings of Add: chain the setters, or assign the fields.
AttentionAttention module: fixed settings, optional bound tensors, one fused call per forward. The module form of ops::attention.
AttentionOptionsSettings of Attention: chain the setters, or assign the fields. The defaults are plain non-causal attention over head-split inputs.
AvgPool1dAverage pooling over a 1-D window; the input is channels-last [N, D1, C]. The module form of ClikaRT::ops::avg_pool1d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AvgPool1dOptionsSettings of AvgPool1d: chain the setters, or assign the fields.
AvgPool2dAverage pooling over a 2-D window; the input is channels-last [N, D1..D2, C]. The module form of ClikaRT::ops::avg_pool2d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AvgPool2dOptionsSettings of AvgPool2d: chain the setters, or assign the fields.
AvgPool3dAverage pooling over a 3-D window; the input is channels-last [N, D1..D3, C]. The module form of ClikaRT::ops::avg_pool3d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
AvgPool3dOptionsSettings of AvgPool3d: chain the setters, or assign the fields.
BatchNormBatch normalization, inference form.
BatchNormOptionsConstruction settings of BatchNorm beyond the channel count (a plain aggregate; assign the fields you need).
BCELossBinary cross entropy between probabilities input and target. The module form of ClikaRT::ops::binary_cross_entropy, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
BCELossOptionsSettings of BCELoss: chain the setters, or assign the fields.
BCEWithLogitsLossBinary cross entropy over logits: the sigmoid fused with the loss in one stable pass. The module form of ClikaRT::ops::binary_cross_entropy_with_logits, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
BCEWithLogitsLossOptionsSettings of BCEWithLogitsLoss: chain the setters, or assign the fields.
CELUContinuously differentiable exponential linear unit: max(0, input) + min(0, alpha * (exp(input / alpha) - 1)). The module form of ClikaRT::ops::celu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
CELUOptionsSettings of CELU: chain the setters, or assign the fields.
CircularPadPads by wrapping around the axis. The module form of ClikaRT::ops::circular_pad, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
CircularPadOptionsSettings of CircularPad: chain the setters, or assign the fields.
ClampClamps every element of input into [min_val, max_val]; an unset bound leaves that side open. The module form of ClikaRT::ops::clamp, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ClampOptionsSettings of Clamp: chain the setters, or assign the fields.
ConstantPadPads with a constant value. The module form of ClikaRT::ops::constant_pad, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ConstantPadOptionsSettings of ConstantPad: chain the setters, or assign the fields.
ConvN-dimensional convolution module (1-D/2-D/3-D by the weight's rank).
ConvTransposeN-dimensional transposed-convolution module (learnable upsampling).
CosineSimilarityCosine similarity of x1 and x2 along dim, with broadcasting. The module form of ClikaRT::ops::cosine_similarity, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
CosineSimilarityOptionsSettings of CosineSimilarity: chain the setters, or assign the fields.
CrossEntropyLossCross entropy over class logits with the class axis last: the log-softmax followed by the negative log-likelihood. The module form of ClikaRT::ops::cross_entropy, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
CrossEntropyLossOptionsSettings of CrossEntropyLoss: chain the setters, or assign the fields.
DivDivides input by other elementwise with broadcasting, with an optional quotient rounding and fused activation. The module form of ClikaRT::ops::div, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
DivOptionsSettings of Div: chain the setters, or assign the fields.
DropoutThe inference form of dropout: the input passes through unchanged. The runtime serves trained models, where dropout is disabled, so no element is dropped and nothing is rescaled; p is kept so the module reads and exports like the layer it stands for.
DropoutOptionsSettings of Dropout: chain the setter, or assign the field.
ELUExponential linear unit: scale * input for input > 0, scale * alpha * (exp(input_scale * input) - 1) otherwise. The module form of ClikaRT::ops::elu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ELUOptionsSettings of ELU: chain the setters, or assign the fields.
EmbeddingEmbedding-table module: gathers rows of its table by token index, the module form of ops::embedding.
FlattenMerges the inclusive dimension range [start_dim, end_dim] into one dimension; a view when the layout permits. The module form of ClikaRT::ops::flatten, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
FlattenOptionsSettings of Flatten: chain the setters, or assign the fields.
FoldSums sliding windows [N, L, prod(kernel_size), C] back into a channels- last image [N, output_size.., C]; overlapping windows add. The module form of ClikaRT::ops::fold, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
FoldOptionsSettings of Fold: chain the setters, or assign the fields.
GeGLUGeGLU over a concatenated [*, 2d] gate-then-up input, emitting [*, d]: gelu(gate) * up. The module form of ClikaRT::ops::geglu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
GeGLUOptionsSettings of GeGLU: chain the setters, or assign the fields.
GELUGaussian error linear unit: input * Phi(input), the exact erf form by default or the tanh approximation. The module form of ClikaRT::ops::gelu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
GELUOptionsSettings of GELU: chain the setters, or assign the fields.
GLUGated linear unit: splits the input in half along dim and multiplies the first half by the sigmoid of the second; the size of dim must be even and the output halves it. The module form of ClikaRT::ops::glu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
GLUOptionsSettings of GLU: chain the setters, or assign the fields.
GroupNormGroup normalization.
GroupNormOptionsConstruction settings of GroupNorm beyond the group and channel counts (a plain aggregate; assign the fields you need).
HardshrinkHard shrinkage: zeroes every element within [-lambd, lambd] and keeps the rest. The module form of ClikaRT::ops::hardshrink, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
HardshrinkOptionsSettings of Hardshrink: chain the setters, or assign the fields.
HardsigmoidPiecewise-linear sigmoid approximation: clamp(input / 6 + 1 / 2, 0, 1). The module form of ClikaRT::ops::hardsigmoid, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
HardsigmoidOptionsHardsigmoid has no settings; the struct keeps every module in this header constructible the same way.
HardswishPiecewise-linear swish approximation: input * clamp(input / 6 + 1 / 2, 0, 1). The module form of ClikaRT::ops::hardswish, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
HardswishOptionsHardswish has no settings; the struct keeps every module in this header constructible the same way.
HardtanhClamps every element to [min_val, max_val]. The module form of ClikaRT::ops::hardtanh, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
HardtanhOptionsSettings of Hardtanh: chain the setters, or assign the fields.
HookOutcomeWhat one call of a custom operator's output_shapes or compute reported back to the runtime: Status::Ok with empty text on success, otherwise the status and text of the exception the hook raised. Plain data on purpose: an exception raised in application code is caught by the application's own C++ runtime (the entry points of detail::ModuleHookTable) and only this record crosses into the library. Not something a module author fills in by hand: the entry points in detail::ModuleHookTable produce it.
HuberLossHalf the squared error below delta, a linear penalty above it, between input and target. The module form of ClikaRT::ops::huber_loss, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
HuberLossOptionsSettings of HuberLoss: chain the setters, or assign the fields.
IdentityReturns its input unchanged: it keeps a slot in a container where a module is expected (a skipped normalization, a disabled projection) without changing the data flow.
IdentityOptionsIdentity has no settings; the struct keeps every module in this header constructible the same way.
InstanceNormInstance normalization.
InstanceNormOptionsConstruction settings of InstanceNorm beyond the channel count (a plain aggregate; assign the fields you need).
KLDivLossKullback-Leibler divergence of target from input, where input holds log-probabilities. The module form of ClikaRT::ops::kl_div, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
KLDivLossOptionsSettings of KLDivLoss: chain the setters, or assign the fields.
KVCacheThe serving-side KV cache: per-layer storage rows (attention K/V, conv windows, recurrent slabs) over a slot-indexed batch.
KVCacheConfigConstruction parameters shared by every cache strategy. num_kv_heads is the KV head count (grouped-query attention: KV heads ≤ query heads); the cache stores K/V as [seq_len, num_kv_heads, head_dim] per (slot, layer).
KVLayerSpecOne decoder layer's storage row, the per-layer entry of the spec vector KVCache::make consumes. Two orthogonal axes: Kind says WHAT the row stores, Layout says HOW a token-indexed row is laid out.
KVQuantSpecPer-side KV quantization scheme (quantize-on-append). Three families:ELEMENT-CODED (block == KVBlockScheme::None, the default): the side stores CODE bytes at dtype (Int8 / UInt8 / Float8_E4M3 / Float8_E5M2) under the per-tensor scale (+ optional integer zero_point; fp8 codes are scale-only; one there is rejected at construction). BLOCK-QUANTIZED (block set): the side stores packed blocks with inline scales (see KVBlockScheme); dtype / scale / zero_point must stay at their defaults. PER-TOKEN (per_token true): the online scheme; every appended row encodes with its own absmax-derived scale, stored in a scale plane the cache allocates beside its buffer. Element-coded over the symmetric byte codes (Int8 / Float8_E4M3 / Float8_E5M2); scale / zero_point must stay at their defaults (the plane owns the scales), and block must stay None. Either way the scheme rides every tensor the cache hands out: the attention op encodes appended rows through it and decodes attended rows back.
L1LossMean absolute error between input and target. The module form of ClikaRT::ops::l1_loss, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
L1LossOptionsSettings of L1Loss: chain the setters, or assign the fields.
LayerNormLayer normalization module.
LeakyReLUReLU with a small slope on the negative side: input for input > 0, negative_slope * input otherwise. The module form of ClikaRT::ops::leaky_relu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
LeakyReLUOptionsSettings of LeakyReLU: chain the setters, or assign the fields.
LinearAffine transformation of the trailing dimension: y=xW⊤+by = x W^\top + b.
LinearOptionsConstruction options: everything beyond the two feature counts. Every field is optional / defaulted; the defaults produce a plain packing Linear.
LoadOptionsOptions for the load_state_dict overloads (a plain aggregate; assign the fields you need). The same options serve BOTH load paths (an in-memory NamedTensors dict and a lazy checkpoint container), so a model accounts its weight names identically whichever way it loads. strict gates the completeness report; the allow-lists carve per-name exceptions out of it without turning the whole check off:allow_missing: declared slot names the checkpoint may legitimately omit, e.g. a tied head's lm_head.weight, absent from checkpoints that share it with the embedding and bound by weight-sharing instead. allow_unexpected: checkpoint names no slot declares that are fine to leave unread, e.g. a checkpoint shipping a redundant copy of a tied head's weight. weight_residency: a load-scope residency override. When set, every module in the tree that admits a residency choice adopts this value before the checkpoint binds, as if it had been constructed with it, so one load call pins the whole model's posture without touching the model's own construction defaults. A module whose class pins its residency keeps the pin; the load logs one summary line with how many modules adopted the override and how many kept a class-pinned residency. Packed on a weight whose layout cannot pack still refuses readably at bind. Absent (the default) changes nothing. pack_on_load: pack every bound weight at the end of THIS load, so a serving process pays the one-time pack at load time instead of on its first forward. It is the same pack the first forward would run (once, thread-safe), driven through the tree's own pack hooks from the root, so a composite that fuses or folds at pack time does that work here too. After it each packed weight has exactly one resident copy and still reads back through named_parameters() / state_dict(). Default off: the first forward packs.
LogSigmoidLogarithm of the sigmoid, computed stably: -softplus(-input). The module form of ClikaRT::ops::log_sigmoid, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
LogSigmoidOptionsLogSigmoid has no settings; the struct keeps every module in this header constructible the same way.
LogSoftmaxLogarithm of the softmax along dim, computed stably in one pass. The module form of ClikaRT::ops::log_softmax, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
LogSoftmaxOptionsSettings of LogSoftmax: chain the setters, or assign the fields.
MaximumElementwise maximum of input and other with broadcasting; a NaN in either operand yields NaN. The module form of ClikaRT::ops::maximum, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MaximumOptionsMaximum has no settings; the struct keeps every module in this header constructible the same way.
MaxPool1dMax pooling over a 1-D window; the input is channels-last [N, D1, C]. The module form of ClikaRT::ops::max_pool1d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MaxPool1dOptionsSettings of MaxPool1d: chain the setters, or assign the fields.
MaxPool2dMax pooling over a 2-D window; the input is channels-last [N, D1..D2, C]. The module form of ClikaRT::ops::max_pool2d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MaxPool2dOptionsSettings of MaxPool2d: chain the setters, or assign the fields.
MaxPool3dMax pooling over a 3-D window; the input is channels-last [N, D1..D3, C]. The module form of ClikaRT::ops::max_pool3d, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MaxPool3dOptionsSettings of MaxPool3d: chain the setters, or assign the fields.
MinimumElementwise minimum of input and other with broadcasting; a NaN in either operand yields NaN. The module form of ClikaRT::ops::minimum, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MinimumOptionsMinimum has no settings; the struct keeps every module in this header constructible the same way.
MishMish activation: input * tanh(softplus(input)). The module form of ClikaRT::ops::mish, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MishOptionsMish has no settings; the struct keeps every module in this header constructible the same way.
ModuleBase class of every ClikaRT module: an ownership tree of parameters, buffers, and child modules.
ModuleDictA named set of modules registered under their keys, in insertion order, so a checkpoint names their weights heads.cls.weight. A dict has no forward of its own: the owner looks a module up and calls it.
ModuleHooksThe entry points a module captures when it is constructed, one per virtual the runtime invokes from its own frames (see Module::Module() and detail::ModuleHookTable). A Result-returning hook reports a caught exception as the failed Result; the others report it as a HookOutcome. Filled by Module(); a module author never constructs one.
ModuleListAn indexed list of modules registered under their positions ("0", "1", ...), so a checkpoint names their weights layers.3.weight. A list has no forward of its own: the owner iterates it.
MoEDense mixture-of-experts module: a router picks top-k experts per token, each expert runs its gated FFN, and the outputs combine under the router weights.
MoeOptionsMoE configuration beyond the required top-k. Every field is optional; the defaults are the runtime's (SoftmaxTopK routing, renormalized top-k weights, SwiGLU with the Interleaved gate‖up layout, erf gelu).
MSELossMean squared error between input and target. The module form of ClikaRT::ops::mse_loss, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MSELossOptionsSettings of MSELoss: chain the setters, or assign the fields.
MulMultiplies input by other elementwise with broadcasting, then the optional fused activation. The module form of ClikaRT::ops::mul, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
MulOptionsSettings of Mul: chain the setters, or assign the fields.
NLLLossNegative log-likelihood over log-probabilities with the class axis last. The module form of ClikaRT::ops::nll_loss, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
NLLLossOptionsSettings of NLLLoss: chain the setters, or assign the fields.
OutputQuantStatic output quantization for a Linear's product: requantize the output with scale / zero_point along axis, at out_dtype.
PagedParamsShared block-pool geometry for the cache's paged rows. block_size must be a power of two in [16, 512] (a kernel addressing contract); num_blocks is the pool capacity; max_blocks_per_seq caps one sequence's block-table row.
PairwiseDistanceThe p-norm of x1 - x2 + eps over the last axis, with broadcasting of the leading dimensions. The module form of ClikaRT::ops::pairwise_distance, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
PairwiseDistanceOptionsSettings of PairwiseDistance: chain the setters, or assign the fields.
PixelShuffleMoves channels into the spatial axes: [N, H, W, C * r * r] to [N, H * r, W * r, C] for the factor r. The module form of ClikaRT::ops::pixel_shuffle, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
PixelShuffleOptionsSettings of PixelShuffle: chain the setters, or assign the fields.
PixelUnshuffleMoves spatial samples into the channels: [N, H * r, W * r, C] to [N, H, W, C * r * r] for the factor r. The module form of ClikaRT::ops::pixel_unshuffle, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
PixelUnshuffleOptionsSettings of PixelUnshuffle: chain the setters, or assign the fields.
PowRaises input to exponent elementwise with broadcasting. The module form of ClikaRT::ops::pow, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
PowOptionsPow has no settings; the struct keeps every module in this header constructible the same way.
PReLUParametric ReLU.
PReLUOptionsConstruction settings of PReLU beyond the parameter count (a plain aggregate; assign the fields you need).
QConvStatically-quantized convolution module: quantized weight AND activation-quantized compute where the backend serves the scheme, Conv's static-quant sibling (construct with a QTensor weight).
QConvWoQWeight-only-quantized convolution module: the weight rests quantized (a QTensor) and decodes inside the kernels; compute runs at the activation dtype, Conv's weight-only sibling.
QEmbeddingQuantized embedding-table module: gathers rows of a packed table by token index and decodes them to float.
QEmbeddingOptionsSettings of QEmbedding: chain the setters, or assign the fields.
QLinearStatically-quantized linear module: quantized weight AND activation-quantized compute where the backend serves the scheme, Linear's static-quant sibling (same forward surface).
QLinearWoQWeight-only-quantized linear module: the weight rests quantized (a QTensor) and decodes inside the matmul kernels; compute runs at the activation dtype, Linear's weight-only sibling (same forward surface).
QMoEWoQWeight-only-quantized mixture-of-experts module: the stacked [E, ...] expert weights rest quantized and decode inside the per-expert kernels, MoE's weight-only sibling (same routing surface and options).
ReflectionPadPads by mirroring about the edge, without repeating the edge sample. The module form of ClikaRT::ops::reflect_pad, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ReflectionPadOptionsSettings of ReflectionPad: chain the setters, or assign the fields.
ReGLUReGLU over a concatenated [*, 2d] gate-then-up input, emitting [*, d]: relu(gate) * up. The module form of ClikaRT::ops::reglu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ReGLUOptionsReGLU has no settings; the struct keeps every module in this header constructible the same way.
ReLURectified linear unit: max(0, input). The module form of ClikaRT::ops::relu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ReLU6ReLU capped at 6: clamp(input, 0, 6). The module form of ClikaRT::ops::relu6, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ReLU6OptionsReLU6 has no settings; the struct keeps every module in this header constructible the same way.
ReLUOptionsReLU has no settings; the struct keeps every module in this header constructible the same way.
ReplicationPadPads by repeating the edge sample. The module form of ClikaRT::ops::replicate_pad, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ReplicationPadOptionsSettings of ReplicationPad: chain the setters, or assign the fields.
RMSNormRoot-mean-square normalization module.
RotaryEmbeddingRotary position embedding module: the angle tables built once at make, the rotation applied per call. The module form of ops::rotary_embedding.
RotaryEmbeddingOptionsSettings of RotaryEmbedding: chain the setters, or assign the fields. The defaults are the plain rotary embedding of most decoder checkpoints: base 10000, no context scaling, half-split pairs.
SELUScaled exponential linear unit: ELU with the fixed self-normalizing constants (alpha about 1.6733, scale about 1.0507). The module form of ClikaRT::ops::selu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SELUOptionsSELU has no settings; the struct keeps every module in this header constructible the same way.
SequentialRuns its steps in order, each fed the previous step's output.
SigmoidLogistic sigmoid: 1 / (1 + exp(-input)). The module form of ClikaRT::ops::sigmoid, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SigmoidOptionsSigmoid has no settings; the struct keeps every module in this header constructible the same way.
SiLUSigmoid linear unit (swish): input * sigmoid(input). The module form of ClikaRT::ops::silu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SiLUOptionsSiLU has no settings; the struct keeps every module in this header constructible the same way.
SmoothL1LossSquared error below beta, absolute error above it, between input and target. The module form of ClikaRT::ops::smooth_l1_loss, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SmoothL1LossOptionsSettings of SmoothL1Loss: chain the setters, or assign the fields.
SoftmaxSoftmax along dim, computed stably. The module form of ClikaRT::ops::softmax, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SoftmaxOptionsSettings of Softmax: chain the setters, or assign the fields.
SoftminSoftmax of the negated input along dim: small values weigh most. The module form of ClikaRT::ops::softmin, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SoftminOptionsSettings of Softmin: chain the setters, or assign the fields.
SoftplusSmooth ReLU: log(1 + exp(beta * input)) / beta, switching to the exact linear input where beta * input exceeds threshold. The module form of ClikaRT::ops::softplus, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SoftplusOptionsSettings of Softplus: chain the setters, or assign the fields.
SoftshrinkSoft thresholding: moves every element toward zero by lambd, to zero within [-lambd, lambd]. The module form of ClikaRT::ops::softshrink, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SoftshrinkOptionsSettings of Softshrink: chain the setters, or assign the fields.
SoftsignSoftsign activation: input / (1 + |input|). The module form of ClikaRT::ops::softsign, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SoftsignOptionsSoftsign has no settings; the struct keeps every module in this header constructible the same way.
SubSubtracts other, scaled by alpha, from input elementwise with broadcasting: input - alpha * other, then the optional fused activation. The module form of ClikaRT::ops::sub, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SubOptionsSettings of Sub: chain the setters, or assign the fields.
SwiGLUSwiGLU over a concatenated [*, 2d] gate-then-up input, emitting [*, d]: (clamp(up, -limit, limit) + beta) * G * sigmoid(alpha * G) with G = min(gate, limit); the defaults reduce to silu(gate) * up. The module form of ClikaRT::ops::swiglu, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
SwiGLUOptionsSettings of SwiGLU: chain the setters, or assign the fields.
TanhElementwise hyperbolic tangent. The module form of ClikaRT::ops::tanh, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
TanhOptionsTanh has no settings; the struct keeps every module in this header constructible the same way.
ThresholdKeeps every element above threshold and replaces the rest with value. The module form of ClikaRT::ops::threshold, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ThresholdOptionsSettings of Threshold: chain the setters, or assign the fields.
UnflattenSplits the dimension dim into sizes, whose product must equal its extent; a view when the layout permits. The module form of ClikaRT::ops::unflatten, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
UnflattenOptionsSettings of Unflatten: chain the setters, or assign the fields.
UnfoldExtracts sliding windows from a channels-last input [N, D1..Dn, C] into [N, L, prod(kernel_size), C], one row per window placement. The module form of ClikaRT::ops::unfold, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
UnfoldOptionsSettings of Unfold: chain the setters, or assign the fields.
UpsampleResamples the spatial dimensions to size or by scale_factor (exactly one given); the input is channels-last and its rank picks the spatial variant. The two settings share one type, so the module is constructed from its options: Upsample(UpsampleOptions().scale_factor({2, 2})). The module form of ClikaRT::ops::interpolate, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
UpsampleOptionsSettings of Upsample: chain the setters, or assign the fields.
WhereElementwise select with broadcasting: input where condition holds, other elsewhere. The module form of ClikaRT::ops::where, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
WhereOptionsWhere has no settings; the struct keeps every module in this header constructible the same way.
ZeroPadPads with zeros. The module form of ClikaRT::ops::constant_pad with value fixed to 0.0, which documents the full contract (shapes, dtypes, errors); construct once, then call it like a function.
ZeroPadOptionsSettings of ZeroPad: chain the setters, or assign the fields.

Enumerations​

enum KVBlockScheme​

enum class KVBlockScheme : uint8_t

Block-quantized KV storage selector: the block-wise online schemes whose scales ride INLINE in the cache's own block bytes (no external scale to supply; the append encodes each block as it writes). Q8_0 stores 8-bit codes + one f16 scale per 32 elements (8.5 bits/element at rest); NVFP4 stores 4-bit float codes + one UE4M3 sub-scale per 16 elements over 64-element blocks (4.5 bits/element at rest). head_dim must be a whole number of the scheme's blocks. The wider block-scheme taxonomy is reachable by name through KVCacheConfig::kv_cache_scheme; this enum names only the common selections.

EnumeratorDescription
Noneelement-coded (the dtype/scale fields below apply)
Q8_032-element blocks, int8 codes, inline f16 scales
NVFP464-element blocks, fp4 codes, inline per-16 UE4M3 sub-scales

Declared in ClikaRT/nn/kv_cache.h, line 110

enum KVCacheMode​

enum class KVCacheMode : uint8_t

The cache's serving mode: what every KVLayerSpec::Layout::Auto row resolves to. A composition with no token-indexed rows serves identically under both (the mode describes the cache's token-indexed rows; a state-only stack has nothing to page).

EnumeratorDescription
Continuousper-slot buffers (preallocated or grow-on-demand)
Pagedtoken-indexed rows share the cache's block pool

Declared in ClikaRT/nn/kv_cache.h, line 152

enum LinearKind​

enum class LinearKind : std::uint8_t

The SERVING POSTURE of a Linear-family module: what the bound weight is and how forward serves it. Reported by Linear::kind(); the runtime is built without RTTI, so this type tag is the one sanctioned runtime introspection. It names the posture, NOT the class identity: a plain Linear reports Dense until a quantized checkpoint payload binds onto its declared slot, then WeightQuantized; the subclasses report their pinned posture constantly.

EnumeratorDescription
Densedense weight, dense matmul
WeightQuantizedquantized-at-rest weight, float activations
StaticQuantquantized weight AND quantized activations (QLinear)

Declared in ClikaRT/nn/linear.h, line 116

enum WeightResidency​

enum class WeightResidency : std::uint8_t

How a bound weight lives at rest.

EnumeratorDescription
PackedPack once into the backend kernel layout; the pack is the resident copy.
BorrowedNever pack a private copy; forward serves off the shared storage. A derived runtime form the call needs (e.g. a quantized weight's compute form) is retained once when it fits the device's free-memory budget, and rebuilt per call when it does not.
AutoPack only when the copy fits the device's free-memory budget.
PerCallNever retain a derived runtime form on this module, the memory-pressure fallback made deterministic: forward serves off the shared storage and rebuilds any derived form per call, regardless of the free-memory budget. Minimum per-module retention at a per-call rebuild cost; the runtime may still serve repeated calls from its own bounded, process-wide working set, which this choice does not pin.

Declared in ClikaRT/nn/weight_residency.h, line 15

Variables​

kAutoPackMaxFreeFraction​

double kAutoPackMaxFreeFraction = 0.10

Auto packs when the head copy stays at or under this share of the device's currently-free memory: small enough that downstream pool sizing (which claims its own fraction of what remains after load) barely notices, large enough to admit every real vocab table on a workstation card. A device that reports no free-memory figure resolves to Borrowed, never a blind pack. The resolution is logged either way.

Declared in ClikaRT/nn/linear.h, line 107

ClikaRT/nn/attention.h​

#include <ClikaRT/nn/attention.h>

Multi-head attention as an nn::Module leaf over the runtime's fused attention. The settings (head split, causality, scale, soft-cap, sliding window) are fixed at make; the optional bound tensors, an attention mask, a learned per-head softmax sink and rotary tables, live in the module as registered slots and ride every forward(query, key, value) as inputs of the fused attention. The module holds no projection weights: the query / key / value projections and the output projection are Linear modules the caller composes around it. This class serves the dense form (every key and value passed per call); a form attending over a KVCache is a later addition.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/batch_norm.h​

#include <ClikaRT/nn/batch_norm.h>

BatchNorm: per-channel normalization with STORED statistics, the inference form of batch normalization over channels-last input ([N, D1..Dn, C], the channel axis last), exposed as an nn::Module leaf. Unlike the stateless ops::batch_norm free function, a BatchNorm binds its weight / bias and its running_mean / running_var ONCE and reuses them every forward.

The lifecycle (uniform across the weight-bearing nn modules): make(num_features, ...) declares storage-free slots under their checkpoint names; set_weights(...) / set_running_stats(...) (positional) or load_state_dict(...) (a checkpoint, by dotted name) bind them; the FIRST forward builds (once, thread-safe); LoadOptions::pack_on_load pays the build at load time. The module holds each weight exactly once and never re-packs it, so named_parameters() / named_buffers() / state_dict() are complete before and after any forward.

Slots and their checkpoint names: parameters weight (scale) and bias (shift) when affine; buffers running_mean, running_var and num_batches_tracked. The counter is an Int64 scalar the normalization never reads: it is declared with the value 0 so a checkpoint written by torch binds strictly; a checkpoint without it loads with LoadOptions::allow_missing = {"num_batches_tracked"}.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/containers.h​

#include <ClikaRT/nn/containers.h>

Module containers and two stateless utility modules: Sequential, ModuleList, ModuleDict, Identity, Dropout. Header-only: a container holds no weights of its own, the children it registers enumerate under their dotted names (0.weight, encoder.1.bias), and every forward runs in the calling code, one child after the other.

The classes follow the shape of every stateless module: an options struct where there are settings, forward_impl returning a Result, and forward / operator() unwrapping it the way every public method does.

ClikaRT/nn/generated/activation.h​

#include <ClikaRT/nn/generated/activation.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Activation functions as modules: each class applies one nonlinearity to its input, elementwise or along one axis, with the settings fixed at construction.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/distance.h​

#include <ClikaRT/nn/generated/distance.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Distances and similarities between two tensors as modules.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/elementwise.h​

#include <ClikaRT/nn/generated/elementwise.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Elementwise arithmetic as modules: each class combines its operands the way the matching ClikaRT::ops function does (broadcasting and type promotion included), with the settings fixed at construction.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/flatten.h​

#include <ClikaRT/nn/generated/flatten.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Dimension merging and splitting as modules; the results are views when the layout permits.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/fold.h​

#include <ClikaRT/nn/generated/fold.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Sliding-window extraction (Unfold) and its inverse (Fold) as modules over channels-last input.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/loss.h​

#include <ClikaRT/nn/generated/loss.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Losses as modules: forward(input, target) with the reduction and the per-class weights fixed at construction.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/padding.h​

#include <ClikaRT/nn/generated/padding.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Padding as modules: the pad widths are (low, high) pairs in axis order, the first pair padding the leading axis; fewer pairs than the input rank pad the leading axes only, and a negative width crops.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/pixelshuffle.h​

#include <ClikaRT/nn/generated/pixelshuffle.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Pixel shuffle as modules over channels-last input: channels move to and from the spatial axes by a fixed factor.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/pooling.h​

#include <ClikaRT/nn/generated/pooling.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Pooling as modules over channels-last input ([N, D1..Dn, C], the channel axis last): one class per window rank, the window geometry fixed at construction. A geometry list with one value broadcasts to every spatial dimension.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/generated/upsampling.h​

#include <ClikaRT/nn/generated/upsampling.h>

Machine-generated; do not hand-edit. Regenerated with each release of the public surface.

Resampling as a module over channels-last input: the target size or the scale factors and the interpolation mode are fixed at construction.

Every class derives from nn::Module, keeps its settings in a <Class>Options struct (chainable setters; a single positional setting also constructs it implicitly) and runs one ClikaRT::ops function. forward_impl returns a Result; forward and operator() unwrap it the way every public method does. The modules hold no weights, so construct one where the model is built and call it every step.

ClikaRT/nn/group_norm.h​

#include <ClikaRT/nn/group_norm.h>

GroupNorm: normalization over groups of channels of channels-last input ([N, D1..Dn, C], the channel axis last), exposed as an nn::Module leaf. The statistics come from the input itself, per (sample, group) over the group's channels and the spatial dimensions.

The lifecycle (uniform across the weight-bearing nn modules): make(num_groups, num_channels, ...) declares storage-free slots under their checkpoint names; set_weights(...) (positional) or load_state_dict(...) (a checkpoint, by dotted name) binds them; the FIRST forward builds (once, thread-safe); LoadOptions::pack_on_load pays the build at load time. The module holds each weight exactly once and never re-packs it, so named_parameters() / state_dict() are complete before and after any forward.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/instance_norm.h​

#include <ClikaRT/nn/instance_norm.h>

InstanceNorm: normalization per (sample, channel) over the spatial dimensions of channels-last input ([N, D1..Dn, C], the channel axis last), exposed as an nn::Module leaf. By default the statistics come from the input itself; a module made with track_running_stats folds stored running_mean / running_var instead (the inference form).

The lifecycle (uniform across the weight-bearing nn modules): make(num_features, ...) declares storage-free slots under their checkpoint names; set_weights(...) / set_running_stats(...) (positional) or load_state_dict(...) (a checkpoint, by dotted name) bind them; the FIRST forward builds (once, thread-safe); LoadOptions::pack_on_load pays the build at load time. The module holds each weight exactly once and never re-packs it, so named_parameters() / named_buffers() / state_dict() are complete before and after any forward.

Slots and their checkpoint names: parameters weight (scale) and bias (shift) when affine; buffers running_mean, running_var and the Int64 counter num_batches_tracked when track_running_stats (the counter is declared with the value 0 and never read; a checkpoint without it loads with LoadOptions::allow_missing).

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/kv_cache.h​

#include <ClikaRT/nn/kv_cache.h>

A key/value cache for autoregressive decoding, the public face of the runtime's IKVCache. It holds per-layer K/V history across decode steps so attention reads the whole context while each step only computes the new tokens. Bind a layer's keys(l)/values(l) straight into ops::group_query_attention_varlen as BOTH past_key/past_value AND out_present_key/out_present_value; the op then appends this step's post-RoPE K/V into the cache buffer IN PLACE (no realloc, no copy), which is the mechanism that makes per-token decode fast.

One cache serves every decoder composition through a per-layer spec vector, on two orthogonal axes:

  • Kind, WHAT a row stores: token-indexed K/V (AttentionKV, full or windowed via window), or fixed-size serving state (ConvState, RecurrentState, HybridState: slot-resident slabs the recurrence ops bind directly; no token indexing, so layout does not apply).
  • Layout, HOW a token-indexed row is laid out: Continuous (a per-slot buffer, preallocated or grow-on-demand) or Paged (blocks from the cache's shared pool, handed out by reserve / prepare_step). Auto (the default) follows the cache's serving mode, so one spec vector serves both serving shapes unchanged.

Construction is ONE call: KVCache::make(config, layer_specs); the config carries the cache-wide defaults, the serving mode, and (when any row lays out paged) the shared pool geometry in config.paged. Move-only.

​

Prefix caching (paged rows): the contract​

A cache with paged rows can reuse the K/V of already-seen prompt prefixes, so a repeated prefix (a system prompt, a chat transcript resubmitted next turn) skips its prefill compute entirely. The rules:

  • Opt-in, per session. Caching happens ONLY through two calls: admit(slot, prompt_ids) at session start (binds the longest already-cached prefix, and enrolls the prompt's own full blocks in the cache as they fill) and retire(slot, full_ids) at clean finish (keeps the session's blocks, generated tokens included, addressable for future admits). Skip both and the paged cache does no hashing and no caching, zero overhead.
  • Opt-out, two levels. Per session: finish via evict_batch (the cancel/error path). It caches nothing NEW: the session's blocks that no hash names (the partial tail, its generated tokens) return to the pool. Its sealed prefix blocks (full prompt blocks an admit enrolled) are PARKED, not dropped: they stay addressable to a later admit of the same prompt, which reports them as cached tokens. A parked block is free capacity (reclaimed under pool pressure like any cached block). To drop every cached block, reset the cache. Entirely: never call admit/retire.
  • Memory-neutral. The pool's footprint is fixed at construction (num_blocks × block_size × paged rows × heads × head_dim × dtype); caching never allocates beyond it and never pins: a cached block is FREE memory that happens to retain its bytes, reclaimed least-recently-used the moment a live sequence needs a block. A cache entry can therefore disappear under pool pressure; admit reports fewer (or zero) cached tokens.
  • What it saves is prefill compute, not storage. admit returns num_cached; forward only prompt_ids[num_cached:]. Concurrent sessions sharing a prefix also share the physical blocks (copy-on-write on divergence), which REDUCES live block usage.
  • State rows ride along. On a cache mixing paged rows with state rows, retire snapshots each state row's block-boundary checkpoint keyed by the prefix content hash, and admit adopts a cached prefix only when a snapshot matches the hit length exactly; otherwise it binds nothing and returns 0 (shared K/V over the wrong recurrent state is never served).
  • Controls. num_blocks bounds how much history can stay cached opportunistically (size it above the live working set to leave cache headroom); block_size sets hit granularity (a hit needs a full identical block; smaller blocks match finer, at more block-table rows); retire vs evict_batch decides per session what enters the cache.
  • Multimodal sessions declare their media boundary. Sharing is keyed on token ids, and the K/V under a media placeholder comes from NON-token inputs (pixels, audio) the ids cannot identify, so a session whose prompt carries media passes its FIRST media position as addressable_len on admit AND retire. Caching then covers only the pure-text prefix before it: blocks at or past the boundary are never shared or kept, and text-only sessions (the default) pay nothing.

ClikaRT/nn/linear.h​

#include <ClikaRT/nn/linear.h>

Linear: the public face of the runtime's packed matmul family, exposed as an nn::Module leaf and the BASE of its posture hierarchy (QLinearWoQ pins a quantized-at-rest weight; QLinear adds static output quantization). Unlike the stateless ops::matmul / ops::linear free functions (which re-pack the weight on every call), a Linear binds its weight ONCE and packs it into the backend's kernel layout; every forward reuses that packed weight. This is the pack-once path that makes per-token decode fast; build one Linear per projection at load, reuse it every step.

Construction is ONE factory with two symmetric overloads, and the factory IS the posture dispatch:

  • make(weight[, bias][, options]): construct FROM tensors: geometry, dtype and device all read off the weight; declared and bound in one call. The returned shared_ptr<Linear>'s dynamic type follows what was given: options.output_quant set ⇒ a QLinear (this wins even over a quantized weight; quantized weight + output requant IS the static-quant posture); else weight.is_quantized() ⇒ a QLinearWoQ; else a plain Linear. Model code holds the base pointer and never branches.
  • make(in_features, out_features[, options]): declare-then-bind for checkpoint flows: storage-free weight / bias slots are declared under their canonical names (bias presence rides options.bias; placement rides options.device); set_weights (tensors in hand) or load_state_dict (by dotted name) binds them, and the first forward packs (once, thread-safe). options.output_quant set returns a QLinear here too. Otherwise the object is a plain Linear whose SERVING POSTURE follows the payload that later binds: a dense payload serves through the dense matmul, a quantized one through the weight-only-quantized matmul; kind() reports which. LoadOptions::pack_on_load pays the pack at load time instead of on the first forward.

Dtype: adoption over declaration. A declared slot ADOPTS the bound payload's dtype (a weight keeps its checkpoint dtype, never a cast at rest). Until a payload binds, an unpinned slot HAS no dtype: it enumerates through named_parameters() as DataType::Undefined. LinearOptions::dtype PINS a dtype instead: the slots declare at it and every bind casts the payload to it. to(dtype), called before or after any bind, casts the slots and pins that dtype for every subsequent re-bind. One seam to know: a checkpoint RE-load (load_state_dict onto a module whose slot already holds a payload) re-binds at the slot's CURRENT dtype; re-bind through set_weights when adoption of the new payload's dtype is wanted.

Weight orientation: the canonical dense layout is the HuggingFace [out_features, in_features]. set_weights also accepts the transposed [in_features, out_features], validated against the declared feature counts (the ambiguous square DENSE case is read as [out, in]).

A QUANTIZED payload states its own orientation, and every posture and entrance (Linear, QLinear, QLinearWoQ; make(Tensor), set_weights, load_state_dict, share_weight) reads it by ONE law, in this order:

  1. a declared group axis (a grouped scheme): groups ride the input features;
  2. a per-channel scale axis: one scale per OUTPUT feature, so quant_axis 0 names an [out, in] payload and 1 an [in, out] one (a scale whose length fits only the other axis is refused);
  3. an n-major-canonical scheme (the plane-addressed schemes): [out, in]. A payload none of these describe follows the declared feature counts (a non-square shape fits exactly one order; a block stream bound in the transposed [out, in] order is read as such and routed by the residency law); past the counts, a block stream (a load_gguf / MX / NVFP4 payload, the one carrying a logical_shape) reads in the loader's [in, out] order, and a plain code tensor from the tensor alone reads the dense contract, [out, in], a square one included, the same reading a square dense weight takes. Ahead of every rung, a payload whose FACTORY fixes its orientation carries that statement on its own params (make_quantized_fp8 adopts codes as [in, out], make_quantized_fp8_blocked as the checkpoint's [out, in]), so a square fp8 payload binds right from the tensor alone; a caller holding a square [in, out] payload of any other kind states it through a per-channel scale axis along its output features. One thing is refused with a sentence at make / set_weights (never at the first forward): evidence that contradicts the declared counts. Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/module.h​

#include <ClikaRT/nn/module.h>

ClikaRT::nn::Module: the PyTorch-nn.Module-style base every model and layer subclasses. It is BOTH the parameter/submodule container (so named_parameters() produces the same canonical dotted names, in registration order, as PyTorch, i.e. the keys in a HuggingFace model.safetensors) AND the user-space custom-op seam: a leaf module overrides output_shapes() + compute() with its own kernel logic and calls dispatch(), which runs that kernel through the runtime (eager and tracing alike).

Authoring:

  • Composite: subclass, register_module(...) / register_parameter(...) in the ctor, and define your own forward(...) (any signature) composing ops:: / child modules.
  • Leaf custom op: subclass, override output_shapes() + compute(), and have forward(...) call this->dispatch({inputs...}). Modules are held via std::shared_ptr (register_module stores one); a leaf that calls dispatch() must itself be owned by a shared_ptr.

ClikaRT/nn/moe.h​

#include <ClikaRT/nn/moe.h>

A bound, weight-packing Mixture-of-Experts block over DENSE expert weights, the public face of the runtime's fused MoE (router top-k selection + per-expert gated FFN + the weighted combine, in ONE op). Like Linear vs ops::matmul, a MoE binds its stacked expert weights ONCE and packs them into the backend's kernel layout; every forward reuses the pack. Build one per MoE layer at load, reuse it every step. Weight-only-QUANTIZED experts are the sibling module QMoEWoQ (exactly the Linear vs QLinearWoQ split).

The lifecycle (uniform across every weight-bearing nn module):

  1. make(experts, hidden, intermediate, top_k, options, ...), the ONE constructor: declares the storage-free expert / bias slots.
  2. set_weights(...) or load_state_dict(...) binds them.
  3. forward(x, router_logits) packs on first call (once, thread-safe), then serves; LoadOptions::pack_on_load pays the pack at load time.

The caller computes the per-token router logits itself (families differ in how: a plain projection, a normalized/scaled one) and hands them to forward beside the activations; routing (MoeRouting), the gated activation, and the gate‖up layout (ops::fused::SwigluFusion) are configuration.

Expert weights are stacked 3-D tensors (the HF "experts as one parameter" layout): gate_up_experts [E, F·I, H] (F=2 fused gate‖up, F=1 with a separate gate_experts) and down_experts [E, H, I].

ClikaRT/nn/prelu.h​

#include <ClikaRT/nn/prelu.h>

PReLU: the parametric ReLU, a ReLU whose negative-side slope is a learned weight (one slope for all channels, or one per channel), exposed as an nn::Module leaf. Unlike the stateless ops::prelu free function, a PReLU holds its weight and reuses it every forward.

The lifecycle: make(num_parameters, ...) declares the weight parameter [num_parameters] FILLED with options.init, so the module is usable at once; set_weights(...) (positional) or load_state_dict(...) (a checkpoint, by dotted name) replaces it; the FIRST forward builds (once, thread-safe); LoadOptions::pack_on_load pays the build at load time. The module holds the weight exactly once and never re-packs it, so named_parameters() / state_dict() are complete before and after any forward.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/qembedding.h​

#include <ClikaRT/nn/qembedding.h>

A quantized embedding table as an nn::Module leaf: the packed weight stays quantized at rest and every row decodes at gather. The module form of ops::q_embedding, with the weight owned by the module.

The weight is the nn.Embedding layout, slot weight, [num_embeddings, embedding_dim], bound as a QTensor (the packed payload carrying its scheme; the scales and zero points ride the payload). The lifecycle is the one every weight-bearing nn module follows: make(...) declares the slot (or binds it, for the from-weight overload); set_weights(...) / load_state_dict(...) bind; the FIRST forward builds the lookup once (thread-safe) and hands the registry slot over to the packed form, which is the single resident copy; the slot then enumerates as its reconstruction, so named_parameters() and state_dict() stay total. LoadOptions::pack_on_load pays the build at load time.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/rotary_embedding.h​

#include <ClikaRT/nn/rotary_embedding.h>

Rotary position embedding (RoPE) as an nn::Module leaf. The module owns the angle tables: make(dim, max_positions, options) builds the cos / sin planes once, through ops::generate_rotary_cache, and registers them as the buffers cos and sin (each [max_positions, dim / 2] Float32), so named_buffers(), state_dict() and to(device) all see them. forward(input, position_ids) rotates one tensor by position and forward_qk(query, key, position_ids) rotates the pair with one angle gather. One module serves every layer that shares the geometry.

Held via std::shared_ptr (an nn::Module leaf); copy/move are pinned by the base.

ClikaRT/nn/weight_residency.h​

#include <ClikaRT/nn/weight_residency.h>

ClikaRT::nn::WeightResidency: how a bound weight lives at rest. The load call decides it, through nn::LoadOptions::weight_residency; this enum is that option's type. A module picks its own posture only where its weight format leaves one choice (a shared-storage tie, a quantized pack), and the option overrides every posture a module leaves open.