Skip to main content

ClikaRT::nn::QLinear

class

Header: ClikaRT/nn/qlinear.h
Inherits: ClikaRT::nn::Linear

QLinear: the static-quant (QDQ) member of the Linear family: BOTH operands are quantized. The weight binds quantized (codes + scheme) and packs into the backend's integer kernel layout on the first forward; the activation arrives ALREADY quantized (from calibration or a preceding ops::quantize; this module never quantizes it). For the weight-only flavor (float activations against a quantized weight) a plain Linear (which serves whatever binds) or the QLinearWoQ pin is the home.

This is the family's genuine API subclass: it carries the output-requant state. Output dtype is conditional on it:

  • requant BOUND (options.output_quant at make, or set_output_quantization later): the product requantizes to the bound scheme and forward returns a QUANTIZED tensor carrying it.
  • NOT bound: the dequant path; forward returns a FLOAT tensor at the weight scheme's float target dtype. Linear::make returns this class whenever options.output_quant is set (that wins even over a quantized weight; quantized weight + output requant IS the static-quant posture); kind() reports StaticQuant constantly.

The weight's orientation is the PAYLOAD's own, read by the family's one quantized-orientation law (linear.h): a per-channel scale axis names the output features, a grouped scheme's group axis the input features; past those, the declared counts decide a non-square shape, a block stream's logical shape reads in the loader's [in, out] order, and a per-tensor code tensor reads the HuggingFace [out_features, in_features] from the tensor alone, a square one included; a payload whose factory fixes its orientation (make_quantized_fp8: [in, out]) carries that statement and is read by it first. The registry slot declares [out_features, in_features], and a payload in either logical order binds onto it; the pack serves the resolved orientation. A dtype to is Unsupported: the weight stays quantized at rest. Lifecycle, movement, and slot rules are the family's (linear.h).

Statically-quantized linear module: quantized weight AND activation-quantized compute where the backend serves the scheme, Linear's static-quant sibling (same forward surface).

The module owns its weights: construct with make(...), or declare shapes and bind a checkpoint via load_state_dict. The first forward packs the weight (or LoadOptions::pack_on_load packs it at the end of the load); the pack is the one resident copy, and named_parameters() / state_dict() read the weight back from it; to(dtype) restores, casts, and re-packs on the next forward.

Static member functions​

make(Tensor, OptionalTensor, LinearOptions)​

static std::shared_ptr<QLinear> make(
    Tensor weight,
    OptionalTensor bias = {},
    LinearOptions options = {}
)

Construct FROM a quantized weight (a static factory, so a bad config is a clean error): geometry (the payload's own orientation, by the law in linear.h), dtype and device all read off the payload; declared AND bound in one call (an optional bias [out] binds beside it, at its own dtype: the bias follows the EXIT's domain, one law for QLinear and QConv on every backend: an integer bias is a quantized value that adds in the integer accumulator at a_scale·w_scale; a float bias is a real value that adds at a float exit as is, and quantizes onto the accumulator grid, round-half-to-even, only where the exit is the raw Int32 accumulator; a bias slot exists exactly when bias is passed). options.output_quant binds the output requant in the same call. The QTensor overload is the typed convenience over the same factory. Raises ClikaRT::Error on a dense payload; the static-quant weight is quantized by definition.

Declared in ClikaRT/nn/qlinear.h, line 74

make(QTensor, OptionalTensor, LinearOptions)​

static std::shared_ptr<QLinear> make(
    QTensor weight,
    OptionalTensor bias = {},
    LinearOptions options = {}
)

Construct from a quantized weight in either logical order (its own orientation, by the law in linear.h; + optional bias and the shared LinearOptions). Raises when the scheme cannot be served, or when the payload's evidence contradicts a declared geometry.

Declared in ClikaRT/nn/qlinear.h, line 83

make(int64_t, int64_t, LinearOptions)​

static std::shared_ptr<QLinear> make(
    std::int64_t in_features,
    std::int64_t out_features,
    LinearOptions options = {}
)

Construct FROM a quantized weight (a static factory, so a bad config is a clean error): geometry (the payload's own orientation, by the law in linear.h), dtype and device all read off the payload; declared AND bound in one call (an optional bias [out] binds beside it, at its own dtype: the bias follows the EXIT's domain, one law for QLinear and QConv on every backend: an integer bias is a quantized value that adds in the integer accumulator at a_scale·w_scale; a float bias is a real value that adds at a float exit as is, and quantizes onto the accumulator grid, round-half-to-even, only where the exit is the raw Int32 accumulator; a bias slot exists exactly when bias is passed). options.output_quant binds the output requant in the same call. The QTensor overload is the typed convenience over the same factory. Raises ClikaRT::Error on a dense payload; the static-quant weight is quantized by definition.

Declared in ClikaRT/nn/qlinear.h, line 100

Member functions​

~QLinear()​

~QLinear() override

Releases the module's packed weights.

Declared in ClikaRT/nn/qlinear.h, line 107

kind()​

virtual LinearKind kind() const noexcept override

StaticQuant, constantly; the posture is this class's definition.

Declared in ClikaRT/nn/qlinear.h, line 110

set_weights(QTensor, OptionalTensor)​

void set_weights(QTensor weight, OptionalTensor bias = {})

Bind the declared slots positionally: the QUANTIZED weight (its logical shape in either orientation, read by the orientation law in linear.h and validated against the declared feature counts) and, when the module was declared with one, bias ([out]). A DENSE weight is refused. The quantized payload lands as-is (no dtype cast). Re-binding after a pack drops the pack. The QTensor overload is the typed convenience. Raises ClikaRT::Error on a dense payload, a geometry mismatch, or a bias without a declared bias slot.

Declared in ClikaRT/nn/qlinear.h, line 123

set_output_quantization()​

void set_output_quantization(
    Tensor scale,
    OptionalTensor zero_point = {},
    std::int64_t quant_axis = -1,
    std::optional<DataType> out_dtype = std::nullopt
)

Optionally bind the OUTPUT quantization: the product requantizes to scale / zero_point (per-tensor affine) and forward returns a quantized tensor whose code dtype is out_dtype (defaulting to the zero-point's dtype when one is bound, else Int8). Calling it again re-binds (and drops any pack); never calling it keeps the float (dequant) output path. Raises ClikaRT::Error on undefined scale.

Declared in ClikaRT/nn/qlinear.h, line 136

forward(QTensor)​

Tensor forward(QTensor x) const

Apply the bound quantized weight to the QUANTIZED activation: x [*, in] must carry a quantization scheme (from calibration or a preceding ops::quantize) -> [*, out] (float, or quantized when an output requant is bound). The QTensor overload is the typed convenience. The first call packs (once, thread-safe); every later call reuses the pack. Raises ClikaRT::Error on a float activation, a shape mismatch, or a still-fake slot.

Declared in ClikaRT/nn/qlinear.h, line 152

set_weights(Tensor, OptionalTensor)​

void set_weights(Tensor weight, OptionalTensor bias = {})

Bind the declared slots positionally: the QUANTIZED weight (its logical shape in either orientation, read by the orientation law in linear.h and validated against the declared feature counts) and, when the module was declared with one, bias ([out]). A DENSE weight is refused. The quantized payload lands as-is (no dtype cast). Re-binding after a pack drops the pack. The QTensor overload is the typed convenience. Raises ClikaRT::Error on a dense payload, a geometry mismatch, or a bias without a declared bias slot.

Declared in ClikaRT/nn/qlinear.h, line 122

forward(Tensor)​

Tensor forward(Tensor x) const

Apply the bound weight: x [*, in] -> [*, out] (then bias + activation if configured). The first call packs the weight into the backend kernel layout (once, thread-safe); every later call reuses the pack: the dense matmul or the weight-only-quantized matmul, per the bound payload.

Throws

  • ClikaRT::Error: on a shape/dtype mismatch or a still-fake slot.

Declared in ClikaRT/nn/qlinear.h, line 151

Protected member functions​

initialize_impl()​

virtual Result<void> initialize_impl() override

The pack hook (linear.h) over the static-quant pack, plus the posture pin at the gate: a dense payload that reached the slot by name is refused rather than packed.

Declared in ClikaRT/nn/qlinear.h, line 163