Skip to main content

QLinearWoQ

QLinearWoQ(in_features, out_features, bias=True, device=None, dtype=None, *, activation=None)

Weight-only-quantized linear: y = x @ dequant(weight).T + bias with the weight resting quantized. A quantized payload binds as is (its scheme travels with it); a dense payload is refused.

Args: in_features: the size of each input sample. out_features: the size of each output sample. bias: declare the (floating-point) bias slot. device: where the layer's tensors live; None declares on the cpu. dtype: the activation dtype the layer computes in; None adopts the bound payload's logical dtype. activation: a fused epilogue applied to every output (a name such as "silu" or an :class:~clika_runtime.Activation value); None for none. How the bound weight lives at rest is the weight_residency option of the load that binds it.

__init__​

__init__(self, in_features: 'int', out_features: 'int', bias: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, activation: 'ActivationLike' = None) -> 'None'

Initialize self. See help(type(self)) for accurate signature.

extra_repr​

extra_repr(self) -> 'str'

extra_repr() -> str

One line of per-class detail for :meth:__repr__; a layer prints its geometry here (in_features=64, out_features=256).

forward​

forward(self, input: 'Tensor') -> 'Tensor'

forward(input) -> Tensor

Floating-point [*, in_features] to [*, out_features].

from_weights​

from_weightsfrom_weights(weight, bias=None, *, activation=None) -> QLinearWoQ

from_weights(weight, bias=None, *, activation=None) -> QLinearWoQ

Construct FROM a quantized weight (crt.ops.quantize's QTensor): geometry, dtype, and device read off the payload; an optional bias binds beside it.

set_weights​

set_weights(self, weight: 'QTensor', bias: 'Tensor | None' = None) -> 'None'

set_weights(weight, bias=None) -> None

Bind the quantized weight (its scheme travels with it) and the optional floating-point bias; a dense weight is refused.