QLinearWoQ
QLinearWoQ(in_features, out_features, bias=True, device=None, dtype=None, *, activation=None)
Weight-only-quantized linear: y = x @ dequant(weight).T + bias with
the weight resting quantized. A quantized payload binds as is (its
scheme travels with it); a dense payload is refused.
Args:
in_features: the size of each input sample.
out_features: the size of each output sample.
bias: declare the (floating-point) bias slot.
device: where the layer's tensors live; None declares on the
cpu.
dtype: the activation dtype the layer computes in; None adopts
the bound payload's logical dtype.
activation: a fused epilogue applied to every output (a name such
as "silu" or an :class:~clika_runtime.Activation
value); None for none. How the bound weight lives at rest
is the weight_residency option of the load that binds it.
__init__
__init__(self, in_features: 'int', out_features: 'int', bias: 'bool' = True, device: 'Device | str | None' = None, dtype: 'DtypeLike | None' = None, *, activation: 'ActivationLike' = None) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
extra_repr
extra_repr(self) -> 'str'
extra_repr() -> str
One line of per-class detail for :meth:__repr__; a layer prints
its geometry here (in_features=64, out_features=256).
forward
forward(self, input: 'Tensor') -> 'Tensor'
forward(input) -> Tensor
Floating-point [*, in_features] to [*, out_features].
from_weights
from_weightsfrom_weights(weight, bias=None, *, activation=None) -> QLinearWoQ
from_weights(weight, bias=None, *, activation=None) -> QLinearWoQ
Construct FROM a quantized weight (crt.ops.quantize's
QTensor): geometry, dtype, and device read off the payload; an
optional bias binds beside it.
set_weights
set_weights(self, weight: 'QTensor', bias: 'Tensor | None' = None) -> 'None'
set_weights(weight, bias=None) -> None
Bind the quantized weight (its scheme travels with it) and the optional floating-point bias; a dense weight is refused.