Skip to main content

LoadOptions

Load-time knobs shared by every loader.

``device`` places the weights (a string like "cuda:0" works; "auto" or
None is the automatic placement: the best available accelerator, and a
load the chosen accelerator cannot take or run falls back down the
selection order, ending at the CPU, as the CLI's does; a device you name
never falls back); ``stream``
instead loads them on that stream; ``devices`` declares a multi-device
set the family places across. ``dtype`` selects the compute dtype (None
keeps the checkpoint's): a dense checkpoint's weights are cast to it at
load and the activations and the KV cache run at it, a quantized
checkpoint decodes to it, and on an image or video generation pipeline it
is the denoising transformer's dtype; ``component_dtypes`` names a
pipeline's other components by their ``model_index.json`` slot
(``{"text_encoder": "float32", "vae": "bfloat16"}``; a single model
refuses it); ``dtype_cast_eps`` checks every weight a load casts before
any weight byte moves to the device (the weight cast to the target and
back must differ from its stored values by at most this value, else the
load refuses naming the weight; None, the default, checks nothing);
``max_seq`` is the context window (0 = auto: the largest window that fits
the device's free memory beside the weights, bounded by the checkpoint's
own, as the CLI's default; a modern checkpoint's full window, e.g. 40960,
would make a chat/serving engine pre-build a KV pool of tens of seconds
and gigabytes). Pass ``keep_max_seq=True`` for the checkpoint's full
window, or a positive ``max_seq`` for an exact one, clamped to the
checkpoint's; the two together are refused. ``kv_cache`` ``kv_cache``
names the KV-cache storage every pipeline
over the model runs on ("paged" or "continuous"; None keeps the
one-session default, continuous, which the first call builds once);
``kv_quant`` names an online quantized KV-cache scheme; ``max_active`` is
the requests the model serves at once, one knob for every family (a chat
model's decode sessions, the KV cache's slot count; a speech model's
decoder replicas over one encoder; a served engine's admitted requests),
0 = one;
``model_args`` carries family-defined arguments. ``revision``,
``token``, ``cache_dir`` and ``offline`` govern hub access; ``weights``
selects one weight option of a multi-option repo ("Q6_K", a glob, a
file). ``progress(done, total)`` and ``file_progress(info)`` report the
download; return False from either to cancel.

__init__​

__init__(self, device: 'clika_runtime.Device | str | None' = None, stream: 'clika_runtime.Stream | None' = None, devices: 'Sequence[clika_runtime.Device | str]' = (), dtype: 'clika_runtime.dtype | str | None' = None, component_dtypes: 'Mapping[str, clika_runtime.dtype | str] | None' = None, dtype_cast_eps: 'float | None' = None, max_seq: 'int' = 0, keep_max_seq: 'bool' = False, kv_cache: 'str | None' = None, kv_quant: 'str' = '', model_args: 'Mapping[str, Any] | None' = None, max_active: 'int' = 0, max_kv_cache_bytes: 'int | None' = None, revision: 'str' = 'main', token: 'str | None' = None, cache_dir: 'str | None' = None, offline: 'bool' = False, weights: 'str | None' = None, progress: 'Progress | None' = None, file_progress: 'FileProgress | None' = None) -> None

Initialize self. See help(type(self)) for accurate signature.

hub_kwargs​

hub_kwargs(self) -> 'dict[str, Any]'

The hub-access subset, as keyword arguments.

to_bound​

to_bound(self) -> 'Any'

The model library's own options object carrying these values.