---
title: "LoadOptions"
sidebar_label: "LoadOptions"
description: "The clika_runtime.modelverse.models LoadOptions class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

Load-time knobs shared by every loader.

    ``device`` places the weights (a string like "cuda:0" works; "auto" or
    None is the automatic placement: the best available accelerator, and a
    load the chosen accelerator cannot take or run falls back down the
    selection order, ending at the CPU, as the CLI's does; a device you name
    never falls back); ``stream``
    instead loads them on that stream; ``devices`` declares a multi-device
    set the family places across. ``dtype`` selects the compute dtype (None
    keeps the checkpoint's): a dense checkpoint's weights are cast to it at
    load and the activations and the KV cache run at it, a quantized
    checkpoint decodes to it, and on an image or video generation pipeline it
    is the denoising transformer's dtype; ``component_dtypes`` names a
    pipeline's other components by their ``model_index.json`` slot
    (``{"text_encoder": "float32", "vae": "bfloat16"}``; a single model
    refuses it); ``dtype_cast_eps`` checks every weight a load casts before
    any weight byte moves to the device (the weight cast to the target and
    back must differ from its stored values by at most this value, else the
    load refuses naming the weight; None, the default, checks nothing);
    ``max_seq`` is the context window (0 = auto: the largest window that fits
    the device's free memory beside the weights, bounded by the checkpoint's
    own, as the CLI's default; a modern checkpoint's full window, e.g. 40960,
    would make a chat/serving engine pre-build a KV pool of tens of seconds
    and gigabytes). Pass ``keep_max_seq=True`` for the checkpoint's full
    window, or a positive ``max_seq`` for an exact one, clamped to the
    checkpoint's; the two together are refused. ``kv_cache`` ``kv_cache``
    names the KV-cache storage every pipeline
    over the model runs on ("paged" or "continuous"; None keeps the
    one-session default, continuous, which the first call builds once);
    ``kv_quant`` names an online quantized KV-cache scheme; ``max_active`` is
    the requests the model serves at once, one knob for every family (a chat
    model's decode sessions, the KV cache's slot count; a speech model's
    decoder replicas over one encoder; a served engine's admitted requests),
    0 = one;
    ``model_args`` carries family-defined arguments. ``revision``,
    ``token``, ``cache_dir`` and ``offline`` govern hub access; ``weights``
    selects one weight option of a multi-option repo ("Q6_K", a glob, a
    file). ``progress(done, total)`` and ``file_progress(info)`` report the
    download; return False from either to cancel.
    

## `__init__`

```python
__init__(self, device: 'clika_runtime.Device | str | None' = None, stream: 'clika_runtime.Stream | None' = None, devices: 'Sequence[clika_runtime.Device | str]' = (), dtype: 'clika_runtime.dtype | str | None' = None, component_dtypes: 'Mapping[str, clika_runtime.dtype | str] | None' = None, dtype_cast_eps: 'float | None' = None, max_seq: 'int' = 0, keep_max_seq: 'bool' = False, kv_cache: 'str | None' = None, kv_quant: 'str' = '', model_args: 'Mapping[str, Any] | None' = None, max_active: 'int' = 0, max_kv_cache_bytes: 'int | None' = None, revision: 'str' = 'main', token: 'str | None' = None, cache_dir: 'str | None' = None, offline: 'bool' = False, weights: 'str | None' = None, progress: 'Progress | None' = None, file_progress: 'FileProgress | None' = None) -> None
```

Initialize self.  See help(type(self)) for accurate signature.

## `hub_kwargs`

```python
hub_kwargs(self) -> 'dict[str, Any]'
```

The hub-access subset, as keyword arguments.

## `to_bound`

```python
to_bound(self) -> 'Any'
```

The model library's own options object carrying these values.
