Skip to main content

clika_runtime.modelverse functions

check​

check(source: 'str', device: 'str | None' = None, max_seq: 'int | None' = None, image: 'tuple[int, int] | None' = None, *, dtype: 'clika_runtime.dtype | str | None' = None, devices: 'Sequence[tuple[str, int]]' = (), revision: 'str' = 'main', cache_dir: 'str | None' = None, token: 'str | None' = None, offline: 'bool' = False, weights: 'str | None' = None) -> 'list[FitVerdict]'

One :class:FitVerdict per device for source (a hub repo id, a hub URL, a local model directory or a local .gguf file).

device judges one device: a label the devices listing prints (CPU:0), its --device spelling (cpu, vulkan:1), an api name alone (cuda, every device of that api), or a label from devices below; None judges every device. max_seq is the sequence length judged (None: the model's own context length). dtype judges the model at that load-time dtype (a dense checkpoint's weights priced at its width, the cast the load applies, and its cache at it; None: the stored dtype). image is the (width, height) an image generation pipeline is judged at (None: its weights alone, with the largest square image that fits reported). devices names (label, total_bytes) pairs to judge instead of this machine's devices (a phone's memory, say). weights selects one option of a multi-option GGUF repository (a tag, a glob, a file name); offline refuses any network access.

device_memory​

device_memory(device: "'_clika_runtime.Device | str'") -> 'dict[str, object]'

A device's memory as the runtime reads it: device, total_bytes, free_bytes, runtime_bytes (the runtime's own share) and active_bytes (the pool bytes in use). device is a :class:clika_runtime.Device or its string form ("cpu", "cuda:0").

hardware​

hardware() -> 'dict[str, object]'

The inventory of this machine's compute, the document the devices verb prints: runtime_version, the cpu cores, and backends, each with api, compiled, available, device_count, unavailable_reason and its devices (index, name, total_memory_bytes).

info​

info(sources: 'str | Sequence[str]', *, weights: 'str | None' = None, revision: 'str | None' = None, cache_dir: 'str | None' = None, token: 'str | None' = None, offline: 'bool' = False) -> 'dict[str, object]'

The catalog document of sources: one source or a sequence of them, each a hub repo id, a hub URL, a local model directory or a local .gguf file.

A repository the hub tags as a quantization of another joins that base's entry; a finetune, a merge or an adapter is an entry of its own naming its parent. A source that cannot be read lands in the document's refused list with its reason, and the document still carries the rest; the call raises only when no source could be read. weights narrows a multi-option GGUF repository to one option (a tag, a glob, a file name); revision pins the hub revision (None: the default branch); offline refuses any network access.

info_schema​

info_schema() -> 'dict[str, object]'

The JSON Schema (draft 2020-12) every :func:info document satisfies: every key required, no other key admitted, the closed word lists (requires, precision, format, role, parameters_source) as enums.

list_models​

list_models() -> 'list[_ext.ModelEntry]'

Every registered model family (ModelEntry objects: metadata, the config and GGUF architecture names the family claims, what its runnable factories serve (served), whether it is runnable, and what an identity-only registration is).

load​

load(source: 'str', *, task: 'str | None' = None, options: 'LoadOptions | None' = None, **load_options: 'Any') -> 'Any'

Load the model at source (a local directory, a hub repo id, or a .gguf file) through the door its family calls for.

task names the door explicitly (the model-hub task vocabulary, "text-generation", "automatic-speech-recognition", ...); without it the family's declared modalities decide. The remaining keyword arguments are :class:LoadOptions fields (device, dtype, max_seq, revision, cache_dir, token, offline, ...).

pipeline​

pipeline(task: 'str', model: 'str | Any | None' = None, *, options: 'LoadOptions | None' = None, stage_observer: 'Any' = None, **load_options: 'Any') -> 'Pipeline'

The callable for task over model: a source to load (a local directory, a hub repo id, a .gguf file) or an already loaded model of the task's kind. The keyword arguments are :class:LoadOptions fields (device, dtype, max_seq, revision, ...). An unknown task raises ValueError naming the eleven. stage_observer installs the model's stage observer (set_stage_observer on the loaded model: the stages of every call, on their own threads) where the model's kind reports stages.

serve​

serve(model_or_source: 'str | GenerativeModel | ImageGenerationModel | VideoGenerationModel | Any', host: 'str' = '127.0.0.1', port: 'int' = 8000, *, task: 'str' = 'text-generation', start: 'bool' = True, ready_timeout: 'float' = 30.0, **options: 'Any') -> 'Server'

Serve a text-generation, a text-to-image or a text-to-video model over the OpenAI-compatible API and return the running :class:Server.

model_or_source is a loaded model, an engine, or a source to load (a local directory, a hub repo id, a .gguf file); task picks the load door for a source string: "text-generation" (the default), "text-to-image" or "text-to-video". The remaining keyword arguments are the load options (device, dtype, max_seq, revision, cache_dir, token, ...) and the server options (model_id, max_active, max_queued, default_budget, enable_web_ui, enable_cors, log_request_timing, max_upload_bytes, read_timeout_seconds, api_key, the key every request must carry when it is set, and codec, the video door's clip codec). start=False returns the server unstarted; otherwise it is started and probed ready within ready_timeout seconds.

snapshot_download​

snapshot_download(source: 'str', **options: 'Any') -> 'str'

Download (or find in the cache) the snapshot of source through the registry's download door and return its local directory. Options are those of :func:snapshot; the companions the family names are fetched beside it, and the license line is logged.