clika_runtime.modelverse functions
check
check(source: 'str', device: 'str | None' = None, max_seq: 'int | None' = None, image: 'tuple[int, int] | None' = None, *, dtype: 'clika_runtime.dtype | str | None' = None, devices: 'Sequence[tuple[str, int]]' = (), revision: 'str' = 'main', cache_dir: 'str | None' = None, token: 'str | None' = None, offline: 'bool' = False, weights: 'str | None' = None) -> 'list[FitVerdict]'
One :class:FitVerdict per device for source (a hub repo id, a
hub URL, a local model directory or a local .gguf file).
device judges one device: a label the devices listing prints
(CPU:0), its --device spelling (cpu, vulkan:1), an api name
alone (cuda, every device of that api), or a label from devices
below; None judges every device. max_seq is the sequence length
judged (None: the model's own context length). dtype judges the
model at that load-time dtype (a dense checkpoint's weights priced at
its width, the cast the load applies, and its cache at it; None: the
stored dtype). image is the
(width, height) an image generation pipeline is judged at (None:
its weights alone, with the largest square image that fits reported).
devices names (label, total_bytes) pairs to judge instead of this
machine's devices (a phone's memory, say). weights selects one option
of a multi-option GGUF repository (a tag, a glob, a file name);
offline refuses any network access.
device_memory
device_memory(device: "'_clika_runtime.Device | str'") -> 'dict[str, object]'
A device's memory as the runtime reads it: device, total_bytes,
free_bytes, runtime_bytes (the runtime's own share) and
active_bytes (the pool bytes in use). device is a
:class:clika_runtime.Device or its string form ("cpu",
"cuda:0").
hardware
hardware() -> 'dict[str, object]'
The inventory of this machine's compute, the document the devices
verb prints: runtime_version, the cpu cores, and backends, each
with api, compiled, available, device_count,
unavailable_reason and its devices (index, name,
total_memory_bytes).
info
info(sources: 'str | Sequence[str]', *, weights: 'str | None' = None, revision: 'str | None' = None, cache_dir: 'str | None' = None, token: 'str | None' = None, offline: 'bool' = False) -> 'dict[str, object]'
The catalog document of sources: one source or a sequence of them,
each a hub repo id, a hub URL, a local model directory or a local
.gguf file.
A repository the hub tags as a quantization of another joins that base's
entry; a finetune, a merge or an adapter is an entry of its own naming
its parent. A source that cannot be read lands in the document's
refused list with its reason, and the document still carries the
rest; the call raises only when no source could be read. weights
narrows a multi-option GGUF repository to one option (a tag, a glob, a
file name); revision pins the hub revision (None: the default
branch); offline refuses any network access.
info_schema
info_schema() -> 'dict[str, object]'
The JSON Schema (draft 2020-12) every :func:info document satisfies:
every key required, no other key admitted, the closed word lists
(requires, precision, format, role, parameters_source)
as enums.
list_models
list_models() -> 'list[_ext.ModelEntry]'
Every registered model family (ModelEntry objects: metadata, the
config and GGUF architecture names the family claims, what its runnable
factories serve (served), whether it is runnable, and what an
identity-only registration is).
load
load(source: 'str', *, task: 'str | None' = None, options: 'LoadOptions | None' = None, **load_options: 'Any') -> 'Any'
Load the model at source (a local directory, a hub repo id, or a
.gguf file) through the door its family calls for.
task names the door explicitly (the model-hub task vocabulary,
"text-generation", "automatic-speech-recognition", ...); without
it the family's declared modalities decide. The remaining keyword
arguments are :class:LoadOptions fields (device, dtype,
max_seq, revision, cache_dir, token, offline, ...).
pipeline
pipeline(task: 'str', model: 'str | Any | None' = None, *, options: 'LoadOptions | None' = None, stage_observer: 'Any' = None, **load_options: 'Any') -> 'Pipeline'
The callable for task over model: a source to load (a local
directory, a hub repo id, a .gguf file) or an already loaded model of the
task's kind. The keyword arguments are :class:LoadOptions fields
(device, dtype, max_seq, revision, ...). An unknown task
raises ValueError naming the eleven. stage_observer installs the
model's stage observer (set_stage_observer on the loaded model: the
stages of every call, on their own threads) where the model's kind
reports stages.
serve
serve(model_or_source: 'str | GenerativeModel | ImageGenerationModel | VideoGenerationModel | Any', host: 'str' = '127.0.0.1', port: 'int' = 8000, *, task: 'str' = 'text-generation', start: 'bool' = True, ready_timeout: 'float' = 30.0, **options: 'Any') -> 'Server'
Serve a text-generation, a text-to-image or a text-to-video model over
the OpenAI-compatible API and return the running :class:Server.
model_or_source is a loaded model, an engine, or a source to load (a
local directory, a hub repo id, a .gguf file); task picks the load
door for a source string: "text-generation" (the default),
"text-to-image" or "text-to-video". The remaining keyword
arguments are the load options (device, dtype, max_seq,
revision, cache_dir, token, ...) and the server options
(model_id, max_active, max_queued, default_budget,
enable_web_ui, enable_cors, log_request_timing,
max_upload_bytes, read_timeout_seconds, api_key, the key
every request must carry when it is set, and codec, the video
door's clip codec). start=False returns
the server unstarted; otherwise it is started and probed ready within
ready_timeout seconds.
snapshot_download
snapshot_download(source: 'str', **options: 'Any') -> 'str'
Download (or find in the cache) the snapshot of source through the
registry's download door and return its local directory. Options are
those of :func:snapshot; the companions the family names are fetched
beside it, and the license line is logged.