GenerativeModel
A loaded text-generation model.
``generate`` returns the reply to one prompt (or to a conversation
passed as ``messages``); ``stream_generate`` yields the reply piece by
piece as the model produces it. Decode knobs ride keyword arguments over
the checkpoint's defaults (``max_new_tokens``, ``temperature``,
``top_p``, ...; see :class:`GenerationConfig`). ``report`` holds how the
last generation ended.
assistant_model (property)
The attached assistant model, or None.
compute_dtype (property)
The dtype the model computes at (Undefined when the checkpoint's own applies).
defaults (property)
The checkpoint's decode defaults.
eos_token_ids (property)
The ids that end a sequence.
has_chat_template (property)
True when the tokenizer ships a chat template.
kv_cache_mode (property)
The KV-cache storage the load named ("paged" or "continuous").
None when the load named nothing: the streaming pipeline and
generate then run one continuous slot, the one-session default.
local_dir (property)
The resolved snapshot directory the model loaded from.
max_seq (property)
The context window in tokens.
reasoning_mode (property)
How the reasoning channel behaves: "never" (no channel),
"always" (no off switch), "toggleable" (on by default;
enable_thinking=False turns it off), "opt_in" (off by default;
enable_thinking=True turns it on) or "adaptive" (the model
decides for itself).
report (property)
How the last generation ended (a GenerationReport), or None before the first one.
reports (property)
How each reply of the last :meth:generate_batch ended, one
:class:ChatOutcome per prompt in prompt order (empty before the
first batch).
residency (property)
What the loaded model holds: the weights part the load left on each
device (_Loaded.residency) plus the pipeline part, the KV cache
the model's streaming pipeline built (kv_cache_bytes 0 before the
first request builds it), in the one document shape.
supports_audio (property)
supports_speculative_decoding (property)
True when a request may draft through prompt lookup on this model
(prompt_lookup_num_tokens= on :meth:generate, :meth:chat and
a session's send): the family scores every fed position in one
step and keeps the whole history. The reply's report then carries
the tally.
supports_thinking (property)
True when the model produces a reasoning channel: one a request
switches (enable_thinking= / reasoning_effort= on
:meth:generate, :meth:chat and :meth:ChatSession.send), one it
always opens, or one it opens when it decides to; read from the
checkpoint's chat template at load. reasoning_mode says which.
supports_tools (property)
True when the model takes tool definitions and its replies' calls are read: a chat template whose call syntax a registered format reads.
supports_video (property)
supports_vision (property)
template_renders_tools (property)
True when the chat template itself renders offered tools into the prompt (otherwise the format's tools block joins the system turn).
tokenizer (property)
The checkpoint's tokenizer (a clika_runtime tokenizer).
tool_call_format (property)
The format the model's tool calls are read with ("hermes", "mistral", "llama3_json", "pythonic", "qwen_xml", "generic"), probed from the checkpoint's chat template at load.
vocab_size (property)
__init__
__init__(self, handle: 'Any') -> 'None'
Initialize self. See help(type(self)) for accurate signature.
chat
chat(self, messages: 'Sequence[Mapping[str, Any]]', *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'
The assistant's reply to a conversation of {"role", "content"}
turns: the text when stream is False, an iterator of text pieces
when it is True. tools offers tool definitions for the turn
(:meth:ChatSession.send has the shape of the reply); keyword
arguments overlay the decode policy for this call
(enable_thinking=False switches a thinking model's channel off).
The turns run on the model's chat engine (built on the first call and
kept), so consecutive calls reuse its KV cache and prefix cache; the
conversation itself is the caller's (pass the whole history each
time), or use :meth:chat_session for one that keeps it. report
afterwards holds the reply's :class:ChatOutcome. Attached media ride
the last user turn as the chat template's parts (the layout
generate and the CLI produce); a marker hand-written into the text
would count as a second one.
chat_session
chat_session(self, **options: 'Any') -> "'ChatSession'"
A new :class:ChatSession on this model (its own serving engine;
system= seeds the conversation, max_active= sizes the engine).
conversation_text
conversation_text(self, messages: 'Sequence[Mapping[str, Any]]') -> 'str'
The text the model reads for a conversation: the chat template's
rendering when the checkpoint ships one, else the turns flattened as
role: content lines with the assistant's turn opened.
generate
generate(self, prompt: 'str | Sequence[str] | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), max_active: 'int | None' = None, **overrides: 'Any') -> 'str | list[str]'
Generate the reply to prompt (one user message), to
messages (a conversation of {"role", "content"} turns), or to
a list of prompts (one reply per prompt, in order).
A model with a chat template gets it applied; a model without one
reads the prompt as raw text. images ([H, W, 3] tensors),
audio ([S] float mono tensors) and videos ((frames,
timestamps) pairs) ride the user turn. Keyword arguments overlay the
checkpoint's decode defaults (max_new_tokens=64,
temperature=0.7, enable_thinking=False to switch a thinking
model's channel off for this request, ...). report afterwards says how the reply
ended. Requests run through the model's serving pipeline, built on
the first request and reused by every later one.
A list of prompts runs as one batch through the model's serving
engine (see :meth:generate_batch): the prompts decode together,
step by step, instead of one after another; max_active caps how
many decode at once (default: all of them).
generate_batch
generate_batch(self, prompts: 'Sequence[str]', *, config: 'GenerationConfig | Mapping[str, Any] | None' = None, max_active: 'int | None' = None, **overrides: 'Any') -> 'list[str]'
Generate one reply per prompt, the prompts decoding TOGETHER.
Every prompt becomes a session on one continuous-batching engine over
this model (the engine chat_session rides), all of them submitted
before any reply is read, so each decode step serves every prompt
still generating instead of the prompts taking turns. max_active
caps how many decode at once (default: every prompt; a smaller cap
queues the rest and admits them as others finish). The replies come
back in prompt order; reports holds one :class:ChatOutcome per
prompt and report the last. An empty list answers an empty list.
render
render(self, messages: 'Sequence[Mapping[str, Any]]', *, add_generation_prompt: 'bool' = True, tools: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, **overrides: 'Any') -> 'str'
The prompt text the chat template renders for messages
([{"role": ..., "content": ...}, ...]; an assistant turn may carry
tool_calls and a tool turn a tool_call_id). tools
offers tool definitions (the OpenAI envelope or the bare function
object): the template renders them, or a template without a tools
slot gets the format's tools block in the system turn. config and
the keyword overrides name a request's policy: the render then folds
its reasoning switch into the template variables the way the request
would (enable_thinking=False renders a toggleable model's block
closed; an unknown level, or an off switch on a model without one,
raises ValueError); with tools alone, or a policy, the model's
default policy folds the same way.
set_assistant
set_assistant(self, model: "'GenerativeModel | None'") -> 'None'
Attach model as this model's assistant (assisted generation):
a second generative model sharing this one's vocabulary and device,
with a context window at least this one's (refused otherwise, naming
the mismatch); load(..., assistant_model=) does this at the load.
A request's num_assistant_tokens then drafts with it, and its
report carries the tally. None detaches. Set before the first
generation.
set_stage_observer
set_stage_observer(self, callback: 'StageObserver | None') -> 'None'
Report every request's pipeline stages (tokenize, prefill,
decode once per request, detokenize with the final text,
vision when a tower runs) to callback, called with a
:class:StageEvent on the stage's own thread; it must return at
once. The pipeline reads it when it is built, so set it before the
first request. None detaches.
stream_generate
stream_generate(self, prompt: 'str | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), skip_special_tokens: 'bool' = True, decode: 'bool' = True, **overrides: 'Any') -> 'Iterator[str] | Iterator[list[int]]'
Generate as generate does, yielding the reply as the model
produces it: decoded text pieces (the default), or with
decode=False the token ids themselves, one list per batch the
model hands over. Attached media ride the (last) user turn as the
chat template's parts, the layout generate and the CLI produce;
a marker hand-written into the text would count as a second one.
The model's thread produces token ids into a queue; this iterator
drains it, so the calling thread never blocks the model. A text piece
is not a token: it may carry several tokens, or the decoder may hold
a piece back until a multi-byte character completes, so per-token
timing, logprob or id consumers take the id stream and decode it
themselves (tokenizer.streaming_decoder is the decoder the text
stream uses). Closing the iterator early stops delivery; the
generation itself runs to completion. report is set once the
reply is complete. The model's serving pipeline is built on the
first request (that call pays the build once) and reused after.