Skip to main content

GenerativeModel

A loaded text-generation model.

``generate`` returns the reply to one prompt (or to a conversation
passed as ``messages``); ``stream_generate`` yields the reply piece by
piece as the model produces it. Decode knobs ride keyword arguments over
the checkpoint's defaults (``max_new_tokens``, ``temperature``,
``top_p``, ...; see :class:`GenerationConfig`). ``report`` holds how the
last generation ended.

assistant_model (property)​

The attached assistant model, or None.

compute_dtype (property)​

The dtype the model computes at (Undefined when the checkpoint's own applies).

defaults (property)​

The checkpoint's decode defaults.

eos_token_ids (property)​

The ids that end a sequence.

has_chat_template (property)​

True when the tokenizer ships a chat template.

kv_cache_mode (property)​

The KV-cache storage the load named ("paged" or "continuous").

None when the load named nothing: the streaming pipeline and generate then run one continuous slot, the one-session default.

local_dir (property)​

The resolved snapshot directory the model loaded from.

max_seq (property)​

The context window in tokens.

reasoning_mode (property)​

How the reasoning channel behaves: "never" (no channel), "always" (no off switch), "toggleable" (on by default; enable_thinking=False turns it off), "opt_in" (off by default; enable_thinking=True turns it on) or "adaptive" (the model decides for itself).

report (property)​

How the last generation ended (a GenerationReport), or None before the first one.

reports (property)​

How each reply of the last :meth:generate_batch ended, one :class:ChatOutcome per prompt in prompt order (empty before the first batch).

residency (property)​

What the loaded model holds: the weights part the load left on each device (_Loaded.residency) plus the pipeline part, the KV cache the model's streaming pipeline built (kv_cache_bytes 0 before the first request builds it), in the one document shape.

supports_audio (property)​

supports_speculative_decoding (property)​

True when a request may draft through prompt lookup on this model (prompt_lookup_num_tokens= on :meth:generate, :meth:chat and a session's send): the family scores every fed position in one step and keeps the whole history. The reply's report then carries the tally.

supports_thinking (property)​

True when the model produces a reasoning channel: one a request switches (enable_thinking= / reasoning_effort= on :meth:generate, :meth:chat and :meth:ChatSession.send), one it always opens, or one it opens when it decides to; read from the checkpoint's chat template at load. reasoning_mode says which.

supports_tools (property)​

True when the model takes tool definitions and its replies' calls are read: a chat template whose call syntax a registered format reads.

supports_video (property)​

supports_vision (property)​

template_renders_tools (property)​

True when the chat template itself renders offered tools into the prompt (otherwise the format's tools block joins the system turn).

tokenizer (property)​

The checkpoint's tokenizer (a clika_runtime tokenizer).

tool_call_format (property)​

The format the model's tool calls are read with ("hermes", "mistral", "llama3_json", "pythonic", "qwen_xml", "generic"), probed from the checkpoint's chat template at load.

vocab_size (property)​

__init__​

__init__(self, handle: 'Any') -> 'None'

Initialize self. See help(type(self)) for accurate signature.

chat​

chat(self, messages: 'Sequence[Mapping[str, Any]]', *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'

The assistant's reply to a conversation of {"role", "content"} turns: the text when stream is False, an iterator of text pieces when it is True. tools offers tool definitions for the turn (:meth:ChatSession.send has the shape of the reply); keyword arguments overlay the decode policy for this call (enable_thinking=False switches a thinking model's channel off).

The turns run on the model's chat engine (built on the first call and kept), so consecutive calls reuse its KV cache and prefix cache; the conversation itself is the caller's (pass the whole history each time), or use :meth:chat_session for one that keeps it. report afterwards holds the reply's :class:ChatOutcome. Attached media ride the last user turn as the chat template's parts (the layout generate and the CLI produce); a marker hand-written into the text would count as a second one.

chat_session​

chat_session(self, **options: 'Any') -> "'ChatSession'"

A new :class:ChatSession on this model (its own serving engine; system= seeds the conversation, max_active= sizes the engine).

conversation_text​

conversation_text(self, messages: 'Sequence[Mapping[str, Any]]') -> 'str'

The text the model reads for a conversation: the chat template's rendering when the checkpoint ships one, else the turns flattened as role: content lines with the assistant's turn opened.

generate​

generate(self, prompt: 'str | Sequence[str] | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), max_active: 'int | None' = None, **overrides: 'Any') -> 'str | list[str]'

Generate the reply to prompt (one user message), to messages (a conversation of {"role", "content"} turns), or to a list of prompts (one reply per prompt, in order).

A model with a chat template gets it applied; a model without one reads the prompt as raw text. images ([H, W, 3] tensors), audio ([S] float mono tensors) and videos ((frames, timestamps) pairs) ride the user turn. Keyword arguments overlay the checkpoint's decode defaults (max_new_tokens=64, temperature=0.7, enable_thinking=False to switch a thinking model's channel off for this request, ...). report afterwards says how the reply ended. Requests run through the model's serving pipeline, built on the first request and reused by every later one.

A list of prompts runs as one batch through the model's serving engine (see :meth:generate_batch): the prompts decode together, step by step, instead of one after another; max_active caps how many decode at once (default: all of them).

generate_batch​

generate_batch(self, prompts: 'Sequence[str]', *, config: 'GenerationConfig | Mapping[str, Any] | None' = None, max_active: 'int | None' = None, **overrides: 'Any') -> 'list[str]'

Generate one reply per prompt, the prompts decoding TOGETHER.

Every prompt becomes a session on one continuous-batching engine over this model (the engine chat_session rides), all of them submitted before any reply is read, so each decode step serves every prompt still generating instead of the prompts taking turns. max_active caps how many decode at once (default: every prompt; a smaller cap queues the rest and admits them as others finish). The replies come back in prompt order; reports holds one :class:ChatOutcome per prompt and report the last. An empty list answers an empty list.

render​

render(self, messages: 'Sequence[Mapping[str, Any]]', *, add_generation_prompt: 'bool' = True, tools: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, **overrides: 'Any') -> 'str'

The prompt text the chat template renders for messages ([{"role": ..., "content": ...}, ...]; an assistant turn may carry tool_calls and a tool turn a tool_call_id). tools offers tool definitions (the OpenAI envelope or the bare function object): the template renders them, or a template without a tools slot gets the format's tools block in the system turn. config and the keyword overrides name a request's policy: the render then folds its reasoning switch into the template variables the way the request would (enable_thinking=False renders a toggleable model's block closed; an unknown level, or an off switch on a model without one, raises ValueError); with tools alone, or a policy, the model's default policy folds the same way.

set_assistant​

set_assistant(self, model: "'GenerativeModel | None'") -> 'None'

Attach model as this model's assistant (assisted generation): a second generative model sharing this one's vocabulary and device, with a context window at least this one's (refused otherwise, naming the mismatch); load(..., assistant_model=) does this at the load. A request's num_assistant_tokens then drafts with it, and its report carries the tally. None detaches. Set before the first generation.

set_stage_observer​

set_stage_observer(self, callback: 'StageObserver | None') -> 'None'

Report every request's pipeline stages (tokenize, prefill, decode once per request, detokenize with the final text, vision when a tower runs) to callback, called with a :class:StageEvent on the stage's own thread; it must return at once. The pipeline reads it when it is built, so set it before the first request. None detaches.

stream_generate​

stream_generate(self, prompt: 'str | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), skip_special_tokens: 'bool' = True, decode: 'bool' = True, **overrides: 'Any') -> 'Iterator[str] | Iterator[list[int]]'

Generate as generate does, yielding the reply as the model produces it: decoded text pieces (the default), or with decode=False the token ids themselves, one list per batch the model hands over. Attached media ride the (last) user turn as the chat template's parts, the layout generate and the CLI produce; a marker hand-written into the text would count as a second one.

The model's thread produces token ids into a queue; this iterator drains it, so the calling thread never blocks the model. A text piece is not a token: it may carry several tokens, or the decoder may hold a piece back until a multi-byte character completes, so per-token timing, logprob or id consumers take the id stream and decode it themselves (tokenizer.streaming_decoder is the decoder the text stream uses). Closing the iterator early stops delivery; the generation itself runs to completion. report is set once the reply is complete. The model's serving pipeline is built on the first request (that call pays the build once) and reused after.