---
title: "GenerativeModel"
sidebar_label: "GenerativeModel"
description: "The clika_runtime.modelverse GenerativeModel class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

A loaded text-generation model.

    ``generate`` returns the reply to one prompt (or to a conversation
    passed as ``messages``); ``stream_generate`` yields the reply piece by
    piece as the model produces it. Decode knobs ride keyword arguments over
    the checkpoint's defaults (``max_new_tokens``, ``temperature``,
    ``top_p``, ...; see :class:`GenerationConfig`). ``report`` holds how the
    last generation ended.
    

## `assistant_model` (property)

The attached assistant model, or ``None``.

## `compute_dtype` (property)

The dtype the model computes at (Undefined when the checkpoint's own applies).

## `defaults` (property)

The checkpoint's decode defaults.

## `eos_token_ids` (property)

The ids that end a sequence.

## `has_chat_template` (property)

True when the tokenizer ships a chat template.

## `kv_cache_mode` (property)

The KV-cache storage the load named ("paged" or "continuous").

None when the load named nothing: the streaming pipeline and
``generate`` then run one continuous slot, the one-session default.

## `local_dir` (property)

The resolved snapshot directory the model loaded from.

## `max_seq` (property)

The context window in tokens.

## `reasoning_mode` (property)

How the reasoning channel behaves: ``"never"`` (no channel),
``"always"`` (no off switch), ``"toggleable"`` (on by default;
``enable_thinking=False`` turns it off), ``"opt_in"`` (off by default;
``enable_thinking=True`` turns it on) or ``"adaptive"`` (the model
decides for itself).

## `report` (property)

How the last generation ended (a GenerationReport), or None before
the first one.

## `reports` (property)

How each reply of the last :meth:`generate_batch` ended, one
:class:`ChatOutcome` per prompt in prompt order (empty before the
first batch).

## `residency` (property)

What the loaded model holds: the weights part the load left on each
device (``_Loaded.residency``) plus the ``pipeline`` part, the KV cache
the model's streaming pipeline built (``kv_cache_bytes`` 0 before the
first request builds it), in the one document shape.

## `supports_audio` (property)



## `supports_speculative_decoding` (property)

True when a request may draft through prompt lookup on this model
(``prompt_lookup_num_tokens=`` on :meth:`generate`, :meth:`chat` and
a session's ``send``): the family scores every fed position in one
step and keeps the whole history. The reply's ``report`` then carries
the tally.

## `supports_thinking` (property)

True when the model produces a reasoning channel: one a request
switches (``enable_thinking=`` / ``reasoning_effort=`` on
:meth:`generate`, :meth:`chat` and :meth:`ChatSession.send`), one it
always opens, or one it opens when it decides to; read from the
checkpoint's chat template at load. ``reasoning_mode`` says which.

## `supports_tools` (property)

True when the model takes tool definitions and its replies' calls
are read: a chat template whose call syntax a registered format
reads.

## `supports_video` (property)



## `supports_vision` (property)



## `template_renders_tools` (property)

True when the chat template itself renders offered tools into the
prompt (otherwise the format's tools block joins the system turn).

## `tokenizer` (property)

The checkpoint's tokenizer (a clika_runtime tokenizer).

## `tool_call_format` (property)

The format the model's tool calls are read with ("hermes",
"mistral", "llama3_json", "pythonic", "qwen_xml", "generic"), probed
from the checkpoint's chat template at load.

## `vocab_size` (property)



## `__init__`

```python
__init__(self, handle: 'Any') -> 'None'
```

Initialize self.  See help(type(self)) for accurate signature.

## `chat`

```python
chat(self, messages: 'Sequence[Mapping[str, Any]]', *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'
```

The assistant's reply to a conversation of ``{"role", "content"}``
turns: the text when ``stream`` is False, an iterator of text pieces
when it is True. ``tools`` offers tool definitions for the turn
(:meth:`ChatSession.send` has the shape of the reply); keyword
arguments overlay the decode policy for this call
(``enable_thinking=False`` switches a thinking model's channel off).

The turns run on the model's chat engine (built on the first call and
kept), so consecutive calls reuse its KV cache and prefix cache; the
conversation itself is the caller's (pass the whole history each
time), or use :meth:`chat_session` for one that keeps it. ``report``
afterwards holds the reply's :class:`ChatOutcome`. Attached media ride
the last user turn as the chat template's parts (the layout
``generate`` and the CLI produce); a marker hand-written into the text
would count as a second one.

## `chat_session`

```python
chat_session(self, **options: 'Any') -> "'ChatSession'"
```

A new :class:`ChatSession` on this model (its own serving engine;
``system=`` seeds the conversation, ``max_active=`` sizes the engine).

## `conversation_text`

```python
conversation_text(self, messages: 'Sequence[Mapping[str, Any]]') -> 'str'
```

The text the model reads for a conversation: the chat template's
rendering when the checkpoint ships one, else the turns flattened as
``role: content`` lines with the assistant's turn opened.

## `generate`

```python
generate(self, prompt: 'str | Sequence[str] | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), max_active: 'int | None' = None, **overrides: 'Any') -> 'str | list[str]'
```

Generate the reply to ``prompt`` (one user message), to
``messages`` (a conversation of ``{"role", "content"}`` turns), or to
a list of prompts (one reply per prompt, in order).

A model with a chat template gets it applied; a model without one
reads the prompt as raw text. ``images`` ([H, W, 3] tensors),
``audio`` ([S] float mono tensors) and ``videos`` ((frames,
timestamps) pairs) ride the user turn. Keyword arguments overlay the
checkpoint's decode defaults (``max_new_tokens=64``,
``temperature=0.7``, ``enable_thinking=False`` to switch a thinking
model's channel off for this request, ...). ``report`` afterwards says how the reply
ended. Requests run through the model's serving pipeline, built on
the first request and reused by every later one.

A list of prompts runs as one batch through the model's serving
engine (see :meth:`generate_batch`): the prompts decode together,
step by step, instead of one after another; ``max_active`` caps how
many decode at once (default: all of them).

## `generate_batch`

```python
generate_batch(self, prompts: 'Sequence[str]', *, config: 'GenerationConfig | Mapping[str, Any] | None' = None, max_active: 'int | None' = None, **overrides: 'Any') -> 'list[str]'
```

Generate one reply per prompt, the prompts decoding TOGETHER.

Every prompt becomes a session on one continuous-batching engine over
this model (the engine ``chat_session`` rides), all of them submitted
before any reply is read, so each decode step serves every prompt
still generating instead of the prompts taking turns. ``max_active``
caps how many decode at once (default: every prompt; a smaller cap
queues the rest and admits them as others finish). The replies come
back in prompt order; ``reports`` holds one :class:`ChatOutcome` per
prompt and ``report`` the last. An empty list answers an empty list.

## `render`

```python
render(self, messages: 'Sequence[Mapping[str, Any]]', *, add_generation_prompt: 'bool' = True, tools: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, **overrides: 'Any') -> 'str'
```

The prompt text the chat template renders for ``messages``
(``[{"role": ..., "content": ...}, ...]``; an assistant turn may carry
``tool_calls`` and a ``tool`` turn a ``tool_call_id``). ``tools``
offers tool definitions (the OpenAI envelope or the bare function
object): the template renders them, or a template without a tools
slot gets the format's tools block in the system turn. ``config`` and
the keyword overrides name a request's policy: the render then folds
its reasoning switch into the template variables the way the request
would (``enable_thinking=False`` renders a toggleable model's block
closed; an unknown level, or an off switch on a model without one,
raises ``ValueError``); with tools alone, or a policy, the model's
default policy folds the same way.

## `set_assistant`

```python
set_assistant(self, model: "'GenerativeModel | None'") -> 'None'
```

Attach ``model`` as this model's assistant (assisted generation):
a second generative model sharing this one's vocabulary and device,
with a context window at least this one's (refused otherwise, naming
the mismatch); ``load(..., assistant_model=)`` does this at the load.
A request's ``num_assistant_tokens`` then drafts with it, and its
``report`` carries the tally. ``None`` detaches. Set before the first
generation.

## `set_stage_observer`

```python
set_stage_observer(self, callback: 'StageObserver | None') -> 'None'
```

Report every request's pipeline stages (``tokenize``, ``prefill``,
``decode`` once per request, ``detokenize`` with the final text,
``vision`` when a tower runs) to ``callback``, called with a
:class:`StageEvent` on the stage's own thread; it must return at
once. The pipeline reads it when it is built, so set it before the
first request. ``None`` detaches.

## `stream_generate`

```python
stream_generate(self, prompt: 'str | None' = None, *, messages: 'Sequence[Mapping[str, Any]] | None' = None, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), skip_special_tokens: 'bool' = True, decode: 'bool' = True, **overrides: 'Any') -> 'Iterator[str] | Iterator[list[int]]'
```

Generate as ``generate`` does, yielding the reply as the model
produces it: decoded text pieces (the default), or with
``decode=False`` the token ids themselves, one list per batch the
model hands over. Attached media ride the (last) user turn as the
chat template's parts, the layout ``generate`` and the CLI produce;
a marker hand-written into the text would count as a second one.

The model's thread produces token ids into a queue; this iterator
drains it, so the calling thread never blocks the model. A text piece
is not a token: it may carry several tokens, or the decoder may hold
a piece back until a multi-byte character completes, so per-token
timing, logprob or id consumers take the id stream and decode it
themselves (``tokenizer.streaming_decoder`` is the decoder the text
stream uses). Closing the iterator early stops delivery; the
generation itself runs to completion. ``report`` is set once the
reply is complete. The model's serving pipeline is built on the
first request (that call pays the build once) and reused after.
