ChatSession
A conversation with a text-generation model that keeps its own history.
::
with model.chat_session(system="You are terse.") as chat:
print(chat.send("hi"))
for piece in chat.send("and now?", stream=True):
print(piece, end="", flush=True)
Each ``send`` appends the user turn, runs the whole conversation through
the model's chat engine (a continuous-batching serving pipeline the
session owns), and appends the reply. Turns run one after another on the
session's own thread; closing a streamed reply early cancels its decode
at the next step. ``report`` holds the last reply's :class:`ChatOutcome`.
Tools: ``tools=`` (on the session, or per turn) offers tool definitions
(the OpenAI envelope ``{"type": "function", "function": {...}}`` or the
bare function object). A turn offered tools speaks the user-facing
channel: ``send`` returns the reply's visible text, a streamed reply
yields its text pieces and then one :class:`ToolCall` per call the reply
made, ``tool_calls`` holds them, and the history's assistant turn carries
them as ``tool_calls`` in the message shape (``report.finish`` reads
``"tool_calls"``). :meth:`add_tool_result` appends a tool's answer as a
tool turn, and :meth:`respond` runs the conversation on from there. A
model that takes no tools (``supports_tools`` is False) refuses an offer
with ValueError.
card (property)
What the session's engine advertises about the model (a ChatCard).
closed (property)
engine (property)
The session's chat engine (shareable with :func:serve).
report (property)
How the last reply ended (a ChatOutcome), or None before the first.
tool_calls (property)
The tool calls the last reply made (empty when none, or when the turn offered no tools).
__init__
__init__(self, model: 'GenerativeModel', *, system: 'str | None' = None, model_id: 'str' = '', max_active: 'int' = 1, max_queued: 'int' = 256, default_budget: 'int' = 256, tools: 'Sequence[Mapping[str, Any]] | None' = None) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
add_tool_result
add_tool_result(self, call_id: 'str', content: 'str', *, name: 'str' = '') -> 'None'
Append a tool's answer to the history as a tool turn
({"role": "tool", "tool_call_id": call_id, "content": content},
plus name when given); :meth:respond then runs the model on.
close
close(self) -> 'None'
Cancel the live turn at its next decode step, drop queued ones, and release the engine; later turns raise ValueError.
complete
complete(self, messages: 'Sequence[Mapping[str, Any]]', *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'
The assistant's reply to an explicit conversation, leaving the
session's own history untouched (the reply's shape is :meth:send's).
reset
reset(self, *, keep_system: 'bool' = True) -> 'None'
Forget the conversation (the system turn stays unless told otherwise).
respond
respond(self, *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'
Run the conversation on from where it stands (after
:meth:add_tool_result, or a history edited by hand) and append the
reply; the reply's shape is :meth:send's.
send
send(self, content: 'str', *, stream: 'bool' = False, config: 'GenerationConfig | Mapping[str, Any] | None' = None, images: 'Sequence[clika_runtime.Tensor]' = (), audio: 'Sequence[clika_runtime.Tensor]' = (), videos: 'Sequence[tuple[clika_runtime.Tensor, clika_runtime.Tensor]]' = (), tools: 'Sequence[Mapping[str, Any]] | None' = None, **overrides: 'Any') -> 'str | Iterator[str | ToolCall]'
Add a user turn and return the assistant's reply: the text, or an
iterator of text pieces when stream is True. The reply joins the
history once complete (a streamed reply when its iterator finishes).
tools offers tool definitions for this turn (None: the session's);
the reply then speaks the user-facing channel, a streamed one yields
a :class:ToolCall per call after its text, and tool_calls holds
the calls. Keyword arguments overlay the decode policy for this turn
(enable_thinking=False switches a thinking model's channel off;
reasoning_effort= picks a level).