//clika-runtime/io.clika.modelverse/GenerativeModel
GenerativeModel
[common]
class GenerativeModel : PreTrainedModel
A text model loaded from its files on disk, on the CPU.
Open it with open from a snapshot directory (the layout a Hugging Face download leaves: config.json, the tokenizer files and the weights) or from one GGUF file. Read capabilities to learn what it takes. Send a conversation with chat; the reply streams through the listener on the model's own worker thread, one request at a time per model, and the returned handle cancels it. close frees the weights; a request sent after it ends in GenerationListener.onError.
The chat template is the checkpoint's: the turns are rendered through it, so a system turn, the user turns and the earlier assistant turns reach the model as its publisher intended.
generate is the blocking form of chat, answering the reply's text with report filled in; chatSession keeps a conversation over the model, its tools channel included; render shows the prompt the template writes; and setStageObserver hears the pipeline's stages as they run.
Types
| Name | Summary |
|---|---|
| Companion | [common] object Companion |
Properties
| Name | Summary |
|---|---|
| capabilities | [common] val capabilities: Capabilities What the model takes and gives, read when it opened. |
| config | [common] open override val config: PretrainedConfig The checkpoint's identity, read when the model opened. |
| report | [common] var report: GenerationReport? The last reply's report, from generate, chat or a ChatSession over this model; null before the first reply ends. |
Functions
| Name | Summary |
|---|---|
| chat | [common] fun chat(messages: List<ChatMessage>, config: GenerationConfig, listener: GenerationListener): GenerationHandle Stream a reply to messages under config. Returns at once; the worker runs the request and calls listener on its thread. A second chat while one runs waits for the first to end. After close the listener receives GenerationListener.onError and nothing else. |
| chatSession | [common] fun chatSession(system: String? = null, tools: List<ToolDef> = emptyList()): ChatSession A conversation over this model, opened with system as its system turn when given and tools offered on every turn (a turn may offer its own). |
| close | [common] open fun close() |
| generate | [common] fun generate(prompt: String? = null, messages: List<ChatMessage>? = null, config: GenerationConfig = capabilities.defaults, images: List<ByteArray> = emptyList(), tools: List<ToolDef> = emptyList(), toolChoice: ToolChoice = ToolChoice.Auto, stageObserver: StageObserver? = null): String One reply, blocking: prompt as the user turn (with images), or messages as the conversation; exactly one of the two. A model with a chat template renders it, one without reads the prompt as text. Answers the reply's visible text; report then says how it ended and what it cost. tools are offered under toolChoice, and the calls the reply made are on the report. stageObserver hears this call's stages alone, beside the model's own (setStageObserver). A failure throws ModelverseException with the library's message (an offer of tools to a model that takes none among them); a close while the reply runs ends it as FinishReason.CANCELED with the text so far. |
| render | [common] fun render(messages: List<ChatMessage>, tools: List<ToolDef> = emptyList()): String The prompt the chat template writes for messages with tools offered, as a request would send it (the generation prompt appended); a model without a template answers the turns' text joined by newlines. |
| residency | [common] open override fun residency(): ResidencyReport Free the model. A running request is canceled first and the call waits for it to return; called from inside a listener callback, the free runs right after that callback's request ends instead. Idempotent. |
| setStageObserver | [common] fun setStageObserver(observer: StageObserver?) Hear every request's pipeline stages through observer (StageEvent), on the stage's own thread; null detaches it. Takes effect at once, for a request already running too; an event already in flight may still reach the observer it replaced. |