Skip to main content

//clika-runtime/io.clika.modelverse/LoadOptions

LoadOptions

[common]
data class LoadOptions(val contextLength: Int = 0, val maxActive: Int = 0, val thinking: Boolean? = null, val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions

The load-time choices a caller makes for a generative model. Every field has the library's own default; a zero or a null keeps it.

Constructors​

LoadOptions[common]
constructor(contextLength: Int = 0, maxActive: Int = 0, thinking: Boolean? = null, revision: String = "main", cacheDir: String? = null, token: String? = null, offline: Boolean = false, device: String? = null, weights: String? = null, dtype: String? = null)

Properties​

NameSummary
cacheDir[common]
open override val cacheDir: String? = null
contextLength[common]
val contextLength: Int = 0
The context window in tokens: the longest prompt plus reply the model serves. 0 is auto: the largest window that fits the device's free memory beside the weights, bounded by the checkpoint's own; a value names the window, clamped to the checkpoint's. A smaller window shrinks the key-value cache the load allocates, which is what lets a phone load a model whose full window would not fit. A speech model reads it the same way (its decoder's window comes down, never up).
device[common]
open override val device: String? = null
dtype[common]
open override val dtype: String? = null
The precision the weights load at: "float16" or "bfloat16" casts a dense checkpoint's float weights at load (about half the memory of a float32 checkpoint; the cast is best effort, and the stored precision keeps the model as trained) and selects a quantized checkpoint's compute dtype; null keeps the checkpoint's own. A load that reads neither meaning refuses naming the loads that do. The same dtype the command line's --dtype reads.
maxActive[common]
val maxActive: Int = 0
The requests the model serves at once: a text model's decode sessions (the key-value cache's slot count a chat session or a served model opens), a speech model's admitted requests. 0 keeps one, the single-session footprint; a phone app keeps it there.
offline[common]
open override val offline: Boolean = false
revision[common]
open override val revision: String
thinking[common]
val thinking: Boolean? = null
The thinking channel of a model whose Capabilities.thinking is ThinkingControl.TOGGLE: true renders the prompt with thinking on, false with it off, null keeps the checkpoint's default. A model with no such switch ignores it; a request may override it per call through GenerationConfig.thinking.
token[common]
open override val token: String? = null
weights[common]
open override val weights: String? = null
The weight option of a repository that ships several (a GGUF repository with many quantizations): a quantization tag (Q4_K_M), a repo-relative .gguf file name, or a glob over the option names. Null picks nothing: a repository with one option needs none, one with several refuses naming them. The same weights the command line's --weights reads.