//clika-runtime/io.clika.modelverse/LoadOptions
LoadOptions
[common]
data class LoadOptions(val contextLength: Int = 0, val maxActive: Int = 0, val thinking: Boolean? = null, val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions
The load-time choices a caller makes for a generative model. Every field has the library's own default; a zero or a null keeps it.
Constructors
| LoadOptions | [common] constructor(contextLength: Int = 0, maxActive: Int = 0, thinking: Boolean? = null, revision: String = "main", cacheDir: String? = null, token: String? = null, offline: Boolean = false, device: String? = null, weights: String? = null, dtype: String? = null) |
Properties
| Name | Summary |
|---|---|
| cacheDir | [common] open override val cacheDir: String? = null |
| contextLength | [common] val contextLength: Int = 0 The context window in tokens: the longest prompt plus reply the model serves. 0 is auto: the largest window that fits the device's free memory beside the weights, bounded by the checkpoint's own; a value names the window, clamped to the checkpoint's. A smaller window shrinks the key-value cache the load allocates, which is what lets a phone load a model whose full window would not fit. A speech model reads it the same way (its decoder's window comes down, never up). |
| device | [common] open override val device: String? = null |
| dtype | [common] open override val dtype: String? = null The precision the weights load at: "float16" or "bfloat16" casts a dense checkpoint's float weights at load (about half the memory of a float32 checkpoint; the cast is best effort, and the stored precision keeps the model as trained) and selects a quantized checkpoint's compute dtype; null keeps the checkpoint's own. A load that reads neither meaning refuses naming the loads that do. The same dtype the command line's --dtype reads. |
| maxActive | [common] val maxActive: Int = 0 The requests the model serves at once: a text model's decode sessions (the key-value cache's slot count a chat session or a served model opens), a speech model's admitted requests. 0 keeps one, the single-session footprint; a phone app keeps it there. |
| offline | [common] open override val offline: Boolean = false |
| revision | [common] open override val revision: String |
| thinking | [common] val thinking: Boolean? = null The thinking channel of a model whose Capabilities.thinking is ThinkingControl.TOGGLE: true renders the prompt with thinking on, false with it off, null keeps the checkpoint's default. A model with no such switch ignores it; a request may override it per call through GenerationConfig.thinking. |
| token | [common] open override val token: String? = null |
| weights | [common] open override val weights: String? = null The weight option of a repository that ships several (a GGUF repository with many quantizations): a quantization tag ( Q4_K_M), a repo-relative .gguf file name, or a glob over the option names. Null picks nothing: a repository with one option needs none, one with several refuses naming them. The same weights the command line's --weights reads. |