Skip to main content

//clika-runtime/io.clika.modelverse/Capabilities

Capabilities

[common]
data class Capabilities(val inputModalities: List<String>, val outputModalities: List<String>, val hasChatTemplate: Boolean, val contextLength: Int, val maxImagesPerRequest: Int, val thinking: ThinkingControl, val defaults: GenerationConfig, val family: String, val modelType: String, val device: String, val voices: List<String> = emptyList(), val languages: List<String> = emptyList(), val needsReference: Boolean = false, val speakKnobs: List<SpeakKnob> = emptyList(), val visionTasks: List<String> = emptyList(), val labels: List<String> = emptyList(), val openVocabulary: Boolean = false, val embeddingDim: Int = 0, val defaultImageWidth: Int = 0, val defaultImageHeight: Int = 0, val defaultFrames: Int = 0, val defaultFps: Double = 0.0, val metricDepth: Boolean = false, val supportsRepetitionKnobs: Boolean = false, val supportsTools: Boolean = false, val toolCallFormat: String = "", val templateRendersTools: Boolean = false)

What a loaded model can take and give, read from the checkpoint's documents and the family's registration when the model opens. A consumer greys out the controls a model cannot serve from this and nothing else.

Constructors​

Capabilities[common]
constructor(inputModalities: List<String>, outputModalities: List<String>, hasChatTemplate: Boolean, contextLength: Int, maxImagesPerRequest: Int, thinking: ThinkingControl, defaults: GenerationConfig, family: String, modelType: String, device: String, voices: List<String> = emptyList(), languages: List<String> = emptyList(), needsReference: Boolean = false, speakKnobs: List<SpeakKnob> = emptyList(), visionTasks: List<String> = emptyList(), labels: List<String> = emptyList(), openVocabulary: Boolean = false, embeddingDim: Int = 0, defaultImageWidth: Int = 0, defaultImageHeight: Int = 0, defaultFrames: Int = 0, defaultFps: Double = 0.0, metricDepth: Boolean = false, supportsRepetitionKnobs: Boolean = false, supportsTools: Boolean = false, toolCallFormat: String = "", templateRendersTools: Boolean = false)

Properties​

NameSummary
contextLength[common]
val contextLength: Int
The context window the model loaded with, in tokens.
defaultFps[common]
val defaultFps: Double = 0.0
defaultFrames[common]
val defaultFrames: Int = 0
The frame count and the frame rate a video-generation request of zeros takes (a video-generation model; 0 otherwise).
defaultImageHeight[common]
val defaultImageHeight: Int = 0
defaultImageWidth[common]
val defaultImageWidth: Int = 0
The picture size an image-generation request of 0 by 0 takes (an image-generation model; 0 otherwise).
defaults[common]
val defaults: GenerationConfig
The sampling defaults the checkpoint ships (generation_config.json); a zero temperature decodes greedily.
device[common]
val device: String
The device the weights landed on, as the compute API's name (CPU; an accelerator among several carries its ordinal, CUDA:1).
embeddingDim[common]
val embeddingDim: Int = 0
The width of the vectors an embedding model answers; 0 for every other model.
family[common]
val family: String
The model family as the library names it (qwen, llama).
hasChatTemplate[common]
val hasChatTemplate: Boolean
True when the chat route renders a template: the checkpoint's own, or its family's.
inputModalities[common]
val inputModalities: List<String>
The modalities the model's requests may carry, in the library's own spellings: text, audio, vision, video, bounding_boxes, embedding, depth, segmentation, ranking, keypoints. A model whose inputs carry vision takes images in its turns (ChatMessage.images); maxImagesPerRequest says how many one request may carry.
labels[common]
val labels: List<String>
A detector's own label table, in id order; empty for an open-vocabulary detector (its labels are the request's) and for every other model.
languages[common]
val languages: List<String>
The language codes a speech model declares; empty when it declares none.
maxImagesPerRequest[common]
val maxImagesPerRequest: Int
The most images one request may carry across all of its turns: the model's image slots for a model that takes images, 0 for one that does not. A request past it is refused before any decode.
metricDepth[common]
val metricDepth: Boolean = false
True when a depth model's values are meters; false for relative depth (larger is closer).
modelType[common]
val modelType: String
The checkpoint's own type (qwen3, llama).
needsReference[common]
val needsReference: Boolean = false
True when a text-to-speech model conditions every utterance on a reference voice (TtsModel.speak with a VoiceReference); a request without one is refused and ends in SpeakListener.onError. False for a model that speaks its own voice, and for every model that is not a text-to-speech model.
openVocabulary[common]
val openVocabulary: Boolean = false
True for a detector that looks for the phrases a request names (DetectOptions.labels).
outputModalities[common]
val outputModalities: List<String>
The modalities a reply carries, in the same spellings (text).
speakKnobs[common]
val speakKnobs: List<SpeakKnob>
The synthesis knobs a text-to-speech model reads, each with the range the binding accepts and the value that keeps the family's own default; empty for every other model.
supportsRepetitionKnobs[common]
val supportsRepetitionKnobs: Boolean = false
True when the library carries the repetition knobs (GenerationConfig.repetitionPenalty, GenerationConfig.noRepeatNgramSize); false refuses any value but their defaults.
supportsTools[common]
val supportsTools: Boolean = false
True when the model takes tool definitions and its replies' calls are read: it has a chat template whose call syntax a known format reads.
templateRendersTools[common]
val templateRendersTools: Boolean = false
True when the chat template renders the offered tools itself (a tools block in the prompt).
thinking[common]
val thinking: ThinkingControl
The reasoning controls the model admits.
toolCallFormat[common]
val toolCallFormat: String
The call format stamped at the load: the first of the family's formats the chat template writes (hermes, mistral, llama3_json, pythonic, qwen_xml), else generic, the fallback reader, a model without a template included; whether the model takes tools is supportsTools.
visionTasks[common]
val visionTasks: List<String>
The tasks a vision model serves, as the VisionTask constants spell them (VisionTask.DETECT, VisionTask.DEPTH, VisionTask.OCR, VisionTask.EMBED, VisionTask.CLASSIFY); a VisionModel.detect on a model whose list lacks the task ends in VisionListener.onError. Empty for every other model.
voices[common]
val voices: List<String>
The preset voices a text-to-speech model names; empty for a model that clones the reference voice it is given, and for every model that is not a speech model.