Skip to main content

//clika-runtime/io.clika.modelverse/VisionModel

VisionModel

[common]
class VisionModel : PreTrainedModel

A vision model loaded from its files on disk, on the CPU: a detector, a depth estimator, a text reader, or an embedding model (a model that embeds text and pictures in one space also classifies against labels).

Open it with open from a snapshot directory; capabilities says which tasks the model serves (Capabilities.visionTasks). Hand it a picture with the task's call; the result arrives through the listener on the model's own worker thread, one request at a time per model. close frees the weights.

Types​

NameSummary
Companion[common]
object Companion

Properties​

NameSummary
capabilities[common]
val capabilities: Capabilities
What the model takes and gives, read when it opened.
config[common]
open override val config: PretrainedConfig
The checkpoint's identity, read when the model opened.

Functions​

NameSummary
classify[common]
fun classify(image: ImageInput, labels: List<String>, listener: VisionListener): VisionHandle
Score image against labels on a model that embeds text and pictures in one space; VisionListener.onScores gets one score per label, descending.
close[common]
open fun close()
depth[common]
fun depth(image: ImageInput, listener: VisionListener): VisionHandle
Estimate the depth of every pixel of image; VisionListener.onDepth gets the map at the picture's resolution.
detect[common]
fun detect(image: ImageInput, options: DetectOptions, listener: VisionListener): VisionHandle
Find the objects in image under options; VisionListener.onDetections gets the boxes in the picture's pixels, score-descending, then VisionListener.onDone. A model without VisionTask.DETECT, labels given to a detector with its own table, or none given to an open-vocabulary one end in VisionListener.onError.
embed[common]
fun embed(image: ImageInput, listener: VisionListener): VisionHandle
Embed image; VisionListener.onEmbedding gets its vector of Capabilities.embeddingDim values.
recognizeText[common]
fun recognizeText(image: ImageInput, options: OcrOptions, listener: VisionListener): VisionHandle
Read the text written in image; VisionListener.onText gets it whole.
residency[common]
open override fun residency(): ResidencyReport
Free the model. A running request is canceled first and the call waits for it to return; called from inside a listener callback, the free runs right after that callback's request ends instead. Idempotent.