//clika-runtime/io.clika.modelverse/VisionModel
VisionModel
[common]
class VisionModel : PreTrainedModel
A vision model loaded from its files on disk, on the CPU: a detector, a depth estimator, a text reader, or an embedding model (a model that embeds text and pictures in one space also classifies against labels).
Open it with open from a snapshot directory; capabilities says which tasks the model serves (Capabilities.visionTasks). Hand it a picture with the task's call; the result arrives through the listener on the model's own worker thread, one request at a time per model. close frees the weights.
Types
| Name | Summary |
|---|---|
| Companion | [common] object Companion |
Properties
| Name | Summary |
|---|---|
| capabilities | [common] val capabilities: Capabilities What the model takes and gives, read when it opened. |
| config | [common] open override val config: PretrainedConfig The checkpoint's identity, read when the model opened. |
Functions
| Name | Summary |
|---|---|
| classify | [common] fun classify(image: ImageInput, labels: List<String>, listener: VisionListener): VisionHandle Score image against labels on a model that embeds text and pictures in one space; VisionListener.onScores gets one score per label, descending. |
| close | [common] open fun close() |
| depth | [common] fun depth(image: ImageInput, listener: VisionListener): VisionHandle Estimate the depth of every pixel of image; VisionListener.onDepth gets the map at the picture's resolution. |
| detect | [common] fun detect(image: ImageInput, options: DetectOptions, listener: VisionListener): VisionHandle Find the objects in image under options; VisionListener.onDetections gets the boxes in the picture's pixels, score-descending, then VisionListener.onDone. A model without VisionTask.DETECT, labels given to a detector with its own table, or none given to an open-vocabulary one end in VisionListener.onError. |
| embed | [common] fun embed(image: ImageInput, listener: VisionListener): VisionHandle Embed image; VisionListener.onEmbedding gets its vector of Capabilities.embeddingDim values. |
| recognizeText | [common] fun recognizeText(image: ImageInput, options: OcrOptions, listener: VisionListener): VisionHandle Read the text written in image; VisionListener.onText gets it whole. |
| residency | [common] open override fun residency(): ResidencyReport Free the model. A running request is canceled first and the call waits for it to return; called from inside a listener callback, the free runs right after that callback's request ends instead. Idempotent. |