//clika-runtime/io.clika.modelverse
Package-level declarations
Types
| Name | Summary |
|---|---|
| AutoConfig | [common] object AutoConfig The identity of the checkpoint at source, from its configuration alone. |
| AutomaticSpeechRecognitionPipeline | [common] class AutomaticSpeechRecognitionPipeline : Pipeline Speech recognition over an SttModel: a clip's mono 16-bit samples answer its transcript. |
| AutoModel | [common] object AutoModel Load any registered model by name: the family's declared modalities pick the door, or task names it (Task). The answer is the door's handle, a PreTrainedModel. A door this binding serves no handle for yet (image segmentation) refuses with UNSUPPORTED, naming the task. |
| AutoModelForCausalLM | [common] object AutoModelForCausalLM A text-generation model (chat included) by name. |
| AutoModelForDepthEstimation | [common] object AutoModelForDepthEstimation A monocular depth model by name. |
| AutoModelForImageTextToText | [common] object AutoModelForImageTextToText A multimodal chat model (an image, audio or video in the prompt) by name. |
| AutoModelForObjectDetection | [common] object AutoModelForObjectDetection A closed-set object detector by name. |
| AutoModelForSeq2SeqLM | [common] object AutoModelForSeq2SeqLM A translation (sequence-to-sequence) model by name. |
| AutoModelForSpeakerDiarization | [common] object AutoModelForSpeakerDiarization A speaker-diarization model (who spoke when, over a clip) by name. |
| AutoModelForSpeechSeq2Seq | [common] object AutoModelForSpeechSeq2Seq A speech-to-text model by name. |
| AutoModelForTextEmbedding | [common] object AutoModelForTextEmbedding A text-embedding model (one vector per text) by name. |
| AutoModelForTextRanking | [common] object AutoModelForTextRanking A reranking model (one relevance score per document against a query) by name. |
| AutoModelForTextToImage | [common] object AutoModelForTextToImage An image-generation model (pictures out of a prompt) by name. |
| AutoModelForTextToVideo | [common] object AutoModelForTextToVideo A video-generation model (a clip out of a prompt) by name. |
| AutoModelForTextToWaveform | [common] object AutoModelForTextToWaveform A text-to-speech model by name. |
| AutoModelForZeroShotObjectDetection | [common] object AutoModelForZeroShotObjectDetection An open-vocabulary detector (a text query) by name. |
| AutoProcessor | [common] object AutoProcessor The input processor of the checkpoint at source: the composite of its tokenizer and media processors, from the snapshot's processor files (a hub source fetches them without the weights). |
| AutoTokenizer | [common] object AutoTokenizer The tokenizer of the checkpoint at source (a local directory or a hub repo id; a hub source fetches the configuration and tokenizer files alone), the same tokenizer a loaded model of the checkpoint uses. |
| BackendReport | [common] data class BackendReport(val api: String, val compiled: Boolean, val available: Boolean, val deviceCount: Int, val unavailableReason: String, val devices: List<DeviceReport>) One compute backend of the runtime, as the hardware inventory lists it. |
| BenchCacheMode | [common] enum BenchCacheMode : Enum<BenchCacheMode> The key-value cache the bench's decoders run on: the paged pool (the serving default, whose prefix cache the admittedTok and retires columns report) or one window per session. |
| BenchCell | [common] data class BenchCell(val isl: Long, val osl: Long, val lateIsl: Long = 0) One cell of the sweep (Modelverse.bench): isl prompt tokens in, osl tokens decoded, both exact (the prompts are synthesized in the model's own vocabulary at that length and the decode is pinned to the count). A positive lateIsl adds one more session with a prompt of that many tokens that joins once every regular session has streamed its first token, the shape of a long prompt arriving while others decode; the row then reports what its prefill cost the running sessions ( itlMaxMs, itlP99Ms) and what the prompt waited itself (lateTtftMs). The command line spells a cell ISL:OSL or ISL:OSL+LATE. |
| BenchColumns | [common] object BenchColumns The column names of a bench report's table, the library's own spellings: the columns of BenchReport.table, the keys of BenchRow.cells and the header of the CSV text. |
| BenchConfigKeys | [common] object BenchConfigKeys The keys of a bench report's configuration document (BenchReport.config), the library's own spellings. |
| BenchOptions | [common] data class BenchOptions(val cells: List<BenchCell>? = null, val concurrency: List<Int>? = null, val warmup: Int? = null, val iters: Int? = null, val cacheMode: BenchCacheMode? = null, val seed: Long? = null, val kvQuant: String? = null, val prefillChunkTokens: Long? = null, val stepTokenBudget: Long? = null, val sharedPrefixTokens: Long? = null, val promptLookupNumTokens: Long? = null, val uniquePrompts: Boolean? = null, val profile: Boolean? = null, val profilePipeline: Boolean? = null, val profileDir: String? = null, val reportPrefix: String? = null, val maxKvCacheBytes: Long? = null) The knobs of Modelverse.bench, the command line's bench verb's and the Python modelverse.bench.run's: every field null keeps the library's own default (the one the command line prints in its --help), so a BenchOptions() runs the library's single-cell sweep at concurrency 1. |
| BenchReport | [common] class BenchReport One bench's record (Modelverse.bench), the library's own BenchReport: table, the runtime's columnar frame with one row per (cell, concurrency) in the sweep's order under the columns of BenchColumns (the same frame the command line's bench verb prints as CSV and writes as PREFIX.csv; toCsvString renders it); config, the run's configuration document under the keys of BenchConfigKeys (the model, the device, the cache mode and the knobs, the machine's hardware inventory, the checkpoint's on-disk bytes and the loaded model's resident bytes; what the verb writes as PREFIX.json); profileSummary, the per-op text of a profiled run (empty otherwise); and rows, the table read into typed BenchRows. The table and the document are native objects: close releases them (use {}), and the typed rows stay readable after it. |
| BenchRow | [common] data class BenchRow(val isl: Long, val osl: Long, val lateIsl: Long, val concurrency: Int, val status: BenchRowStatus, val reason: String, val notes: String, val firstCallMs: Double, val ttftMs: Double, val lateTtftMs: Double, val prefillTokS: Double, val prefillChunkMsP50: Double, val prefillChunkMsMax: Double, val itlP50Ms: Double, val itlP99Ms: Double, val itlMaxMs: Double, val decodeTokS: Double, val speedupVsC1: Double, val batchingEfficiency: Double, val effectiveIsl: Long, val admittedTok: Long, val retires: Long, val peakActiveBytes: Long, val numAllocs: Long, val numInflightParkWaits: Long, val acceptedPredictionTokens: Long, val rejectedPredictionTokens: Long, val speculativeSteps: Long, val cells: Map<String, Any?>) One row of a bench report, typed: one (cell, concurrency) of the sweep. status says what the row is; the metrics of a row that is not BenchRowStatus.OK are zeros, never measurements, and reason says why in the library's words. The times are milliseconds, the rates tokens per second. cells carries every column of the row as the table holds it, by the names of BenchColumns: a Long for an integer column, a Double for a float one, a String for a text one, null for an empty cell; a column a later library adds is there before it has a typed field. The rows are read off BenchReport.table once, so a row outlives the report's close. |
| BenchRowStatus | [common] enum BenchRowStatus : Enum<BenchRowStatus> What a report row is: a measurement, a cell the loaded window cannot fit, or a run whose pipeline failed. |
| Box | [common] data class Box(val left: Float, val top: Float, val right: Float, val bottom: Float) A rectangle in the input picture's pixels, top-left origin, left and top the inclusive edges. |
| Capabilities | [common] data class Capabilities(val inputModalities: List<String>, val outputModalities: List<String>, val hasChatTemplate: Boolean, val contextLength: Int, val maxImagesPerRequest: Int, val thinking: ThinkingControl, val defaults: GenerationConfig, val family: String, val modelType: String, val device: String, val voices: List<String> = emptyList(), val languages: List<String> = emptyList(), val needsReference: Boolean = false, val speakKnobs: List<SpeakKnob> = emptyList(), val visionTasks: List<String> = emptyList(), val labels: List<String> = emptyList(), val openVocabulary: Boolean = false, val embeddingDim: Int = 0, val defaultImageWidth: Int = 0, val defaultImageHeight: Int = 0, val defaultFrames: Int = 0, val defaultFps: Double = 0.0, val metricDepth: Boolean = false, val supportsRepetitionKnobs: Boolean = false, val supportsTools: Boolean = false, val toolCallFormat: String = "", val templateRendersTools: Boolean = false) What a loaded model can take and give, read from the checkpoint's documents and the family's registration when the model opens. A consumer greys out the controls a model cannot serve from this and nothing else. |
| ChatMessage | [common] data class ChatMessage(val role: String, val content: String, val images: List<ByteArray> = emptyList(), val toolCalls: List<ToolCall> = emptyList(), val toolCallId: String = "", val name: String = "") One turn of a conversation. role is system, user or assistant; content the text; images the encoded image files (PNG or JPEG bytes) the turn carries, decoded by the library and placed in the turn where the model's template puts them. A model that takes no images refuses a turn with one, and a request carrying more than Capabilities.maxImagesPerRequest across its turns is refused; either request ends in GenerationListener.onError. An assistant turn that called tools carries them as toolCalls; a tool's result is a ROLE_TOOL turn naming the call it answers (toolCallId). |
| ChatSession | [common] class ChatSession A conversation over a GenerativeModel: the turns so far (messages), a send that appends the user turn and the reply, a respond that replies to the turns as they stand, the tools channel (tools, toolCalls, addToolResult) and the last reply's report. The whole conversation goes to the model on every turn, rendered through the checkpoint's chat template; the model's cache serves the earlier turns (GenerationReport.cachedPromptTokens). One session speaks at a time, on the model's own worker thread; a second session over the same model waits its turn. The session owns no weights: close ends it and leaves the model open. Built by GenerativeModel.chatSession. |
| CheckOptions | [common] data class CheckOptions(val device: String? = null, val maxSeq: Long? = null, val imageWidth: Int = 0, val imageHeight: Int = 0, val devices: List<FitDevice> = emptyList(), val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val weights: String? = null, val dtype: String? = null) : SourceOptions The knobs of Modelverse.check, the command line's check verb's and the Python modelverse.check's. device is the device judged: a label the devices listing prints (CPU:0), the runtime's spelling (cpu, vulkan:1), an api name alone (cuda, every device of that api), or a label of devices; null judges every device. maxSeq is the sequence length judged, null the model's own context length. imageWidth and imageHeight are the image an image generation pipeline is judged at (0 = its weights alone, with the largest square image that fits reported). devices names devices to judge instead of this machine's (a phone's memory, a step's remaining budget); empty takes this machine's inventory. The hub knobs (revision, cacheDir, token, offline, weights) are the ones every load reads; nothing is fetched beyond the model's small documents and its file listing, and nothing runs. |
| DepthEstimationPipeline | [common] class DepthEstimationPipeline : Pipeline Depth estimation over a VisionModel: a picture answers its depth map. |
| DepthMap | [common] data class DepthMap(val width: Int, val height: Int, val values: FloatArray, val metric: Boolean, val min: Float, val max: Float) A depth value per pixel of the input picture, row by row (width × height values): meters when metric, else relative depth where a larger value is closer; min and max are the map's own range. |
| Detection | [common] data class Detection(val label: String, val labelId: Long, val score: Float, val box: Box) One detected object: the label (empty when the model's table names none for the id), its id, the score in 0 to 1, and the box in the input picture's pixels. |
| DetectOptions | [common] data class DetectOptions(val maxResults: Int = DEFAULT_MAX_RESULTS, val scoreThreshold: Float = DEFAULT_SCORE_THRESHOLD, val labels: List<String> = emptyList()) How a detection runs. The engine floors the scores at scoreThreshold, keeps the strongest maxResults and answers them score-descending. |
| DeviceMemory | [common] data class DeviceMemory(val device: String, val totalBytes: Long, val freeBytes: Long, val runtimeBytes: Long, val activeBytes: Long) A device's memory as the runtime reads it (Modelverse.deviceMemory): the physical total, what is free, the runtime's own share and the pool bytes in use. |
| DeviceReport | [common] data class DeviceReport(val index: Int, val name: String, val totalMemoryBytes: Long) One compute device of a backend, as the hardware inventory lists it. |
| DiarizationModel | [common] class DiarizationModel : PreTrainedModel A speaker-diarization model: a clip in, who spoke when out, as SpeakerSegments in time order (the voice-activity-detection kind). |
| DiarizeHandle | [common] interface DiarizeHandle A speaker-diarization request in flight. |
| DiarizeListener | [common] interface DiarizeListener The calls a speaker-diarization request makes on the model's worker thread, in order: onSegments once with who spoke when, in time order, then onDone; or onError alone when the request fails. |
| DiarizeOptions | [common] data class DiarizeOptions(val streaming: Boolean = false, val streamingMode: String? = null, val threshold: Float = 0.5f) How a diarization runs. streaming runs the clip chunk by chunk with the speaker cache, as a live stream does, streamingMode naming the latency mode (the checkpoint's default when null); threshold in (0, 1) is the activity a frame needs to count for a speaker. |
| DiarizeReport | [common] data class DiarizeReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The facts of one speaker-diarization request. |
| DownloadEvent | [common] data class DownloadEvent(val file: String, val index: Int, val count: Int, val done: Long, val total: Long, val planBytesDone: Long, val planBytesTotal: Long, val isWeight: Boolean) One step of a download (DownloadListener): the file transferring (repo-relative), its position in the plan so far (index of count; the plan grows as the resolve names its missing files), its bytes, and the plan's bytes done and total (planBytesTotal is 0 while a planned file's size is unknown). A cached file never joins the plan. |
| DownloadListener | [common] fun interface DownloadListener Hears a download's progress (Modelverse.snapshot): one call per transfer step of each file the resolve fetches, on the calling thread, with the file's bytes and the plan's. Return false to stop the download: the call then refuses with the hub's own cancel status; an exception the listener lets escape stops it the same way. A snapshot served from the cache reports nothing, since a cached file never joins the plan. |
| EmbeddingHandle | [common] interface EmbeddingHandle A text-embedding request in flight. |
| EmbeddingListener | [common] interface EmbeddingListener The calls a text-embedding request makes on the model's worker thread, in order: onEmbeddings once with one vector per text (in the texts' order), then onDone; or onError alone when the request fails. |
| EmbeddingModel | [common] class EmbeddingModel : PreTrainedModel A text-embedding model: one vector per text, the same length for every text (Capabilities.embeddingDim), so two texts compare by the cosine of their vectors. |
| EmbeddingReport | [common] data class EmbeddingReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The facts of one text-embedding request. |
| FeatureExtractionPipeline | [common] class FeatureExtractionPipeline : Pipeline Text embeddings over an EmbeddingModel: a text answers its vector, a list of texts one vector each, in order ( feature-extraction and its alias sentence-similarity). |
| FinishReason | [common] enum FinishReason : Enum<FinishReason> How a generation ended. |
| FitColumns | [common] object FitColumns The column names of FitReport.table, the library's own spellings. |
| FitDevice | [common] data class FitDevice(val label: String, val totalBytes: Long, val cores: Int = 0) One device to judge a model against: label as the devices listing prints it (CPU:0, Vulkan:1), or any name the caller chooses for a device that is not this machine (a phone, or the memory left for one step of a pipeline after the earlier steps took theirs); totalBytes its memory; cores the CPU cores it decodes with (0 = unknown, no many-cores bonus in the speed estimate). |
| FitGrade | [common] enum FitGrade : Enum<FitGrade> How comfortably a model's need sits in a device's memory, read from the memory share at the admitted context: at most 60 percent is PERFECT, at most 85 GOOD, at most 98 MARGINAL, above it TOO_TIGHT (no room left for the allocator), and a need no context of the device holds is DOES_NOT_FIT. word is the library's own spelling, the one the command line's check verb prints. |
| FitReport | [common] class FitReport The answer of Modelverse.check: verdicts, one per device judged, in the request's order; and table, the runtime's Table the command line's check verb prints (one row per verdict, sorted by grade then memory share, the columns of FitColumns; toCsvString renders it). The table is a native object: close releases it (use {}); the verdicts stay. |
| FitVerdict | [common] data class FitVerdict(val source: String, val device: String, val deviceBytes: Long, val weightBytes: Long, val kvBytesFixed: Long, val kvBytesPerToken: Long, val kvDtype: String, val contextLength: Long, val maxSeq: Long, val activationBytes: Long, val needBytes: Long, val fits: Boolean, val maxSeqFits: Long, val maxImageFits: Int, val reason: String, val judged: Boolean, val grade: FitGrade?, val memorySharePercent: Int, val parameters: Long, val weightEncoding: String, val tokSEstimate: Double, val tokSMeasured: Double) Does a model fit a device, judged from its metadata alone, the library's own FitVerdict. The law, every term a field: |
| GeneratedImage | [common] data class GeneratedImage(val width: Int, val height: Int, val rgba8888: ByteArray) One generated picture: width by height pixels, four bytes each in the order red, green, blue, alpha (the alpha opaque). |
| GeneratedImages | [common] data class GeneratedImages(val images: List<GeneratedImage>, val seed: Long) The pictures of one request and the seed they were drawn from. |
| GeneratedVideo | [common] data class GeneratedVideo(val frames: List<GeneratedImage>, val fps: Double, val seed: Long, val audioPcm16: ShortArray?, val audioSampleRate: Int) One generated clip: its frames in order (each a GeneratedImage), the fps they play at, the seed they were drawn from, and the sound as 16-bit mono samples at audioSampleRate hertz where the family answers one (audioPcm16 null and the rate 0 otherwise). |
| GenerationConfig | [common] data class GenerationConfig(val maxTokens: Int = 0, val temperature: Float = 0.0f, val topK: Int = 0, val topP: Float = 1.0f, val repetitionPenalty: Float = 1.0f, val thinking: Boolean? = null, val noRepeatNgramSize: Int = 0, val noRepeatNgramWindow: Int = 0) The decode policy of one request. A zero maxTokens or topK keeps the model's own default for that knob (Capabilities.defaults); temperature and topP apply as given, so a caller that wants the model's defaults for them copies them from Capabilities.defaults. |
| GenerationHandle | [common] interface GenerationHandle A running generation. cancel may be called from any thread: the decode stops at its next step, the listener receives the text streamed so far, and GenerationListener.onDone arrives with FinishReason.CANCELED. Calling it after the generation ended is a no-op. |
| GenerationListener | [common] interface GenerationListener The stream of one generation. The calls arrive in order on the model's own worker thread, never on the thread that called GenerativeModel.chat: onToken and onThinking as text is produced, then exactly one of onDone or onError. A listener that needs the main thread posts to it. |
| GenerationReport | [common] data class GenerationReport(val finish: FinishReason, val promptTokens: Int, val completionTokens: Int, val timeToFirstTokenMs: Float, val prefillTokensPerSecond: Float, val decodeTokensPerSecond: Float, val latencyMs: Float, val accelerator: String, val cachedPromptTokens: Int = 0, val text: String = "", val reasoning: String = "", val visibleText: String = "", val toolCalls: List<ToolCall> = emptyList(), val toolsCalled: Boolean = false) The counts and timings of one generation, delivered once through GenerationListener.onDone. |
| GenerativeModel | [common] class GenerativeModel : PreTrainedModel A text model loaded from its files on disk, on the CPU. |
| HardwareReport | [common] data class HardwareReport(val runtimeVersion: String, val physicalCores: Int, val logicalCores: Int, val backends: List<BackendReport>) The inventory of this machine's compute (Modelverse.hardware): the runtime's version, the CPU's cores and every backend with its devices. |
| ImageFeatureExtractionPipeline | [common] class ImageFeatureExtractionPipeline : Pipeline Image embeddings over a VisionModel: a picture answers its vector. |
| ImageGenerationHandle | [common] interface ImageGenerationHandle An image-generation request in flight. |
| ImageGenerationListener | [common] interface ImageGenerationListener The calls an image-generation request makes on the model's worker thread, in order: onImages once with the pictures and the seed the model drew, then onDone; or onError alone when the request fails. |
| ImageGenerationModel | [common] class ImageGenerationModel : PreTrainedModel An image-generation model: a prompt in, pictures out (the text-to-image kind), each as the RGBA pixels an Android bitmap takes. |
| ImageGenerationOptions | [common] data class ImageGenerationOptions(val negativePrompt: String? = null, val width: Int = 0, val height: Int = 0, val steps: Int = 0, val guidanceScale: Float? = null, val seed: Long = -1L, val count: Int = 1) How pictures are generated. A knob left at its zero keeps the model's own default: width and height in pixels (the family's size law applies; see Capabilities.defaultImageWidth and Capabilities.defaultImageHeight), steps the denoising steps, guidanceScale the prompt's pull (null keeps the model's), seed the draw (-1 lets the model pick one; the answer names it), count the number of pictures, negativePrompt what to steer away from (null for none). |
| ImageGenerationReport | [common] data class ImageGenerationReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The facts of one image-generation request. |
| ImageInput | [common] sealed interface ImageInput A picture handed to a vision model: the bitmap's pixels, or the encoded file. |
| ImageToTextPipeline | [common] class ImageToTextPipeline : Pipeline Text recognition over a VisionModel: a picture answers the text read off it. |
| LabelScore | [common] data class LabelScore(val label: String, val score: Float) One label with its score against the picture (the cosine of the two embeddings), descending in a result. |
| LoadOptions | [common] data class LoadOptions(val contextLength: Int = 0, val maxActive: Int = 0, val thinking: Boolean? = null, val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions The load-time choices a caller makes for a generative model. Every field has the library's own default; a zero or a null keeps it. |
| Modelverse | [common] object Modelverse The entry of the Modelverse binding: loads the three native libraries in the one order that works, then answers the library's version. |
| ModelverseException | [common] class ModelverseException(message: String, val codeName: String = "", cause: Throwable? = null) A Modelverse failure, in words a person can act on: a model directory that holds no checkpoint the library serves, a weight file the runtime cannot decode, a request past the context window, a library that did not load. |
| ObjectDetectionPipeline | [common] class ObjectDetectionPipeline : Pipeline Detection over a VisionModel: a picture answers its detections; the candidateLabels name the phrases an open-vocabulary detector looks for. |
| OcrOptions | [common] data class OcrOptions(val maxNewTokens: Int = 0) How text recognition runs. |
| Pipeline | [common] sealed class Pipeline A task pipeline: one call that runs a loaded model's door end to end and answers the result, under the task's own name (Modelverse.pipeline with a task string: text-generation, automatic-speech-recognition, text-to-speech, translation, image-to-text, object-detection, depth-estimation, image-feature-extraction, zero-shot-image-classification, feature-extraction, text-ranking, voice-activity-detection, text-to-image, text-to-video, and their aliases). Each pipeline's invoke blocks on the caller's thread for the model's reply; a failure throws ModelverseException with the model's message. The pipeline owns its model: close closes it. The handle's own doors stay reachable through model for a caller that streams. |
| PretrainedConfig | [common] data class PretrainedConfig(val family: String, val modelType: String, val architecture: String, val vendor: String, val description: String, val licenseUrl: String, val pipelineTags: List<String>, val inputModalities: List<String>, val outputModalities: List<String>, val isMoe: Boolean, val isMultimodal: Boolean, val hasTokenizer: Boolean, val hasImageProcessor: Boolean, val hasAudioProcessor: Boolean, val hasVideoProcessor: Boolean, val numHiddenLayers: Int, val hiddenSize: Int, val numAttentionHeads: Int, val numKeyValueHeads: Int, val slidingWindow: Long, val numExperts: Int, val runnable: Boolean, val identityOnlyPurpose: String) A checkpoint's identity as the registry resolves it from its configuration alone, before any weight is read: the family that serves it, its model type and architecture, the family's description, the modalities it takes and gives, the variant's facts, what its processor carries, and whether the family runs in this binding (runnable; identityOnlyPurpose says what an identity-only family is for). |
| PreTrainedModel | [common] sealed interface PreTrainedModel A loaded model of any door: its config (the identity the registry resolved for its source) and close. The handles are GenerativeModel, SttModel, TtsModel, VisionModel and TranslateModel. |
| RegisteredFamily | [common] data class RegisteredFamily(val family: String, val vendor: String, val runnable: Boolean, val modelTypes: List<String>, val architectures: List<String>, val ggufArchitectures: List<String>, val inputModalities: List<String>, val outputModalities: List<String>, val description: String = "", val licenseUrl: String = "", val pipelineTags: List<String> = emptyList(), val identityOnlyPurpose: String = "") One model family the model library registers, as Modelverse.registry lists it: the identity keys a checkpoint is matched by, whether the family carries a runnable path, and its modalities. A checkpoint whose model_type or architectures[0] (a config.json) or general.architecture (a GGUF header) is in no runnable family's keys has no template in this library and will be refused at load; a key that is listed names a template, and the load may still decline a checkpoint on its own facts (a weight format the runtime does not decode, a model type the family registers only to refuse). The per-release supported-models document is the measured verdict. |
| RerankerModel | [common] class RerankerModel : PreTrainedModel A reranking model: a query and documents in, one relevance score per document out, in the documents' order (a higher score is a better match; the caller sorts). |
| RerankHandle | [common] interface RerankHandle A reranking request in flight. |
| RerankListener | [common] interface RerankListener The calls a reranking request makes on the model's worker thread, in order: onScores once with one relevance score per document (in the documents' order), then onDone; or onError alone when the request fails. |
| RerankReport | [common] data class RerankReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The facts of one reranking request. |
| ResidencyPart | [common] data class ResidencyPart(val name: String, val device: String, val weightsBytes: Long, val kvCacheBytes: Long, val otherBytes: Long, val totalBytes: Long) What one part of a loaded model holds on one device (ResidencyReport). |
| ResidencyReport | [common] data class ResidencyReport(val parts: List<ResidencyPart>, val totalBytes: Long) What a loaded model holds on each device, as the runtime's pools report it (PreTrainedModel.residency). The weights are the pool growth the load left on the device; two loads running at once on one device are not told apart by that read, so load one at a time where the number matters. |
| ServeOptions | [common] data class ServeOptions(val host: String = "127.0.0.1", val port: Int = 8000, val modelId: String = "", val maxActive: Int = 4, val maxQueued: Int = 256, val defaultBudget: Int = 256, val enableWebUi: Boolean = true, val enableCors: Boolean = false, val logRequestTiming: Boolean = false, val maxUploadBytes: Long = 0, val readTimeoutSeconds: Int = 0) The server's knobs (Modelverse.serve): the bind address, the advertised model id (the family when empty), the engine's width (requests decoding at once), its queue and its default decode budget, the web page, CORS, the request timing log, and the upload and read limits (0 keeps the library's). |
| Server | [common] class Server The server over a loaded GenerativeModel (Modelverse.serve): the chat API ( /v1/chat/completions, the models listing, /health, the web page when enabled) on its own threads. start listens and returns at once; waitUntilReady probes /health; stop ends the listening and joins the server's threads; close stops and releases the served engine. The model stays open, the caller's, and refuses its own close() while a server holds it: close the server first. Several servers may serve one model. http is the server's own http server, borrowed, for routes and middleware of the app's beside the API's, registered before start. |
| Snapshot | [common] data class Snapshot(val localDir: String, val files: List<String>, val weightsPrimary: String?, val licenseLine: String?, val companions: List<SnapshotCompanion>) What Modelverse.snapshot answers: the directory the model files rest in and the file names in it; the selected weight file on the GGUF lane (null elsewhere); the checkpoint's license line as the command line logs it (null when the facts were unavailable); and the companions, each with its outcome. |
| SnapshotCompanion | [common] data class SnapshotCompanion(val repo: String, val localDir: String, val files: Int, val failure: String, val failureCode: String) One companion of a snapshot: a hub repository the model's family needs beside the source (a codec, say), where it landed, and the hub's refusal when it did not (failure empty means it landed, or was cached). |
| SourceOptions | [common] interface SourceOptions The source knobs every load and resolve reads, the hub ecosystem's own: a source is a local directory, a hub repo id or a .gguf path; revision is the hub revision (a branch, a tag or a commit), cacheDir the hub cache (null for the hub's default), token a hub access token passed through to the download (null for the hub's own token chain), offline answers from the cache alone, and device is the runtime's device spelling ("cpu", "vulkan:0"; null keeps the door's default). A local source ignores the hub knobs. |
| SpeakerSegment | [common] data class SpeakerSegment(val speaker: Int, val startMs: Long, val endMs: Long) One speaker's turn: who (speaker, a 0-based index the model assigns) spoke from startMs to endMs. |
| SpeakHandle | [common] interface SpeakHandle A running synthesis. cancel may be called from any thread. The model library runs an utterance to its end, so a cancel does not shorten the wait: the clip is dropped when the synthesis returns, SpeakListener.onClip is not called, and SpeakListener.onDone arrives with SpeakReport.canceled true. |
| SpeakKnob | [common] data class SpeakKnob(val name: String, val min: Float, val max: Float?, val keepsDefault: Float?) One synthesis knob a text-to-speech model reads (Capabilities.speakKnobs), named as the constants here spell it, with the range the binding accepts. |
| SpeakListener | [common] interface SpeakListener The outcome of one synthesis. The calls arrive on the model's own worker thread: onClip with the utterance, then exactly one of onDone or onError. This release of the model library answers the whole utterance in one call, so onClip is called once; a library that produces the audio in chunks calls it once per chunk, in order, on the same listener. |
| SpeakOptions | [common] data class SpeakOptions(val exaggeration: Float? = null, val temperature: Float? = null, val cfgScale: Float? = null, val steps: Int = 0, val seed: Long = -1, val maxSpeechTokens: Int = 0, val language: String? = null) The synthesis knobs of one request. A null knob, or the value SpeakKnob.keepsDefault names for it, keeps the family's own trained default (the chatterbox family: exaggeration 0.5, temperature 0.6, guidance 1.0, ten refinement steps, no seed, the text-length budget). A value under a knob's SpeakKnob.min is refused before the model runs. |
| SpeakReport | [common] data class SpeakReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The outcome of one synthesis, delivered once through SpeakListener.onDone. |
| SpeechClip | [common] data class SpeechClip(val pcm16: ShortArray, val sampleRate: Int, val durationMs: Long) A synthesized utterance: mono 16-bit samples at sampleRate hertz (the chatterbox family speaks at 24 kHz), durationMs long. |
| StageEvent | [common] data class StageEvent(val kind: StageEventKind, val stage: String, val slot: Int, val tStartNs: Long, val tEndNs: Long, val text: String = "", val isFinal: Boolean = false, val language: String = "", val mediaCount: Int = 0, val samples: FloatArray? = null, val sampleRate: Int = 0) One event of a model's pipeline, reported to a StageObserver on the stage's own thread. stage names the stage ( tokenize, prefill, decode and detokenize on a text model, vision when a tower runs); slot is the request's slot when the stage serves several at once, -1 otherwise; the times are the runtime's monotonic clock in nanoseconds. A payload field is empty, or null, when the kind carries none. |
| StageEventKind | [common] enum StageEventKind : Enum<StageEventKind> What one StageEvent reports. |
| StageObserver | [common] fun interface StageObserver Hears a model's pipeline stages (GenerativeModel.setStageObserver, or one call's through GenerativeModel.generate). onStage runs on the stage's own thread, never the caller's, and must return at once: a slow observer holds that stage and nothing else. An exception it lets escape goes to its thread's uncaught-exception handler. |
| SttModel | [common] class SttModel : PreTrainedModel A speech recognition model loaded from its files on disk, on the CPU. |
| Task | [common] object Task The hub task vocabulary: what AutoModel.fromPretrained takes as a task to name the door explicitly, and what kindFor answers. |
| TextGenerationPipeline | [common] class TextGenerationPipeline : Pipeline Text generation over a GenerativeModel: a prompt answers the generated text (the prompt and the reply together when returnFullText, the reply alone otherwise), a conversation answers its turns plus the assistant's. |
| TextRankingPipeline | [common] class TextRankingPipeline : Pipeline Reranking over a RerankerModel: a query and documents answer one relevance score per document, in the documents' order ( text-ranking and its alias reranking); a higher score is a better match, and the caller sorts. |
| TextToImagePipeline | [common] class TextToImagePipeline : Pipeline Image generation over an ImageGenerationModel: a prompt answers its pictures and the seed they were drawn from ( text-to-image). |
| TextToSpeechPipeline | [common] class TextToSpeechPipeline : Pipeline Speech synthesis over a TtsModel: a text answers one clip, the utterance's chunks joined. |
| TextToVideoPipeline | [common] class TextToVideoPipeline : Pipeline Video generation over a VideoGenerationModel: a prompt answers its clip and the seed it was drawn from ( text-to-video). |
| ThinkingControl | [common] enum ThinkingControl : Enum<ThinkingControl> The reasoning controls a model admits, as its chat template declares them. |
| ToolCall | [common] data class ToolCall(val id: String, val name: String, val argumentsJson: String) One call a reply made: the function's name and its arguments as JSON text (argumentsJson); id names the call for its result (ChatSession.addToolResult), empty when the model's format wrote none (a ChatSession mints call_<n>). |
| ToolChoice | [common] sealed class ToolChoice How the model may use the tools a request offers. |
| ToolDef | [common] data class ToolDef(val name: String, val description: String = "", val parametersJson: String = "{}") One tool a request offers: the function's name, its description and its parameters as a JSON schema document (parametersJson, an object; {} takes no arguments). The chat template renders the offer, and the reply's calls are read against the names. |
| TranscribeHandle | [common] interface TranscribeHandle A running transcription. cancel may be called from any thread: a clip transcribed in several windows stops at the next window boundary, and TranscribeListener.onDone arrives with TranscribeReport.canceled true. |
| TranscribeListener | [common] interface TranscribeListener The stream of one transcription. The calls arrive in order on the model's own worker thread: onSegment for each transcribed span, then exactly one of onDone or onError. |
| TranscribeOptions | [common] data class TranscribeOptions(val language: String? = null, val translate: Boolean = false, val targetLanguage: String? = null) How a clip is transcribed: language names the spoken language as an ISO code ( en), null lets the model detect it; targetLanguage names the language of the text, null keeps the spoken one (a transcription) and a code asks a model that translates as it transcribes for the text in it; a model that cannot translate refuses a target that differs from the spoken language, naming what it lacks, and the request ends in TranscribeListener.onError. translate is the English shorthand: true reads as targetLanguage = "en" where no target is named. |
| TranscribeReport | [common] data class TranscribeReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean, val language: String? = null) The outcome of one transcription, delivered once through TranscribeListener.onDone. |
| Transcript | [common] data class Transcript(val text: String, val report: TranscribeReport) A clip's transcript with the report of the request that made it (AutomaticSpeechRecognitionPipeline.transcribe). |
| TranscriptSegment | [common] data class TranscriptSegment(val startMs: Long, val endMs: Long, val text: String) One transcribed span, with its position in the clip in milliseconds. |
| TranslateHandle | [common] interface TranslateHandle A running translation. cancel may be called from any thread; the model library runs the batch to its end, the translations are dropped when it returns, and TranslateListener.onDone arrives with TranslateReport.canceled true. |
| TranslateListener | [common] interface TranslateListener The outcome of one translation request. The calls arrive on the model's own worker thread: onTranslations with one translation per input text, in order, then exactly one of onDone or onError. |
| TranslateLoadOptions | [common] data class TranslateLoadOptions(val contextLength: Int = 0, val modelArgs: Map<String, String> = emptyMap(), val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions The load-time choices a caller makes for a translation model. |
| TranslateModel | [common] class TranslateModel : PreTrainedModel A translation model loaded from its files on disk, on the CPU: an encoder-decoder model that translates between the languages it names (Capabilities.languages). A chat model that translates by instruction is a GenerativeModel, not this. |
| TranslateOptions | [common] data class TranslateOptions(val maxNewTokens: Int = 0) How a translation runs. |
| TranslateReport | [common] data class TranslateReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The outcome of one translation request, delivered once through TranslateListener.onDone. |
| Translation | [common] data class Translation(val text: String, val sourceTokens: Int, val generatedTokens: Int, val stoppedAtBudget: Boolean) One translated text. stoppedAtBudget is true when the translation stopped at its token budget instead of the model's end token, so it may be incomplete. |
| TranslationPipeline | [common] class TranslationPipeline : Pipeline Translation over a TranslateModel: texts answer their translations, in order. |
| TtsLoadOptions | [common] data class TtsLoadOptions(val contextLength: Int = 0, val scratchDir: <Error class: unknown class>? = null, val modelArgs: Map<String, String> = emptyMap(), val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions The load-time choices a caller makes for a text-to-speech model. |
| TtsModel | [common] class TtsModel : PreTrainedModel A text-to-speech model loaded from its files on disk, on the CPU. |
| VideoGenerationHandle | [common] interface VideoGenerationHandle A video-generation request in flight. |
| VideoGenerationListener | [common] interface VideoGenerationListener The calls a video-generation request makes on the model's worker thread, in order: onVideo once with the clip and the seed the model drew, then onDone; or onError alone when the request fails. |
| VideoGenerationModel | [common] class VideoGenerationModel : PreTrainedModel A video-generation model: a prompt in, a clip out (the text-to-video kind), its frames each as the RGBA pixels an Android bitmap takes, with the clip's sound where the family answers one. |
| VideoGenerationOptions | [common] data class VideoGenerationOptions(val negativePrompt: String? = null, val width: Int = 0, val height: Int = 0, val frames: Int = 0, val fps: Float = 0.0f, val steps: Int = 0, val guidanceScale: Float? = null, val seed: Long = -1L) How a clip is generated. A knob left at its zero keeps the model's own default: width and height in pixels and frames the frame count (the family's size and frame laws apply; see Capabilities.defaultImageWidth, Capabilities.defaultImageHeight and Capabilities.defaultFrames), fps the frame rate (Capabilities.defaultFps), steps the denoising steps, guidanceScale the prompt's pull (null keeps the model's), seed the draw (-1 lets the model pick one; the answer names it), negativePrompt what to steer away from (null for the family's own). |
| VideoGenerationReport | [common] data class VideoGenerationReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The facts of one video-generation request. |
| VisionHandle | [common] interface VisionHandle A running vision request. cancel may be called from any thread. The model library runs a request to its end, so a cancel does not shorten the wait: the result is dropped when the model returns, no result callback is made, and VisionListener.onDone arrives with VisionReport.canceled true. |
| VisionListener | [common] interface VisionListener The outcome of one vision request. The calls arrive on the model's own worker thread: the one result callback the request's task answers with (a detection request calls onDetections, a depth request onDepth, a text reading onText, an embedding onEmbedding, a classification onScores), then exactly one of onDone or onError. The result callbacks default to nothing, so a listener implements the one its request needs. |
| VisionLoadOptions | [common] data class VisionLoadOptions(val modelArgs: Map<String, String> = emptyMap(), val revision: String = "main", val cacheDir: String? = null, val token: String? = null, val offline: Boolean = false, val device: String? = null, val weights: String? = null, val dtype: String? = null) : SourceOptions The load-time choices a caller makes for a vision model. |
| VisionModel | [common] class VisionModel : PreTrainedModel A vision model loaded from its files on disk, on the CPU: a detector, a depth estimator, a text reader, or an embedding model (a model that embeds text and pictures in one space also classifies against labels). |
| VisionReport | [common] data class VisionReport(val latencyMs: Float, val accelerator: String, val canceled: Boolean) The outcome of one vision request, delivered once through VisionListener.onDone. |
| VisionTask | [common] object VisionTask The tasks a vision model may serve, the spellings Capabilities.visionTasks carries. |
| VoiceActivityDetectionPipeline | [common] class VoiceActivityDetectionPipeline : Pipeline Speaker diarization over a DiarizationModel: a clip answers who spoke when, in time order ( voice-activity-detection, the hub's name for the kind). |
| VoiceReference | [common] sealed interface VoiceReference The voice a text-to-speech request clones: a clip of the speaker, mono, a few seconds long (a voice-cloning family reads about the first ten seconds for the generator and the whole clip, silence trimmed, for the speaker embedding). Any sample rate; the library resamples. |
| ZeroShotImageClassificationPipeline | [common] class ZeroShotImageClassificationPipeline : Pipeline Zero-shot image classification over a dual-tower VisionModel: a picture and the candidate labels answer each label's score. |
Functions
| Name | Summary |
|---|---|
| load | [android] @JvmOverloads fun Modelverse.load(context: Context, license: String? = null) The Android entry: pin the runtime's cache root under the app's cache directory and place license through ClikaRtAndroid.load, then Modelverse.load. One call readies the binding on Android; a second call is a no-op. license is the runtime's license credential (the CLIKA1-... text, or the path of a file holding it); null leaves the process environment as it is. |