Language bindings
ClikaRT is a C++ library, and every other language reaches it through a binding compiled against the same public C++ headers a C++ program includes. There is no separate C layer in between: what a binding can do is exactly what the C++ API does, at one of two levels.
The two levels
| Level | Languages | What it covers |
|---|---|---|
| FULL | C++, Python | every public header: the runtime (tensors, operators, devices and streams, model loading, graphs and transforms, tokenizers and processors, serving) and the Modelverse model library |
| INFERENCE | Kotlin | what runs, serves and measures a ready model: loading it from a file or a hub snapshot, the registry and its model cards, generation and chat with streaming and cancel, the tokenizers and processors the model needs, image, audio and video input, the task pipelines, serving, the benchmark and the fit check of a model, the device, hardware and memory facts, the host utilities (JSON, tables, templates, regular expressions), errors, logging, the profiler and progress, and the Android entry |
A FULL language binds every public header. An INFERENCE language binds what runs a ready model and declines the rest by design: building or editing a model, the nn modules, graph queries and transforms, tracing and compiling your own code, and ONNX export are FULL-level work, done from C++ or Python. A model you author reaches an app as a served model or a compiled graph, never as Kotlin source.
What each binding carries
The rows are the surfaces a program reaches for; a cell names the member where the answer is partial.
| Surface | C++ | Python | Kotlin |
|---|---|---|---|
| tensors, operators, dtypes, in-place forms | yes | yes | yes (Ops, one function per operator) |
| devices, streams, execution scopes (synchronous, tracing, placement) | yes | yes | yes |
nn modules (Linear, Conv, KVCache, fused projections, load_state_dict) | yes | yes | no |
| ONNX: open a model, compile it, optimize, run by position or name | yes | yes | yes (OnnxModel, ModelGraph) |
graph query and edit, transforms, trace and compile of your own code, ONNX export | yes | yes | no |
readers: safetensors, .npy, GGUF, images, audio, video | yes | yes | images, audio, video and .npy (Io, VideoReader); no safetensors or GGUF reader |
| tokenizer, chat template, streaming decode | yes | yes | yes |
| image, audio and video processors | yes | yes | yes |
the serving runtime: FunctionModel, Executor, Pipeline, batching | yes | yes | yes |
| HTTP client and server, JSON, regular expressions, templates, tables | yes | yes | yes |
| downloading a model from the hub | hub::snapshot (the model library) | mv.snapshot_download | Modelverse.snapshot, and inside fromPretrained |
| the model library: generate, chat, serve, pipelines | yes | yes | yes |
the benchmark and the fit check of a model (bench, check) | yes | yes | yes (BenchReport, FitReport, maxContextLength) |
| speech to text, text to speech, vision, translation | yes | yes | yes, as the handles SttModel, TtsModel, VisionModel and TranslateModel |
| text embedding and reranking | yes | yes | no |
| image and video generation | yes | yes | no |
PyTorch interoperation (torch.compile backend, DLPack exchange) | no | yes | no |
the Android entry (ClikaRtAndroid.load, memory-pressure trim, logcat) | no | no | yes |
What every binding promises
- The same names. The INFERENCE binding carries the model library's own object names (
AutoConfig,AutoModel,AutoProcessor,AutoTokenizer,fromPretrained,generate,pipeline(task)), so a model loads, generates and serves under the names the C++ and Python surfaces use. - The same errors. A failure arrives as a typed error carrying the status, the stable code name and the message, in each language's own error type.
- The same version law. A binding is built against one release of the runtime and refuses to load another, naming both versions.
- One runtime library per process. The Python wheel carries the runtime library once for both products; the Kotlin artifact's Android variant carries it too, and a desktop JVM program takes it from the release archive's
lib/. Nothing ships a second copy.
The artifacts
Every language ships one artifact that carries both products, the runtime and the Modelverse model library. Each is a download of the platform (Download the ClikaRT SDK), taken from the same release.
| Language | Level | How you get it |
|---|---|---|
| C++ | FULL | the release archive for your platform: the headers, the libraries, find_package(ClikaRT CONFIG) and find_package(Modelverse CONFIG) |
| Python | FULL | the clika-runtime wheel for your CPython version and platform, installed from the file (pip install <wheel>), with Modelverse inside as clika_runtime.modelverse |
| Kotlin | INFERENCE | the io.clika:clika-runtime Maven artifact in clika-runtime-maven-<version>.zip: one coordinate with an Android AAR variant and a desktop JVM jar variant, carrying io.clika.runtime and io.clika.modelverse. The AAR carries the runtime libraries, so an Android app needs nothing else; the desktop jar carries the classes and the JNI bridges and loads the runtime libraries from the release archive of the same platform, its lib/ named on java.library.path |
Where to go next
- C++: the API reference, every public namespace, class and function.
- Python: Use ClikaRT from Python, the wheel's shape from the NumPy boundary to models as
nn.Module. - Kotlin: Deploy to mobile in the tutorial, Package ClikaRT in an Android app, and the Modelverse part Use it from code.