# CLIKA documentation, full text
> ClikaRT is the inference runtime, ClikaRT CLI the command line interface that infers, serves and benchmarks popular models on it (built on Modelverse, the CLIKA model library), and the Platform the service that runs benchmarks on devices and issues the runtime licenses. Every page below is a plain address on this site; the site also offers a self-contained offline copy (the "Offline Docs" button).
Every page of `https://docs.clika.io/llms.txt`, in its order; each begins with its title, its one-line description and its address.
---
# ClikaRT
ClikaRT is CLIKA's inference runtime, a C++ library with Python and Kotlin bindings, that runs models on CPUs, GPUs and other hardware accelerators through one public API.
Source: https://docs.clika.io/clikart.md
ClikaRT is the CLIKA inference runtime: a C++ library, with Python and Kotlin bindings, that loads models and runs them on CPUs, GPUs and other hardware accelerators through one public API. Link one library and include one header, or import one package, and the same code runs on every backend the runtime ships.
## Why ClikaRT
For the AI developer. The API is PyTorch-shaped C++, and the same shape reaches Python in full and Kotlin at the inference level through the [bindings](bindings.md). Tensors move with `.to(device)`, and the built-in operators chain the way you expect (`.reshape()`, `.relu()`, `.matmul()`). What PyTorch does not give you is where this runs. The same code covers CPU, CUDA, Vulkan and Metal (mobile GPU via Vulkan, Apple silicon via Metal), and the backend is picked at run time, not at build time. Library built-ins provide the convenience of PyTorch, but you can implement whatever you want, as close to the hardware as you like; the library is built with freedom as a first-class citizen. Serving, tokenizers and pipelines live in the same library, so a model checkpoint becomes a served endpoint.
For the backend engineer. No AI background is required to serve a model well. Pull a packaged model from [Modelverse](/modelverse), load it, and hand requests to the serving runtime; batching, device placement and precision are the runtime's job, not model configuration details you need to understand. The library behaves like normal C++, which inference stacks usually do not. Integration is one `find_package`, one link target, C++17, and no Python in the ship path.
For the embedded developer. Small targets are a first-class platform. A complete Android deployment is the Kotlin artifact's Android AAR, roughly 140 MB to download (the core library, its Vulkan backend and the model library inside), and the same code runs on a phone's GPU, on Apple silicon and on Jetson, from the same bundle with no device-specific build. The device carries no Python and no package manager, only the libraries your binary links.
For the defense, healthcare and finance developer. Everything runs on your hardware. User data never leaves the device, and the self-contained bundle installs and builds where there is no network at all.
For the business. Nothing external to install, resolve, or keep in sync. The bundle is self-contained, and every third-party component it embeds is listed in its `licenses/` directory. There is no dependency tree to audit, no third-party library update that breaks your team's product, and no copyleft surprise for legal. Hardware stays a choice, not a commitment. TensorRT runs only on NVIDIA GPUs; ClikaRT runs the same models on NVIDIA, AMD, Intel and Qualcomm GPUs and on plain CPUs, years-old hardware included. Run your state-of-the-art models on older or legacy machines to reduce inference and cloud costs.
## What you get
- **Tensors and operators.** `ClikaRT::Tensor` plus the `ClikaRT::ops` library: element-wise math, reductions, convolutions, attention, indexing. This is the operator set a model needs.
- **Backends.** CPU always; CUDA, Vulkan and Metal where the platform has them. Accelerator backends load on demand; unavailable ones are absent.
- **Quantized weights.** Weight-only quantization schemes (`nn::QLinearWoQ`, GGUF block formats, FP8, NF4 and more) decode inside the kernels.
- **Model loading.** safetensors, GGUF, ONNX and NumPy through `ClikaRT::io`; tokenizers, chat templates and processors alongside.
- **Serving.** A runtime with sessions, continuous batching and pipelines; an HTTP server and client; a CLI framework.
- **Bindings.** Two levels over the one C++ public surface: Python (a wheel) binds all of it; Kotlin (a desktop library and an Android AAR) binds what runs a ready model, under the model library's own names. [Language bindings](bindings.md) says what each level promises.
## Where to go next
- [Getting started](getting-started/index.md): install the bundle and write your first program.
- [System requirements](system-requirements.md): supported platforms and hardware.
- [Examples](examples.md): the public example projects, one topic each.
- [API reference](api/index.md): every public namespace, class and function.
---
# API reference
The ClikaRT public C++ API, by namespace and class.
Source: https://docs.clika.io/clikart/api.md
The public API of ClikaRT is the `ClikaRT::` namespace, declared by the headers under `ClikaRT/`. Include `ClikaRT/clika_rt.h` for everything, or the individual headers named on each page.
## Namespaces
| Name | Description |
| --- | --- |
| [`ClikaRT::cli`](./ClikaRT/cli/index.md) | |
| [`ClikaRT::device`](ClikaRT/device-namespace/index.md) | |
| [`ClikaRT::dtype`](./ClikaRT/dtype/index.md) | The dtype-system helpers: classification predicates, size arithmetic, and the readable name. [`DataType`](./index.md#DataType) itself stays at [`ClikaRT`](./index.md)`::` (the one name every signature spells); everything ABOUT a dtype lives here. |
| [`ClikaRT::encoding`](./ClikaRT/encoding/index.md) | |
| [`ClikaRT::env`](./ClikaRT/env/index.md) | |
| [`ClikaRT::graph`](./ClikaRT/graph/index.md) | |
| [`ClikaRT::http`](./ClikaRT/http/index.md) | |
| [`ClikaRT::io`](./ClikaRT/io/index.md) | |
| [`ClikaRT::json`](./ClikaRT/json/index.md) | |
| [`ClikaRT::logging`](./ClikaRT/logging/index.md) | |
| [`ClikaRT::nn`](./ClikaRT/nn/index.md) | |
| [`ClikaRT::ops`](./ClikaRT/ops/index.md) | |
| [`ClikaRT::placement`](./ClikaRT/placement/index.md) | |
| [`ClikaRT::platform`](./ClikaRT/platform/index.md) | |
| [`ClikaRT::processor`](./ClikaRT/processor/index.md) | |
| [`ClikaRT::progress`](./ClikaRT/progress/index.md) | |
| [`ClikaRT::quant`](./ClikaRT/quant/index.md) | The quantized-checkpoint import taxonomy: the parsed `quantization_config` facts and the routing enums the split factories consume. |
| [`ClikaRT::random`](./ClikaRT/random/index.md) | |
| [`ClikaRT::regex`](./ClikaRT/regex/index.md) | |
| [`ClikaRT::runtime`](./ClikaRT/runtime/index.md) | |
| [`ClikaRT::spec`](./ClikaRT/spec/index.md) | |
| [`ClikaRT::tables`](./ClikaRT/tables/index.md) | |
| [`ClikaRT::threading`](./ClikaRT/threading/index.md) | |
| [`ClikaRT::tokenizer`](./ClikaRT/tokenizer/index.md) | |
## Classes
| Name | Description |
| --- | --- |
| [`Device`](./ClikaRT/Device.md) | A compute device: a backend API plus a device index (e.g. the `1` in CUDA device 1). Default-constructs to the whole-machine host CPU. |
| [`EagerScope`](./ClikaRT/EagerScope.md) | While alive, forces EAGER execution on the calling thread: every op runs its kernel and produces a real value before dispatch returns, even when nested inside a [`TracingScope`](./ClikaRT/TracingScope.md). Use it to guard a region that must compute concrete values (a metric, a running average, a control-flow decision read back to the host) so a caller who wrapped the surrounding code in a [`TracingScope`](./ClikaRT/TracingScope.md) cannot turn those ops into un-materialized placeholders. Outside any scope execution is already eager, so this is a no-op there. Non-copyable, non-movable. |
| [`Error`](./ClikaRT/Error.md) | Thrown by the throwing API overloads and by [Result::value\_or\_throw()](./ClikaRT/Result.md#value_or_throw). |
| [`FakeTensor`](./ClikaRT/FakeTensor.md) | A value type (no out-of-line surface): it holds a `vector<`[`spec::SymInt`](./ClikaRT/spec/SymInt.md)`>`, a dtype, a copyable [`Stream`](./ClikaRT/Stream.md), and a flag, all ABI-stable, so it crosses by layout. |
| [`NamedTensors`](./ClikaRT/NamedTensors.md) | An insertion-agnostic `name → `[`Tensor`](./ClikaRT/Tensor.md) map. Copyable and movable (a copy shares each entry's storage; the tensors are refcount handles). A moved-from [NamedTensors](./ClikaRT/NamedTensors.md) is empty: every read answers empty / false / 0, and [`set`](./ClikaRT/NamedTensors.md#set) fills it again (a dict moved into `load_state_dict` stays a usable object). |
| [`ProfileSession`](./ClikaRT/ProfileSession.md) | A profiling session: start one or more captures, then export the results. Move-only. The session must outlive any in-flight async work it captured (workers that finish after a capture ends still record into it). |
| [`ProfileSummary`](./ClikaRT/ProfileSummary.md) | The counters a session accumulated, as plain values: the typed form of [`ProfileSession::summary_text()`](./ClikaRT/ProfileSession.md#summary_text) and of the `summary` object in `report.json`, field for field, so a consumer reads numbers instead of parsing text. Durations are nanoseconds; counts are events recorded inside the session's captures. |
| [`QTensor`](./ClikaRT/QTensor.md) | A quantized weight, viewed with its scheme. A value type over a refcounted payload handle; copying a [`QTensor`](./ClikaRT/QTensor.md) never copies weight bytes. |
| [`Result`](./ClikaRT/Result.md) | Holds either a success value of type T or a ([Status](./index.md#Status), message) failure. Move-only. Check [ok()](./ClikaRT/Result.md#ok) before reading [value()](./ClikaRT/Result.md#value), or use [value\_or\_throw()](./ClikaRT/Result.md#value_or_throw). \[\[nodiscard\]\]: a call that returns a [`Result`](./ClikaRT/Result.md) and discards it swallows the failure it may carry; the compiler warns at such a site; consume it ([`unwrap`](./index.md#unwrap), [`CLIKART_CHECK`](./index.md#CLIKART_CHECK), a named read) or discard deliberately with a `(void)` cast. |
| [`Result\`](./ClikaRT/Result.void.md) | Outcome of a fallible operation that yields no value on success (e.g. registering a route, writing a file). Check [`ok()`](./ClikaRT/Result.void.md#ok), or call [`value_or_throw()`](./ClikaRT/Result.void.md#value_or_throw) to raise [`ClikaRT::Error`](./ClikaRT/Error.md) on failure in your own TU. \[\[nodiscard\]\]: a statement-position effect call that ignores its result swallows the failure; the compiler warns; consume it ([`CLIKART_CHECK`](./index.md#CLIKART_CHECK), [`unwrap`](./index.md#unwrap), a named read) or discard deliberately with a `(void)` cast. |
| [`Scalar`](./ClikaRT/Scalar.md) | A scalar operand that keeps its KIND across the boundary: a bool, an integer (any width, stays integral), or a double. What it buys over a bare double parameter: an integer scalar against an integer tensor stays in the integer domain (`ops::sub(2, int64_tensor)` is Int64, exact at any magnitude), where a double would force weak-float promotion. Every ctor is implicit on purpose. Pass the bare literal. Distinct from [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md): a [`Scalar`](./ClikaRT/Scalar.md) never carries a tensor, which is what keeps a two-tensor call (`ops::add(a, b)`) unambiguous against the tensor-first overloads. |
| [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md) | A ClikaRT-level value type (NOT op-layer machinery): one optional slot that carries a scalar (a bool, an integer or a double) OR a tensor. Every ctor is implicit on purpose. Pass a bare `bool`, `double`, an integer, a [`Tensor`](./ClikaRT/Tensor.md), or `std::nullopt` straight to a [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md) parameter; spelling the wrap at a call site (`ops::ScalarOrTensor(x)`) is redundant noise. [`ops::ScalarOrTensor`](./ClikaRT/ops/ScalarOrTensor.md) remains a valid spelling via the alias below. |
| [`Span`](./ClikaRT/Span.md) | A non-owning view over [`size()`](./ClikaRT/Span.md#size) contiguous elements of `T`, the C++17 stand-in for `std::span` (see the file note above). It never allocates, copies, or owns: the viewed memory must outlive every copy of the view. `T`'s constness is the access law: [`Span`](./ClikaRT/Span.md)`` reads, [`Span`](./ClikaRT/Span.md)`` writes through to the caller's memory. |
| [`Stream`](./ClikaRT/Stream.md) | A handle to a device execution stream. Cheap to copy (it is an identifier, not the stream's resources). |
| [`StreamOrDevice`](./ClikaRT/StreamOrDevice.md) | A value type: it holds a copyable [`Stream`](./ClikaRT/Stream.md) handle and/or a [`Device`](./ClikaRT/Device.md), so it crosses the ABI by layout; [`resolve_impl`](./ClikaRT/StreamOrDevice.md#resolve) is its one library entry (the placement law lives in the runtime). |
| [`SynchronousScope`](./ClikaRT/SynchronousScope.md) | While alive, every op dispatched on the calling thread COMPLETES before the dispatch call returns, whichever stream runs it: the moment a call returns, its result is ready. Thread-local (it changes nothing for other threads and throttles no stream), so it opens anywhere, on a busy lane included; it trades away the throughput that asynchronous pipelining buys, for as long as it lives. Scopes nest. Non-copyable, non-movable. |
| [`SynchronousStreamScope`](./ClikaRT/SynchronousStreamScope.md) | While alive, every op dispatched on the calling thread WAITS for completion before the dispatch call returns, and `stream` runs at most one task at a time, so the moment a dispatch returns, the result is ready ([`Stream::query_idle()`](./ClikaRT/Stream.md#query_idle) is true). Deterministic, ideal for debugging / per-op stepping; it trades away the throughput that async pipelining buys. The stream's prior limit is restored on destruction. Open it only on a quiescent stream (synchronize first). Non-copyable, non-movable. |
| [`Tensor`](./ClikaRT/Tensor.md) | |
| [`TracingScope`](./ClikaRT/TracingScope.md) | While alive, the calling thread builds a LAZY graph: ops return `Unscheduled` placeholder tensors carrying lineage, and NO kernel runs. Call [`Tensor::synchronize()`](./ClikaRT/Tensor.md#synchronize) to materialize the graph in one pass. Outside the scope execution is eager (the default). Non-copyable, non-movable. |
## Enumerations
### enum Status {#Status}
`enum class` **`Status`** `:` `std::int32_t`
Outcome of an operation.
| Enumerator | Value | Description |
| --- | --- | --- |
| `Ok` | `0` | |
| `InvalidArgument` | | |
| `NotFound` | | |
| `Unsupported` | | |
| `Internal` | | |
| `Unavailable` | | The work did not run: it was shed under load, ran past its deadline, or was canceled. Retry later (HTTP 503). |
Declared in `ClikaRT/common/result.h`, line 93
### enum DataType {#DataType}
`enum class` **`DataType`** `:` `std::int32_t`
The element format of a tensor's values.
`int32`-backed for a stable ABI. The set covers booleans, signed/unsigned integers (including sub-byte widths), the IEEE-style floats, and the narrow floating formats used by quantized models (FP8 / FP6 / FP4). `Undefined` is the unset value.
ABI note: the numeric values are part of the public ABI: new types append at the end; existing ones never reorder or drop.
| Enumerator | Value | Description |
| --- | --- | --- |
| `Undefined` | `0` | |
| `Bool` | | |
| `Int2` | | |
| `Int4` | | |
| `Int8` | | |
| `Int16` | | |
| `Int32` | | |
| `Int64` | | |
| `UInt1` | | |
| `UInt2` | | |
| `UInt4` | | |
| `UInt8` | | |
| `UInt16` | | |
| `UInt32` | | |
| `UInt64` | | |
| `Float16` | | |
| `BFloat16` | | |
| `Float32` | | |
| `Float64` | | |
| `Float8_E4M3` | | |
| `Float8_E5M2` | | |
| `Float8_E4M3FNUZ` | | |
| `Float8_E5M2FNUZ` | | |
| `Float8_E8M0` | | |
| `Float6_E2M3` | | |
| `Float6_E3M2` | | |
| `Float4_E2M1` | | |
Declared in `ClikaRT/compute/data_type.h`, line 21
## Type aliases
### using OptionalTensor {#OptionalTensor}
A tensor argument that may be omitted. It is just [`Tensor`](./ClikaRT/Tensor.md); pass a default [`Tensor`](./ClikaRT/Tensor.md)`{}` (an undefined tensor) to mean "not provided". The alias documents, at a signature, that the parameter is optional (e.g. a norm's `weight`/`bias`, an attention `mask`). Mirrors the role of `c10::optional<`[`Tensor`](./ClikaRT/Tensor.md)`>` in ATen, using the undefined-tensor sentinel [ClikaRT](./index.md) already carries.
True when `code_name` (an [`Error::code_name()`](./ClikaRT/Error.md#code_name) or [`Result::code_name()`](./ClikaRT/Result.md#code_name)) names a capability decline: a case the runtime does not serve, whether the operation itself, the device, or the dtype, memory layout, quantization scheme, source and destination pair, geometry or parameter value it was asked for. It is the one failure a caller may serve another way (another device, another route); every other failure belongs to the operation, and so does a backend that runs nothing at all (`BACKEND_UNAVAILABLE`, `BACKEND_NOT_LOADED`, `BACKEND_DRIVER_ABSENT`). False for an empty or unknown name; names compare exactly, case included.
The coarse [`Status`](./index.md#Status) a failure named `code_name` carries: the `status()` the runtime reports beside that `code_name()`, so a binding or a log reader can class a failure from its name alone. [`Status::Internal`](./index.md#Status) for an empty or unknown name; names compare exactly, case included.
The rvalue identity: the argument's VALUE (one move), never a reference into it. A reference-returning arm would hand a range-for or an `auto&&` binding the storage of a temporary that dies at the end of the range-init full-expression; returning the value makes that binding own what it reads.
[Stream](./ClikaRT/Stream.md) a tensor's `to_string()` summary (shape, dtype, device, first values) to an `ostream`, so `std::cout << t << '\n'` works. Header-only; forwards to the (infallible) `to_string()`; reads values for any dtype, integers included.
Declared in `ClikaRT/compute/tensor.h`, line 971
## Platform notes {#platform_notes}
### macOS: JIT acceleration and the Hardened Runtime {#macos_jit}
On macOS, [ClikaRT](./index.md)'s CPU runtime can accelerate some workloads by compiling specialized kernels at run time. The host application must be allowed to map executable memory: an application built with the Hardened Runtime needs the `com.apple.security.cs.allow-jit` entitlement. Without it, [ClikaRT](./index.md) detects the restriction at startup and serves its standard kernels; results are identical, and only the acceleration is unavailable.
## `ClikaRT/cli/cli.h` {#cli-h}
```cpp
#include
```
Umbrella for the public command-line parser ([`ClikaRT::cli`](./ClikaRT/cli/index.md)): typed options / flags / positionals, subcommand routing, display + mutually-exclusive groups, auto help/usage, env-var fallbacks, and shell-completion generation. Every fallible call returns `Result`; the parser never throws, never exits, and never mutates argv.
## `ClikaRT/clika_rt.h` {#clika_rt-h}
```cpp
#include
```
[ClikaRT](./index.md) public API umbrella header: one include for the whole public C++ API, the [`ClikaRT`](./index.md)`::` namespace with one sub-namespace per module, and the one link target `ClikaRT::ClikaRT` (`find_package(ClikaRT CONFIG)`). `AGENTS.md` beside these headers is the lookup index: what each module gives, which header to open for a task, and the gotchas.
The modules, each under its own directory and namespace: `compute/` (`Tensor`, `Device`, `Stream`, the dtypes, the execution scopes, and `ops::`, the operator library), `nn/` (the module tree a model is built from, the packed leaves, the KV cache), `graph/` (`ModelGraph`, trace, compile, the queries and transforms), `io/` (checkpoints, media, the ONNX file), `runtime/` (the serving runtime: nodes, executors, pipelines), `tokenizer/`, `processor/`, `json/`, `template/`, `regex/`, `http/`, `cli/`, `progress/`, `profiler/`, `tables/`, `threading/`, `logging/`, `encoding/`, `platform/` and `common/` (the result type, `Span`, the environment names, the version).
Three laws every program meets first. Every fallible call has two spellings: the distribution's default build returns values and raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure, a build with the `Result` surface returns `Result`; [`ClikaRT::unwrap`](./index.md#unwrap), [`CLIKART_TRY`](./index.md#CLIKART_TRY) and [`CLIKART_CHECK`](./index.md#CLIKART_CHECK) read the same under both, and a failure's `code_name()` is the name to branch on ([`common/result.h`](./index.md#result-h)). Operators are asynchronous: an `ops::` call queues its work and returns, and a host read (`to_string`, `item`, `item_as_vec`) waits for the value ([`compute/stream.h`](./ClikaRT/Stream.md), [`compute/scope.h`](./index.md#scope-h)). Every process that runs an operator or a model needs the license credential, `CLIKA_RT_LICENSE` or the per-user file, before its first call ([`common/env_vars.h`](./ClikaRT/env/index.md#env_vars-h)).
## `ClikaRT/common/macros.h` {#macros-h}
```cpp
#include
```
### Macros
#### \#define CLIKART\_IS\_WINDOWS {#CLIKART_IS_WINDOWS}
## `ClikaRT/common/result.h` {#result-h}
```cpp
#include
```
Value-returned result type for the [ClikaRT](./index.md) API. Holds a value on success or a Status + message on failure. Every fallible public method returns `Result` and never throws on its own. You turn a failure into an exception on the calling side with `value_or_throw()` (which raises [`ClikaRT::Error`](./ClikaRT/Error.md)), or inspect `ok()` / `status()` / `message()` and never pay for exceptions at all.
Works with exceptions disabled: a consumer compiling with `-fno-exceptions` (or defining `CLIKART_NO_EXCEPTIONS`) gets a `value_or_throw()` that reports the failure to stderr and `std::abort()`s instead of throwing; the inspecting API (`ok()` / `status()` / `message()`) is unaffected.
── The error-handling vocabulary: each name's role ─────────────────────
Two LAYERS live in this header and they are not duplicates:
The LIBRARY's own boundary shims (not for consumer code):
- [`CLIKART_RESULT(T)`](./index.md#CLIKART_RESULT) / [`CLIKART_UNWRAP(expr)`](./index.md#CLIKART_UNWRAP): how the public headers' inline wrappers adapt the library's `Result` to the built error-handling shape ([`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE)). Consumer code never writes these.
- [`CLIKART_INPLACE_RESULT(T)`](./index.md#CLIKART_INPLACE_RESULT) / [`CLIKART_INPLACE_UNWRAP(expr, self)`](./index.md#CLIKART_INPLACE_UNWRAP): the same adaptation for the IN-PLACE wrappers (the write-through ops and mutating tensor methods): under the value-returning surface they keep the historical `T&` return (the mutated `self`/`out`); under the `Result`-returning surface they return the impl's `Result` straight through. [`CLIKART_CHECK`](./index.md#CLIKART_CHECK) is the caller's mode-stable spelling for these calls.
The caller vocabulary (each spelling compiles and behaves identically whichever shape the library was built with):
- `ClikaRT::unwrap(x)`: read a value; extracts a `Result` (raising on failure) and forwards anything else unchanged, so the same call-site text serves both shapes.
- Implicit extraction: `T x = fn(...);` / `g(fn(...))` compiles under BOTH shapes: a TEMPORARY `Result` converts to `T`, raising on failure exactly as `unwrap`. Call sites written against the value-returning surface keep compiling when a build opts into the `Result` surface, so a codebase adopts explicit handling gradually rather than all at once. Rvalue-only (a NAMED `Result` is read through `ok()`/`value()`/`unwrap`), and excluded for `bool` payloads (`operator bool` tests OKNESS everywhere, never the payload; read a bool payload through `unwrap`/`value()`). Note `auto x = fn(...);` still binds the `Result` itself; spell the type (or `unwrap`) to extract.
- [`CLIKART_TRY(expr)`](./index.md#CLIKART_TRY): the CAPTURE idiom, "hand me data, never throw": yields a `Result` whatever happens, catching [`ClikaRT::Error`](./ClikaRT/Error.md) AND any other exception (nothing propagates).
- [`CLIKART_TRY_OR_RETURN(var, expr)`](./index.md#CLIKART_TRY_OR_RETURN) / [`CLIKART_CHECK(expr)`](./index.md#CLIKART_CHECK): the PROPAGATE idiom for `Result`-returning consumer functions: on failure they RETURN the error (status + message + code\_name verbatim) to the enclosing function's caller; only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted; foreign exceptions pass through untouched.
### Macros
#### \#define CLIKART\_DETAIL\_HAS\_CXXABI {#CLIKART_DETAIL_HAS_CXXABI}
CLIKART\_TRY: opt back INTO Result-style error handling.
The public API returns values directly and raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure. When you would rather inspect a `Result` than catch, wrap the call:
```text
Result r = CLIKART_TRY(ops::matmul(a, b));
if (!r.ok()) { log(r.message()); return; }
use(r.value());
```
Yields a `Result` where `T` is the (decayed) type the expression produces: `Result` for a `void` call, `Result` for a `Tensor`- or `Tensor&`-returning one. An expression that already produces a `Result` passes through as that same `Result` (never `Result>`), so the idiom reads identically whichever error-handling shape the library was built with. The expression is evaluated exactly once. A [`ClikaRT::Error`](./ClikaRT/Error.md) is captured with its status; any other exception is captured as `Status::Internal`; nothing propagates. This is the inverse of the internal unwrap-or-propagate idiom: it CATCHES exceptions at the call site and hands you data.
Declared in `ClikaRT/common/result.h`, line 581
#### \#define CLIKART\_RESULT {#CLIKART_RESULT}
`#define` **`CLIKART_RESULT(...)`** `__VA_ARGS__`
CLIKART\_RESULT / CLIKART\_UNWRAP: the two-mode public boundary shape.
A fallible public method delegates to a `Result`-returning `*_impl` in the library and adapts that `Result` for the caller. These macros pick the adaptation at build time from [`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE) (the [`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE) CMake option, baked into `build_info.h`); the SAME header source compiles both ways, no per-method edit to flip. [`CLIKART_UNWRAP`](./index.md#CLIKART_UNWRAP) **supplies its own `return`**, so the wrapper body is just the unwrap; one macro serves both a value and a `void` boundary (a `void`-returning function may `return` a `void` expression):
[CLIKART\_RESULT(Tensor)](./index.md#CLIKART_RESULT) to(Device d) const \{ [CLIKART\_UNWRAP(to\_impl(d))](./index.md#CLIKART_UNWRAP); \} [CLIKART\_RESULT(void)](./index.md#CLIKART_RESULT) write(...) const \{ [CLIKART\_UNWRAP(write\_impl(...))](./index.md#CLIKART_UNWRAP); \}
- **CLIKART\_USE\_RESULT\_TYPE == 0** (default): the return type is the bare value `T` (`void`), and the unwrap is `return (expr).value_or_throw();`; the wrapper raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure, caller-side. Byte-behaviour-identical to the historical value-returning surface.
- **CLIKART\_USE\_RESULT\_TYPE == 1**: the return type is `Result` (`Result`) and the unwrap is `return (expr);`; the impl's `Result` passes straight through; the wrapper never throws and the caller inspects `.ok()` / `.status()` / `.value()`.
[`CLIKART_RESULT`](./index.md#CLIKART_RESULT) is variadic so a comma-bearing type ([`CLIKART_RESULT`](./index.md#CLIKART_RESULT)`(std::array)`, [`CLIKART_RESULT`](./index.md#CLIKART_RESULT)`(std::vector)`) is not split into two macro arguments. Boundaries that transform the impl's result before returning it (an in-place op returning `Tensor&` / `*this`, a templated readback, a contained callback) keep their explicit `value_or_throw()`; they are not `Result`-shaped and are the deliberate carve-outs (same family as the user-callback types).
CLIKART\_INPLACE\_RESULT / CLIKART\_INPLACE\_UNWRAP: the boundary shims for the IN-PLACE wrapper family (a write-through op returning its `out`, a mutating tensor method returning `*this`). The impl twins all return `Result`; the wrapper manufactures the reference:
inline [CLIKART\_INPLACE\_RESULT(Tensor)](./index.md#CLIKART_INPLACE_RESULT) relu\_(Tensor\& self) \{ [CLIKART\_INPLACE\_UNWRAP(impl::relu\_(self), self)](./index.md#CLIKART_INPLACE_UNWRAP); \}
- **CLIKART\_USE\_RESULT\_TYPE == 0** (default): the return type is `T&` and the unwrap is `(expr).value_or_throw(); return self;`, the historical reference-returning surface, raising [`ClikaRT::Error`](./ClikaRT/Error.md) caller-side on failure.
- **CLIKART\_USE\_RESULT\_TYPE == 1**: the return type is `Result` and the unwrap is `return (expr);`; the impl's `Result` passes straight through (the caller already holds the buffer it handed in).
The caller's mode-stable spelling for these calls is [`CLIKART_CHECK(...)`](./index.md#CLIKART_CHECK); reference-chaining is a value-surface-only idiom.
CLIKART\_TRY\_OR\_RETURN: the PROPAGATE idiom's value form, for `Result`-returning consumer functions.
```text
Result mean_of(Tensor t) {
CLIKART_TRY_OR_RETURN(auto m, ClikaRT::ops::mean(t));
return m.item();
}
```
Evaluates `expr` exactly once; on success declare-assigns (or assigns) `var` from the value; on a LIBRARY failure returns the error (status, message, and the fine `code_name` verbatim, through the three-argument `Result` constructor) to the enclosing function's caller. Only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted; a foreign exception passes through untouched. The SAME text serves both error-handling shapes: under the value-returning surface `expr` yields the bare value (a failure throws and is converted here); under the `Result`-returning surface `expr` yields a `Result` that passes through unflattened. Error timing is the settle-time law (see `unwrap`): failures surface at the read site, and a settled failure re-reports identically on re-read. Expands to multiple statements; use it in statement context (never as an unbraced `if` body), one per source line. With exceptions disabled a failure already aborted inside the library, so a reached call only ever sees success.
Declared in `ClikaRT/common/result.h`, line 673
#### \#define CLIKART\_CHECK {#CLIKART_CHECK}
`#define` **`CLIKART_CHECK(expr)`** `do { \ auto clikart_check_state_ = \ ::ClikaRT::detail::propagate_capture([&]() { return (expr); }); \ if (!clikart_check_state_.ok()) \ return {clikart_check_state_.status(), \ clikart_check_state_.message(), \ clikart_check_state_.code_name()}; \ } while (false)`
CLIKART\_CHECK: the PROPAGATE idiom's effect form: `expr` is run for its effect (an in-place op, a write, a registration) and any LIBRARY failure returns the error (status/message/`code_name` verbatim) to the enclosing function's caller; only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted, foreign exceptions pass through. The mode-stable spelling for the in-place `Tensor&`-returning ops ([`CLIKART_CHECK(ops::relu_(x))`](./index.md#CLIKART_CHECK)`;`) and for any `Result` call under the `Result`-returning surface. Single statement (safe as an `if` body); settle-time + sticky re-read semantics as on `unwrap`.
Declared in `ClikaRT/common/result.h`, line 691
## `ClikaRT/common/version.h` {#version-h}
```cpp
#include
```
[ClikaRT](./index.md) version information.
## `ClikaRT/compute/scope.h` {#scope-h}
```cpp
#include
```
RAII execution-mode scopes. [ClikaRT](./index.md) executes asynchronously by default; ops dispatch and return, and the work runs later. These scopes change that for the calling thread, for their lifetime: force every op to complete before dispatch returns (deterministic / debug), or switch to lazy graph-building (tracing).
## `ClikaRT/http/http.h` {#http-h}
```cpp
#include
```
HTTP umbrella: the client ([`client.h`](./ClikaRT/http/DownloadOptions.md)) and the server ([`server.h`](./ClikaRT/http/index.md#server-h)) together. Include this to get both sides; include the role header directly ([`ClikaRT/http/client.h`](./ClikaRT/http/DownloadOptions.md) or [`ClikaRT/http/server.h`](./ClikaRT/http/index.md#server-h)) to name the side you use.
## `ClikaRT/nn/nn.h` {#nn-h}
```cpp
#include
```
[ClikaRT](./index.md) public NN-module umbrella; include this one header for the whole `nn` surface. New public modules are added HERE (and only here); consumers and [`clika_rt.h`](./index.md#clika_rt-h) never enumerate the individual headers.
A translation unit that defines a class deriving from `nn::Module` is compiled with run-time type information disabled, as the library is.
## `ClikaRT/profiler/profiler.h` {#profiler-h}
```cpp
#include
```
Capture a profile of [ClikaRT](./index.md) compute work and export it: a Chrome trace you can open in `chrome://tracing` / Perfetto, a per-op summary, or the raw report JSON.
While a capture is active, every op dispatched on the capturing thread (and on the streams it drives) is timed automatically; you do not instrument individual calls. Work that other threads dispatch is outside the capture: a serving pipeline runs its nodes on its own executor threads, so a capture opened around a pipeline request records only the capture's own markers. Profile a pipeline through the `clikart-cli` CLI's `bench --profile`, which captures where the nodes run; a model driven on the calling thread (a detector, a transcriber, an embedder) profiles through this session directly. The shape is RAII:
```text
ClikaRT::ProfileSession session;
{
auto capture = session.start_capture("decode");
// ... run ops / drive streams ...
} // capture ends here
session.save("profile_out"); // chrome_trace.json + summary.txt + ...
printf("%s\n", ClikaRT::unwrap(session.summary_text()).c_str());
const ClikaRT::ProfileSummary s = ClikaRT::unwrap(session.summary()); // the same numbers, typed
printf("%zu allocations\n", s.total_allocations);
```
---
# Kotlin API
The ClikaRT Kotlin binding, documented from its sources.
Source: https://docs.clika.io/clikart/api-kt.md
The [`clika-runtime`](./clika-runtime/index.md) module, package by package.
| Package |
| --- |
| [`io.clika.modelverse`](./clika-runtime/io.clika.modelverse/index.md) |
| [`io.clika.runtime`](./clika-runtime/io.clika.runtime/index.md) |
---
# clika_runtime
The clika_runtime module, introspected from the installed wheel.
Source: https://docs.clika.io/clikart/api-py.md
clika-runtime: the Python package of the ClikaRT on-device inference runtime.
Importing this package loads the runtime library shipped in-package (under
``clika_runtime/lib/``). The Modelverse model library ships in the same
package as :mod:`clika_runtime.modelverse` (``import clika_runtime.modelverse
as mv``); importing it loads the runtime first and resolves it from here.
The top level publishes :class:`Tensor`, :class:`Size`, :class:`Device`,
:class:`Stream`, the dtype objects (``float32``,
``bfloat16``, ...), the tensor factories (``tensor``, ``zeros``,
``randn``, ...), every operator of :mod:`clika_runtime.ops` as a
free function (``clika_runtime.matmul(a, b)``), the placement and
execution scopes (``device``, ``stream``, ``synchronous``, ``tracing``,
``eager``, ``meta_init``), ``eval`` / ``async_eval`` / ``synchronize``,
``save`` / ``load`` / ``load_with_metadata``, ``compile`` / ``trace``, the device functions, and the
exception classes. Modes are strings (``approximate="tanh"``); the typed
enumerations stay under ``clika_runtime._core.ops``. Models are built with
:mod:`clika_runtime.nn`; :mod:`clika_runtime.torch` (imported on first use,
never as a side effect) exchanges tensors and modules with PyTorch.
Execution is asynchronous by default: an operator returns at once and its
work rides the calling thread's current lane on the input's device. A
value settles when it is read (``numpy()``, ``item()``, ``tolist()``,
``print``), when ``eval(*trees)`` or ``synchronize()`` is called, or inside
a ``synchronous()`` region where every operator completes before it
returns. ``tracing()`` records operations instead of running them until
``eval`` materializes them; ``eager()`` restores the default inside a
tracing region. Reading the value of a traced or storage-free tensor raises
:class:`ClikaRTError`.
Every process that runs an operator or a model needs the license credential
CLIKA issued for the project: ``CLIKA_RT_LICENSE`` in the environment (the
``CLIKA1-...`` text, or the path of a file holding it), set before this
package is imported, else the per-user file the ``clikart-license-init``
console script writes once. Without one the first call raises
:class:`ClikaRTError` with the code name ``LICENSE_FAILED``; every error
carries a stable ``code_name`` beside its message, and the code name is the
thing to branch on (:mod:`clika_runtime.errors`).
A model from a model hub loads through :mod:`clika_runtime.modelverse`:
``mv.AutoModelForCausalLM.from_pretrained(repo_id)`` downloads the snapshot
into the hub cache and loads it, ``mv.snapshot_download(repo_id)`` downloads
without loading, and a local directory is a source everywhere a repository
id is. A model written here from the operators starts from
:mod:`clika_runtime.nn` (``nn.Module``, ``nn.Linear``, ``nn.KVCache``) and
the fused attention operator; the documentation's examples walk one end
to end.
| Name | Kind |
| --- | --- |
| [`ClikaRTError`](./ClikaRTError.md) | class |
| [`CompiledFunction`](./CompiledFunction.md) | class |
| [`CompiledModule`](./CompiledModule.md) | class |
| [`Device`](./Device.md) | class |
| [`DeviceProperties`](./DeviceProperties.md) | class |
| [`InternalError`](./InternalError.md) | class |
| [`InvalidArgumentError`](./InvalidArgumentError.md) | class |
| [`NotFoundError`](./NotFoundError.md) | class |
| [`OutOfMemoryError`](./OutOfMemoryError.md) | class |
| [`QTensor`](./QTensor.md) | class |
| [`Size`](./Size.md) | class |
| [`Stream`](./Stream.md) | class |
| [`Tensor`](./Tensor.md) | class |
| [`TensorSpec`](./TensorSpec.md) | class |
| [`TracedGraph`](./TracedGraph.md) | class |
| [`UnavailableError`](./UnavailableError.md) | class |
| [`UnsupportedError`](./UnsupportedError.md) | class |
| [`dtype`](./dtype.md) | class |
| [`finfo`](./finfo.md) | class |
| [`iinfo`](./iinfo.md) | class |
| [functions](./functions.md) | module functions |
---
# Language bindings
ClikaRT has two API levels: FULL, the whole public surface, for C++ and Python; INFERENCE, everything that runs a ready model, for Kotlin. Every binding compiles against the C++ public headers.
Source: https://docs.clika.io/clikart/bindings.md
ClikaRT is a C++ library, and every other language reaches it through a binding compiled against the same public C++ headers a C++ program includes. There is no separate C layer in between: what a binding can do is exactly what the C++ API does, at one of two levels.
## The two levels
| Level | Languages | What it covers |
|---|---|---|
| **FULL** | C++, Python | every public header: the runtime (tensors, operators, devices and streams, model loading, graphs and transforms, tokenizers and processors, serving) and the Modelverse model library |
| **INFERENCE** | Kotlin | what runs, serves and measures a ready model: loading it from a file or a hub snapshot, the registry and its model cards, generation and chat with streaming and cancel, the tokenizers and processors the model needs, image, audio and video input, the task pipelines, serving, the benchmark and the fit check of a model, the device, hardware and memory facts, the host utilities (JSON, tables, templates, regular expressions), errors, logging, the profiler and progress, and the Android entry |
A FULL language binds every public header. An INFERENCE language binds what runs a ready model and declines the rest by design: building or editing a model, the `nn` modules, graph queries and transforms, tracing and compiling your own code, and ONNX export are FULL-level work, done from C++ or Python. A model you author reaches an app as a served model or a compiled graph, never as Kotlin source.
## What each binding carries
The rows are the surfaces a program reaches for; a cell names the member where the answer is partial.
| Surface | C++ | Python | Kotlin |
| --- | --- | --- | --- |
| tensors, operators, dtypes, in-place forms | yes | yes | yes (`Ops`, one function per operator) |
| devices, streams, execution scopes (synchronous, tracing, placement) | yes | yes | yes |
| `nn` modules (`Linear`, `Conv`, `KVCache`, fused projections, `load_state_dict`) | yes | yes | no |
| ONNX: open a model, compile it, optimize, run by position or name | yes | yes | yes (`OnnxModel`, `ModelGraph`) |
| graph query and edit, transforms, `trace` and `compile` of your own code, ONNX export | yes | yes | no |
| readers: safetensors, `.npy`, GGUF, images, audio, video | yes | yes | images, audio, video and `.npy` (`Io`, `VideoReader`); no safetensors or GGUF reader |
| tokenizer, chat template, streaming decode | yes | yes | yes |
| image, audio and video processors | yes | yes | yes |
| the serving runtime: `FunctionModel`, `Executor`, `Pipeline`, batching | yes | yes | yes |
| HTTP client and server, JSON, regular expressions, templates, tables | yes | yes | yes |
| downloading a model from the hub | `hub::snapshot` (the model library) | `mv.snapshot_download` | `Modelverse.snapshot`, and inside `fromPretrained` |
| the model library: generate, chat, serve, pipelines | yes | yes | yes |
| the benchmark and the fit check of a model (`bench`, `check`) | yes | yes | yes (`BenchReport`, `FitReport`, `maxContextLength`) |
| speech to text, text to speech, vision, translation | yes | yes | yes, as the handles `SttModel`, `TtsModel`, `VisionModel` and `TranslateModel` |
| text embedding and reranking | yes | yes | no |
| image and video generation | yes | yes | no |
| PyTorch interoperation (`torch.compile` backend, DLPack exchange) | no | yes | no |
| the Android entry (`ClikaRtAndroid.load`, memory-pressure trim, logcat) | no | no | yes |
## What every binding promises
- **The same names.** The INFERENCE binding carries the model library's own object names (`AutoConfig`, `AutoModel`, `AutoProcessor`, `AutoTokenizer`, `fromPretrained`, `generate`, `pipeline(task)`), so a model loads, generates and serves under the names the C++ and Python surfaces use.
- **The same errors.** A failure arrives as a typed error carrying the status, the stable code name and the message, in each language's own error type.
- **The same version law.** A binding is built against one release of the runtime and refuses to load another, naming both versions.
- **One runtime library per process.** The Python wheel carries the runtime library once for both products; the Kotlin artifact's Android variant carries it too, and a desktop JVM program takes it from the release archive's `lib/`. Nothing ships a second copy.
## The artifacts
Every language ships one artifact that carries both products, the runtime and the Modelverse model library. Each is a download of the platform ([Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk)), taken from the same release.
| Language | Level | How you get it |
|---|---|---|
| C++ | FULL | the release archive for your platform: the headers, the libraries, `find_package(ClikaRT CONFIG)` and `find_package(Modelverse CONFIG)` |
| Python | FULL | the `clika-runtime` wheel for your CPython version and platform, installed from the file (`pip install `), with Modelverse inside as `clika_runtime.modelverse` |
| Kotlin | INFERENCE | the `io.clika:clika-runtime` Maven artifact in `clika-runtime-maven-.zip`: one coordinate with an Android AAR variant and a desktop JVM jar variant, carrying `io.clika.runtime` and `io.clika.modelverse`. The AAR carries the runtime libraries, so an Android app needs nothing else; the desktop jar carries the classes and the JNI bridges and loads the runtime libraries from the release archive of the same platform, its `lib/` named on `java.library.path` |
## Where to go next
- C++: the [API reference](api/index.md), every public namespace, class and function.
- Python: [Use ClikaRT from Python](how-to/use-clikart-from-python.mdx), the wheel's shape from the NumPy boundary to models as `nn.Module`.
- Kotlin: [Deploy to mobile](getting-started/first-program/06-deploy-to-mobile.mdx) in the tutorial, [Package ClikaRT in an Android app](how-to/package-clikart-in-an-android-app.mdx), and the Modelverse part [Use it from code](/modelverse/getting-started/first-model/use-it-from-code).
---
# Additional examples
The example programs that ship with a release: the runtime's own CMake projects inside every desktop archive, and the examples archive with complete C++, Python, Kotlin and Android programs beside their recorded output.
Source: https://docs.clika.io/clikart/examples.md
Two sets of example programs ship with a release, and both are downloads of the platform ([Get ClikaRT](getting-started/get-clikart.mdx)). The release archive for a desktop platform carries the runtime's own examples under `examples/src`: each is a standalone `find_package(ClikaRT CONFIG)` project on the public API only, each builds its topic up one chapter at a time, starting at `00_hello_world`, and every example directory has its own `README.md` walk-through; this page catalogs that set first. The examples archive, `ClikaRT--examples.tar.xz`, is the other: complete programs in C++, Python and Kotlin, three Android apps and a voice translator, each beside its recorded output, under one `examples/` directory whose `README.md` and `AGENTS.md` index them.
## The runtime's own examples, inside the archive
Read roughly top to bottom; each row assumes a little of the ones above it.
| Example | What it shows | Chapters |
| --- | --- | --- |
| `version` | The smallest consumer: link the bundle, print `GetVersionInfo()` | `00_hello_world` |
| `compute` | The tensor/op engine: tensors and dtypes, device properties, data movement, the `ops::` library, the async model, zero-copy `.to(device)`, hardware probes, a distributed matmul, the profiler, cast chains | `00_hello_world` · `01_tensors` · `02_devices` · `03_data_movement` · `04_operators` · `05_async` · `06_zero_copy` · `07_hardware` · `08_distributed_matmul` · `09_profiler` · `10_cast_chain` |
| `async` | The async execution model in depth: dispatch vs ready, safe host reads, `on_complete`, the synchronous scope, tracing | `00_dispatch_vs_ready` · `01_safe_reads` · `02_on_complete` · `03_sync_scope` · `04_tracing_scope` |
| `runtime` | The serving runtime: nodes and phases, per-session state, continuous batching, pipelines, vision models, the interface layer, a two-stage model, an encoder with a stateful decoder | `00_functional_api` · `01_model_api` · `02_stateful_functional_api` · `03_stateful_model_api` · `04_pipeline` · `05_simple_vision_model` · `06_interface` · `07_two_stage_vision_model` · `08_encoder_and_stateful_decoder` |
| `nn` | Neural-network modules: make with plain counts, bind weights, forward; Linear, Conv, and a KV cache bound straight into the attention op | `00_linear` · `01_conv` · `02_attention_kvcache` |
| `processor` | Image and audio pre-processing: resize, rescale and normalize a generated image; log-mel features from a synthesized waveform | `00_image` · `01_audio` |
| `json` | A small JSON value for payloads: parse, typed getters, build with `operator[]`, `dump` round trip | `00_hello_world` |
| `errors` | The failure vocabulary: `unwrap`, `CLIKART_TRY` capture, `CLIKART_TRY_OR_RETURN` propagation, and `code_name()`, the stable machine-readable failure name | `00_result` |
| `cli` | Typed argument parsing: options and flags, subcommands, validators, shell completion | `00_hello_world` · `01_args` · `02_subcommands` · `03_completion` |
| `logging` | Structured logging: levels, the level filter, named subsystem loggers | `00_hello_world` · `01_levels` · `02_named` |
| `progress` | Terminal progress: bars (bytes, ETA, rate), spinners, multi-bar groups | `00_hello_world` · `01_bytes_eta_rate` · `02_spinner` · `03_group` |
| `templating` | Jinja2-compatible rendering: variables, logic, filters, chat prompts | `00_hello_world` · `01_logic` · `02_filters` · `03_chat_prompts` |
| `tokenizer` | Text to token ids and back: HF `tokenizer.json`, byte offsets, batching and vocab, chat templates | `00_hello_world` · `01_offsets` · `02_batch_and_vocab` · `03_huggingface` |
| `tables` | A columnar dataframe: CSV, dtypes, select/filter/sort, group-by/join, transforms, GPU acceleration, a row cursor | `00_hello_world` · `01_columns_and_dtypes` · `02_csv` · `03_select_filter_sort` · `04_groupby_join` · `05_transform` · `06_acceleration` · `07_rows` |
| `tables_benchmark` | The pandas-vs-ClikaRT performance sibling of `tables` | one standalone project |
| `io` | Loading data and weights: NumPy `.npy`, safetensors, GGUF, images, and the ONNX model stack ([guide](how-to/run-an-onnx-model.mdx)) | `00_hello_world` · `01_safetensors` · `02_gguf` · `03_image` · `04_onnx` |
| `http_server` | An HTTP service: routing, JSON APIs, middleware, a templated site, SSE streaming, an image-upload endpoint that runs compute | `00_hello_world` · `01_routing` · `02_json_api` · `03_middleware` · `04_serve_a_website` · `05_sse` · `06_image_compute` |
| `download` | A download tool combining CLI, progress bar and HTTP client | `00_hello_world` · `01_multi_file` |
| `flash_attention` | A hand-written CUDA kernel on ClikaRT streams: upstream Flash Attention ported, not rewritten, dispatched through the custom-op interface | one project (`main.cpp` + the ported kernel) |
### Build and run
```bash
cmake -S "$CLIKART_BUNDLE_DIR/examples/src/compute" -B build-compute \
-DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build-compute
./build-compute/compute_00_hello_world
```
Or build them all at once from the top-level project: `cmake -S "$CLIKART_BUNDLE_DIR/examples/src" -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"`. The top-level project adapts to the distribution it is built against; it skips an example the distribution cannot build rather than failing.
On platforms where the archive carries pre-built example binaries (`examples/bin`), run them in place; they find the libraries through a relative rpath, with no library paths to set.
## The examples archive
`ClikaRT--examples.tar.xz` extracts to one `examples/` directory. Every program in it runs on its own against a release, and every tutorial and how-to program sits beside its recorded output (`.out`), the text the release printed when the recording was made. `examples/README.md` is the front door, one table per language; `examples/AGENTS.md` is the index written for a reader or an AI assistant: the facts a first run needs, a task table, the commands, the download call in each language, and what each binding carries. Nothing in the archive reaches outside it.
Each language's programs name the one distribution they need, taken from the same release:
| Directory | Programs | Needs | Build and run |
| --- | --- | --- | --- |
| `cpp/` | `clika_rt/`, the engine one topic per directory in chapters; `tutorial/`, the getting-started programs; `howto/`, one directory per how-to page; `modelverse/`, programs over the model library ([the Modelverse catalog](/modelverse/examples)) | the release archive for your platform, extracted, with `CLIKART_BUNDLE_DIR` naming its directory | `cmake -S cpp -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" && cmake --build build`, then the binary under `build/`; `cpp/README.md` |
| `python/` | `clika_rt/`, self-asserting chapters over the wheel, from the NumPy boundary through `nn.Module`, ONNX compile-and-run, tokenizers, processors and tracing; `tutorial/` and `howto/` as the C++ tree has them, with the Python-only pages (a model authored in Python, PyTorch interoperation, pytrees) | the `clika-runtime` wheel for your interpreter and platform, installed with `pip install ` | `python3 `; `python/README.md` |
| `kotlin/` | `clika_rt/`, instrumented tests that run the operator surface and the device facts on a connected Android device; `tutorial/`, the getting-started programs on a desktop JVM; `howto/`, the vision-task programs on a desktop JVM | the Kotlin artifact, `clika-runtime-maven-.zip`, extracted anywhere and named to gradle as `-PclikaRtMavenRepo=`; a desktop run also takes the release archive's `lib/` on `java.library.path` | `gradle connectedAndroidTest -PclikaRtMavenRepo=` in a chapter's directory; `kotlin/README.md` |
| `android/` | three sample apps, one screen each: `hello` loads the runtime and prints the version, the backends and one operator's result; `chat` loads a chat model and streams its replies; `serve` hosts a loaded model for the network from a foreground service | the Kotlin artifact, as above, and an Android SDK | one gradle build beside `shared/`; `android/README.md` |
| `voice_translate/` | a live voice translator as one Compose Multiplatform app for Android and the desktop JVM: speech to text, translation and text to speech over the model library's Kotlin binding | the Kotlin artifact and, per platform, the release archives the build lays out | `voice_translate/README.md` |
Every program here runs compute, so every one of them needs a license credential in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes ([Get ClikaRT](getting-started/get-clikart.mdx#license-credential)); without one a call is refused with the code name `LICENSE_FAILED`. A model program names a checkpoint from the local hub cache or takes one as an argument; `clikart-cli fetch ` downloads it once, with no credential.
The tutorial programs this documentation shows ([Your first program](getting-started/first-program/01-your-first-program.mdx) onward) are the archive's `cpp/tutorial/`, `python/tutorial/` and `kotlin/tutorial/` programs, with the output recorded beside them. [Use ClikaRT from Python](how-to/use-clikart-from-python.mdx) walks the first steps of the Python lane; [Language bindings](bindings.md) says what each binding carries.
---
# First steps
New to ClikaRT? Start here. What the runtime is, how to install it, and the tutorial series.
Source: https://docs.clika.io/clikart/getting-started.md
New to ClikaRT? This section is where to start. It gives enough orientation to hold the whole library in your head, an install you can verify in minutes, and a first program that grows into a real pipeline. Read it in order:
1. **[ClikaRT at a glance](overview.mdx)**: what the runtime is and is not, who it is for, and the five-minute mental model.
2. **[Get ClikaRT](get-clikart.mdx)**: pick your platform, get the download and verify commands.
3. **[Quick install](installation.md)**: prerequisites, the bundle, and two ways to prove it works.
4. **Tutorial series**: six parts, each a complete step, with core concepts explained where they first appear. [Your first program](first-program/01-your-first-program.mdx), [tensors and operators](first-program/02-tensors-and-operators.mdx), [devices and the async model](first-program/03-devices-and-async.mdx), [your first pipeline](first-program/04-your-first-pipeline.mdx) (weights from disk, compute on the best device, results back), [serve it](first-program/05-serve-it.mdx) (the same model behind the serving runtime, answering requests), and [deploy to mobile](first-program/06-deploy-to-mobile.mdx) (the part 4 program on a real phone, unchanged).
5. **[What to read next](next-steps.md)**: where to go once it runs.
## How the ClikaRT docs are layered
- **This section** orients: condensed, in reading order, concepts woven in.
- **[How-to guides](/how-to/index.md)**: problem-oriented recipes, one per "how do I X", with [additional examples](/examples.md) as the end-to-end reading inside it.
- **[API reference](/api/index.md)**: the full reference per language: C++, Python and Kotlin.
[System requirements](/system-requirements.md) sits alongside, for the platform and hardware tables.
---
# Complete installation
Per-platform toolchains, IDE setup, and the full failure catalog. Under construction.
Source: https://docs.clika.io/clikart/getting-started/complete-installation.md
This page will carry the full installation reference. Until it lands, [Quick install](installation.md) covers the fast path and [system requirements](/system-requirements.md) lists the platforms, toolchains and drivers.
Planned contents:
- Per-OS toolchains in depth: Linux distributions, macOS and Xcode, Windows and Visual Studio.
- Platform notes: Windows, macOS, Android.
- IDE setup.
- Adding ClikaRT to an existing CMake project (also planned as a how-to guide).
- The full failure catalog, beyond quick install's four lines.
---
# Deploy to mobile
Cross-compile the part 4 program for Android, push it over adb, and run it on the phone's GPU. The code does not change.
Source: https://docs.clika.io/clikart/getting-started/first-program/deploy-to-mobile.md
**Your code does not change.** The program from [part 4 of the series](04-your-first-pipeline.mdx), byte for byte, compiled for Android and pushed to a phone, prints the same numbers there that it printed on your workstation, on the phone's GPU when it has one and on its CPU otherwise. This page is the complete path: cross-compile, push, run, and the few lines of Kotlin an app adds around it.
You need the bundle (its `android-arm64` dist ships the CPU and Vulkan backends), the Android NDK, and a phone with USB debugging enabled.
## 1. Cross-compile
Same project, same `main.cpp`. The configure line adds the NDK's toolchain file and the ABI; the bundle's CMake package selects the Android dist on its own:
```bash
cmake -S . -B build-android \
-DCMAKE_TOOLCHAIN_FILE="$ANDROID_NDK/build/cmake/android.toolchain.cmake" \
-DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 \
-DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build-android
```
```text
-- ClikaRT 0.6.4: dist android-arm64 (backends: cpu;vulkan)
```
## 2. Push and run
The binary plus the two libraries from the Android dist go to the device; the phone needs nothing else installed:
```bash
adb shell mkdir -p /data/local/tmp/clikart
adb push build-android/hello \
"$CLIKART_BUNDLE_DIR/lib/libClikaRT.so" \
"$CLIKART_BUNDLE_DIR/lib/libClikaRT_vulkan.so" \
/data/local/tmp/clikart/
adb shell "cd /data/local/tmp/clikart && chmod +x hello && \
CLIKA_RT_LICENSE=CLIKA1-... LD_LIBRARY_PATH=. ./hello"
```
The program runs compute, so it needs a license credential like any other ClikaRT program, and a binary under `adb shell` reads it from the environment of the shell that starts it ([Get ClikaRT](../get-clikart.mdx#license-credential)). Push the credential to a file on the device and give `CLIKA_RT_LICENSE` that path when you would rather keep it off the command line.
On a Galaxy S24 Ultra:
```text
y = Tensor(shape=[2, 4], dtype=Float32, device=Vulkan:0, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...])
row means = [4.25, 4.25] (expected 8*1*0.5 + 0.25 = 4.25)
```
The workstation printed `device=CUDA:0`; the phone prints `device=Vulkan:0`. Same program, same numbers. `pick_device()` from [part 3](03-devices-and-async.mdx) found the phone's GPU the same way it found the workstation's; on this phone the part 3 discovery loop lists:
```text
CPU 0: ARM
Vulkan 0: Adreno (TM) 750
```
## 3. From an app: a few lines of Kotlin
ClikaRT is a C++ library, and an Android app reaches it through the app's own native code. Kotlin loads your library and calls your function; ClikaRT stays on the native side:
```kotlin
object Pipeline {
init { System.loadLibrary("pipeline") }
external fun run(): String
}
```
```cpp title="pipeline_jni.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
extern "C" JNIEXPORT jstring JNICALL
Java_com_example_app_Pipeline_run(JNIEnv* env, jobject) {
const Tensor x = Tensor::ones({2, 8}, DataType::Float32);
const Tensor w = Tensor::full({4, 8}, 0.5, DataType::Float32);
const Tensor y = ops::relu(ops::linear(x, w));
return env->NewStringUTF(y.to_string().c_str());
}
```
The CMake side adds one target next to `hello`:
```cmake
add_library(pipeline SHARED pipeline_jni.cpp)
target_link_libraries(pipeline PRIVATE ClikaRT::ClikaRT)
```
Beneath the JNI boundary the native side is the C++ library you built for arm64. Kotlin code can also skip the custom library entirely: the `io.clika:clika-runtime` artifact ships its own JNI bridge and every runtime library, and calls the runtime directly. Python stays on the server and desktop side.
An app has no shell to export a variable from and no per-user license file, so the credential is an argument instead. `ClikaRtAndroid.load(context, license = "CLIKA1-...")` places it for the process before the binding loads. A credential given there replaces one already in the process environment, and the default, `null`, leaves the environment as it is. Ship the credential the way you ship any other secret your app needs at start-up.
What the app ships beside the library, the build levels it declares (`compileSdk` 36, `minSdk` 28, the Android Gradle plugin 9.3.2 on gradle 9.5 or newer), the model library's own `libClikaRT_modelverse.so` when the app runs a model from the library, and the two facts it places before the first call (the CPU worker count `CLIKA_RT_NUM_THREADS`, read once before the first compute, and the cache root) are on [Package ClikaRT in an Android app](../../how-to/package-clikart-in-an-android-app.mdx).
Every platform works this way: one bundle, one toolchain file, the same code. A Jetson needs no cross-compile at all: it is an arm64 Linux machine and runs the linux-arm64 build directly. The [system requirements](/system-requirements.md) table lists the platforms and their backends. The tutorial ends here; [what to read next](../next-steps.md).
---
# Devices and the async model
Discover the machine's devices, place tensors on them, and understand dispatch vs ready, and why host reads are always safe.
Source: https://docs.clika.io/clikart/getting-started/first-program/devices-and-async.md
The same program from parts 1-2 runs unchanged on a GPU. Data placement is the only new ingredient. This part adds device discovery and `.to(device)`, then explains the execution model behind every `ops::` call you have made so far.
## The program
```cpp title="main.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::device::ComputeAPI;
namespace device = ClikaRT::device;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
// The best device this machine has, probed at run time. An unavailable
// backend is a fact, not an error: the predicate answers false.
namespace {
Device pick_device() {
if (device::is_cuda_available()) return Device::cuda();
if (device::is_vulkan_available()) return Device::vulkan();
if (device::is_metal_available()) return Device::metal();
return Device::cpu();
}
} // namespace
int main() {
// What is this machine carrying? enumerate_devices lists the concrete
// handles a backend exposes; get_device_properties describes one.
for (ComputeAPI api : {ComputeAPI::CPU, ComputeAPI::CUDA,
ComputeAPI::Vulkan, ComputeAPI::Metal}) {
for (Device dev : device::enumerate_devices(api)) {
const device::DeviceProperties p = device::get_device_properties(dev);
std::printf("%-7s %d: %s\n",
device::compute_api_name(api), dev.index, p.name.c_str());
}
}
const Device dev = pick_device();
std::printf("running on %s\n", device::compute_api_name(dev.api));
// .to(device) moves data; ops run where their inputs live.
const Tensor a = Tensor::ones({512, 512}, DataType::Float32).to(dev);
// This call DISPATCHES the matmul and returns. The kernel runs in its
// own time; nothing here waits for it.
const Tensor c = ops::matmul(a, a);
// A host read is where the wait lands: it synchronizes first, so you
// never observe unfinished bytes. Every element is 512 (= K).
std::printf("every element = %.0f\n", ops::amax(c).item());
return 0;
}
```
```python title="main.py"
import clika_runtime as crt
def main() -> None:
# The best device this machine has: Device.gpu() probes the available
# accelerators and FALLS BACK to the CPU when none is present; it never
# fails, so the same script runs everywhere.
dev = crt.Device.gpu()
print(f"running on {dev!r}")
# Factories take the device directly; ops run where their inputs live.
a = crt.ones(512, 512, device=dev)
# This call DISPATCHES the matmul and returns. The kernel runs in its
# own time; nothing here waits for it.
c = a @ a
# A host read is where the wait lands: it synchronizes first, so you
# never observe unfinished bytes. Every element is 512 (= K).
print(f"every element = {c.amax().item():.0f}")
if __name__ == "__main__":
main()
```
```kotlin title="Main.kt"
import io.clika.runtime.Backends
import io.clika.runtime.ClikaRt
import io.clika.runtime.ComputeApi
import io.clika.runtime.Device
import io.clika.runtime.Ops
import io.clika.runtime.Tensors
import java.util.Locale
// The best device this machine has, probed at run time. A backend that is
// not in this build, or whose devices do not serve, is a fact the probe
// answers, never an error.
private fun pickDevice(): Device {
for (api in listOf(ComputeApi.CUDA, ComputeApi.VULKAN, ComputeApi.METAL)) {
if (Backends.isAvailable(api) && Backends.deviceStatus(api) == null && Backends.deviceCount(api) > 0) {
return Device(api, 0)
}
}
return Device(ComputeApi.CPU)
}
fun main() {
ClikaRt.load()
// What is this machine carrying? Backends.devices lists the devices a
// backend exposes; Device.properties() describes one.
for (api in listOf(ComputeApi.CPU, ComputeApi.CUDA, ComputeApi.VULKAN, ComputeApi.METAL)) {
if (!Backends.isAvailable(api)) continue
for (dev in Backends.devices(api)) {
println("%-7s %d: %s".format(Locale.ROOT, api.label, dev.index, dev.properties().name))
}
}
val dev = pickDevice()
println("running on ${dev.api.label}")
// Placement is a factory argument; operators run where their inputs live.
val a = Tensors.ones(longArrayOf(512, 512), where = dev)
// This call DISPATCHES the matmul and returns. The kernel runs in its
// own time; nothing here waits for it.
val c = Ops.matmul(a, a)
// A host read is where the wait lands: item() synchronizes first, so
// you never observe unfinished bytes. Every element is 512 (= K).
val m = Ops.amax(c)
println("every element = ${"%.0f".format(Locale.ROOT, m.item())}")
listOf(a, c, m).forEach { it.release() }
}
```
On a machine with an NVIDIA GPU:
```text
CPU 0: AMD
CUDA 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
CUDA 1: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Vulkan 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Vulkan 1: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
running on CUDA
every element = 512
```
The same binary on a CPU-only machine lists only the CPU and runs there; no rebuild, no configuration.
## Devices and backends
A `Device` is a backend API plus a zero-based index. `Device::cuda(1)` is the second CUDA GPU, and `Device::cpu()` is the default everything starts on. Every distribution has the CPU backend compiled in; CUDA, Vulkan and Metal are shared libraries the runtime loads on demand, the first time something asks. `is_backend_available(api)` (and the shorthands `is_cuda_available()` and friends) answers whether that load works on this machine; `enumerate_devices(api)` returns the concrete handles, and an empty list is a valid answer, not an error. This is why the same binary runs everywhere. Absent hardware costs you a branch, not a build configuration.
`.to(device)` returns the tensor moved (a no-op copy if it is already there), and operators run on the device their inputs live on; there is no global "current device" to set.
## Dispatch is not execution
ClikaRT is asynchronous by nature. An `ops::` call **dispatches** work and returns; the kernel runs and the result becomes **ready** in its own time. On the default CPU stream ops happen to run inline, which is why parts 1-2 never confronted this. On an accelerator, or a worker stream made with `Stream::create`, the dispatch returns first and the compute overlaps with your code. `t.synchronize()` blocks until `t`'s pending work is done; `t.on_complete(callback)` is the push-style equivalent, firing when the result is ready.
## Reading across the async boundary
Two rules cover every host read:
1. **Reads wait for you.** `item()`, `item_as_vec()` and `const_data_ptr()` synchronize before handing back bytes. A read issued right after dispatching heavy work blocks until the result is real. The wait moves into the read; it never disappears. You can never observe garbage through the public read surface.
2. **Pointers are device pointers.** `const_data_ptr()` addresses the buffer on the tensor's own device. On a CUDA tensor that is CUDA memory; call `t.to(Device::cpu())` first to read it on the host. (`item` / `item_as_vec` do the host transfer for you.)
So the failure mode is never corruption, it is a surprise stall, a "cheap" read that waited for a matmul. When latency matters, choose where the wait lands: an explicit `synchronize()`, an `on_complete` callback, or a read whose cost you have accepted. The `async` [example project](/examples.md) measures all of this with timers.
Next: [part 4](04-your-first-pipeline.mdx). It loads weights from disk, computes, and reads results back.
---
# Serve it
Wrap part 4's model in the serving runtime: a schema, a lambda serving the forward pass, and an executor answering requests.
Source: https://docs.clika.io/clikart/getting-started/first-program/serve-it.md
Part 4 ended with a program that loads a checkpoint and computes once. This part puts the same model behind the serving runtime: declare its input/output contract, serve the forward pass with a lambda, and hand requests to an executor. The tutorial ends where the overview's pitch ends, a model checkpoint answering requests.
## The program
```cpp title="main.cpp"
#include
#include
#include
#include
#include
namespace rt = ClikaRT::runtime;
using ClikaRT::DataType;
using ClikaRT::NamedTensors;
using ClikaRT::Tensor;
using ClikaRT::spec::TensorSpec;
namespace io = ClikaRT::io;
namespace ops = ClikaRT::ops;
namespace {
// The layer's geometry, the checkpoint's names, and the edge names the schema
// declares and every request and response addresses. A misspelled edge name
// routes nothing, silently, so each is spelled once.
constexpr std::int64_t kFeatures = 8;
constexpr std::int64_t kOutputs = 4;
constexpr std::int64_t kBatch = 2;
constexpr const char* kWeight = "mlp.weight";
constexpr const char* kBias = "mlp.bias";
constexpr const char* kX = "x";
constexpr const char* kY = "y";
} // namespace
int main() {
// The checkpoint from part 4, re-created so this program stands alone.
NamedTensors weights;
weights.set(kWeight, Tensor::full({kOutputs, kFeatures}, 0.5, DataType::Float32));
weights.set(kBias, Tensor::full({kOutputs}, 0.25, DataType::Float32));
const std::string ckpt =
(std::filesystem::temp_directory_path() / "first_program.safetensors").string();
io::save_safetensors(weights, ckpt);
// Load it and declare the model's I/O contract: batches of 8 features
// in, batches of 4 activations out. kDynamicDim leaves the batch open.
const NamedTensors loaded = io::load_safetensors(ckpt);
const Tensor w = loaded.get(kWeight);
const Tensor b = loaded.get(kBias);
rt::ModelSchema schema;
schema.inputs = {TensorSpec{kX, DataType::Float32, {TensorSpec::kDynamicDim, kFeatures}, false}};
schema.outputs = {TensorSpec{kY, DataType::Float32, {TensorSpec::kDynamicDim, kOutputs}, false}};
// Part 4's forward pass, served by a lambda. No subclass needed.
rt::FunctionModel model{schema};
model.on_run_once("run", [&](rt::PhaseContext& ctx) {
const Tensor x = ctx.inputs->get(kX);
ctx.outputs->set(kY, ops::relu(ops::linear(x, w, b)));
});
rt::Executor exec = rt::Executor::create(model); // borrowed; keep `model` alive
// One request in, one response out: the serving loop in miniature.
rt::Request req;
req.inputs.set(kX, Tensor::ones({kBatch, kFeatures}, DataType::Float32));
const rt::Response resp = exec.await(exec.enqueue(std::move(req)));
std::printf("y = %s\n", resp.outputs.get(kY).to_string().c_str());
exec.shutdown();
return 0;
}
```
{/* CERTIFICATION: the arm as shown is verified against the pinned release's
cp313 wheel (tools/tutorial_check.py over the program the page embeds; trace ->
optimize -> finalize -> run, matching its recording). The pipeline executor over graphs
(clika_runtime._core.runtime: Pipeline, Request, Response, enqueue/wait) is
bound but has no published example chapter in the pinned tree (chapters end at
15_trace_and_compile), so the arm stops at the finalized ModelGraph. Extend it
to the executor when the upstream chapter lands. */}
The Python lane serves a traced graph: `crt.trace` captures the forward pass as a `ModelGraph`, and after `optimize()` and `finalize()` the finalized graph is the servable unit. The pipeline executor over graphs is bound but not yet published, so this arm stops at the graph.
```python title="main.py"
import tempfile
import clika_runtime as crt
import clika_runtime.nn.functional as F
def main() -> None:
# The checkpoint from part 4, re-created so this program stands alone.
ckpt = f"{tempfile.gettempdir()}/first_program.safetensors"
crt.io.save_safetensors({
"mlp.weight": crt.full((4, 8), 0.5),
"mlp.bias": crt.full((4,), 0.25),
}, ckpt)
loaded = crt.io.load_safetensors(ckpt)
w, b = loaded["mlp.weight"], loaded["mlp.bias"]
# The traced callable takes and returns LISTS of tensors. crt.trace runs
# it once over data-free stand-ins and captures the operator graph, which
# comes back as recorded: optimize() runs the graph optimizer and
# finalize() readies the graph to serve.
def forward(ins: list) -> list:
return [F.relu(F.linear(ins[0], w, b))]
g = crt.trace(forward, example_inputs=[crt.ones(2, 8)])
g.optimize()
g.finalize()
# A request in, a response out.
(y,) = g.run([crt.ones(2, 8)])
print(f"y = {y}")
if __name__ == "__main__":
main()
```
The Kotlin binding does not carry the serving runtime; the C++ arm is the serving story today.
```text
y = Tensor(shape=[2, 4], dtype=Float32, device=CPU, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...])
```
The same 4.25s part 4 computed, produced this time by an executor answering a request.
## Schemas, lambdas, executors
`ModelSchema` is the model's I/O contract: named `TensorSpec`s with dtype and dims, and `kDynamicDim` leaves a dimension open, so one served model accepts any batch size. `FunctionModel` serves a callable against that schema with no subclass; the lambda reads its inputs and sets its outputs through the `PhaseContext`. `Executor::create` borrows the model (keep it alive) and turns it into a queue: `enqueue` accepts a `Request`, `await` blocks for its `Response`, and `shutdown` drains the queue. Sessions, continuous batching and pipelines build on this same executor; the `runtime` [example project](/examples.md) walks each one.
## One flag for the serving runtime
The serving runtime is built without RTTI, so a target that uses `runtime::` adds one line to part 1's CMake (the bundle's own runtime examples set the same flag):
```cmake
target_compile_options(hello PRIVATE -fno-rtti)
```
Next: [part 6](06-deploy-to-mobile.mdx), the same tutorial program on a real phone.
---
# Tensors and operators
Factories and dtypes, the ops:: library, views, operator sugar, and reading values back to the host.
Source: https://docs.clika.io/clikart/getting-started/first-program/tensors-and-operators.md
Part 1 made one tensor; this part covers the compute vocabulary you will use everywhere: building tensors, transforming them with `ops::`, and reading values back. Same project as [part 1](01-your-first-program.mdx); only `main.cpp` changes.
## The program
```cpp title="main.cpp"
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
int main() {
// Factories build tensors from a shape and a dtype.
const Tensor threes = Tensor::full({2, 3}, 3.0, DataType::Float32);
// from_data copies host bytes into a tensor of the given shape + dtype.
const float host[6] = {0, 1, 2, 3, 4, 5};
const Tensor x = Tensor::from_data(host, {2, 3}, DataType::Float32);
// ops:: free functions return their result directly. Elementwise math
// broadcasts, and a scalar binds wherever a tensor does.
const Tensor y = ops::add(ops::mul(x, 2.0), threes); // y = 2x + 3
// Shape ops are views: the same bytes behind a new layout, no copy.
const Tensor yt = ops::permute(y, {1, 0}); // 2x3 -> 3x2
const Tensor flat = ops::reshape(y, {6});
// Tensor carries operator and method sugar over the same ops, so a
// chain reads like the math it computes.
const Tensor m = (y - 3.0).abs().max();
// Reading back: to_string() for a summary, item() for the single
// element of a one-element tensor, item_as_vec() for a 0-D/1-D tensor.
std::printf("y = %s\n", y.to_string().c_str());
std::printf("y^T = %s\n", yt.to_string().c_str());
std::printf("max|y - 3| = %.0f\n", m.item());
const std::vector v = flat.item_as_vec();
std::printf("flat = [");
for (std::size_t i = 0; i < v.size(); ++i) std::printf("%s%.0f", i ? ", " : "", v[i]);
std::printf("]\n");
return 0;
}
```
```python title="main.py"
import clika_runtime as crt
def main() -> None:
# Factories build tensors directly; dtypes are attributes on the package.
threes = crt.full((2, 3), 3.0)
x = crt.arange(6, dtype=crt.float32).reshape(2, 3)
# Elementwise math reads as operators; scalars broadcast.
y = 2 * x + threes # y = 2x + 3
# Shape methods are views: the same bytes behind a new layout, no copy.
yt = y.permute(1, 0) # 2x3 -> 3x2
flat = y.reshape(-1)
# Method chains: |y - 3| reduced to its global maximum. amax is the
# global reduce; max(dim) is the dim-wise form returning values and
# indices.
m = (y - 3).abs().amax()
# Reading back: repr(t) is the summary, item() reads a scalar, and
# numpy() on a CPU tensor is a zero-copy view when an array is wanted.
print(f"y = {y}")
print(f"y^T = {yt}")
print(f"max|y - 3| = {m.item():.0f}")
v = flat.numpy()
print("flat = [" + ", ".join(f"{e:.0f}" for e in v) + "]")
if __name__ == "__main__":
main()
```
```kotlin title="Main.kt"
import io.clika.runtime.ClikaRt
import io.clika.runtime.Ops
import io.clika.runtime.Tensors
import java.util.Locale
fun main() {
ClikaRt.load()
// Factories build tensors from a shape or from host values; float32 is
// the default dtype, and Tensors.of copies the values in.
val threes = Tensors.full(longArrayOf(2, 3), 3.0)
val x = Tensors.of(floatArrayOf(0f, 1f, 2f, 3f, 4f, 5f), longArrayOf(2, 3))
// Ops carries one function per operator. Elementwise math broadcasts,
// and a number binds wherever a tensor does.
val scaled = Ops.mul(x, 2.0)
val y = Ops.add(scaled, threes) // y = 2x + 3
// Shape operators are views: the same bytes behind a new layout, no copy.
val yt = Ops.permute(y, longArrayOf(1, 0)) // 2x3 -> 3x2
val flat = Ops.reshape(y, longArrayOf(6))
// A chain of operators reads like the math it computes: max|y - 3| = 10.
val d = Ops.sub(y, 3.0)
val ad = Ops.abs(d)
val m = Ops.amax(ad)
// Reading back: summary() renders the shape, dtype, device and values;
// item() reads the one element of a one-element tensor; toFloatArray()
// copies a tensor's values out.
println("y = ${y.summary()}")
println("y^T = ${yt.summary()}")
println("max|y - 3| = ${"%.0f".format(Locale.ROOT, m.item())}")
println("flat = [${flat.toFloatArray().joinToString(", ") { "%.0f".format(Locale.ROOT, it) }}]")
listOf(x, threes, scaled, y, yt, flat, d, ad, m).forEach { it.release() }
}
```
```text
y = Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[3, 5, 7, 9, 11, 13])
y^T = Tensor(shape=[3, 2], dtype=Float32, device=CPU, numel=6, data=[3, 9, 5, 11, 7, 13])
max|y - 3| = 10
flat = [3, 5, 7, 9, 11, 13]
```
## Dtypes
`DataType` names the element format. The everyday set is `Float32`, `Float16`, `BFloat16`, `Float64`, the signed and unsigned integer widths (`Int8` ... `Int64`, `UInt8` ... `UInt64`) and `Bool`; beyond it are the sub-byte integers (`Int4`, `Int2`) and the narrow float families (FP8, FP6, FP4) that quantized models use. `ClikaRT::data_type_name(t.dtype())` prints one; `t.to(DataType::Float16)` casts. Factories take the dtype explicitly. Nothing defaults behind your back.
## Copies are handles
A `Tensor` copy is a cheap reference to the same underlying data, not a deep copy. Writes through one copy are visible through the others, and the data stays alive as long as any copy does. For independent data, build a fresh tensor (a factory or `from_data`, which copies the source bytes and does not retain the pointer).
## The `ops::` library
Every operator is a free function in `ClikaRT::ops`, taking tensors and returning a tensor: elementwise math, reductions, matrix products, convolutions, attention, indexing. This is the operator set a model needs. Shape ops (`reshape`, `permute`, `narrow`) return **views**, metadata over the source's bytes, no copy. For the common ones, `Tensor` adds sugar. Arithmetic operators (`y - 3.0`) and chainable methods (`.abs()`, `.relu()`, `.max()`, `.matmul(...)`) forward to the same `ops::` functions with the same error contract as part 1. The [API reference](/api/index.md) documents every operator.
## Reading values back
Three host-side reads, in increasing weight. `to_string()` is an infallible summary (shape, dtype, device, first values) for logging. `item()` reads the single element of a one-element tensor, typically a reduction result. `item_as_vec()` reads all elements of a 0-D or 1-D tensor, contiguous and dtype-matched. A wrong `T` or a wrong element count raises `ClikaRT::Error`. Every one of these reads is also a synchronization point; [part 3](03-devices-and-async.mdx) explains what that means.
Next: [part 3](03-devices-and-async.mdx), devices, and the asynchrony you have been using without noticing.
---
# Your first pipeline
Load weights and input from disk with ClikaRT::io, compute on the best device, read the results back.
Source: https://docs.clika.io/clikart/getting-started/first-program/your-first-pipeline.md
Part 4 puts the pieces together into the shape every real ClikaRT program has: **load weights and input from disk, place them on a device, compute, read results back**. It is self-contained. Stage 1 writes the files it needs, standing in for a real export.
## The program
```cpp title="main.cpp"
#include
#include
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::NamedTensors;
using ClikaRT::Tensor;
namespace io = ClikaRT::io;
namespace ops = ClikaRT::ops;
namespace {
// The layer's geometry and the checkpoint's names, each spelled once.
constexpr std::int64_t kFeatures = 8; // columns of x, columns of W
constexpr std::int64_t kOutputs = 4; // rows of W, the activations out
constexpr std::int64_t kBatch = 2;
constexpr const char* kWeight = "mlp.weight";
constexpr const char* kBias = "mlp.bias";
Device pick_device() { // from part 3
namespace device = ClikaRT::device;
if (device::is_cuda_available()) return Device::cuda();
if (device::is_vulkan_available()) return Device::vulkan();
if (device::is_metal_available()) return Device::metal();
return Device::cpu();
}
} // namespace
int main() {
const std::filesystem::path dir = std::filesystem::temp_directory_path();
const std::string ckpt = (dir / "first_program.safetensors").string();
const std::string input = (dir / "first_program_input.npy").string();
// ── Stage 1: write the artifacts (a stand-in for a real export) ──
// A checkpoint is a NamedTensors: a name -> tensor map.
NamedTensors weights;
weights.set(kWeight, Tensor::full({kOutputs, kFeatures}, 0.5, DataType::Float32));
weights.set(kBias, Tensor::full({kOutputs}, 0.25, DataType::Float32));
io::save_safetensors(weights, ckpt);
io::save_npy(Tensor::ones({kBatch, kFeatures}, DataType::Float32), input); // two rows of ones
// ── Stage 2: load, compute, read back ──
// Loaders take the target device: the tensors land there directly.
const Device dev = pick_device();
const NamedTensors loaded = io::load_safetensors(ckpt, dev);
const Tensor x = io::load_npy(input, dev);
// One MLP layer: y = relu(x * W^T + b), [2,8] x [4,8]^T -> [2,4].
const Tensor y = ops::relu(ops::linear(x, loaded.get(kWeight), loaded.get(kBias)));
// Per-row mean, then back to the host (the reads from part 3).
const std::vector pooled = ops::mean(y, {1}).item_as_vec();
std::printf("y = %s\n", y.to_string().c_str());
std::printf("row means = [%.2f, %.2f] (expected 8*1*0.5 + 0.25 = 4.25)\n",
pooled[0], pooled[1]);
return 0;
}
```
```python title="main.py"
import tempfile
import clika_runtime as crt
import clika_runtime.nn.functional as F
def main() -> None:
tmp = tempfile.gettempdir()
ckpt = f"{tmp}/first_program.safetensors"
inp = f"{tmp}/first_program_input.npy"
# -- Stage 1: write the artifacts (a stand-in for a real export) --
# A checkpoint is a plain dict: a name -> tensor map.
weights = {
"mlp.weight": crt.full((4, 8), 0.5),
"mlp.bias": crt.full((4,), 0.25),
}
crt.io.save_safetensors(weights, ckpt)
crt.io.save_npy(crt.ones(2, 8), inp) # two rows of ones
# -- Stage 2: load, compute, read back --
# Loaders take the target device: the tensors land there directly.
dev = crt.Device.gpu() # the part 3 probe
loaded = crt.io.load_safetensors(ckpt, device=dev)
x = crt.io.load_npy(inp, device=dev)
# One MLP layer: y = relu(x * W^T + b), [2,8] x [4,8]^T -> [2,4].
# nn.functional (imported as F by convention) carries the stateless
# neural-network operations; the same functions the nn modules call.
y = F.relu(F.linear(x, loaded["mlp.weight"], loaded["mlp.bias"]))
# Per-row mean, then back to the host (the reads from part 3).
pooled = y.mean([1]).to("cpu").numpy()
print(f"y = {y}")
print(f"row means = [{pooled[0]:.2f}, {pooled[1]:.2f}] "
"(expected 8*1*0.5 + 0.25 = 4.25)")
if __name__ == "__main__":
main()
```
The Kotlin binding is an INFERENCE-level binding ([the two kinds](../../bindings.md)): loading a ready model is its job, and the save-side members of this part are FULL-level work, done from C++ or Python. The compute pieces of this part are typed today (`F.linear`, `F.relu`, `F.mean`, `Tensor.to(device)`); the C++ arm is the pipeline story.
On a machine with an NVIDIA GPU (the device line follows what `pick_device()` found):
```text
y = Tensor(shape=[2, 4], dtype=Float32, device=CUDA:0, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...])
row means = [4.25, 4.25] (expected 8*1*0.5 + 0.25 = 4.25)
```
Every element of `y` is `8 * 1 * 0.5 + 0.25 = 4.25`, so both row means print `4.25`, on whatever device the machine offered.
## Model checkpoints are `NamedTensors`
A model's weights travel as a `NamedTensors`, a `name -> tensor` map with `set` / `get` / `size` / `for_each`. `io::save_safetensors` / `io::load_safetensors` round-trip it; `io::load_gguf` returns the same map plus the file's metadata; `io::load_npy` reads a single array. The loaders take the **target device**, so weights stream to where they will be used with no separate `.to()` step. That is the placement lesson from part 3, folded into I/O.
## The pipeline shape
Stage 2 is the skeleton to keep: **discover -> load onto the device -> compute -> read back**. Scaling it up changes the sizes, not the shape (more names in the checkpoint, a deeper stack of `ops::` calls between load and read). The pieces this series did not need live in the [examples](/examples.md). The `runtime` project wraps this exact shape in sessions, continuous batching and pipelines, and `io`'s later chapters cover GGUF and images.
You have a complete, device-portable ClikaRT program. Next: [part 5](05-serve-it.mdx), the same model behind the serving runtime.
---
# Your first program
Link the bundle, print the version, make your first tensor, and meet the error model.
Source: https://docs.clika.io/clikart/getting-started/first-program/your-first-program.md
This tutorial has six parts: four build your first ClikaRT program, one serves it, and one deploys it to mobile. Each part is a complete program that compiles and runs on any machine the runtime supports (no accelerator is needed). Core concepts are explained where they first appear. This part links the bundle, prints the version, and makes one tensor.
The C++ and Kotlin paths assume the bundle is extracted and `CLIKART_BUNDLE_DIR` points at it (see [Quick install](../installation.md)); the Python wheel carries the runtime itself.
Every program here runs compute, so every one of them needs a license credential. `CLIKA_RT_LICENSE` in the environment carries it to the C++ and Kotlin programs and to Python, and `clikart-license-init` stores it once per user account instead; [License the runtime](../../how-to/license-the-runtime.mdx) is the contract. Python has one extra rule, in its tab below.
## The program
```cpp title="main.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
int main() {
std::printf("ClikaRT %s\n", ClikaRT::GetVersionInfo().c_str());
// A 2x3 tensor of ones on the CPU (the default device).
const Tensor a = Tensor::ones({2, 3}, DataType::Float32);
// ops:: are free functions returning the value directly.
const Tensor b = ops::add(a, a);
std::printf("a + a =\n%s\n", b.to_string().c_str());
try {
const Tensor bad = ops::matmul(a, Tensor::ones({5, 7}, DataType::Float32));
} catch (const ClikaRT::Error& e) {
std::printf("failed: %s\n", e.what()); // e.status() has the coarse category
}
return 0;
}
```
```python title="main.py"
import clika_runtime as crt
def main() -> None:
print(f"ClikaRT {crt.version()}")
# A 2x3 tensor of ones on the CPU (the default device). Factories take
# the shape directly; float32 is the default dtype.
a = crt.ones(2, 3)
# Operators read as math; every result is a new tensor.
b = a + a
print(f"a + a =\n{b}")
try:
bad = a @ crt.ones(5, 7)
except RuntimeError as e:
print(f"failed: {e}") # the message ends in the machine-readable [code: NAME]
if __name__ == "__main__":
main()
```
```kotlin title="Main.kt"
import io.clika.runtime.ClikaRt
import io.clika.runtime.ClikaRtException
import io.clika.runtime.Ops
import io.clika.runtime.Tensors
fun main() {
ClikaRt.load()
println("ClikaRT ${ClikaRt.version()}")
// A 2x3 tensor of ones on the CPU (the default placement); float32 is
// the default dtype.
val a = Tensors.ones(longArrayOf(2, 3))
// Every operator is one function of Ops. It returns a fresh tensor and
// leaves its operands live; a tensor is AutoCloseable, so release the
// ones you named, or scope them with `use`.
val b = Ops.add(a, a)
println("a + a =\n${b.summary()}")
try {
Tensors.ones(longArrayOf(5, 7)).use { bad -> Ops.matmul(a, bad) }
} catch (e: ClikaRtException) {
println("failed: ${e.message}") // e.codeName is the machine channel, e.status the category
}
b.release()
a.release()
}
```
## Set up, build, run
The setup is where the languages differ.
The CMake side is three lines of substance (the package and one target).
```cmake title="CMakeLists.txt"
cmake_minimum_required(VERSION 3.19)
project(hello_clikart LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
find_package(ClikaRT CONFIG REQUIRED)
add_executable(hello main.cpp)
target_link_libraries(hello PRIVATE ClikaRT::ClikaRT)
```
```bash
cmake -S . -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build
./build/hello
```
The wheel carries the runtime and its backends; the bundle is not needed on this path. [Get ClikaRT](../get-clikart.mdx#python-wheels) has the wheel to download (on Linux, the flavor follows your NVIDIA driver) and the install command. `numpy` is the array bridge for building tensors from Python data (optional by design, and these pages use it).
```bash
pip install ./clika_runtime---.whl
pip install numpy
export CLIKA_RT_LICENSE=CLIKA1-...
python main.py
```
`import clika_runtime` loads the runtime, so `CLIKA_RT_LICENSE` has to be set before the import runs. A shell `export` does that. From inside the program, assign `os.environ["CLIKA_RT_LICENSE"]` above the import. The wheel also installs `clikart-license-init` as a console script, which stores the credential once per user account and drops the variable entirely.
{/* CERTIFICATION: the Kotlin arm follows the release's Kotlin artifact (clika-runtime-maven-.zip)
and the tutorial's gradle project beside the programs, both of the pinned release. */}
The Kotlin artifact, `clika-runtime-maven-0.6.4.zip`, extracts to a Maven repository directory that the gradle project names through the `clikaRtMavenRepo` property; it ships the Kotlin API and its JNI bridge. `libClikaRT.so` comes from the bundle, so run the JVM with `-Djava.library.path` pointing at the bundle's `lib` directory.
```kotlin title="build.gradle.kts"
// examples/kotlin/tutorial: the Kotlin programs of the getting-started
// tutorial, one file per chapter (NN_.kt), built for a desktop JVM
// against the release's Kotlin artifact, io.clika:clika-runtime (the
// extracted clika-runtime-maven-.zip, named to settings.gradle.kts
// by -PclikaRtMavenRepo=). At run time ClikaRt loads two libraries:
// libClikaRT.so by name from java.library.path, a runtime distribution's
// lib/, and the JNI bridge libclika_rt_jni.so from the artifact's own jar
// (its native// entry, extracted at start):
//
// gradle installDist -PclikaRtMavenRepo=
// java -Djava.library.path=/lib -cp 'build/install/tutorial/lib/*' _01_your_first_programKt
//
// or in one step (CLIKA_RT_LIB_DIR in the environment names /lib too):
//
// gradle run -Pchapter=01_your_first_program -PclikaRtMavenRepo= -PclikaRtLibDir=/lib
//
// 06_deploy_to_mobile.kt is the Android chapter's Kotlin half: it compiles
// here and has no main to run.
plugins {
kotlin("jvm") version "2.3.0"
application
}
java {
toolchain {
languageVersion.set(JavaLanguageVersion.of(17))
}
}
// One compilation: the chapter files beside this build script. Each chapter's
// top-level main lands in its own class, named after the file
// (01_your_first_program.kt -> _01_your_first_programKt).
sourceSets {
main {
kotlin.setSrcDirs(listOf(layout.projectDirectory))
kotlin.include("??_*.kt")
}
}
dependencies {
// The runtime's Kotlin API, the release's artifact (its version read from the repository directory).
implementation("io.clika:clika-runtime:${project.extra["clikaRtVersion"]}")
}
val chapter = providers.gradleProperty("chapter").orElse("01_your_first_program")
val clikaRtLibDir = providers.gradleProperty("clikaRtLibDir")
.orElse(providers.environmentVariable("CLIKA_RT_LIB_DIR"))
.orElse("")
application {
mainClass.set(chapter.map { "_${it}Kt" })
}
tasks.named("run") {
workingDir = projectDir
doFirst {
if (clikaRtLibDir.get().isEmpty()) {
throw GradleException("name the runtime distribution's lib/: -PclikaRtLibDir=/lib or CLIKA_RT_LIB_DIR")
}
}
systemProperty("java.library.path", clikaRtLibDir.get())
}
```
```bash
gradle run -Pchapter=01_your_first_program -PclikaRtMavenRepo="$PWD/clika-runtime-maven-0.6.4" -PclikaRtLibDir="$CLIKART_BUNDLE_DIR/lib"
```
```text
ClikaRT 0.6.4
a + a =
Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[2, 2, 2, 2, 2, 2])
failed: ops::matmul: cannot contract a [2, 3] with b [5, 7]: a's last dim (3) must equal b's dim 0 (5)
```
The version line proves the runtime loaded, and the tensor prints its shape, dtype, device and values, every element `2`. The last line is the failure path, demonstrated at the end of the program and explained below.
## One library, one header
The umbrella header `ClikaRT/clika_rt.h` includes the entire public API (tensors, operators, devices, I/O, the serving runtime). `ClikaRT::ClikaRT` is the only link target. Nothing else from the bundle enters your build; the accelerator backends are shared libraries the runtime loads on its own at run time ([part 3](03-devices-and-async.mdx)).
## Values out, `ClikaRT::Error` on failure
Every fallible call in the public API (factories like `Tensor::ones`, operators like `ops::add`) returns its result directly and raises `ClikaRT::Error` on failure. There are no output parameters and no error codes to check at each call; a shape mismatch or an unavailable device surfaces as an exception where the call was made:
```cpp
try {
const Tensor bad = ops::matmul(a, Tensor::ones({5, 7}, DataType::Float32));
} catch (const ClikaRT::Error& e) {
std::printf("failed: %s\n", e.what()); // e.status() has the coarse category
}
```
```python
try:
bad = a @ crt.ones(5, 7)
except RuntimeError as e:
print(f"failed: {e}") # the message ends in the machine-readable [code: NAME]
```
```kotlin
try {
Tensors.ones(longArrayOf(5, 7)).use { bad -> Ops.matmul(a, bad) }
} catch (e: ClikaRtException) {
println("failed: ${e.message}") // e.codeName is the machine channel, e.status the category
}
```
Catch it where you can act on it. The programs in this series let a failure terminate the process, the right default for a small batch program. The message names the operation and the values it refused (here the two shapes); the stable, machine-readable channel is the code name, which [Handle errors by code](../../how-to/handle-errors-by-code.mdx) covers.
Next: [part 2](02-tensors-and-operators.mdx), tensors and the `ops::` library properly.
---
# Get ClikaRT
Pick your OS, architecture and accelerator, and get the download, verify and extract commands for the archive the CLIKA platform hands out.
Source: https://docs.clika.io/clikart/getting-started/get-clikart.md
import GetClikaRT from '@site/src/components/GetClikaRT';
ClikaRT is a set of downloads on the CLIKA platform: one release archive per platform, the Python wheels, the Kotlin artifact and the examples archive, every file with its SHA-256 shown beside it. Take every one you use from the same release. The release the dialog offers is the deployment's own pin, which may differ from the release these pages are written for (`0.6.4`); a download's own `README.md` is the authority for the release it came from, its commands and file names included. Your platform deployment's **Download ClikaRT** dialog (your name in the top bar, or a project's Overview), the platform CLI (`clika-cli runtime-sdk download`) and the MCP tools all hand out the same files; [Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) is the walk through each surface.
Pick the platform you are deploying to and copy the commands. Each platform has its own archive, named `ClikaRT__-.tar.xz` (`.zip` on Windows); the tree inside is that platform's complete bundle, the runtime and the Modelverse model library together. The dialog shows the archive's SHA-256 as text with a Copy control; paste it in place of `` below.
Then continue with [Quick install](installation.md), which picks up at the extracted bundle. [Complete installation](complete-installation.md) has per-platform depth.
## License credential
ClikaRT runs compute under a license, so every process that runs an operator or a model needs the credential CLIKA issued for your deployment. The credential is the `CLIKA1-...` text your project's license shows as its license key, the same text for an online license key and an offline license bundle; an API key (`clika_rk_...`) is not a credential. [Runtime licenses](/platform/concepts/runtime-licenses) is where a credential comes from, and the download dialog shows your project's key beside the archives.
Put it in the environment variable `CLIKA_RT_LICENSE`, as the credential text or as the path of a file holding it, before the program starts:
```bash
export CLIKA_RT_LICENSE=CLIKA1-... # the credential text
export CLIKA_RT_LICENSE=/etc/clika/clikart.license # or the path of a file holding it
```
Or store it once per user account with `clikart-license-init`, which ships in every archive's `bin/` (`clikart-license-init.exe` on Windows) and as a console script of the `clika-runtime` wheel; every later process under that account reads the file when the variable is not set, and the variable wins when both are present:
```bash
"$CLIKART_BUNDLE_DIR"/bin/clikart-license-init CLIKA1-...
```
Without a valid credential a call is refused with the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a license past its end date. [License the runtime](../how-to/license-the-runtime.mdx) is the whole contract: where each language and packaging puts the credential, and how a program reads a refusal's code name.
## All combinations
The pattern is the same for every platform; only the archive name and the checksum command change. The dialog shows the SHA-256 beside the download button, with the one line that checks the file on your platform:
```bash
echo " ClikaRT_linux_x86_64-0.6.4.tar.xz" | sha256sum -c
tar -xf ClikaRT_linux_x86_64-0.6.4.tar.xz
export CLIKART_BUNDLE_DIR="$PWD/ClikaRT_linux_x86_64-0.6.4"
```
On macOS the check is `shasum -a 256 -c` over the same line. On Windows, in PowerShell: `(Get-FileHash ).Hash -eq ""`, then `tar -xf ` and `$env:CLIKART_BUNDLE_DIR = "$PWD\"`. The platform CLI runs the check itself and refuses a file that does not match.
Per platform, the archive and the driver requirements:
| OS | Architecture | Archive | Accelerators and drivers |
| --- | --- | --- | --- |
| Linux | x86_64 | `ClikaRT_linux_x86_64-.tar.xz` | CPU (no driver) · CUDA (NVIDIA driver only) · Vulkan (Vulkan 1.2 or newer) |
| Linux | arm64 | `ClikaRT_linux_arm64-.tar.xz` | CPU · CUDA · Vulkan, as above |
| Windows | x86_64 | `ClikaRT_windows_x86_64-.zip` | CPU · Vulkan |
| Windows | arm64 | `ClikaRT_windows_arm64-.zip` | CPU · Vulkan |
| macOS | Apple silicon | `ClikaRT_macos_arm64-.tar.xz` | CPU · Metal (ships with macOS) |
| Android | arm64-v8a | `ClikaRT_android_arm64-.tar.xz`, then [Deploy to mobile](first-program/06-deploy-to-mobile.mdx) | CPU · Vulkan |
The two Linux archives carry both CUDA images, `lib/libClikaRT_cuda13x.so` and `lib/libClikaRT_cuda12x.so`, and each image is self-contained: the CUDA runtime and cuBLASLt are inside it. A CUDA machine needs its NVIDIA driver and nothing else, no CUDA toolkit install and no `LD_LIBRARY_PATH`; the runtime loads the image the driver serves (a driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image).
Every desktop archive verifies the same way after extraction: run `examples/bin/version_00_hello_world` (`.exe` on Windows) from the extracted root; it prints the ClikaRT version.
## Python wheels
The `clika-runtime` wheel is the dialog's Python build: pick the platform, the CUDA image (Linux) and the CPython version (3.10 to 3.14), and the download is a zip holding the wheel for that selection, with its SHA-256 shown like the archives'. The wheel carries the runtime and its backends inside the package, and the [Modelverse](/modelverse) model library as `clika_runtime.modelverse`, so the Python path needs no archive and no library paths to set. On Linux (x86_64 and aarch64) the wheel comes in three flavors, told apart by the local version tag; the NVIDIA driver picks the flavor (`nvidia-smi` prints the driver version in its header):
| Version | Contents | Pick it when |
| --- | --- | --- |
| `0.6.4` | every non-CUDA backend, no CUDA image | the machine has no NVIDIA GPU |
| `0.6.4+cu13x` | the non-CUDA backends plus the CUDA 13 image | the NVIDIA driver is major version 580 or newer |
| `0.6.4+cu12x` | the non-CUDA backends plus the CUDA 12 image | the NVIDIA driver is older than 580 |
A CUDA flavor's image is self-contained, as in the archives: the driver is the one requirement, no CUDA toolkit and no `LD_LIBRARY_PATH`. macOS (Apple silicon) and Windows x64 carry no CUDA and ship one unsuffixed wheel each for CPython 3.10 to 3.14; Windows arm64 ships one for CPython 3.11 to 3.14. No wheel declares a requirement on an NVIDIA package. Check the downloaded zip against the SHA-256 the dialog shows, unzip it, and install the wheel inside from the file; `numpy` is the array bridge the tutorial uses, installed separately.
```bash
echo "" | sha256sum -c
unzip
pip install ./clika_runtime-0.6.4+cu13x-cp313-cp313-manylinux_2_28_x86_64.whl
pip install numpy
python -c "import clika_runtime as crt; print(crt.version())"
```
The last line prints `0.6.4`. The platform CLI's `clika-cli runtime-sdk pip-command` prints one `pip install` line against the deployment's own package index instead, which picks the wheel for the interpreter it runs under. [Use ClikaRT from Python](../how-to/use-clikart-from-python.mdx) is the guide to the wheel's surface.
Install these wheels with `pip`. Every wheel is an LZMA-compressed zip, the one compression method that fits the `+cu12x` flavor's CUDA image under the release asset size limit. `pip` reads LZMA on any CPython 3.10 or newer whose build carries the `lzma` module, which the python.org, manylinux, conda and distribution builds do. `uv pip install` does not read LZMA-compressed wheels.
## Modelverse from Python
The same wheel carries [Modelverse](/modelverse), the model library, as the `clika_runtime.modelverse` subpackage: model resolution, generation, pipelines and the serving API. There is no second wheel.
```bash
python -c "import clika_runtime.modelverse as mv; print(mv.__version__)"
```
The line prints `0.6.4`. Importing the subpackage loads the runtime first, then the Modelverse extension, which finds the runtime library the same wheel installed; a runtime and a model library from different releases refuse to import, naming both versions. The `clikart-cli` command-line program is a console script of the same wheel.
## The Kotlin artifact and the examples
Three more downloads sit beside the archives and the wheels in the dialog's listing, taken from the same release:
- `clika-runtime-maven-.zip`, the Kotlin artifact: a Maven repository directory holding `io.clika:clika-runtime` as an Android AAR and a desktop JVM jar under one coordinate. Extract it anywhere and name the directory to gradle; on Android the AAR is the whole runtime, and on the desktop JVM the jar runs beside the release archive of the same platform (its `lib/` on `java.library.path`); [Language bindings](../bindings.md) says what it carries and [Package ClikaRT in an Android app](../how-to/package-clikart-in-an-android-app.mdx) how an app ships it.
- `release_manifest.json`, the release manifest: one entry per published file with the platform, build type and language it serves, the backends inside it, the files it needs beside it, its digest and size, and the pieces of a split archive in join order; a download page or a script reads the release from it instead of from file names. Its `.sha256` sits beside it like every other file.
- `ClikaRT--examples.tar.xz`, the examples archive: complete programs in C++, Python and Kotlin with their recorded output, under one `examples/` directory whose `README.md` and `AGENTS.md` index them. [Additional examples](../examples.md) is the catalog.
[System requirements](/system-requirements.md) is the authority on platforms, drivers and hardware.
---
# Quick install
Install a C++ toolchain, extract the ClikaRT bundle, and prove it works in two commands.
Source: https://docs.clika.io/clikart/getting-started/installation.md
The fast path from nothing to a running ClikaRT program. For per-platform depth, IDE setup and the full failure catalog, see [Complete installation](complete-installation.md).
## 0. Prerequisites
A C++17 compiler and CMake 3.19 or newer:
```bash
sudo apt install build-essential cmake # Debian/Ubuntu
brew install cmake # macOS; compiler: xcode-select --install
winget install Kitware.CMake # Windows; compiler: Visual Studio Build Tools
```
A runtime credential, issued by your platform for the project the program belongs to; step 2 below puts it in place.
## 1. Get the bundle
Extract the archive for your platform ([Get ClikaRT](get-clikart.mdx) has the download and checksum commands) and remember where it is:
```bash
tar -xf ClikaRT_linux_x86_64-.tar.xz
export CLIKART_BUNDLE_DIR="$PWD/ClikaRT_linux_x86_64-"
ls "$CLIKART_BUNDLE_DIR"
```
The `ls` is the success check: `bin/`, `cmake/`, `include/`, `lib/` and `examples/` are among the entries. A command-line-only download from the platform's dialog carries `bin/` and `lib/` and no `cmake/`, `include/` or `examples/`; those come with the C/C++ and Everything downloads, and a composed archive's own `README.md` says exactly what it holds. A new terminal loses the `export`; re-run it there, every command below uses it. On Windows, run these in PowerShell (`tar` is built in since Windows 10).
## 2. Place the license credential
ClikaRT runs compute under a license. Put the `CLIKA1-...` credential your project's license shows in the environment, or store it once under your user account:
```bash
export CLIKA_RT_LICENSE=CLIKA1-... # this shell, and the programs it starts
"$CLIKART_BUNDLE_DIR"/bin/clikart-license-init CLIKA1-... # or once, for this user account
```
[License the runtime](../how-to/license-the-runtime.mdx) has the whole contract: what the credential is, where each language and packaging reads it, and what a refused call carries.
## 3. Run a pre-built example
The fastest proof the bundle works. Desktop archives ship pre-built example binaries; run one in place. Nothing gets installed, and no GPU or driver is needed:
```bash
"$CLIKART_BUNDLE_DIR"/examples/bin/version_00_hello_world
```
```text
ClikaRT 0.6.4
```
On a platform without pre-built binaries, skip to step 4. On a Jetson, extract the `linux_arm64` archive; its binaries run out of the box there, as on any other arm64 Linux machine.
## 4. Build the version example yourself
The same proof through your own toolchain. The three flags, once: `-S` is the source directory, `-B` is the build directory, and `ClikaRT_DIR` tells CMake where the bundle's `cmake/` folder is.
```bash
cmake -S "$CLIKART_BUNDLE_DIR/examples/src/version" -B build-version \
-DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build-version
./build-version/version_00_hello_world
```
```text
ClikaRT 0.6.4
```
## If something failed
- `cmake: command not found`: step 0.
- `tar: xz: Cannot exec` or `xz: command not found`: install `xz-utils` (Debian/Ubuntu) or `xz`.
- `Could not find a package configuration file provided by "ClikaRT"`: `ClikaRT_DIR` is wrong, or the `export` was lost to a new terminal. An error naming a bare `/cmake` path means the variable expanded empty; re-run the `export` from step 1.
- The binary is missing after an MSVC build: it lands in a configuration subdirectory, `build-version\Debug\version_00_hello_world.exe`.
The install works. [Write your first program](first-program/01-your-first-program.mdx).
---
# What to read next
Where the documentation goes after the tutorial: how-to guides, the API reference, Modelverse, system requirements.
Source: https://docs.clika.io/clikart/getting-started/next-steps.md
You installed the bundle and wrote a device-portable program. The rest of the documentation, in a useful reading order:
- **[How-to guides](/how-to/index.md)**: problem-oriented recipes past the tutorial, from quantized weights to HTTP serving. [Additional examples](/examples.md) sit inside it: standalone programs in the bundle, one per subsystem; `compute` and `async` deepen what the tutorial introduced, and `io`, `runtime`, `tokenizer` and `http_server` carry it to a served model.
- **[API reference](/api/index.md)**: every public namespace, class and function, generated from the headers you compile against. `ClikaRT/clika_rt.h` includes them all.
- **[Modelverse](/modelverse)**: the CLIKA model library. Models packaged so ClikaRT loads and runs them as they are, with quantized variants per class of device.
- **[System requirements](/system-requirements.md)**: the platform, toolchain, driver and accelerator tables, for when you leave your development machine.
---
# ClikaRT at a glance
What the ClikaRT runtime is and is not, and the five-minute mental model.
Source: https://docs.clika.io/clikart/getting-started/overview.md
ClikaRT is CLIKA's inference runtime, a C++ library (with Python and Kotlin bindings) that loads models and runs them on CPUs, GPUs, and other hardware accelerators through one public API. It is not a training framework, and not a bundle of separate CUDA, Vulkan and Metal wrappers. You link one library, include one header, and the same code runs on every backend the runtime ships.
## The mental model
Five ideas carry the whole library, in the order you meet them when you build.
1. **Everything arrives in one bundle.** A directory with the public headers, the libraries per platform, a CMake package and the examples. `find_package(ClikaRT CONFIG)` and the target `ClikaRT::ClikaRT` are the entire integration.
2. **Models load from files you already have.** `ClikaRT::io` reads safetensors, GGUF and NumPy checkpoints, plus images and audio, and a model checkpoint arrives as a `name -> tensor` map on the device you name.
```cpp
const NamedTensors weights = io::load_safetensors("model.safetensors");
const Tensor w = weights.get("layer.weight");
```
```python
weights = crt.io.load_safetensors("model.safetensors")
w = weights["layer.weight"]
```
The Kotlin binding carries no checkpoint reader: a model reaches an app through the model library (`AutoModel.fromPretrained`), or as an ONNX model or a compiled graph; tensors and operators carry the `Ops` surface, one function per operator.
3. **`Tensor` is the core building block.** Each tensor is assigned to a device (the CPU by default; `.to(device)` moves it), and `ops::` operators run where their inputs live. The CPU backend is always present; CUDA, Vulkan and Metal load at run time where the machine has them. Everything returns values directly and raises `ClikaRT::Error` on failure.
```cpp
const Tensor a = Tensor::ones({2, 3}, DataType::Float32);
const Tensor b = ops::add(a, a);
std::printf("%s\n", b.to_string().c_str());
```
```python
a = crt.ones(2, 3)
b = a + a
print(b)
```
```kotlin
val a = Tensors.ones(longArrayOf(2, 3))
val b = a + a // operator extensions return a fresh tensor; operands stay live
println(b.summary())
```
4. **Execution is asynchronous by nature.** An `ops::` call dispatches work and returns; reads wait for the result, so you never observe unfinished bytes.
5. **Serving is built in.** The serving runtime adds sessions, continuous batching and pipelines, so the model you loaded answers requests. Declare a schema, serve it with a lambda, drive it through an executor:
```cpp
rt::FunctionModel model{schema};
model.on_run_once("run", [](rt::PhaseContext& ctx) {
ctx.outputs->set("y", ctx.inputs->get("x") * 2.0);
});
rt::Executor exec = rt::Executor::create(model);
const rt::Response resp = exec.await(exec.enqueue(std::move(req)));
```
The Python lane serves a traced graph: `crt.trace` captures a forward pass as a `ModelGraph`, which runs once finalized, and `crt.runtime` carries the `Executor` and `Pipeline` over it. [Part 5](first-program/05-serve-it.mdx) walks it.
The Kotlin binding carries the serving runtime (`FunctionModel`, `Executor`, `Pipeline`) and the model library's server (`Modelverse.serve`), on a phone and on a desktop JVM alike.
The [first program series](first-program/01-your-first-program.mdx) turns these into working programs, explaining each where it first appears; the serving runtime has its own [example project](/examples.md).
## Platforms
Linux (x86_64, arm64), Android (arm64), Windows (x86_64, arm64) and macOS (Apple silicon), one distribution per platform and architecture, and a distribution works out of the box on every machine of its class: the linux-arm64 dist runs on a Jetson the same way it runs on an arm64 server. [System requirements](/system-requirements.md) has the platform and accelerator tables.
---
# How-to guides
Problem-oriented recipes. Each guide answers one "how do I X" with a worked example, verified against the bundle.
Source: https://docs.clika.io/clikart/how-to.md
Practical guides covering common tasks. Each guide answers one concrete "how do I X" with a worked example: real code, built against the bundle, with its real output. Read the [tutorial](../getting-started/first-program/01-your-first-program.mdx) first; the guides assume its ground (tensors, devices, the async model) and go deeper on one problem at a time, in any order. Entries marked as coming are planned and land here as they are written.
Every section carries a C++ and a Python arm, each a program that builds and runs against the release with its output recorded beside it, and a Kotlin note where the binding reaches the surface. Where a surface is the C++ tier's alone (the HTTP server, the processor family, a custom kernel), the tab says so and points at the tab that has it. The Kotlin binding is an INFERENCE-level binding ([the two kinds](../bindings.md)): a Kotlin tab shows what runs a ready model, and where a guide reaches past that level the tab names the member it lacks. Three guides are Python-only because their subject exists in the Python package alone: authoring a model in Python, PyTorch interop, and pytrees. The six graph guides (query, edit, optimize and finalize, write a transform, merge and split, and export to ONNX) carry a C++ and a Python arm, the two APIs that read and change a graph's structure. The licensing guide's arms show where each language puts the runtime credential, not a program with an output.
## Weights and data
- [Load quantized weights from a GGUF file](load-quantized-weights.mdx): read a block-quantized checkpoint and serve it at its on-disk footprint.
- [Load images and audio for inference](load-images-and-audio.mdx): decode files into tensors, resample audio, preprocess with `ops::`.
## Models and graphs
- [Run an ONNX model](run-an-onnx-model.mdx): open or build a graph, compile it into a `ModelGraph`, optimize and finalize it, execute by position or by name.
- [Preprocess inputs with processors](preprocess-with-processors.mdx): the model's own resize/rescale/normalize recipe, or a log-mel front end, from knobs or its config file.
- [Author a model in Python](author-a-model-in-python.mdx): a decoder-only language model as `nn.Module`s from a Hugging Face config and its safetensors, meta init, fused projections, a KV cache and a greedy decode loop, measured against the C++ command line.
- [Query a graph](query-a-graph.mdx): a `ModelGraph`'s nodes, values and edges, search, walks and orders, paths and the critical path, dominance, regions and the cheapest cut, and views across edits.
- [Edit a graph](edit-a-graph.mdx): change a `ModelGraph` before `finalize()`, from rewiring a value's reads and changing what a node reads to renaming nodes and editing the inputs and outputs, then optimize, finalize and run it.
- [Optimize and finalize](optimize-and-finalize.mdx): run the graph optimizer on a `ModelGraph` and read its report, choose its transforms and defaults, place it with `to()`, and `finalize()` it to run.
- [Write a transform](write-a-transform.mdx): functions over the graph edits that `optimize()` runs beside the runtime's transforms, operators built with `add_node`, `insert` and `replace`, and rewrite rules over a `Pattern`.
- [Merge and split](merge-and-split.mdx): join `ModelGraph`s side by side, as a pipeline or across two devices, cut one into parts at its cheapest cut or by a function, extract the operators between values, and merge the parts back.
- [Export a graph to ONNX](export-a-graph-to-onnx.mdx): write a `ModelGraph` as an ONNX model file with `export_onnx`, or build the same model in memory with `OnnxModel::from_graph`; read it back, keep a named dynamic dimension, choose where the weights go and the opset, and meet what the export refuses.
## Text and chat
- [Tokenize text and apply a chat template](tokenize-and-chat-templates.mdx): text to token ids and back, byte offsets, the model's own prompt format, batching for a model, streaming decode for a generation loop.
## Serving
- [Serve a model over HTTP](serve-over-http.mdx): routes, a JSON inference endpoint, and server-sent events for streaming.
## Execution and memory
- [Control asynchronous execution](control-async-execution.mdx): dispatch vs ready, safe host reads, completion callbacks, synchronous and tracing scopes.
- [Trace eager code to graphs](trace-eager-code-to-graphs.mdx): capture a function as a `ModelGraph` with `trace`, optimize, finalize and run it, and wrap hot paths in `compile`.
- [Wrap existing memory without copying](wrap-existing-memory.mdx): tensors over buffers your application already owns.
- [Write a custom operator](write-a-custom-operator.mdx): compose built-ins in an `nn::Module`, or launch your own kernel on the stream's native handle.
- [Structure inputs and outputs as pytrees](pytrees.mdx): nested containers of tensors at every boundary (`compile`, `trace`, `eval`, `save`), the registry for your own classes, key paths and serialization.
## Integration and packaging
- [Add ClikaRT to an existing CMake project](existing-cmake-project.md): `find_package` against the bundle, or `add_subdirectory`, in a project that already builds.
- [Coming from PyTorch or Hugging Face](coming-from-pytorch.mdx): the conventions that differ, each with the one line that bridges it: channels-last and OHWI, integer slots, reads that wait, the accessors, the tokenizer's files, and the generate calls side by side.
- [Use ClikaRT from Python](use-clikart-from-python.mdx): the `clika-runtime` wheel, the NumPy boundary, models as `nn.Module`, errors you can branch on.
- [Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx): `torch.compile(model, backend="clika")`, zero-copy tensor exchange over DLPack, and `from_torch_module` for an eager module tree.
- [Handle errors by code](handle-errors-by-code.mdx): the three channels every failure carries, the stable code name to branch on, the coarse status for policy.
- [License the runtime](license-the-runtime.mdx): where the runtime reads the project's credential, one setting per language and packaging, and the code names a refused call carries.
- [Package ClikaRT in an Android app](package-clikart-in-an-android-app.mdx): the two artifacts and the libraries copied beside them, the build requirements, the license credential, the thread count and the cache root, a model on the phone, and a server that outlives the screen.
## Tools
- Build a command-line model tool (coming): typed argument parsing, progress bars and structured logging in one small tool.
## Complete programs
- [Additional examples](/examples.md): the bundle's standalone projects, one per subsystem, each building its topic up chapter by chapter.
For every public name, the [API reference](/api/index.md).
---
# Author a model in Python
Write a Llama-class decoder as clika_runtime.nn modules, load a Hugging Face checkpoint as one resident copy, fuse its projections, decode through a KV cache with a one-token lookahead, and measure it against the C++ command line.
Source: https://docs.clika.io/clikart/how-to/author-a-model-in-python.md
A model authored in Python over `clika_runtime` runs at the speed of the runtime's operators: Python decides which operator runs next, the kernels do the work, and the loop never reads a value it does not need. This guide walks the wheel's own chapter, `examples/python/howto/author_a_model_in_python/llama/llama_from_scratch.py`, a Llama-class decoder in about six hundred lines that loads a Hugging Face checkpoint, generates greedily and matches the `clikart-cli` command line token for token. Every sample below is an excerpt of that file; the chapter's README holds the command lines and the measured table.
The walk is linear: the configuration, the module tree, the load, the fusions, the cache, the attention step, the decode loop, the measurement. [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module basics this page builds on; [Tokenize text and apply a chat template](tokenize-and-chat-templates.mdx) covers the tokenizer the prompt goes through.
## The configuration comes from config.json
The architecture is a dataclass whose fields carry the checkpoint's own names, so `config.json` fills it directly. `from_dict` keeps the keys the dataclass declares and drops the rest, and `__post_init__` derives what the file leaves implicit (the key/value head count, the head size). The rotary tables come from one operator, with Llama 3's frequency scaling when the checkpoint declares it.
```python title="llama_from_scratch.py (excerpt)"
@dataclasses.dataclass
class LlamaConfig:
hidden_size: int
intermediate_size: int
num_hidden_layers: int
num_attention_heads: int
vocab_size: int
num_key_value_heads: int | None = None
head_dim: int | None = None
rms_norm_eps: float = 1e-5
rope_theta: float = 10000.0
rope_scaling: dict | None = None
tie_word_embeddings: bool = False
max_position_embeddings: int = 8192
eos_token_id: int | list[int] | None = None
def __post_init__(self) -> None:
if self.num_key_value_heads is None:
self.num_key_value_heads = self.num_attention_heads
if self.head_dim is None:
self.head_dim = self.hidden_size // self.num_attention_heads
@classmethod
def from_dict(cls, fields: dict) -> "LlamaConfig":
known = inspect.signature(cls).parameters
return cls(**{name: value for name, value in fields.items() if name in known})
def rotary_tables(self, max_positions: int, device: crt.Device) -> tuple[crt.Tensor, crt.Tensor]:
scaling = self.rope_scaling or {}
if self.rope_llama3:
return crt.generate_rotary_cache(
self.head_dim, max_positions, theta=self.rope_theta, scaling="llama3",
scale=float(scaling["factor"]),
low_freq_factor=float(scaling["low_freq_factor"]),
high_freq_factor=float(scaling["high_freq_factor"]),
original_max_pos=int(scaling["original_max_position_embeddings"]),
device=device,
)
return crt.generate_rotary_cache(self.head_dim, max_positions, theta=self.rope_theta,
device=device)
```
## A module tree with the checkpoint's names
`load_state_dict` binds by dotted name, so the tree spells the checkpoint's names: `model.layers.3.self_attn.q_proj.weight` is the attribute path `model.layers[3].self_attn.q_proj` and its `weight`. Every constructor takes the model dtype and declares each layer at it (the chapter resolves it from `--dtype`, else from the checkpoint's embedding table), so a bind adopts a matching checkpoint tensor as it is and no layer sits at the default dtype beside a bfloat16 checkpoint. The attention module owns the four projections and, after the load, one fused group for the first three.
```python title="llama_from_scratch.py (excerpt)"
class LlamaAttention(nn.Module):
def __init__(self, config: LlamaConfig, dtype: crt.dtype) -> None:
super().__init__()
hidden, heads, kv_heads, head_dim = (config.hidden_size, config.num_attention_heads,
config.num_key_value_heads, config.head_dim)
self.num_heads, self.num_kv_heads, self.head_dim = heads, kv_heads, head_dim
self.q_proj = nn.Linear(hidden, heads * head_dim, bias=False, dtype=dtype)
self.k_proj = nn.Linear(hidden, kv_heads * head_dim, bias=False, dtype=dtype)
self.v_proj = nn.Linear(hidden, kv_heads * head_dim, bias=False, dtype=dtype)
self.o_proj = nn.Linear(heads * head_dim, hidden, bias=False, dtype=dtype)
class LlamaMLP(nn.Module):
def __init__(self, config: LlamaConfig, dtype: crt.dtype) -> None:
super().__init__()
self.gate_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False, dtype=dtype)
self.up_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False, dtype=dtype)
self.down_proj = nn.Linear(config.intermediate_size, config.hidden_size, bias=False, dtype=dtype)
def forward(self, x: crt.Tensor) -> crt.Tensor:
return self.down_proj(crt.swiglu(self.gate_up(x)))
```
`LlamaDecoderLayer` holds one attention, one MLP and the two `nn.RMSNorm` weights; `LlamaModel` holds the `nn.Embedding`, an `nn.ModuleList` of layers and the final `norm`; `LlamaForCausalLM` adds `lm_head` and keeps the model dtype as `self.dtype`. The names are the checkpoint's at every level.
## Build storage-free, then adopt the checkpoint
The checkpoint comes first: `crt.load` reads the safetensors file into tensors on the device, and a tied checkpoint, which carries no `lm_head.weight`, gets the embedding table bound under the head's name too, so one load adopts one object onto both slots. The model dtype is the given `--dtype`, else the embedding table's own. The tree is then built under `crt.device("meta")`, where every parameter is a shape and a dtype with no bytes; `to_empty` gives the slots their placement, and `load_state_dict(state, strict=True, assign=True)` adopts the checkpoint tensors as the model's parameters: no copy, one resident copy of every weight.
```python title="llama_from_scratch.py (excerpt)"
def build_model(config, snapshot, device, max_positions, dtype=None, state=None):
if state is None:
state = load_state(snapshot, device)
state = dict(state)
if config.tie_word_embeddings and LM_HEAD_KEY not in state:
state[LM_HEAD_KEY] = state[EMBED_KEY]
model_dtype = dtype if dtype is not None else state[EMBED_KEY].dtype
if dtype is not None:
state = cast_state(state, dtype)
with crt.device("meta"):
model = LlamaForCausalLM(config, model_dtype)
model.to_empty(device=device)
model.load_state_dict(state, strict=True, assign=True)
del state
model.tie_weights()
model.fuse()
model.model.build_rotary(max_positions, device)
model.eval()
return model
```
`tie_weights` is a check on a tied checkpoint: the head and the embedding must hold one object, and it raises when they do not. After the load, `model.state_dict()` lists the checkpoint's names in the checkpoint's order. A sharded checkpoint loads the same way: `load_state` reads the shards the index file names into one dict.
```python title="check_load.py"
model = build_model(config, snapshot, crt.Device.cpu(), max_positions=4096)
print(model.lm_head.weight is model.model.embed_tokens.weight) # True
print(model.lm_head.weight.dtype) # clika_runtime.bfloat16
print(len(model.state_dict())) # 147
```
## Fuse the projections after the load
Three projections over the same input are one matmul over the concatenated weights. `nn.fuse_linears` returns the group as an `nn.Linear` whose forward emits the parts' outputs side by side; the parts keep their state-dict names and read their rows back from the fused weight, and the group itself lists no parameters, so holding it as an attribute adds nothing to a save. Fuse once the weights are bound and before the first forward.
```python title="llama_from_scratch.py (excerpt)"
def fuse(self) -> None:
self.qkv = nn.fuse_linears(self.q_proj, self.k_proj, self.v_proj, name="qkv")
# in LlamaMLP
def fuse(self) -> None:
self.gate_up = nn.fuse_linears(self.gate_proj, self.up_proj, name="gate_up")
```
The forward splits the fused output where the parts meet:
```python title="llama_from_scratch.py (excerpt)"
qkv = self.qkv(x) # [tokens, (heads + 2 * kv_heads) * head_dim]
q, k, v = crt.split_with_sizes(qkv, [heads * head_dim, kv_heads * head_dim, kv_heads * head_dim], -1)
```
The parts' activations follow one rule: parts that each apply the same pointwise activation keep it, and the group applies it once over the concatenation; parts that apply different ones refuse. A gated `activation=` (`"swiglu"`, `"geglu"`, `"reglu"`, `"situ"`) emits half the concatenated width, the gate and up convention, and needs parts without a bias. Weight-quantized parts (`nn.QLinearWoQ`) join a group too, and share one matmul only under one quantization scheme. A group the runtime cannot fuse (mixed schemes, unequal input widths) serves its parts side by side plus the concatenation, so the output never depends on whether the fusion engaged, only the dispatch count does.
`nn.fuse_convs` does the same for convolutions over one input: the group is an `nn.Conv` whose output is the parts' outputs concatenated along the channel axis. Parts share one convolution when the kernel, stride, dilation, padding and its mode, `groups=1`, the dtype, and the presence of a bias agree. A gated activation refuses there, because it halves the channels.
## The KV cache and the step indices
`nn.KVCache` owns one key row and one value row per layer, sized once from the configuration. `prepare_step` takes the cumulative query lengths, the past length of each sequence and its slot, and returns the `nn.StepIndices` the attention operator reads: `cu_seqlens_q`, `cu_seqlens_k`, `kvcache_start`, `slot_ids` and `max_seqlen_k`, all device tensors except the last.
```python title="llama_from_scratch.py (excerpt)"
def make_cache(config, model, max_positions, device):
kv_config = nn.KVCacheConfig(
num_layers=config.num_hidden_layers, num_kv_heads=config.num_key_value_heads,
head_dim=config.head_dim, kv_dtype=model.model.embed_tokens.weight.dtype,
max_seqs=1, max_tokens_per_seq=max_positions, device=device,
)
return nn.KVCache.make(kv_config, [nn.KVLayerSpec(preallocate=True) for _ in range(config.num_hidden_layers)])
cache = make_cache(config, model, 4096, device)
step = cache.prepare_step([0, len(prompt_ids)], [0], [0]) # the prefill of one sequence
```
Layer `i` reads its rows through `cache.keys(i)` and `cache.values(i)`; the same handles receive the appended keys and values. The program `decode_step.py` beside `check_load.py` makes the two calls a serving loop makes, the prefill of a prompt and one decode step of the token it produced, with nothing else around them.
## The attention step is one operator
Rotation at the positions the cache start implies, the append of the new keys and values into the cache row, and the causal grouped-query attention over the whole row are one call, for the prefill and for every decode step alike. The rows are passed as `past_key` / `past_value` and again as `out_present_key` / `out_present_value`, so the append lands in place.
```python title="llama_from_scratch.py (excerpt)"
attention, _, _ = crt.group_query_attention_varlen(
q, k, v, step.cu_seqlens_q, step.cu_seqlens_k, max_seqlen_k=step.max_seqlen_k,
past_key=past_key, past_value=past_value, kvcache_start=step.kvcache_start,
rope_cos=rope_cos, rope_sin=rope_sin, is_causal=True, rotary_mode="neox",
num_heads=heads, kv_num_heads=kv_heads,
out_present_key=past_key, out_present_value=past_value, slot_ids=step.slot_ids,
)
return self.o_proj(attention)
```
The operator's contract, as the API reference states it:
| Term | Contract |
| --- | --- |
| query, key, value | hidden-folded `[sum of S, heads * head_dim]`; `num_heads` and `kv_num_heads` drive the head split inside the operator |
| `cu_seqlens_q`, `cu_seqlens_k` | the cumulative lengths of the batch's sequences, on the device |
| `max_seqlen_q`, `max_seqlen_k` | the longest query and the longest cached sequence of the batch, passed as numbers (`step.max_seqlen_k` above is a host value) |
| the continuous cache | `[max_seqs, kv_heads, max_seq, head_dim]` per layer, keys and values as two tensors; `nn.KVCache` owns one pair per layer |
| appending in place | the cache row passed as `past_key` / `past_value` and again as `out_present_key` / `out_present_value`; the rotated keys and the values are appended before the attention |
| `kvcache_start` | on a continuous cache a rank-1 `[B]` layout selector whose values are not read; each row's write offset is `cu_seqlens_k[b] - q_len[b]` |
| `slot_ids` | `[B]` Int32, the cache row each batch row appends to and attends from; absent, batch row `b` uses cache row `b` |
| a paged cache | the block pool `[num_blocks, kv_heads, block_size, head_dim]` with a `[B, max_blocks]` Int32 block table; `-1` pads a row's unused tail |
| rotary embedding | in the operator, from `rope_cos` / `rope_sin` at the positions the cache implies; `rotary_mode` picks the pairing |
| `q_norm_gain`, `k_norm_gain` | a per-head RMS norm applied after the rotation, rank-1 `[head_dim]`, together with `qk_norm_eps` (one without the other refuses); a model that normalizes before the rotation calls `qk_rms_norm` on the projections before this operator and passes no gains |
| `head_sink` | `[heads]`, a per-head virtual logit folded into the softmax denominator |
| a read-only attend | `attention_over_cache`: the same cache, nothing appended, rotary applied to the query alone |
## A residual add and the norm after it, fused
Each block adds a residual and normalizes the sum twice. `crt.add_rms_norm` does both in one operator and returns the pair the next block needs: the normalized stream and the raw residual. The closing call takes the norm weight that follows the block, the next layer's `input_layernorm` or the model's final `norm`, so the first `crt.rms_norm` on the embedding is the only standalone norm in the model.
```python title="llama_from_scratch.py (excerpt)"
o = self.self_attn(x, step, past_key, past_value, rope_cos, rope_sin)
x1, h1 = crt.add_rms_norm(o, h, None, None, [self.hidden_size], None,
self.post_attention_layernorm.weight, None, self.eps)
m = self.mlp(x1)
x2, h2 = crt.add_rms_norm(m, h1, None, None, [self.hidden_size], None, next_gain, None, self.eps)
return x2, h2
```
The head reads each sequence's last row from the device, `crt.index_select(x, -2, step.cu_seqlens_q[1:] - 1)`, so no shape is read on the host inside the forward.
## Greedy decoding with a one-token lookahead
Operators return before their work runs, and a host read waits for the value it needs. The loop uses that: after the prefill, each iteration queues the next step on the device first, with the argmax tensor itself as its input, and reads the current token afterwards. The device is never idle while Python reads a token, and every timing is the host clock at a token read.
```python title="llama_from_scratch.py (excerpt)"
cache.reset()
prompt = crt.tensor(list(prompt_ids), dtype=crt.int32, device=device)
length = len(prompt_ids)
logits = model(prompt, cache.prepare_step([0, length], [0], [0]), cache)
token = crt.argmax(logits, dims=[-1], index_dtype=crt.int32) # [1]
crt.async_eval(token)
past = length
ids: list[int] = []
for produced in range(max_new_tokens):
queued = None
if produced + 1 < max_new_tokens:
next_logits = model(token, cache.prepare_step([0, 1], [past], [0]), cache)
queued = crt.argmax(next_logits, dims=[-1], index_dtype=crt.int32)
crt.async_eval(queued)
value = int(token.item())
ids.append(value)
if value in stops or queued is None:
break
token = queued
past += 1
```
`crt.async_eval` submits the queued step without waiting; `token.item()` is the one read per token. The stop ids come from `generation_config.json`, with `config.json` as the fallback, the same source the command line reads.
## Run it and measure it
The prompt goes through the checkpoint's chat template as one user turn with the assistant priming, rendered at the wall clock (`now_epoch_seconds=int(time.time())`), which is how `clikart-cli prompt` renders it, so the two replies compare byte for byte.
```bash
python3 llama_from_scratch.py --model HuggingFaceTB/SmolLM2-135M-Instruct \
--prompt 'Write a long story about a lighthouse keeper who finds a map.' \
--max-new-tokens 64 --device cpu --json reply.json
```
`--json` writes the timings, the memory readings from `clika_runtime.memory_stats(device)` at four points of the run, the rendered prompt and the token ids. The chapter's `measure_decode.py --full` runs the same cell on both sides, alternating, and writes the report:
```bash
python3 measure_decode.py --full \
--cli /path/to/clikart-cli --snapshot /path/to/checkpoint --device cpu \
--isl 128 --osl 64 --iters 5 --warmup 1 --repeats 3 \
--prompt 'Write a long story about a lighthouse keeper who finds a map.' --max-new-tokens 64 \
--out ./llama_cpu
```
The chapter README carries the measured table for this cell, the machine it ran on, how each column is defined, and how to read the spread against the difference between the two rows. A number is a claim about one command on one machine: quote it with its command lines and its cell.
From here: [Structure inputs and outputs as pytrees](pytrees.mdx) for the containers the compile and save boundaries take, [Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx) for a module tree that starts as `torch.nn.Module`, and [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) for capturing a step as a graph.
---
# Coming from PyTorch or Hugging Face
The conventions a reader who knows PyTorch and the Hugging Face libraries meets first: channels-last tensors and OHWI weights, integer slots, reads that wait, the JSON and audio accessors, the tokenizer's inputs, and the generate calls side by side.
Source: https://docs.clika.io/clikart/how-to/coming-from-pytorch.md
{/* CERTIFICATION: every fact on this page is read from the public headers of the pinned release
(nn/conv.h, compute/ops.h, json/json.h, io/io.h, tokenizer/tokenizer.h, threading/threading.h)
and the wheel's generate_like_transformers README; the samples are excerpts, not a compiled program. */}
ClikaRT's Python package reads like PyTorch and its model library loads a checkpoint by the name
the Hugging Face hub gives it, so most of what you know carries over. This page lists the places
where the runtime's convention differs from the one you expect, each with the one line that
bridges it.
## Tensors are channels-last, weights are OHWI
Activations are `[N, spatial..., C]` and a convolution weight is `[out_channels, K..., in_channels / groups]`
(output-channel-first, channels last). PyTorch stores an activation `[N, C, H, W]` and a
convolution weight `[O, C / groups, K, K]`, and every checkpoint exported from it keeps that
order, so a weight is brought to OHWI once, at load, with one permute and a contiguous copy:
```cpp
const Tensor w_ohwi = ops::contiguous(ops::permute(w_oihw, {0, 2, 3, 1})); // 2-D; {0, 2, 1} for 1-D, {0, 2, 3, 4, 1} for 3-D
```
```python
w_ohwi = w_oihw.permute(0, 2, 3, 1).contiguous()
```
`nn::Conv` and `nn.Conv` declare their geometry in that layout, and
[Load images and audio for inference](load-images-and-audio.mdx) decodes a picture straight into
`[H, W, C]`.
## A size is an integer at the operator's boundary
A split size, an index, a sequence length (`max_seqlen_k` of the attention operators) is an
integer slot. The C++ gateway type that takes a number or a tensor accepts a floating-point value
too, and refuses it at run time when the slot is integer-typed, so pass an `int` (or an
`int64_t`), never a `double` cast from one.
## A read is the wait
An operator returns as soon as its work is queued; the kernels run behind it. A host read waits
for the value: `item`, `numpy()`, `tolist()`, printing a tensor, `eval`, `to_string()` and
`item_as_vec()` in C++. Nothing you read is ever an unfinished value. The consequence for a
measurement: time a step at the read of its result, or after `synchronize()` on the producing
stream; the return of the dispatch is not the end of the work.
[Control asynchronous execution](control-async-execution.mdx) walks the model in full.
## JSON, audio and the tokenizer's files
- A JSON document reads an integer through `as_int64()` and a floating-point value through
`as_double()`. `operator[]` on a mutable document creates a missing key (it turns a null into an
object); `at(key)` and `at(index)` read without creating.
- `io::load_audio` returns the samples beside an `AudioInfo` whose `sample_rate`, `channels` and
`frames` describe them (`frames / sample_rate` is the duration in seconds).
- `Tokenizer::from_huggingface` in C++, `Tokenizer.from_file` in Python and `Tokenizer.fromHuggingface`
in Kotlin read a model directory: its `tokenizer.json` with the special ids and the chat
template; a SentencePiece model beside its `tokenizer_config.json`; or a `vocab.json` for a model
no tokenizer class claims.
## The thread count is read once
`CLIKA_RT_NUM_THREADS` sets the CPU worker count and is read before the first compute, so it goes
into the environment before the first operator runs: in the shell, or from the program before it
touches the runtime. An Android app sets it in its own process before it creates any tensor.
## generate, side by side
The model library's Python surface mirrors the `transformers` idioms; the table the wheel's
`examples/python/howto/generate_like_transformers/README.md` carries is the whole map. The rows
a first program needs:
| `transformers` | `clika_runtime.modelverse` |
| --- | --- |
| `AutoModelForCausalLM.from_pretrained(id)` | `AutoModelForCausalLM.from_pretrained(id, device="cpu")` |
| `model.to("cuda")` | `from_pretrained(id, device="cuda:0")`, chosen at load |
| `ids = tok(prompt).input_ids; out = model.generate(ids); tok.decode(out)` | `model.generate(prompt)`, which takes text and returns text |
| `generate(..., do_sample=False)` | `generate(..., temperature=0.0)` |
| `tok.apply_chat_template(messages)` | `model.render(messages)` |
| `apply_chat_template` then `generate` | `model.chat(messages)` |
| `TextIteratorStreamer` on a second thread | `model.stream_generate(prompt)`, an iterator of text pieces |
| `pipeline("text-generation", model=id)` | `pipeline("text-generation", id)` |
A model's weights download inside `from_pretrained` into the Hugging Face hub cache the other
tooling on the machine shares; `mv.snapshot_download(id)` downloads without loading, and a local
directory is a source everywhere a repository id is
([Run fully offline](/modelverse/how-to/run-fully-offline)).
## Where the same idea has another name
| You reach for | Here |
| --- | --- |
| `torch.compile(model)` | `crt.compile(model)`: capture on the first call, replay after ([Trace eager code to graphs](trace-eager-code-to-graphs.mdx)) |
| `torch.no_grad()`, `model.eval()` | no gradients exist; `model.eval()` is accepted and changes nothing |
| `state_dict()` / `load_state_dict()` | the same names and the same dotted keys; `load_state_dict(state, assign=True)` adopts the checkpoint's tensors as the parameters, one resident copy ([Author a model in Python](author-a-model-in-python.mdx)) |
| `torch.device("meta")` | `crt.device("meta")`: shapes and dtypes without bytes, then `to_empty(device=...)` |
| DLPack exchange with torch | `clika_runtime.torch` ([Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx)) |
---
# Control asynchronous execution
Know when dispatched work actually runs, get completion callbacks, and force synchronous or lazy execution with the stream scopes.
Source: https://docs.clika.io/clikart/how-to/control-async-execution.md
An `ops::` call dispatches work and returns; the kernels run behind it. Where the work lands follows one law: a stream you pass is used as given; an operation placed on a bare `Device` resolves to the ambient placement scope's stream for that device when one is set, otherwise to the calling thread's asynchronous stream; and an operation with no placement of its own follows its inputs. Most programs never notice any of this, because reads wait for the result. This guide is for when you need control anyway: measuring where time goes, reacting the moment a result is ready, stepping op by op while debugging, or building a whole graph before running any of it.
Everything below runs on CPU streams, so it behaves the same on any machine; the same rules apply to CUDA, Vulkan and Metal streams. The timing numbers are from one real run and vary with the machine; the ordering they show does not.
## Dispatch is not execution
An `ops::` call returns as soon as the work is queued, and the result knows where it queued: `Tensor::stream()` names the stream the operation rode (under the asynchronous default, `stream().is_default()` is false), which is the stream to query and synchronize. `Tensor::status()` names where a result is in that lifecycle, `Stream::query_idle()` asks a stream without blocking, and `Stream::synchronize()` blocks until everything queued has run. Before dispatching anything, `StreamOrDevice(device).resolve()` reads where a bare-Device placement would land. Host reads (`to_string`, `item`, `item_as_vec`) wait on the producing work themselves, so a read is always safe; what you never observe is unfinished bytes. A timing is taken at a host read of the result, or after `synchronize()` on the producing stream: the return of the dispatch is not the end of the work, and a result nobody reads is not evidence of when it ran.
```cpp title="dispatch_vs_ready.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Stream;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
namespace {
double ms_since(std::chrono::steady_clock::time_point t0) {
return std::chrono::duration(
std::chrono::steady_clock::now() - t0).count();
}
} // namespace
int main() {
// The inputs settle on a stream of their own. The chain below carries no
// placement, so it rides the calling thread's asynchronous lane, and its
// result names that stream: x.stream() is what to query and synchronize.
Stream worker = Stream::create(Device::cpu());
const Tensor a = ops::full({2048, 2048}, 0.001, DataType::Float32, worker);
const Tensor w = ops::full({2048, 2048}, 0.001, DataType::Float32, worker);
worker.synchronize(); // settle the inputs so only the chain is measured
const auto t0 = std::chrono::steady_clock::now();
Tensor x = a;
// 48 dependent 2048 x 2048 products, about 0.8 TFLOP: the lane is still
// at work when the query after the dispatch returns, on any machine.
for (int i = 0; i < 48; ++i) x = ops::matmul(x, w);
const double dispatch_ms = ms_since(t0);
Stream lane = x.stream(); // the stream the chain actually ran on
std::printf("dispatch returned after %.1f ms, stream idle: %s\n",
dispatch_ms, lane.query_idle() ? "yes" : "no");
lane.synchronize();
std::printf("ready after %.1f ms, stream idle: %s\n",
ms_since(t0), lane.query_idle() ? "yes" : "no");
// A host read needs none of the above; it waits on its producer alone.
// (item() reads a single-element tensor, so reduce first.)
std::printf("max(x) = %g\n", ops::amax(x).item());
return 0;
}
```
Every operator returns at once and its work rides the calling thread's lane on the input's device; a read waits for the value it needs. The contract, in full:
1. These settle (they block until the value exists): `numpy()`, `bytes()`, `item()`, `tolist()`, `repr` and `print`, `bool()`, `__array__`, `__dlpack__`, `crt.save`, `crt.eval(*trees)` and `crt.synchronize(target)`; a truthful `shape` / `numel()` / `nbytes` settles only when the producer's shape depends on data.
2. These never settle: `dtype`, `device`, `ndim`, `stride()`, `is_contiguous()`.
3. The scopes are thread-local context managers that also work as decorators: `crt.synchronous()`, `crt.tracing()`, `crt.eager()`, `crt.meta_init()` (the same as `crt.device("meta")`), `crt.device(d)` and `crt.stream(s)`.
4. `crt.async_eval(*trees)` submits the work and returns without waiting, the one-token-lookahead idiom of a decode loop.
5. A data read under `crt.tracing()` or of a storage-free tensor raises `crt.ClikaRTError`; `numpy()` on a device tensor raises `TypeError` (move it to the cpu first); `bool()` of a tensor with more than one element raises `RuntimeError`.
```python title="settle.py"
import numpy as np
import clika_runtime as crt
x = np.ones((64, 64), dtype=np.float32)
t = crt.tensor(x)
y = t
for _ in range(50):
y = crt.add(crt.mul(y, 1.01), 0.01) # fifty operators, dispatched without waiting
print(y.dtype) # clika_runtime.float32 (metadata never settles)
print(float(y.numpy()[0, 0])) # 2.289262533187866 (the first read returns the settled value)
tree = {"a": crt.exp(t), "b": [crt.sum(t), None, "text"], "c": (crt.relu(t),)}
crt.eval(tree, crt.abs(t)) # settles every tensor leaf of every tree, ignores the rest
z = crt.add(crt.mul(t, 2.0), 1.0)
crt.async_eval(z) # submitted, not waited for
print(float(z.numpy()[0, 0])) # 3.0 (3.0, read later)
borrowed = crt.from_numpy(x) # a zero-copy view over the array
borrowed.add_(1.0)
crt.synchronize() # the calling thread's lane; the array carries the write after this
print(x[0, 0]) # 2.0 (2.0)
```
A raw read of a borrowed source array after an in-place operator settles first, as the last four lines do; a fresh `numpy()` call settles on its own.
```text
dispatch returned after 23.5 ms, stream idle: no
ready after 24.2 ms, stream idle: yes
max(x) = 0.00176683
```
The dispatch and the ready walls sit close together on a CPU stream, and the gap between them is the point: dispatch is still not execution (the stream is not idle when the loop returns), but memory on an asynchronous CPU stream follows execution, so a dependent chain is issued one operation ahead of the one executing and the dispatch loop is paced by the work itself. On a CUDA stream the guarantee is the enqueue rather than the run, so the same loop returns in well under a millisecond while the device works behind it. An op placed on a bare `Device` overlaps with the caller; code that needs it synchronous opts in explicitly, with `Stream::default_stream(device)` as the placement or a `SynchronousStreamScope` around the region.
## React the moment a result is ready
`Tensor::on_complete` registers a callback on a result; it fires the moment the producing kernel finishes, on ClikaRT's callback pool while the producer is still running, or inline at registration when the result has already settled. No polling and no blocked thread, the push-style inverse of `synchronize()`. The tensor's storage is kept alive for the callback, which receives it by const reference.
```cpp title="on_complete.cpp"
#include
#include
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Stream;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
namespace {
double ms_since(std::chrono::steady_clock::time_point t0) {
return std::chrono::duration(
std::chrono::steady_clock::now() - t0).count();
}
} // namespace
int main() {
Stream worker = Stream::create(Device::cpu());
const Tensor a = ops::full({1024, 1024}, 0.001, DataType::Float32, worker);
const Tensor w = ops::full({1024, 1024}, 0.001, DataType::Float32, worker);
worker.synchronize();
std::atomic fired{false};
const auto t0 = std::chrono::steady_clock::now();
Tensor x = a;
for (int i = 0; i < 10; ++i) x = ops::matmul(x, w);
std::printf("main: dispatch done at %.1f ms, registering the callback and doing other work\n",
ms_since(t0));
// Fires the moment the producing kernel finishes: on ClikaRT's callback
// pool while the producer is still running, inline at registration when
// the result has already settled. The storage stays alive for the call.
x.on_complete([&](const Tensor& result) {
std::printf("callback: fired after %.1f ms, x = %s\n",
ms_since(t0), result.to_string().c_str());
fired = true;
});
while (!fired) std::this_thread::yield();
return 0;
}
```
The Python package carries no completion callback; the push-style `on_complete` is the C++ arm's. The Python shape of the same idea is `crt.async_eval` to submit without waiting, and a read at the point of use: hand the tensor to whoever consumes it, and that consumer's `numpy()` or `item()` is the wait, on its own thread if the dispatching thread must not block.
```text
main: dispatch done at 10.8 ms, registering the callback and doing other work
callback: fired after 11.8 ms, x = Tensor(shape=[1024, 1024], dtype=Float32, device=CPU, numel=1048576, data=[0.001268, 0.001268, 0.001268, 0.001268, 0.001268, 0.001268, ...])
```
Use it to hand results to a queue, complete a request, or chain host-side work without dedicating a thread to waiting. Keep callbacks short; they share the callback pool.
## Step synchronously while debugging
`SynchronousStreamScope` is RAII: inside it, every op the calling thread dispatches to the stream completes before the call returns, so the program state after each line is exactly what the line computed. Deterministic and slow, which is the right trade while hunting a numeric bug or stepping in a debugger. Open it on a quiescent stream, and pipelining resumes when the scope closes. It is also the region-sized opt-in for code that needs synchronous bare-Device behavior.
```cpp title="sync_scope.cpp"
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Stream;
using ClikaRT::SynchronousStreamScope;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
int main() {
Stream worker = Stream::create(Device::cpu());
// One 4096 x 4096 product, about 137 GFLOP: still at work when the
// query after the dispatch returns, on any machine.
const Tensor a = ops::full({4096, 4096}, 0.001, DataType::Float32, worker);
const Tensor w = ops::full({4096, 4096}, 0.001, DataType::Float32, worker);
worker.synchronize();
// An op with no placement rides the calling thread's asynchronous lane;
// its result names that stream, which is the one to query.
Tensor x = ops::matmul(a, w); // async: dispatched, likely still running
std::printf("no scope : idle after dispatch: %s\n",
x.stream().query_idle() ? "yes" : "no");
x.stream().synchronize();
{
SynchronousStreamScope scope(worker);
x = ops::matmul(a, w); // completes before this line returns
std::printf("sync scope : idle after dispatch: %s\n",
x.stream().query_idle() ? "yes" : "no");
}
return 0;
}
```
`crt.synchronous()` is the region form: inside it every operator the calling thread dispatches completes before the call returns, so a borrowed source array carries an in-place write the moment the line finishes. It works as a decorator too.
```python title="sync_scope.py"
import numpy as np
import clika_runtime as crt
source = np.zeros((8, 8), dtype=np.float32)
borrowed = crt.from_numpy(source)
with crt.synchronous():
borrowed.add_(3.0)
print(source[0, 0]) # 3.0 (3.0, written before the line returned)
result = crt.mul(borrowed, 2.0)
print(float(result.numpy()[0, 0])) # 6.0 (6.0)
@crt.synchronous()
def step(t: crt.Tensor) -> crt.Tensor:
return crt.relu(t)
print(float(step(borrowed).numpy()[0, 0])) # 3.0 (3.0)
```
```text
no scope : idle after dispatch: no
sync scope : idle after dispatch: yes
```
## Build the whole graph first, run it once
`TracingScope` flips the calling thread the other way, to lazy: ops return `Unscheduled` placeholder tensors carrying lineage and no kernel runs. One `synchronize()` (or any read) materializes the graph leaf-first. Use it to declare a computation in full before spending anything, or to hand the runtime the widest possible scheduling view. The same idea with a reusable artifact is [tracing eager code to a graph](trace-eager-code-to-graphs.mdx): capture a function once as a `ModelGraph` and run it repeatedly, instead of scoping one thread's dispatches.
```cpp title="tracing_scope.cpp"
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Tensor;
using TensorStatus = ClikaRT::Tensor::Status;
using ClikaRT::TracingScope;
namespace ops = ClikaRT::ops;
namespace {
const char* name(TensorStatus s) {
switch (s) {
case TensorStatus::Unscheduled: return "Unscheduled";
case TensorStatus::Evaluated: return "Evaluated";
case TensorStatus::Available: return "Available";
}
return "?";
}
} // namespace
int main() {
const Tensor a = ops::ones({4, 4}, DataType::Float32, Device::cpu());
const Tensor b = ops::ones({4, 4}, DataType::Float32, Device::cpu());
Tensor m, s;
{
TracingScope trace;
m = ops::matmul(a, b); // no kernel runs
s = ops::add(m, a);
std::printf("traced : m=%s s=%s\n", name(m.status()), name(s.status()));
s.synchronize(); // materialize the graph, leaf-first
}
// Outside the scope, ops run eagerly again: a one-element view of the
// realized result and its value readback.
std::printf("realized: m=%s s=%s, s[0] = %g (4 ones dot ones + 1 = 5)\n",
name(m.status()), name(s.status()),
ops::select(s.reshape({-1}), 0, 0).item());
return 0;
}
```
`crt.tracing()` flips the calling thread to lazy: operators return unscheduled tensors that carry their shape and dtype and no value, a data read inside the region raises `crt.ClikaRTError`, and `crt.eval` materializes the whole chain leaf-first. `crt.eager()` restores the default inside a tracing region for the one operator that must run at once.
```python title="tracing_scope.py"
import numpy as np
import clika_runtime as crt
t = crt.tensor(np.ones((8, 8), dtype=np.float32))
with crt.tracing():
y = crt.add(crt.mul(t, 2.0), 1.0)
print(y.shape, y.dtype) # clika_runtime.Size([8, 8]) clika_runtime.float32 ((8, 8) clika_runtime.float32; no kernel has run)
try:
y.numpy()
except crt.ClikaRTError:
print("a value read under tracing raises") # printed: the read raised
with crt.eager():
at_once = crt.mul(t, 2.0)
print(float(at_once.numpy()[0, 0])) # 2.0 (2.0, run at once)
crt.eval(y) # materializes the traced chain
print(float(y.numpy()[0, 0])) # 3.0 (3.0)
```
```text
traced : m=Unscheduled s=Unscheduled
realized: m=Evaluated s=Evaluated, s[0] = 5 (4 ones dot ones + 1 = 5)
```
The four tools compose into one rule of thumb: leave the asynchronous default alone for throughput, read results and let the reads wait, reach for `on_complete` when a thread should not wait, and reserve the two scopes for debugging (synchronous) and up-front graph building (tracing). The bundle's `async` example walks each in its own chapter, including safe-reads patterns this guide leaves implicit.
---
# Edit a graph
Change a ModelGraph before finalize(): rewire the reads of a value, change what a node reads, rename a node, and edit the graph's inputs and outputs. A refused edit changes nothing; then optimize, finalize and run the edited graph.
Source: https://docs.clika.io/clikart/how-to/edit-a-graph.md
{/* Every block is a program under examples//howto/edit_a_graph/: the first
block of each tab is get_a_graph whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output. */}
A `ModelGraph` that you trace or compile can be changed before you call `finalize()`. You can move the reads of a value to another value, remove or bypass a node, change what a node reads, rename a node, and add, remove or rename the graph's inputs and outputs. Each edit checks everything before it changes anything, so a refused edit leaves the graph as it was and says what it refused. The edits are available from C++ and Python, under the same names.
## A graph to edit
A trace returns the graph as built, with every operator as written and nothing optimized or finalized, so it is ready to edit. The model below computes `y = Relu(x) + Neg(x)`: `x` feeds a Relu and a Neg, and one Add reads both. Every example on this page uses it. `label` names a node by its input name or its operator, `labels` lists a graph's nodes, and `producers` lists the nodes a node reads, one per input port.
Edits run before `finalize()`. A finalized graph serves `run()` and refuses every edit with the code name `FAILED_PRECONDITION`; to edit it, trace or compile the model again.
```cpp title="get_a_graph.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Error;
using ClikaRT::Tensor;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::Node;
using ClikaRT::graph::NodeKind;
using ClikaRT::graph::OpCode;
using ClikaRT::graph::Value;
namespace ops = ClikaRT::ops;
namespace {
// y = Relu(x) + Neg(x): x feeds a Relu and a Neg, both read by one Add.
std::vector model(const std::vector& inputs) {
const Tensor rectified = ops::relu(inputs[0]);
const Tensor negated = ops::neg(inputs[0]);
return {ops::add(rectified, negated)};
}
// A node's label: an input's name, an operator's code.
std::string label(const Node& node) {
return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code()));
}
// The labels of `nodes`, separated by spaces.
std::string labels(const std::vector& nodes) {
std::string out;
for (const Node& node : nodes) out += (out.empty() ? "" : " ") + label(node);
return out;
}
// The labels of the nodes that produce what `node` reads, one per input port.
std::string producers(const Node& node) {
std::string out;
for (const Value& value : node.inputs()) out += (out.empty() ? "" : " ") + label(*value.producer());
return out;
}
// A trace returns the graph as built: every operator as written, nothing optimized or finalized.
ModelGraph trace_model() {
const std::vector signature = {{"x", DataType::Float32, {2, 3}}};
const std::vector outputs = {"y"};
return ClikaRT::graph::trace(model, signature, "edit", outputs);
}
} // namespace
int main() {
const ModelGraph graph = trace_model();
const Node add = graph.find_nodes(OpCode::Add).front();
std::printf("%s | %s\n", labels(graph.nodes()).c_str(), producers(add).c_str()); // x Relu Neg Add | Relu Neg
std::printf("%s %s\n", graph.output_names().front().c_str(), graph.is_finalized() ? "true" : "false"); // y false
// Edits run before finalize(): a finalized graph refuses every one of them.
ModelGraph done = trace_model();
done.finalize();
try {
done.rename_output("y", "total");
} catch (const Error& error) {
std::printf("%s\n", error.code_name().c_str()); // FAILED_PRECONDITION
}
return 0;
}
```
```python title="get_a_graph.py"
import clika_runtime as crt
from clika_runtime.graph import NodeKind, OpCode
def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
x = inputs[0]
return [crt.relu(x) + crt.neg(x)] # y = Relu(x) + Neg(x)
def label(node: crt.graph.Node) -> str:
return node.name if node.kind == NodeKind.Input else node.op_code.name
def labels(graph: crt.graph.ModelGraph) -> list[str]:
return [label(node) for node in graph.nodes()]
def producers(node: crt.graph.Node) -> list[str]:
return [label(value.producer()) for value in node.inputs]
# A trace returns the graph as built: every operator as written, nothing optimized or finalized.
graph = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph
(add,) = graph.find_nodes(OpCode.Add)
print(labels(graph), producers(add)) # ['x', 'Relu', 'Neg', 'Add'] ['Relu', 'Neg']
print(graph.output_names(), graph.is_finalized()) # ['y'] False
# Edits run before finalize(): a finalized graph refuses every one of them.
done = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph
done.finalize()
try:
done.rename_output("y", "total")
except crt.InvalidArgumentError as error:
print(error.code_name) # FAILED_PRECONDITION
```
Edit a graph when a model needs a change its source does not make: an operator to remove before deployment, a value to expose or feed, or a name a caller binds.
Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, `label`, `labels`, `producers` and the trace).
## A refused edit changes nothing
Every edit checks its arguments and the graph's rules before it changes anything, so a refused edit leaves the nodes, the inputs, the outputs and the constants as they were. Its message names the call and what it refused. The code name is `INVALID_ARGUMENT` for an argument the edit cannot take, and `FAILED_PRECONDITION` for a finalized graph. In C++ the edit throws a `ClikaRT::Error` (a `Result` carries it in a build without exceptions); in Python it raises `InvalidArgumentError`, whose `code_name` says which. Here `rename_node` refuses an input and names the call that renames one.
```cpp title="refused_edit.cpp"
int main() {
ModelGraph graph = trace_model();
const std::string before = labels(graph.nodes());
try {
graph.rename_node(*graph.node("x"), "features"); // an input is renamed with rename_input()
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// INVALID_ARGUMENT | rename_node: 'x' is a graph input; rename it with rename_input()
}
std::printf("%s %s\n", labels(graph.nodes()) == before ? "true" : "false",
graph.input_names().front().c_str()); // true x
return 0;
}
```
```python title="refused_edit.py"
before = labels(graph)
try:
graph.rename_node(graph.node("x"), "features") # an input is renamed with rename_input()
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | rename_node: 'x' is a graph input; rename it with rename_input()
print(labels(graph) == before, graph.input_names()) # True ['x']
```
Branch on the code name and report the message; [Handle errors by code](handle-errors-by-code.mdx) covers the channels every failure carries.
## Rewire the reads of a value
`replace_all_uses_with(old_value, replacement)` moves every read of a value to another one: each operator input it feeds, and each graph output it returns, which keeps its name. `remove_node(node)` removes an operator that nothing reads, and refuses one that is still read, naming the reader to rewire or remove first. `bypass_node(node)` hands a node's readers the value on its input (the first input and output, unless you name the ports) and removes the node. A value takes another's place only with the same dtype and dims.
```cpp title="rewire.cpp"
int main() {
ModelGraph graph = trace_model();
const Node relu = graph.find_nodes(OpCode::Relu).front();
const Node neg = graph.find_nodes(OpCode::Neg).front();
const Node add = graph.find_nodes(OpCode::Add).front();
graph.replace_all_uses_with(*neg.output(0), *graph.node("x")->output(0)); // the Add reads x where it read Neg(x)
graph.remove_node(neg); // nothing reads the Neg now
std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu Add
graph.bypass_node(relu); // the Add reads the Relu's input, x, and the Relu goes
std::printf("%s | %s\n", labels(graph.nodes()).c_str(), producers(add).c_str()); // x Add | x x
return 0;
}
```
```python title="rewire.py"
(relu,) = graph.find_nodes(OpCode.Relu)
(neg,) = graph.find_nodes(OpCode.Neg)
(add,) = graph.find_nodes(OpCode.Add)
graph.replace_all_uses_with(neg.output(0), graph.node("x").output(0)) # the Add reads x where it read Neg(x)
graph.remove_node(neg) # nothing reads the Neg now
print(labels(graph)) # ['x', 'Relu', 'Add']
graph.bypass_node(relu) # the Add reads the Relu's input, x, and the Relu goes
print(labels(graph), producers(add)) # ['x', 'Add'] ['x', 'x']
```
Rewire a graph to remove an operator a deployment does not need, or to let readers take a value the graph already computes.
## Change what a node reads
`set_input(node, port, value)` gives one input port another value; the port must read a value already. `set_constant_input(node, port, tensor)` binds a new constant holding the tensor's bytes to the port: the graph keeps a handle to the bytes, with no copy, and places them with the graph at `finalize()`. `add_constant(tensor)` makes a constant that nothing reads yet, for `set_input` or `replace_all_uses_with` to wire in, and `finalize()` drops a constant that nothing reads by then. A weight an operator holds itself, such as a compiled MatMul's, stays as it is, and an edit that would replace it refuses.
```cpp title="node_inputs.cpp"
int main() {
ModelGraph graph = trace_model();
const Node relu = graph.find_nodes(OpCode::Relu).front();
const Node neg = graph.find_nodes(OpCode::Neg).front();
const Node add = graph.find_nodes(OpCode::Add).front();
graph.set_input(add, 1, *relu.output(0)); // port 1 reads Relu(x) in place of Neg(x)
graph.remove_node(neg); // which nothing reads now
std::printf("%s\n", producers(add).c_str()); // Relu Relu
// Port 1 reads a constant holding the tensor's bytes.
graph.set_constant_input(add, 1, Tensor::full({2, 3}, 1.0, DataType::Float32));
std::printf("%s %zu\n", add.input(1)->is_constant() ? "true" : "false", graph.constants().size()); // true 1
const Value half = graph.add_constant(Tensor::full({2, 3}, 0.5, DataType::Float32)); // nothing reads it yet
graph.set_input(add, 1, half); // port 1 reads it, and the ones, read by nothing, go
std::printf("%s %zu\n", *add.input(1) == half ? "true" : "false", graph.constants().size()); // true 1
return 0;
}
```
```python title="node_inputs.py"
(relu,) = graph.find_nodes(OpCode.Relu)
(neg,) = graph.find_nodes(OpCode.Neg)
(add,) = graph.find_nodes(OpCode.Add)
graph.set_input(add, 1, relu.output(0)) # port 1 reads Relu(x) in place of Neg(x)
graph.remove_node(neg) # which nothing reads now
print(producers(add)) # ['Relu', 'Relu']
graph.set_constant_input(add, 1, crt.ones(2, 3)) # port 1 reads a constant holding the tensor's bytes
print(add.input(1).is_constant(), len(graph.constants())) # True 1
half = graph.add_constant(crt.full((2, 3), 0.5)) # a constant nothing reads yet
graph.set_input(add, 1, half) # port 1 reads it, and the ones, read by nothing, go
print(add.input(1) == half, len(graph.constants())) # True 1
```
Change a node's inputs to feed a constant of your own into the graph, or to point an operator at another producer.
## Rename a node
`rename_node(node, name)` renames an operator, and its outputs keep their names; an input is renamed with `rename_input` instead. A view of the old name is gone afterwards, so look the node up again with `node(name)`.
```cpp title="rename.cpp"
int main() {
ModelGraph graph = trace_model();
graph.rename_node(graph.find_nodes(OpCode::Add).front(), "total"); // views of the old name are gone afterwards
const Node total = *graph.node("total"); // so find the node again by its new one
std::printf("%s | %s | %s\n", label(total).c_str(), producers(total).c_str(),
graph.output_names().front().c_str()); // Add | Relu Neg | y
return 0;
}
```
```python title="rename.py"
(add,) = graph.find_nodes(OpCode.Add)
graph.rename_node(add, "total") # views of the old name are gone afterwards
total = graph.node("total") # so find the node again by its new one
print(label(total), producers(total), graph.output_names()) # Add ['Relu', 'Neg'] ['y']
```
Rename nodes to give them the names your own tools and reports use.
## Edit the graph's inputs and outputs
`add_input(spec)` adds a graph input with the spec's name, dtype and dims (a dynamic dim takes its size at `run()`) and returns its input node, whose value the other edits take. `add_output(value, name)` returns a value under a new output name, and `remove_output(name)` stops returning one, while the node that computes it stays. `rename_input(name, new_name)` and `rename_output(name, new_name)` rename an input and an output. Input and output names stay unique, and a name a KV cache layer binds stays as it is.
```cpp title="graph_io.cpp"
int main() {
ModelGraph graph = trace_model();
const Node relu = graph.find_nodes(OpCode::Relu).front();
const Node neg = graph.find_nodes(OpCode::Neg).front();
const Node add = graph.find_nodes(OpCode::Add).front();
const Node bias = graph.add_input({"bias", DataType::Float32, {2, 3}}); // a new input, bound at run()
graph.set_input(add, 1, *bias.output(0)); // y = Relu(x) + bias
graph.remove_node(neg);
graph.add_output(*relu.output(0), "rectified"); // Relu(x) is returned too
graph.rename_input("x", "features");
graph.rename_output("y", "total");
const std::vector ins = graph.input_names();
const std::vector outs = graph.output_names();
std::printf("%s %s | %s %s\n", ins[0].c_str(), ins[1].c_str(), outs[0].c_str(), outs[1].c_str());
// features bias | total rectified
graph.remove_output("rectified"); // the Relu that computed it stays: the Add reads it
std::printf("%s | %s\n", graph.output_names().front().c_str(), producers(add).c_str()); // total | Relu bias
return 0;
}
```
```python title="graph_io.py"
(relu,) = graph.find_nodes(OpCode.Relu)
(neg,) = graph.find_nodes(OpCode.Neg)
(add,) = graph.find_nodes(OpCode.Add)
bias = graph.add_input(crt.TensorSpec("bias", crt.float32, [2, 3])) # a new input, bound at run()
graph.set_input(add, 1, bias.output(0)) # y = Relu(x) + bias
graph.remove_node(neg)
graph.add_output(relu.output(0), "rectified") # Relu(x) is returned too
graph.rename_input("x", "features")
graph.rename_output("y", "total")
print(graph.input_names(), graph.output_names()) # ['features', 'bias'] ['total', 'rectified']
graph.remove_output("rectified") # the Relu that computed it stays: the Add reads it
print(graph.output_names(), producers(add)) # ['total'] ['Relu', 'bias']
```
Edit the inputs and outputs to expose an intermediate value, to feed a value from outside the graph, or to match the names a caller binds.
## Optimize, finalize and run the edited graph
After the edits, `optimize()` runs the graph optimizer when you want it, and `finalize()` makes the graph runnable, after which it refuses edits. `run()` binds the inputs in `input_names()` order, the new input included. Each program checks the result against a reference it computes itself; the Python program also imports numpy as `np` for that.
```cpp title="after_the_edits.cpp"
int main() {
ModelGraph graph = trace_model();
const Node neg = graph.find_nodes(OpCode::Neg).front();
const Node add = graph.find_nodes(OpCode::Add).front();
const Node bias = graph.add_input({"bias", DataType::Float32, {2, 3}});
graph.set_input(add, 1, *bias.output(0)); // y = Relu(x) + bias
graph.remove_node(neg);
graph.optimize(); // optional: the graph optimizer runs over the edited graph
graph.finalize(); // the graph runs from here on, and refuses edits
const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F};
const std::vector results = graph.run(
{Tensor::from_data(x.data(), {2, 3}, DataType::Float32), Tensor::full({2, 3}, 0.5, DataType::Float32)});
const std::vector y = results.front().reshape({-1}).item_as_vec();
bool matches = y.size() == x.size();
for (std::size_t i = 0; matches && i < x.size(); ++i) {
matches = y[i] == (x[i] > 0.0F ? x[i] : 0.0F) + 0.5F; // the reference, Relu(x) + 0.5, by hand
}
for (std::size_t i = 0; i < y.size(); ++i) std::printf("%s%g", i == 0 ? "" : " ", y[i]);
std::printf("\n%s\n", matches ? "true" : "false");
// 1.5 0.5 3.5 0.5 5.5 0.5
// true
return 0;
}
```
```python title="after_the_edits.py"
(neg,) = graph.find_nodes(OpCode.Neg)
(add,) = graph.find_nodes(OpCode.Add)
bias = graph.add_input(crt.TensorSpec("bias", crt.float32, [2, 3]))
graph.set_input(add, 1, bias.output(0)) # y = Relu(x) + bias
graph.remove_node(neg)
graph.optimize() # optional: the graph optimizer runs over the edited graph
graph.finalize() # the graph runs from here on, and refuses edits
x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32)
b = np.full((2, 3), 0.5, np.float32)
(y,) = graph.run([crt.tensor(x), crt.tensor(b)])
print(y.numpy().tolist()) # [[1.5, 0.5, 3.5], [0.5, 5.5, 0.5]]
print(np.array_equal(y.numpy(), np.maximum(x, 0) + b)) # True (numpy states the reference)
```
Finalize once the graph has every edit it needs, since a finalized graph takes none.
## Views across edits
A `Node`, `Value` or `Edge` view taken before an edit finds its node, value or edge again by its key and keeps answering, and it refuses once an edit removed or renamed what it names, as [Views across edits](query-a-graph.mdx#views-across-edits) on the Query a graph page shows.
---
# Add ClikaRT to an existing CMake project
Link the bundle into a project that already builds, with find_package or add_subdirectory, and the one linker flag serving-runtime consumers need.
Source: https://docs.clika.io/clikart/how-to/existing-cmake-project.md
Your application already builds, and you want it to call ClikaRT. The integration is two lines in the CMake file you already have: declare the package, link one target. This guide adds ClikaRT to an existing app both ways (the prebuilt bundle via `find_package`, the bundle source tree via `add_subdirectory`), then covers the two things integrations trip on: where the libraries are found at run time, and the RTTI flag the serving runtime requires.
ClikaRT needs C++17 or later; if your project predates `set(CMAKE_CXX_STANDARD 17)`, add it.
## Link the bundle with find_package
Say the existing project is a telemetry service with one executable:
```cmake title="CMakeLists.txt (before)"
cmake_minimum_required(VERSION 3.19)
project(telemetry CXX)
set(CMAKE_CXX_STANDARD 17)
add_executable(telemetry src/main.cpp src/collect.cpp)
```
Two lines make ClikaRT available to it. `find_package(ClikaRT CONFIG)` loads the package from the extracted bundle, and the imported target `ClikaRT::ClikaRT` carries the include paths, the libraries, and their link order:
```cmake title="CMakeLists.txt (after)"
cmake_minimum_required(VERSION 3.19)
project(telemetry CXX)
set(CMAKE_CXX_STANDARD 17)
find_package(ClikaRT CONFIG REQUIRED)
add_executable(telemetry src/main.cpp src/collect.cpp)
target_link_libraries(telemetry PRIVATE ClikaRT::ClikaRT)
```
CMake finds the package through `ClikaRT_DIR`, pointing at the bundle's `cmake/` directory:
```bash
cmake -S . -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build
```
`-DClikaRT_DIR` is a cache variable, so it is remembered after the first configure. To avoid passing it at all, add the bundle to `CMAKE_PREFIX_PATH` (in a toolchain file, a CMake preset, or the environment) and `find_package` finds it there. A wrong or unset path fails the configure with `Could not find a package configuration file provided by "ClikaRT"`; the fix is always the same, point `ClikaRT_DIR` at `/cmake`.
## Where the libraries are found at run time
The package stamps each consumer binary with an rpath to the bundle's platform `lib/` directory, so a freshly built binary runs in place with no library paths to set, `LD_LIBRARY_PATH` included. That rpath names the bundle's absolute path on the build machine. For binaries that ship to other machines, copy the libraries your binary needs from the bundle's `lib/` next to it (or into your package's lib directory) and set your own relative rpath, for example `$ORIGIN/../lib`; the bundle's prebuilt examples ship exactly that way.
## Serving-runtime consumers mirror the library's -fno-rtti
Code that only uses tensors, `ops::`, `io::`, tokenizers, or the HTTP pieces builds with your project's existing flags. Code that uses the serving runtime (`runtime::FunctionModel`, `runtime::Model`, the executor) must be compiled with `-fno-rtti`, matching how the library builds. With RTTI left on, the target fails at link with:
```text
undefined reference to `typeinfo for ClikaRT::runtime::Model'
```
Scope the flag to the targets that touch `runtime::`:
```cmake
target_compile_options(telemetry PRIVATE -fno-rtti)
```
The bundle's own examples set the same flag; it is the one compile-option requirement in the integration.
## Check the bundle into your tree with add_subdirectory
For a monorepo that keeps the bundle in the repository (or fetches it into the tree), the bundle root is also a CMake subproject. It defines the same `ClikaRT::ClikaRT` target with the same options, so consumers cannot tell the difference:
```cmake
add_subdirectory(third_party/clikart)
add_executable(telemetry src/main.cpp src/collect.cpp)
target_link_libraries(telemetry PRIVATE ClikaRT::ClikaRT)
```
No `ClikaRT_DIR` and no configure-time flag are involved; the path in the tree is the whole wiring. Prefer `find_package` when the bundle lives outside the repository (a shared install, a CI cache), `add_subdirectory` when it lives inside.
## Verify
A two-line call in your existing code proves the link end to end:
```cpp title="src/main.cpp (excerpt as a standalone check)"
#include
#include "ClikaRT/clika_rt.h"
int main() {
std::printf("ClikaRT %s\n", ClikaRT::GetVersionInfo().c_str());
std::printf("cuda=%d vulkan=%d\n",
ClikaRT::device::is_cuda_available(), ClikaRT::device::is_vulkan_available());
return 0;
}
```
```text
ClikaRT 0.6.4
cuda=1 vulkan=1
```
The availability flags reflect the machine, not the build: the same binary prints different backends on different hardware, which is the point. From here, [the tutorial](../getting-started/first-program/01-your-first-program.mdx) covers the API the newly linked target now reaches, and [Quick install](../getting-started/installation.md) has the bundle layout.
---
# Export a graph to ONNX
Write a ModelGraph as an ONNX model file, or build the same model in memory: read it back through OnnxModel::open and compile, keep a named dynamic dimension, choose where the weights go and the opset, and meet what the export refuses.
Source: https://docs.clika.io/clikart/how-to/export-a-graph-to-onnx.md
{/* Every block is a program under examples//howto/export_a_graph_to_onnx/: the first
block of each tab is get_the_graph whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output. */}
`ModelGraph::export_onnx` writes a graph as a runnable ONNX model file, and `io::OnnxModel::from_graph` builds the same model in memory without writing anything. The model keeps the graph's input and output names, dtypes and dimensions. A dynamic dimension keeps its name, so two inputs that share one still share it when the model is read back, and the graph's weights become the model's initializers. `io::OnnxModel::open` and `compile()` read the file back into a graph that returns the same values. The calls are available from C++ and Python, under the same names.
## The graph to export
The examples on this page export one elementwise model, `y = Relu(x * w + b)`, over an input `x` of shape [batch, 37] whose first dimension is dynamic and named `batch`. The model captures `w` and `b`, two vectors of 37 values, so the traced graph holds them as its two weights. `input(batch)` builds an input whose row `i` holds `i + 1` in every column, and `row_sums` prints the sum of each output row, so two graphs that return the same values print the same line. Every value on this page is a multiple of 1/8, which float32 holds exactly. `print_specs` prints an input or an output with a dynamic dimension under its name, and the other helpers name a temporary directory for the files a program writes, list its files and print a refusal. In Python the programs import `clika_runtime` as `crt`, and `ModelGraph` and `Transform` from `clika_runtime.graph`. There `rows(batch)` builds the input, `slots` prints the `(name, dtype, dims)` tuples that `inputs()` and `outputs()` return, with a dynamic dimension as -1, and a value's `spec().dim_names` holds the dimension names.
A trace returns the graph as built, with every operator as written and nothing optimized or finalized, and `export_onnx` writes it in that state. This program prints the graph's input and output, then runs it once after `finalize()`. Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, the model and the helpers).
```cpp title="get_the_graph.cpp"
#include
#include
#include
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Error;
using ClikaRT::Tensor;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::OnnxExportOptions;
using ClikaRT::io::OnnxModel;
using ClikaRT::spec::TensorSpec;
namespace fs = std::filesystem;
namespace ops = ClikaRT::ops;
namespace {
constexpr std::int64_t kWidth = 37; // the length of every row
// y = Relu(x * w + b) over x [batch, 37], whose first dim is dynamic and named batch. The model captures
// w and b, two [37] vectors, so the graph holds them as its two weights: w[j] = (j - 18) / 8, from -2.25
// to 2.25, and b = 0.5. A trace returns the graph as built: every operator as written, nothing optimized
// or finalized.
ModelGraph trace_model() {
const Tensor w = ops::div(ops::sub(ops::arange(0, kWidth, 1, DataType::Float32), 18), 8);
const Tensor b = Tensor::full({kWidth}, 0.5, DataType::Float32);
const auto model = [w, b](const std::vector& inputs) -> std::vector {
return {ops::relu(ops::add(ops::mul(inputs[0], w), b))};
};
const std::vector signature = {
TensorSpec("x", DataType::Float32, {TensorSpec::kDynamicDim, kWidth}, false, {"batch", ""})};
const std::vector outputs = {"y"};
return ClikaRT::graph::trace(model, signature, "elementwise", outputs);
}
// x [batch, 37], whose row i holds i + 1 in every column.
Tensor input(std::int64_t batch) { return ops::cumsum(Tensor::ones({batch, kWidth}, DataType::Float32), 0); }
// The sum of each row of y, separated by spaces.
std::string row_sums(const Tensor& y) {
std::string out;
for (const float value : ops::sum(y, {1}).item_as_vec()) {
char text[32];
std::snprintf(text, sizeof(text), "%g", static_cast(value));
out += (out.empty() ? "" : " ") + std::string(text);
}
return out;
}
// Each spec on its own line after `label`, as " []", a dynamic dim written as its name.
void print_specs(const char* label, const std::vector& specs) {
for (const TensorSpec& spec : specs) {
std::string dims;
for (std::size_t d = 0; d < spec.dims.size(); ++d) {
const bool named = d < spec.dim_names.size() && !spec.dim_names[d].empty();
dims += (dims.empty() ? "" : ", ") + (named ? spec.dim_names[d] : std::to_string(spec.dims[d]));
}
std::printf("%s %s %s [%s]\n", label, spec.name.c_str(), ClikaRT::dtype::data_type_name(spec.dtype),
dims.c_str());
}
}
// An empty directory for the files a program writes, under the system's temporary directory.
fs::path fresh_directory(const std::string& name) {
const fs::path dir = fs::temp_directory_path() / ("export_a_graph_to_onnx_" + name);
fs::remove_all(dir);
fs::create_directories(dir);
return dir;
}
// The names of the files in dir, sorted and separated by spaces, or "none".
std::string files_in(const fs::path& dir) {
std::vector names;
for (const fs::directory_entry& entry : fs::directory_iterator(dir)) {
names.push_back(entry.path().filename().string());
}
std::sort(names.begin(), names.end());
std::string out;
for (const std::string& name : names) out += (out.empty() ? "" : " ") + name;
return out.empty() ? "none" : out;
}
// A refusal as " | ".
void print_refusal(const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); }
} // namespace
int main() {
ModelGraph graph = trace_model();
print_specs("input", graph.inputs());
print_specs("output", graph.outputs());
// input x Float32 [batch, 37]
// output y Float32 [batch, 37]
graph.finalize(); // a finalized graph runs
std::printf("%s\n", row_sums(graph.run({input(2)}).front()).c_str());
// 31.625 52.5
return 0;
}
```
```python title="get_the_graph.py"
import tempfile
from pathlib import Path
import clika_runtime as crt
from clika_runtime.graph import ModelGraph, Transform
WIDTH = 37 # the length of every row
def trace_model() -> ModelGraph:
# y = relu(x * w + b) over x [batch, 37], whose first dim is dynamic and named batch. The model captures
# w and b, two [37] vectors, so the graph holds them as its two weights: w[j] = (j - 18) / 8, from -2.25
# to 2.25, and b = 0.5. A trace returns the graph as built: every operator as written, nothing optimized
# or finalized.
w = crt.div(crt.sub(crt.arange(0, WIDTH, dtype=crt.float32), 18), 8)
b = crt.full((WIDTH,), 0.5, dtype=crt.float32)
signature = [crt.TensorSpec("x", crt.float32, [-1, WIDTH], dim_names=["batch", ""])]
return crt.trace(lambda inputs: [crt.relu(crt.add(crt.mul(inputs[0], w), b))], signature,
output_names=["y"]).graph
def rows(batch: int) -> crt.Tensor:
# x [batch, 37], whose row i holds i + 1 in every column.
return crt.cumsum(crt.ones(batch, WIDTH), 0)
def row_sums(y: crt.Tensor) -> str:
# The sum of each row of y, separated by spaces.
return " ".join(f"{value:g}" for value in crt.sum(y, [1]).tolist())
def slots(listed: list[tuple[str, object, tuple[int, ...]]]) -> str:
# (name, dtype, dims) slots as "name [dims]", separated by " | "; a dynamic dim reads -1.
return " | ".join(f"{name} {list(dims)}" for name, _, dims in listed)
def files_in(folder: Path) -> list[str]:
# The names of the files in folder, sorted.
return sorted(path.name for path in folder.iterdir())
def refusal(error: crt.ClikaRTError) -> str:
# A refusal as " | ".
return f"{error.code_name} | {error}"
graph = trace_model()
print(slots(graph.inputs()), "->", slots(graph.outputs()))
# x [-1, 37] -> y [-1, 37]
print(graph.node("x").output(0).spec().dim_names) # the dynamic dim keeps its name
# ['batch', '']
graph.finalize() # a finalized graph runs
print(row_sums(graph.run([rows(2)])[0]))
# 31.625 52.5
```
## Export the graph and read it back
`graph.export_onnx(path)` writes the model to `path`. `OnnxModel::open(path)` reads the file, and `compile()` turns it into a `ModelGraph` again, as built, so `finalize()` readies it to run ([Run an ONNX model](run-an-onnx-model.mdx) covers that side). The read-back graph has the same inputs and outputs, and its dynamic dimension carries the name `batch`, so it runs at any batch size. At batch 3 and at batch 5 the two graphs return the same sums.
In Python, `export_onnx` takes the path as a `str` or any `os.PathLike`, such as a `pathlib.Path`, and its options as keyword arguments. `OnnxModel.open` takes the path as a `str`.
```cpp title="export_and_read_back.cpp"
int main() {
const fs::path dir = fresh_directory("export_and_read_back");
const std::string path = (dir / "elementwise.onnx").string();
ModelGraph graph = trace_model();
graph.export_onnx(path); // writes the graph as traced
// OnnxModel::open reads the file, and compile() makes it a graph again: the same inputs and outputs,
// with the dynamic dim under its name.
ModelGraph back = OnnxModel::open(path).compile();
print_specs("input", back.inputs());
print_specs("output", back.outputs());
// input x Float32 [batch, 37]
// output y Float32 [batch, 37]
graph.finalize();
back.finalize();
for (const std::int64_t batch : {3, 5}) { // the dynamic dim takes any size
const Tensor x = input(batch);
std::printf("%s | %s\n", row_sums(graph.run({x}).front()).c_str(),
row_sums(back.run({x}).front()).c_str());
}
// 31.625 52.5 73.75 | 31.625 52.5 73.75
// 31.625 52.5 73.75 95 116.375 | 31.625 52.5 73.75 95 116.375
fs::remove_all(dir);
return 0;
}
```
```python title="export_and_read_back.py"
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / "elementwise.onnx"
graph = trace_model()
graph.export_onnx(path) # writes the graph as traced; a str or any os.PathLike names the file
# OnnxModel.open reads the file, and compile() makes it a graph again: the same inputs and outputs, with
# the dynamic dim under its name.
back = crt.io.OnnxModel.open(str(path)).compile()
print(slots(back.inputs()), "->", slots(back.outputs()))
# x [-1, 37] -> y [-1, 37]
print(back.node("x").output(0).spec().dim_names)
# ['batch', '']
graph.finalize()
back.finalize()
for batch in (3, 5): # the dynamic dim takes any size
x = rows(batch)
print(row_sums(graph.run([x])[0]), "|", row_sums(back.run([x])[0]))
# 31.625 52.5 73.75 | 31.625 52.5 73.75
# 31.625 52.5 73.75 95 116.375 | 31.625 52.5 73.75 95 116.375
```
## Build the model in memory
`OnnxModel::from_graph(graph)` builds, in memory, the model `export_onnx` writes, and writes nothing. The result is an `OnnxModel` like one that `open` returns: inspect its inputs, outputs and initializers, edit it, `optimize()` it, `compile()` it, or `save()` it when you want the file. It reads the graph's weights only when it is saved or compiled, so it holds no copy of them, and it keeps them alive for as long as it lives. Here the model reports the two weights as its initializers, and the graph it compiles to returns the same sums. In Python, `OnnxModel.from_graph` is a static method with the same keyword options, and `num_initializers` is a property.
```cpp title="in_memory.cpp"
int main() {
ModelGraph graph = trace_model();
// OnnxModel::from_graph builds the model export_onnx writes, in memory: nothing is written.
OnnxModel model = OnnxModel::from_graph(graph);
print_specs("input", model.inputs());
print_specs("output", model.outputs());
std::printf("%zu initializers\n", model.num_initializers());
// input x Float32 [batch, 37]
// output y Float32 [batch, 37]
// 2 initializers
// It compiles like any other model, and save() writes it when you want the file.
ModelGraph back = model.compile();
graph.finalize();
back.finalize();
const Tensor x = input(2);
std::printf("%s | %s\n", row_sums(graph.run({x}).front()).c_str(), row_sums(back.run({x}).front()).c_str());
// 31.625 52.5 | 31.625 52.5
const fs::path dir = fresh_directory("in_memory");
model.save((dir / "elementwise.onnx").string());
std::printf("%s\n", files_in(dir).c_str());
// elementwise.onnx
fs::remove_all(dir);
return 0;
}
```
```python title="in_memory.py"
graph = trace_model()
# OnnxModel.from_graph builds the model export_onnx writes, in memory: nothing is written.
model = crt.io.OnnxModel.from_graph(graph)
print(slots(model.inputs()), "->", slots(model.outputs()))
print(model.num_initializers, "initializers")
# x [-1, 37] -> y [-1, 37]
# 2 initializers
# It compiles like any other model, and save() writes it when you want the file.
back = model.compile()
graph.finalize()
back.finalize()
x = rows(2)
print(row_sums(graph.run([x])[0]), "|", row_sums(back.run([x])[0]))
# 31.625 52.5 | 31.625 52.5
with tempfile.TemporaryDirectory() as tmp:
model.save(str(Path(tmp) / "elementwise.onnx"))
print(files_in(Path(tmp)))
# ['elementwise.onnx']
```
## Choose where the weights go
`OnnxExportOptions::save_as_external_data` decides where the weights go when the model is written. `true` puts them in a `_data` file beside the model, `false` keeps them inside the model file, and unset, the default, keeps them inside until they pass 2 GiB. `OnnxModel::open` finds the data file beside the model. Here the data file holds the two weights, 296 bytes, and the model read from it returns the same sums. In Python the option is `save_as_external_data=True`, `False` or `None`.
```cpp title="where_the_weights_go.cpp"
int main() {
const fs::path dir = fresh_directory("where_the_weights_go");
ModelGraph graph = trace_model();
// save_as_external_data = true writes the weights to a _data file beside the model.
OnnxExportOptions beside;
beside.save_as_external_data = true;
graph.export_onnx((dir / "beside.onnx").string(), beside);
// false keeps them in the model file, and so does the default until they pass 2 GiB.
OnnxExportOptions inside;
inside.save_as_external_data = false;
graph.export_onnx((dir / "inside.onnx").string(), inside);
graph.export_onnx((dir / "default.onnx").string());
std::printf("%s\n", files_in(dir).c_str());
// beside.onnx beside_data default.onnx inside.onnx
std::printf("beside_data: %llu bytes\n", static_cast(fs::file_size(dir / "beside_data")));
// beside_data: 296 bytes: the two [37] float32 weights
// OnnxModel::open finds the data file beside the model.
ModelGraph back = OnnxModel::open((dir / "beside.onnx").string()).compile();
back.finalize();
std::printf("%s\n", row_sums(back.run({input(2)}).front()).c_str());
// 31.625 52.5
fs::remove_all(dir);
return 0;
}
```
```python title="where_the_weights_go.py"
with tempfile.TemporaryDirectory() as tmp:
folder = Path(tmp)
graph = trace_model()
# save_as_external_data=True writes the weights to a _data file beside the model.
graph.export_onnx(folder / "beside.onnx", save_as_external_data=True)
# False keeps them in the model file, and so does None, the default, until they pass 2 GiB.
graph.export_onnx(folder / "inside.onnx", save_as_external_data=False)
graph.export_onnx(folder / "default.onnx")
print(files_in(folder))
# ['beside.onnx', 'beside_data', 'default.onnx', 'inside.onnx']
print("beside_data:", (folder / "beside_data").stat().st_size, "bytes")
# beside_data: 296 bytes: the two [37] float32 weights
# OnnxModel.open finds the data file beside the model.
back = crt.io.OnnxModel.open(str(folder / "beside.onnx")).compile()
back.finalize()
print(row_sums(back.run([rows(2)])[0]))
# 31.625 52.5
```
## Set the opset and the other options
`opset_version` is the `ai.onnx` opset the model imports, and 0 takes the library's default. `contrib_ops` lets the model use the `com.microsoft` operators, such as GroupQueryAttention and MatMulNBits, and imports that domain only when an operator uses it. `optimize` simplifies the written model, composing pairs of transposes that undo each other and dropping the nodes no output reads. `producer_name` is the producer the model records, and empty takes the library's. `OnnxModel::from_graph` takes the same options, and an operator the chosen opset cannot state is refused, naming the node. In Python the options are the keyword arguments of the same names.
```cpp title="options.cpp"
int main() {
const fs::path dir = fresh_directory("options");
const std::string path = (dir / "opset17.onnx").string();
ModelGraph graph = trace_model();
OnnxExportOptions options;
options.opset_version = 17; // the ai.onnx opset the model imports; 0 takes the library's
options.producer_name = "my-exporter"; // the producer the model records; empty takes the library's
graph.export_onnx(path, options);
std::printf("opset %lld\n", static_cast(OnnxModel::open(path).opset()));
// opset 17
// contrib_ops = false keeps the model to the ai.onnx operators, and optimize = false writes it as
// lowered, without simplifying it.
OnnxExportOptions plain;
plain.contrib_ops = false;
plain.optimize = false;
ModelGraph back = OnnxModel::from_graph(graph, plain).compile();
back.finalize();
std::printf("%s\n", row_sums(back.run({input(2)}).front()).c_str());
// 31.625 52.5
fs::remove_all(dir);
return 0;
}
```
```python title="options.py"
graph = trace_model()
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / "opset17.onnx"
# opset_version is the ai.onnx opset the model imports (0 takes the library's), and producer_name the
# producer the model records (empty takes the library's).
graph.export_onnx(path, opset_version=17, producer_name="my-exporter")
print("opset", crt.io.OnnxModel.open(str(path)).opset)
# opset 17
# contrib_ops=False keeps the model to the ai.onnx operators, and optimize=False writes it as lowered,
# without simplifying it.
back = crt.io.OnnxModel.from_graph(graph, contrib_ops=False, optimize=False).compile()
back.finalize()
print(row_sums(back.run([rows(2)])[0]))
# 31.625 52.5
```
## A graph that rewrites a buffer in place
An ONNX model is a function of its inputs, so it cannot state a buffer that a graph changes in place. A trace of a step that writes a captured cache through an in-place operation surfaces the written cache as an extra output, `state_out_0`, and `export_onnx` refuses that graph with `NOT_IMPLEMENTED_FOR_PARAM`, naming the output that carries the update. A graph given a KV cache by `attach_kv_cache()` is refused the same way. Export the graph before `attach_kv_cache()`, or trace the step with the cache as an input and its update as an output. That graph exports, and its model reads the cache in and writes the update out. In Python the refusal raises `crt.UnsupportedError`.
```cpp title="state_buffer.cpp"
int main() {
const fs::path dir = fresh_directory("state_buffer");
const std::string path = (dir / "step.onnx").string();
// A step that adds x into a cache it captures, in place. The trace surfaces the written cache as the
// output state_out_0, and every run rewrites the cache's buffer, which ONNX cannot state.
const Tensor cache = Tensor::zeros({2, kWidth}, DataType::Float32);
const auto in_place = [cache](const std::vector& inputs) -> std::vector {
Tensor written = cache; // a handle on the live buffer
ops::add_(written, inputs[0]);
return {ops::relu(written)};
};
const std::vector x = {TensorSpec("x", DataType::Float32, {2, kWidth})};
const std::vector y = {"y"};
ModelGraph writes = ClikaRT::graph::trace(in_place, x, "step", y);
print_specs("output", writes.outputs());
// output y Float32 [2, 37]
// output state_out_0 Float32 [2, 37]
try {
writes.export_onnx(path);
} catch (const Error& error) {
print_refusal(error);
}
// NOT_IMPLEMENTED_FOR_PARAM | export_onnx: the graph rewrites a state buffer (its update is the output
// 'state_out_0') in place, which ONNX cannot state; export the graph before attach_kv_cache, or trace
// the step with the cache as an input and its update as an output
std::printf("files: %s\n", files_in(dir).c_str());
// files: none
// The same step with the cache as an input and its update as an output exports.
const auto explicit_cache = [](const std::vector& inputs) -> std::vector {
const Tensor updated = ops::add(inputs[1], inputs[0]); // cache + x
return {ops::relu(updated), updated};
};
const std::vector x_and_cache = {TensorSpec("x", DataType::Float32, {2, kWidth}),
TensorSpec("cache", DataType::Float32, {2, kWidth})};
const std::vector y_and_cache = {"y", "cache_out"};
ModelGraph step = ClikaRT::graph::trace(explicit_cache, x_and_cache, "step", y_and_cache);
step.export_onnx(path);
const OnnxModel model = OnnxModel::open(path);
print_specs("input", model.inputs());
print_specs("output", model.outputs());
// input x Float32 [2, 37]
// input cache Float32 [2, 37]
// output y Float32 [2, 37]
// output cache_out Float32 [2, 37]
fs::remove_all(dir);
return 0;
}
```
```python title="state_buffer.py"
cache = crt.zeros(2, WIDTH)
def in_place(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
# A step that adds x into the cache it captures, in place. The trace surfaces the written cache as the
# output state_out_0, and every run rewrites the cache's buffer, which ONNX cannot state.
crt.add_(cache, inputs[0])
return [crt.relu(cache)]
def explicit_cache(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
# The same step with the cache as an input and its update as an output.
updated = crt.add(inputs[1], inputs[0]) # cache + x
return [crt.relu(updated), updated]
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / "step.onnx"
writes = crt.trace(in_place, [crt.TensorSpec("x", crt.float32, [2, WIDTH])], output_names=["y"]).graph
print(writes.output_names())
# ['y', 'state_out_0']
try:
writes.export_onnx(path)
except crt.UnsupportedError as error:
print(refusal(error))
# NOT_IMPLEMENTED_FOR_PARAM | export_onnx: the graph rewrites a state buffer (its update is the output
# 'state_out_0') in place, which ONNX cannot state; export the graph before attach_kv_cache, or trace
# the step with the cache as an input and its update as an output
print(files_in(Path(tmp)))
# []
signature = [crt.TensorSpec("x", crt.float32, [2, WIDTH]), crt.TensorSpec("cache", crt.float32, [2, WIDTH])]
step = crt.trace(explicit_cache, signature, output_names=["y", "cache_out"]).graph
step.export_onnx(path)
model = crt.io.OnnxModel.open(str(path))
print(slots(model.inputs()), "->", slots(model.outputs()))
# x [2, 37] | cache [2, 37] -> y [2, 37] | cache_out [2, 37]
```
## What the export refuses
A refusal raises `ClikaRT::Error`, whose `code_name()` names the code, and writes nothing. Besides a buffer written in place, the export refuses:
- an operator with no ONNX form, naming the node and the operator (`NOT_IMPLEMENTED`). A custom operator is one by nature, since its kernel is your own code ([Write a custom operator](write-a-custom-operator.mdx) builds one);
- an opset the library does not know (`INVALID_ARGUMENT`);
- a call while `optimize()`, `finalize()`, `attach_kv_cache()`, `to()` or an edit runs on the graph (`FAILED_PRECONDITION`, naming the call). This program makes one from inside a transform that `optimize()` runs ([Write a transform](write-a-transform.mdx)).
`OnnxModel::from_graph` refuses the same way.
In Python a refusal raises a subclass of `crt.ClikaRTError` whose `code_name` names the code: `crt.UnsupportedError` for a code that starts with `NOT_IMPLEMENTED`, and `crt.InvalidArgumentError` for `INVALID_ARGUMENT` and `FAILED_PRECONDITION`. A path of another type raises `TypeError`. Python has no surface for writing a custom kernel, so the custom operator's refusal is the C++ tab's, and the Python tab shows the others.
```cpp title="refusals.cpp"
// A custom operator, y = 2x: its own shape rule and its own kernel, run on the host.
class TimesTwo final : public ClikaRT::nn::Module {
public:
static std::shared_ptr make() { return std::shared_ptr(new TimesTwo()); }
Tensor forward(const Tensor& x) const { return this->dispatch({&x, 1})[0]; }
std::vector output_shapes(
ClikaRT::Span inputs) const override {
return {ClikaRT::FakeTensor(inputs[0].shape(), inputs[0].dtype(), inputs[0].stream())};
}
void compute(ClikaRT::Span inputs, ClikaRT::Span outputs) const override {
const float* x = static_cast(inputs[0].const_data_ptr());
float* y = static_cast(outputs[0].mutable_data_ptr());
const std::int64_t n = inputs[0].numel();
for (std::int64_t i = 0; i < n; ++i) y[i] = 2.0F * x[i];
}
private:
TimesTwo() = default;
};
int main() {
const fs::path dir = fresh_directory("refusals");
const std::string path = (dir / "refused.onnx").string();
// An operator with no ONNX form is refused by name. A custom operator is one by nature: its kernel is
// your own code.
const std::shared_ptr twice = TimesTwo::make();
const auto doubled = [twice](const std::vector& inputs) -> std::vector {
return {twice->forward(inputs[0])};
};
const std::vector x = {TensorSpec("x", DataType::Float32, {2, kWidth})};
const std::vector y = {"y"};
ModelGraph custom = ClikaRT::graph::trace(doubled, x, "doubled", y);
try {
custom.export_onnx(path);
} catch (const Error& error) {
print_refusal(error);
}
// NOT_IMPLEMENTED | export_onnx: node 'Custom_1' (Custom) has no ONNX form: it is a user-defined operator
// An opset the library does not know.
ModelGraph graph = trace_model();
OnnxExportOptions unknown;
unknown.opset_version = 99;
try {
graph.export_onnx(path, unknown);
} catch (const Error& error) {
print_refusal(error);
}
// INVALID_ARGUMENT | no compatible ONNX IR version for the given opset version
// A call while another call holds the graph: here export_onnx from inside one of optimize()'s
// transforms.
const auto export_inside = [&path](ModelGraph& held) -> ClikaRT::Result {
try {
held.export_onnx(path);
} catch (const Error& error) {
print_refusal(error);
}
return {};
};
ClikaRT::graph::OptimizeOptions only_this;
only_this.transforms = std::vector{
ClikaRT::graph::Transform::from_function("export_inside", export_inside)};
graph.optimize(only_this);
// FAILED_PRECONDITION | export_onnx: optimize() is running on this graph; call export_onnx() after it
// returns
std::printf("files: %s\n", files_in(dir).c_str());
// files: none: a refusal writes nothing
fs::remove_all(dir);
return 0;
}
```
```python title="refusals.py"
graph = trace_model()
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / "refused.onnx"
# An opset the library does not know.
try:
graph.export_onnx(path, opset_version=99)
except crt.InvalidArgumentError as error:
print(refusal(error))
# INVALID_ARGUMENT | no compatible ONNX IR version for the given opset version
# A call while another call holds the graph: here export_onnx from inside one of optimize()'s
# transforms.
def export_inside(held: ModelGraph) -> None:
try:
held.export_onnx(path)
except crt.InvalidArgumentError as error:
print(refusal(error))
graph.optimize([Transform("export_inside", export_inside)])
# FAILED_PRECONDITION | export_onnx: optimize() is running on this graph; call export_onnx() after it
# returns
# A path of another type.
try:
graph.export_onnx(3)
except TypeError as error:
print(error)
# export_onnx(): path is a str or an os.PathLike, not int
print(files_in(Path(tmp)))
# []: a refusal writes nothing
```
---
# Handle errors by code
Read a failure's three channels, branch on the stable code name, and switch policy on the coarse status, in every language.
Source: https://docs.clika.io/clikart/how-to/handle-errors-by-code.md
Your program needs to react to a runtime failure without parsing prose. Every ClikaRT failure carries the same three channels, in every language: a **coarse status** (the policy switch: retry, reject, fail), a **human message** (diagnostics for a log line), and a **stable code name** (the fine-grained machine channel, e.g. `TIMED_OUT`, `BACKEND_NOT_LOADED`). The rule this page exists for: branch on the code name, never on the message text. A failure you can act on (a bad shape, an option out of range, a missing file) reads as plain text in every build and names the operation and the values it refused. A message of the form `E` means a fault inside the runtime: its value is specific to the build that produced it, so it is nothing to record, compare or branch on; report it verbatim with the runtime version. The code name is machine-readable in every build.
## Read the three channels
The program provokes one failure (a shape-mismatched matmul) and reads everything the error carries.
```cpp title="three_channels.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Error;
using ClikaRT::Status;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
int main() {
const Tensor a = Tensor::ones({3, 4}, DataType::Float32);
const Tensor bad = Tensor::ones({5, 5}, DataType::Float32);
try {
const Tensor y = ops::matmul(a, bad);
} catch (const Error& e) {
// The three channels of every runtime failure.
std::printf("status: %d\n", static_cast(e.status()));
std::printf("message: %s\n", e.what());
std::printf("code name: %s\n", e.code_name().c_str());
// Branch on the code NAME: stable in every build, and the only
// one of the three to compare against. A failure you can act on,
// like this one, states the shapes in plain text in every build; a
// message of the form E means a defect inside the runtime,
// and its value is specific to the build that produced it.
if (e.code_name() == "INVALID_ARGUMENT")
std::printf("-> reject this request, keep serving\n");
// The coarse status is the POLICY switch.
switch (e.status()) {
case Status::Unavailable: /* shed under load: retry later */ break;
case Status::InvalidArgument: /* caller bug: answer 400 */ break;
default: /* fail loudly */ break;
}
}
return 0;
}
```
```python title="three_channels.py"
import numpy as np
import clika_runtime as crt
def code_name_of(err: BaseException) -> str:
"""The machine-readable code name a runtime error carries ('' if none):
every ClikaRT exception exposes it as ``code_name``."""
return getattr(err, "code_name", "")
def main() -> None:
a = crt.tensor(np.ones((3, 4), dtype=np.float32))
bad = crt.tensor(np.ones((5, 5), dtype=np.float32))
try:
a @ bad
except RuntimeError as e:
# Python raises a typed exception per status: str(e) is the message
# alone, and the stable code name rides the exception's code_name.
print(f"message: {e}")
print(f"code name: {code_name_of(e)}")
# Branch on the code NAME: stable in every build, and the only
# one of the three to compare against. A failure you can act on,
# like this one, states the shapes in plain text in every build; a
# message of the form E means a defect inside the runtime,
# and its value is specific to the build that produced it.
if code_name_of(e) == "INVALID_ARGUMENT":
print("-> reject this request, keep serving")
if __name__ == "__main__":
main()
```
```text
status: 1
message: ops::matmul: cannot contract a [3, 4] with b [5, 5]: a's last dim (4) must equal b's dim 0 (5)
code name: INVALID_ARGUMENT
-> reject this request, keep serving
```
Status `1` is `InvalidArgument`, the value of the C++ `Status` enum, whose numbering is append-only. The message names the operation and the two shapes it refused, in a release build as in a debug build. Python prints the status by name rather than by number (`status: InvalidArgument`) and is otherwise identical: the exception carries `.status`, its text and `.code_name`, and its class is the status, so a caller can switch on the type instead of reading the name.
## Branch on the code name, never the message
The message exists for a human reading a log. An argument mistake reads as a sentence that names the operation and the values, in every build, but its wording can change between versions, so it is not a contract. A message of the form `E` is a runtime fault code: report it verbatim together with the runtime version; it is not an identifier to record, compare or branch on. The code name is the contract: a stable, append-only name for the fine-grained status the failure originated with, readable in every build. An empty code name is meaningful too: the failure did not originate inside the runtime (the C++ `Error` from your own code, a client-side load failure), so there is nothing machine-readable to branch on.
## A placement failure names its reason
One message with a fixed shape: placing work on a device whose backend did not come up fails with the reason spelled out, `cannot allocate on device : `. The `` names what actually happened: the backend library is not shipped beside the runtime, it failed to load (with the loader's cause), or it loaded and brought up no device. The code name stays the branch point; the message is now worth logging as is. On a machine without the Metal backend:
```cpp title="backend_reason.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Error;
using ClikaRT::Tensor;
int main() {
try {
const Tensor t = Tensor::ones({2, 2}, DataType::Float32, Device::metal());
std::printf("placed: %s\n", t.to_string().c_str());
} catch (const Error& e) {
std::printf("status: %d\n", static_cast(e.status()));
std::printf("message: %s\n", e.what());
std::printf("code name: %s\n", e.code_name().c_str());
}
return 0;
}
```
```text
status: 4
message: cannot allocate on device Metal:0: The metal backend could not load: not shipped in this install. The metal backend is not part of this ClikaRT build.
code name: BACKEND_NOT_LOADED
```
The same reason text reaches every binding through its language's error channel, and `clikart-cli` prints it when `--device` names a backend that cannot come up.
{/* CERTIFICATION: both captures are real output of the program shown, run
under the pinned release's clika_runtime cp313 manylinux wheel (sha256 checked
against the digest published with it) in a fresh venv on linux x86_64, with no
credential in the environment and no per-user license file present. The refusal texts are the
runtime's own and are quoted unchanged. */}
## A licensing refusal carries `LICENSE_FAILED`
ClikaRT runs compute under a license ([License the runtime](license-the-runtime.mdx) is where the credential goes). A call refused for licensing carries the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a credential past its end date, and the runtime writes one line to the console naming the refusal. On a process started with no credential at all:
```text
[ClikaRT] [error] clikart license invalid — no ClikaRT license is configured for this process
```
```python
import numpy as np
import clika_runtime as crt
t = crt.tensor(np.ones((2, 3), dtype=np.float32))
try:
(t + t).numpy()
except crt.ClikaRTError as e:
print(type(e).__name__, e.status, e.code_name, sep=" | ")
print(e)
```
```text
InternalError | Internal | LICENSE_FAILED
no ClikaRT license is configured for this process — a license key comes from the Licenses page of your project on the CLIKA platform; set CLIKA_RT_LICENSE to the key or to a file holding it, or install it with the clikart-license-init tool
```
The message names where the credential comes from and both places it can go, so it is worth logging as is. Licensing is where this page's rule earns its keep: the coarse status reads `Internal`, whose row in the table below sends a reader to file a report, and an unset variable is not something to file. Branch on `LICENSE_FAILED` before you consult the status.
Making a tensor is not an operator, so a tensor still builds without a credential; the refusal arrives at the first operator. A process that has to survive a missing credential checks for one before it starts computing, or catches the refusal at its first operator.
## Switch policy on the coarse status
The code name says what happened; the coarse status says what KIND of thing happened, which is usually all a policy needs:
| Status | Meaning | Usual reaction |
| --- | --- | --- |
| `InvalidArgument` | the caller's request cannot be right | reject it (a 400 in [the serving guide](serve-over-http.mdx)) |
| `NotFound` | a named thing does not exist | reject or fall back (a 404) |
| `Unsupported` | this build or backend cannot do it | fail fast at startup, not per request |
| `Internal` | the runtime's own invariant broke | log everything, file it |
| `Unavailable` | transient: the work was shed under load | retry later (a 503) |
## Without exceptions
C++ callers that would rather not pay for exceptions read the same three channels off a `Result`: `ok()`, `status()`, `message()`, `code_name()`. `CLIKART_TRY(...)` captures a throwing call as a `Result` where a failure is an expected outcome; [the serving guide](serve-over-http.mdx) uses it for exactly that on its 400 paths.
The serving-side `nn::KVCache` is a worked example of the `Result` surface: its per-layer accessors (`keys_impl`, `values_impl`, `conv_state_impl`, `recurrent_state_impl`) return a failed `Result` for a layer index outside `[0, num_layers())` instead of crashing, and `num_layers()` reports the count without allocating:
```cpp title="kv_bounds.cpp"
#include
#include
using ClikaRT::Result;
using ClikaRT::Tensor;
namespace nn = ClikaRT::nn;
int main() {
// The smallest real cache: two full-attention layers, one slot.
nn::KVCacheConfig config;
config.num_layers = 2;
config.num_kv_heads = 1;
config.head_dim = 4;
config.max_seqs = 1;
config.max_tokens_per_seq = 8;
const nn::KVLayerSpec specs[2] = {}; // AttentionKV rows by default
const nn::KVCache cache = nn::KVCache::make_impl(config, specs).value_or_throw();
// The per-layer accessors return a failed Result for a layer outside
// [0, num_layers()); no exception, no crash.
std::printf("layers: %d\n", cache.num_layers());
const Result in_range = cache.keys_impl(1);
const Result out_range = cache.keys_impl(2);
std::printf("keys(1): ok=%s\n", in_range.ok() ? "true" : "false");
std::printf("keys(2): ok=%s, message: %s, code name: %s\n",
out_range.ok() ? "true" : "false",
out_range.message().c_str(), out_range.code_name().c_str());
return 0;
}
```
```text
layers: 2
keys(1): ok=true
keys(2): ok=false, message: kv cache layer 2 is out of range: the cache has 2 layer(s), code name: INVALID_ARGUMENT
```
[Tutorial part 1](../getting-started/first-program/01-your-first-program.mdx) introduces the error model this page operationalizes; the python examples' errors chapter walks the same contract with an oracle per claim.
---
# License the runtime
Where ClikaRT and Modelverse read the credential your platform issues, per language and packaging, and what a refused call looks like.
Source: https://docs.clika.io/clikart/how-to/license-the-runtime.md
Every ClikaRT runtime needs a credential your platform issued for the project it ships under, and Modelverse, which runs on the same runtime, needs the same one. [Runtime licenses](/platform/concepts/runtime-licenses) is the concept page: why the credential exists, [the two kinds](/platform/concepts/runtime-licenses#the-two-kinds) (an online key, an offline bundle), the entitlements it carries, and [its lifecycle](/platform/concepts/runtime-licenses#lifecycle). This page is the runtime side of that loop: where the runtime looks for the credential, how to put it there from each language and packaging, and what a refused call carries.
## Where the runtime reads the credential
The runtime looks in three places, in this order, and uses the first one it finds:
1. The environment variable `CLIKA_RT_LICENSE`, holding either the credential text itself or the path of a file that holds it.
2. The environment variable `CLIKA_LICENSE`, in the same two forms.
3. The per-user file `~/.clika/runtime/license` (on Windows `%USERPROFILE%\.clika\runtime\license`), which `clikart-license-init ` writes once; the tool also takes the credential from a file with `--from-file `, and prompts for it on standard input when given neither, which keeps it out of your shell history. It ships in the bundle's `bin/` and as a console script of the `clika-runtime` Python wheel.
The tool writes the file readable by its owner alone and prints the file's path, the credential's kind (an online license's carrier bundle, which the platform validates on first use, or a signed offline bundle with its project and its expiry) and the permissions it set. It refuses to replace a file that is already there, so a second run asks for `--force` rather than overwriting what a machine is running on.
The credential is the `CLIKA1-...` license key or license bundle your project's Licenses page reveals. A bare `clika_rk_...` API key is not a runtime credential and is refused everywhere, the license-init tool included.
```bash
export CLIKA_RT_LICENSE=/path/to/credential # a file holding the credential
export CLIKA_RT_LICENSE="$(cat /path/to/credential)" # or the credential text itself
"$CLIKART_BUNDLE_DIR"/bin/clikart-license-init # or the per-user file, once per user
```
Put it in place before the program starts.
## Set it from each language
A C++ program carries no code for the credential. The runtime reads the process environment, so set the variable in the shell, the service unit or the launcher that starts the program, or run `clikart-license-init` once on the machine and rely on the per-user file.
```bash
export CLIKA_RT_LICENSE=/path/to/credential
./my_program
```
The `clika-runtime` wheel loads the runtime at import, so the variable is set before `import clika_runtime`: in the shell that starts Python, or in `os.environ` at the top of the program.
```python title="licensed.py"
import os
os.environ["CLIKA_RT_LICENSE"] = "/path/to/credential" # before the import: the wheel loads the runtime here
import clika_runtime as crt
x = crt.tensor([1.0, 2.0, 3.0])
```
The wheel also installs `clikart-license-init` as a console script. After `clikart-license-init ` has written the per-user file, the program needs no variable at all.
On Android there is no home directory for a per-user file and no `bin/` for the tool, so the credential arrives through the loader itself. `ClikaRtAndroid.load` takes it as its `license` argument; an explicit value wins over an ambient variable, and `null`, the default, leaves the environment alone.
```kotlin title="LicensedApp.kt"
import android.app.Application
import io.clika.runtime.ClikaRtAndroid
class LicensedApp : Application() {
override fun onCreate() {
super.onCreate()
// readCredential() is the app's own: it returns the text the platform
// revealed, from wherever the app keeps its secrets.
ClikaRtAndroid.load(this, license = readCredential())
}
}
```
`ClikaRtAndroid.load` also hands the runtime the app's Context, which is how the license identifies the device from inside an app (the identity Android scopes to the app). A command the app launches instead, such as the runtime's own command line, has no Context: set `XDG_CACHE_HOME` (or `HOME`) to a directory of the app in the child's environment, and the runtime keeps a per-installation identity there. A credential with a hardware seat budget refuses a process that can identify neither the device nor its installation.
On a desktop JVM, `ClikaRt.load()` needs nothing extra: the runtime reads the process environment, as in the C++ tab.
## What a refusal looks like
A call the runtime may not serve fails through the three channels every runtime failure carries ([Handle errors by code](handle-errors-by-code.mdx)): the coarse status, a message, and the stable code name. The code name is `LICENSE_EXPIRED` when the credential's lifetime has run out and `LICENSE_FAILED` for every other refusal. It reads the same in every language, so branch on it, never on the message. A credential that lacks a grant refuses the call that needs it, with the code the [entitlements](/platform/concepts/runtime-licenses#entitlements) section lists.
The coarse status on a licensing refusal reads `Internal`, whose usual policy is to log everything and file a report, which is the wrong reaction to an unset variable. Read the code name first; [Handle errors by code](handle-errors-by-code.mdx) shows the captured refusal and the branch.
## Online and offline
An online key needs a route from the runtime to the platform; an offline bundle is verified locally against Clika's signature and needs none. Which to issue, and how each is revoked, is the concept page's [Lifecycle](/platform/concepts/runtime-licenses#lifecycle) section.
## Containers and CI
Pass the credential through the environment of the service, never through the image. A `docker run -e`, an orchestrator secret exposed as an environment variable, or a CI job's secret variable all reach the runtime as the variable above; a credential written into an image layer travels with every copy of that image. Keep it out of logs the same way: log a refusal by its code name, and never echo the variable's value.
```bash
docker run --rm -e CLIKA_RT_LICENSE="$CLIKA_RT_LICENSE" my-image
```
Issuing, revealing, rotating and revoking a credential is platform work: [Runtime licenses](/platform/concepts/runtime-licenses) is the concept and [the licensing commands](/platform/cli/licensing) drive it from the command line. [Quick install](../getting-started/installation.md) lists the credential among ClikaRT's prerequisites, and [Modelverse's quick install](/modelverse/getting-started/installation) among its own. `clikart-cli` runs on this runtime and reads the credential from the same places.
---
# Load images and audio for inference
Decode image files into [H, W, C] tensors and audio into resampled waveforms, then make them model-ready with ops.
Source: https://docs.clika.io/clikart/how-to/load-images-and-audio.md
Vision and audio models consume tensors, and `ClikaRT::io` gets you there from the files themselves: `load_image` decodes the common encoded formats (JPEG, PNG, GIF, BMP, TGA, PSD, HDR, PNM, PIC) into an `[H, W, C]` tensor, and `load_audio` decodes WAV, FLAC, MP3 and OGG (Vorbis) into a `Float32` waveform, resampling and remixing to the layout you name. Both also take an in-memory byte span, so the same calls serve an upload handler as well as a file path. A file outside those sets is refused by name (`Status::Unsupported`, code `NOT_IMPLEMENTED`): `image decode: photo.webp: WebP is not supported; this build decodes JPEG, PNG, ...`; a damaged file, or a file of the other family, is refused as `InvalidArgument` with the file and the container it carries named, so a batch with one bad input says which one. The write direction ships too: `encode_audio` and `save_audio` turn a waveform back into a WAV, and `save_image` writes PNG by default or JPEG, BMP and TGA when the path's extension asks for them ([Write audio back](#write-audio-back)).
The runs below use a small RGB PNG (`input.png`, 8x6, a color gradient) and a half-second stereo WAV (`clip.wav`, 8 kHz, a 440 Hz tone on the left channel); any image or audio file of yours works the same.
## Decode an image and make it model-ready
`peek_image` reads only the header, the cheap way to validate and route before decoding. `load_image` decodes to `[H, W, C]` with 8-bit pixels; `requested_channels` forces a channel count (3 collapses an alpha channel or expands grayscale, so one call normalizes mixed inputs). The rest of the preprocessing every vision model wants (float, scale, `CHW`, batch dimension) is three ops.
```cpp title="image_to_input.cpp"
int main() {
const std::string input = input_png();
const io::ImageInfo info = io::peek_image(input);
std::printf("header: %lldx%lld, %lld channel(s)\n", static_cast(info.width),
static_cast(info.height), static_cast(info.channels));
const Tensor img = io::load_image(input, /*requested_channels=*/3);
std::printf("decoded: %s\n", img.to_string().c_str());
// Model-ready: float in [0,1], channels first, leading batch dim.
const Tensor x = ops::unsqueeze(
ops::permute(img.to(DataType::Float32) * (1.0 / 255.0), {2, 0, 1}), 0);
std::printf("input: %s\n", x.to_string().c_str());
std::printf("mean pixel value: %.4f\n", ops::mean(x).item());
return 0;
}
```
The image and audio decoders are part of the C++ API today; the C++ tab
shows both. From python, a decoded image or waveform enters as a numpy
array through `crt.tensor(array)`; the processor module then carries the
resize/normalize and feature steps.
```text
header: 8x6, 3 channel(s)
decoded: Tensor(shape=[6, 8, 3], dtype=UInt8, device=CPU, numel=144, data=[0, 0, 128, 32, 0, 128, ...])
input: Tensor(shape=[1, 3, 6, 8], dtype=Float32, device=CPU, numel=144, data=[0, 0.1255, 0.251, 0.3765, 0.502, 0.6275, ...])
mean pixel value: 0.4444
```
The decoded tensor is ordinary compute-engine currency from that point on: `.to(device)` moves it, and a model's forward takes it as-is. For bytes already in memory (an HTTP upload, an asset in an archive), pass a `Span` instead of the path; the `http_server` example's `06_image_compute` chapter runs exactly that flow behind an upload endpoint.
## Decode audio at the model's sample rate
Audio checkpoints are trained at a fixed sample rate and channel count, and `load_audio` meets them at the file: `target_sample_rate` resamples and `target_channels` remixes during the decode, so the tensor that comes out is already the model's input layout. The result is an `AudioData`: `samples` (`Float32`, `[frames]` mono or `[frames, channels]`) plus the effective `AudioInfo`.
```cpp title="audio_to_input.cpp"
int main() {
const std::string clip = clip_wav();
// Whatever the file's native layout, ask for 16 kHz mono.
const io::AudioData audio = io::load_audio(clip, /*target_sample_rate=*/16000,
/*target_channels=*/1);
std::printf("decoded: %d Hz, %d channel(s), %lld frames (%.2f s)\n",
audio.info.sample_rate, audio.info.channels,
static_cast(audio.info.frames),
static_cast(audio.info.frames) / audio.info.sample_rate);
std::printf("samples: %s\n", audio.samples.to_string().c_str());
// Loudness check: RMS over the waveform.
const Tensor rms = ops::sqrt(ops::mean(audio.samples * audio.samples));
std::printf("rms = %.4f\n", rms.item());
return 0;
}
```
This step is part of the C++ API today; the C++ tab shows it.
```text
decoded: 16000 Hz, 1 channel(s), 8000 frames (0.50 s)
samples: Tensor(shape=[8000], dtype=Float32, device=CPU, numel=8000, data=[0.0008917, 0.02581, 0.06115, 0.09333, 0.1175, 0.1378, ...])
rms = 0.1295
```
Passing 0 for either target keeps the file's native rate or channel count, and `peek_audio`'s `AudioInfo` answers "what would this decode produce" without decoding. From here the waveform feeds a feature front end (`ops::` has the FFT-adjacent reductions and windowed views a spectrogram needs) or goes straight into a model that consumes raw samples.
## Write audio back
The write direction is two calls: `encode_audio(samples, sample_rate)` returns the PCM16 WAV bytes of a waveform, and `save_audio(samples, sample_rate, path)` writes the same bytes to a file. A `Float16`, `BFloat16`, `Float32` or `Float64` waveform converts to 16-bit PCM; `Int16` is written as is. `[frames]` is mono and `[frames, channels]` is interleaved multi-channel, the same layouts the decoder produces, so a model's synthesized speech goes to disk with one call and an HTTP handler returns the encoded bytes without touching a file. The path's extension is the format request: `.wav` (or no extension) writes; a name asking for a container the writer does not produce (`reply.mp3`, `reply.ogg`) is refused by name with `Status::Unsupported` and nothing is written, so WAV bytes never land under a foreign suffix.
```cpp title="audio_write.cpp"
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace io = ClikaRT::io;
namespace ops = ClikaRT::ops;
int main() {
// Half a second of a 440 Hz tone at 16 kHz: t = frame / rate,
// s = 0.5 * sin(2 pi f t), a Float32 waveform in [-1, 1].
const int rate = 16000;
const Tensor t = ops::arange(0.0, 8000.0, 1.0, DataType::Float32) * (1.0 / rate);
const Tensor tone = ops::sin(t * (2.0 * 3.14159265358979 * 440.0)) * 0.5;
// encode_audio returns the PCM16 WAV bytes; save_audio writes the same
// bytes to a file.
const std::vector wav_bytes = io::encode_audio(tone, rate);
std::printf("encoded: %zu bytes of RIFF/WAVE\n", wav_bytes.size());
io::save_audio(tone, rate, "tone.wav");
// Round trip: decode what was written and check the layout survived.
const io::AudioData back = io::load_audio("tone.wav", rate, /*target_channels=*/1);
std::printf("reloaded: %d Hz, %d channel(s), %lld frames\n",
back.info.sample_rate, back.info.channels,
static_cast(back.info.frames));
const Tensor rms = ops::sqrt(ops::mean(back.samples * back.samples));
std::printf("rms = %.4f\n", rms.item());
return 0;
}
```
```python title="audio_write.py"
import math
import clika_runtime as crt
# Half a second of a 440 Hz tone at 16 kHz, a Float32 waveform in [-1, 1].
rate = 16000
t = crt.arange(0.0, 8000.0, 1.0) * (1.0 / rate)
tone = crt.sin(t * (2.0 * math.pi * 440.0)) * 0.5
# encode_audio returns the PCM16 WAV bytes; save_audio writes them to a file.
wav_bytes = crt.io.encode_audio(tone, rate)
print(f"encoded: {len(wav_bytes)} bytes of RIFF/WAVE")
crt.io.save_audio(tone, rate, "tone.wav")
# Round trip: decode what was written and check the layout survived.
# load_audio returns (samples, sample_rate, channels, frames).
samples, sample_rate, channels, frames = crt.io.load_audio("tone.wav", rate, 1)
print(f"reloaded: {sample_rate} Hz, {channels} channel(s), {frames} frames")
rms = crt.sqrt((samples * samples).mean())
print(f"rms = {rms.item():.4f}")
```
```text
encoded: 16044 bytes of RIFF/WAVE
reloaded: 16000 Hz, 1 channel(s), 8000 frames
rms = 0.3535
```
The byte count is the 8000 frames as 16-bit PCM plus the 44-byte RIFF header, and the round-trip RMS is the tone's own (0.5 amplitude over root two).
Decoding gives you pixels and samples; the model-specific half (resize, normalize, log-mel) is [the processors guide](preprocess-with-processors.mdx). Weights travel the same road in [the GGUF guide](load-quantized-weights.mdx); the bundle's `io` example walks NumPy and safetensors round-trips in its earlier chapters.
---
# Load quantized weights from a GGUF file
Read a block-quantized checkpoint, keep the weights packed, and serve them through QLinearWoQ at their on-disk footprint.
Source: https://docs.clika.io/clikart/how-to/load-quantized-weights.md
You have a block-quantized GGUF checkpoint and want to run it without inflating it to dense floats. ClikaRT loads the file as-is: `io::load_gguf` returns the tensors plus the file's metadata, quantized entries keep their packed bytes, and `nn::QLinearWoQ` runs the matmul off the packed form. The checkpoint's on-disk footprint is its in-memory footprint.
The programs below use [SmolLM2-135M-Instruct](https://huggingface.co/bartowski/SmolLM2-135M-Instruct-GGUF) (Apache-2.0, 145 MB in Q8_0), small enough to download in a minute. Any `.gguf` file works; nothing here depends on the architecture.
```bash
curl -LO "https://huggingface.co/bartowski/SmolLM2-135M-Instruct-GGUF/resolve/main/SmolLM2-135M-Instruct-Q8_0.gguf"
```
Each program is a complete `main.cpp`; build them like any bundle consumer ([tutorial part 1](../getting-started/first-program/01-your-first-program.mdx) has the four-line CMake project).
## Load the file and read its metadata
`io::load_gguf` returns a `GgufModel`: the weights as a `NamedTensors` map and the file's metadata as one `Json` object. The loader takes the target device (CPU by default), and the tensor payloads are mmap-backed, so loading is cheap.
```cpp title="inspect_metadata.cpp"
int main() {
const io::GgufModel model = io::load_gguf(howto::gguf_fixture());
std::printf("tensors %zu\n", model.tensors.size());
std::printf("metadata %zu keys\n", model.metadata.size());
for (const char* key : {"general.architecture", "general.size_label",
"llama.block_count", "llama.embedding_length"}) {
std::printf(" %-24s = %s\n", key, model.metadata.at(key).dump(0).c_str());
}
return 0;
}
```
```python title="inspect_metadata.py"
def main() -> None:
tensors, metadata_json = crt.io.load_gguf(gguf_path())
metadata = json.loads(metadata_json)
print(f"tensors {len(tensors)}")
print(f"metadata {len(metadata)} keys")
for key in ("general.architecture", "general.size_label",
"llama.block_count", "llama.embedding_length"):
print(f" {key:<24} = {json.dumps(metadata[key])}")
if __name__ == "__main__":
main()
```
```text
tensors 272
metadata 37 keys
general.architecture = "llama"
general.size_label = "135M"
llama.block_count = 30
llama.embedding_length = 576
```
The metadata carries everything the file knows about itself: architecture, hyperparameters, tokenizer configuration (`tokenizer.chat_template` included; [the chat-template guide](tokenize-and-chat-templates.mdx) picks that up). `at(key)` raises `ClikaRT::Error` on a missing key; probe with `contains` when a key is optional.
## What a quantized weight is in memory
A quantized entry rides a plain `Tensor` whose element data is the packed block stream, exactly as it sits in the file. `is_quantized()` separates those from the dense entries (norms and embeddings stay floating point in most files). `quantized_view` names what the payload is: the packed bytes, the scheme, and the logical element shape they encode.
```cpp title="inspect_weights.cpp"
int main() {
const io::GgufModel model = io::load_gguf(howto::gguf_fixture());
int quantized = 0, dense = 0;
model.tensors.for_each([&](std::string_view, const Tensor& t) {
if (t.is_quantized()) ++quantized; else ++dense;
});
std::printf("%d quantized, %d dense\n", quantized, dense);
const QTensor q = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight"));
std::printf("scheme %s\n", q.scheme.c_str());
std::printf("logical [%lld, %lld]\n",
static_cast(q.logical_shape[0]), static_cast(q.logical_shape[1]));
std::printf("payload %s\n", q.payload.to_string().c_str());
const Tensor norm = model.tensors.get("blk.0.attn_norm.weight");
std::printf("dense %s\n", norm.to_string().c_str());
return 0;
}
```
```python title="inspect_weights.py"
def main() -> None:
tensors, _ = crt.io.load_gguf(gguf_path())
quantized = sum(1 for t in tensors.values() if t.is_quantized)
print(f"{quantized} quantized, {len(tensors) - quantized} dense")
q = crt.quantized_view(tensors["blk.0.ffn_up.weight"])
print(f"scheme {q.scheme}")
print(f"logical [{q.logical_shape[0]}, {q.logical_shape[1]}]")
print(f"payload {q.payload}")
print(f"dense {tensors['blk.0.attn_norm.weight']}")
if __name__ == "__main__":
main()
```
```text
211 quantized, 61 dense
scheme GGUF_Q8_0
logical [576, 1536]
payload Tensor(shape=[1536, 612], dtype=UInt8, device=CPU, numel=940032, quantized=true, mmap=checkpoint:SmolLM2-135M-Instruct-Q8_0.gguf+33749344, data=[232, 27, 234, 24, 66, 230, ...])
dense Tensor(shape=[576], dtype=Float32, device=CPU, numel=576, mmap=checkpoint:SmolLM2-135M-Instruct-Q8_0.gguf+31866976, data=[0.01398, 0.0238, -0.01978, -0.03027, -0.01965, -0.03516, ...])
```
Two facts to keep. The **logical shape is `[in_features, out_features]`**, the row-contiguous dimension first; that is the orientation every quantized consumer below expects. The **payload is UInt8 `[rows, row_bytes]`**: for Q8_0, each row of 576 elements packs into blocks of 32 (one fp16 scale plus 32 int8 codes each), 34 bytes per block.
## Serve it packed with QLinearWoQ
`nn::QLinearWoQ` is a `Linear` over a quantized weight. The weight stays packed for the module's lifetime; the first `forward` reshapes the payload once into the backend's kernel layout, and every later call runs the quantized-weight matmul off that. Nothing is ever materialized dense.
The lifecycle is the same as every weight-bearing `nn` module: `make(in, out)` declares the slots, `set_weights` binds the loaded tensor, `forward` runs.
```cpp title="serve_packed.cpp"
int main() {
const io::GgufModel model = io::load_gguf(howto::gguf_fixture());
const QTensor w = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight"));
const std::int64_t in = w.logical_shape[0], out = w.logical_shape[1];
const std::shared_ptr ffn_up = QLinearWoQ::make(in, out);
ffn_up->set_weights(w);
const Tensor x = Tensor::full({1, in}, 0.01, DataType::Float32);
const Tensor y = ffn_up->forward(x); // first call packs, later calls reuse
std::printf("y = %s\n", y.to_string().c_str());
return 0;
}
```
```python title="run_quantized.py"
def main() -> None:
tensors, _ = crt.io.load_gguf(gguf_path())
q = crt.quantized_view(tensors["blk.0.ffn_up.weight"])
in_f, out_f = q.logical_shape
# Dequantize to the scheme's float target and run dense math on it.
# Serving the matmul off the PACKED form (no dense copy at rest) is
# done through the quantized modules of the C++ API; the C++ tab
# shows it.
w = crt.dequantize(q)
x = crt.full((1, in_f), 0.01)
y = F.linear(x, w)
print(f"y = {y}")
if __name__ == "__main__":
main()
```
```text
y = Tensor(shape=[1, 1536], dtype=Float32, device=CPU, numel=1536, data=[0.03397, -0.02917, 0.0793, 0.03382, -0.02785, 0.005078, ...])
```
One orientation trap: `QLinearWoQ` consumes the `[in, out]` logical shape that `load_gguf` produces, while dense `Linear` takes the HuggingFace `[out, in]` layout. Bind the GGUF entry as-is; do not transpose. Moving the module (`ffn_up->to(...)`) re-packs the weight on the target device, and a dtype move is refused: the weight stays quantized at rest.
## Inspect a weight by dequantizing
`ops::dequantize` decodes a packed weight into a dense tensor, `[out, in]` row-major. It is the inspection and tooling path, not the serving path; use it to eyeball values or to check a conversion. The program decodes the same weight, checks the packed forward against the dense one, and prints what staying packed saves.
```cpp title="check_dense.cpp"
int main() {
const io::GgufModel model = io::load_gguf(howto::gguf_fixture());
const QTensor w = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight"));
const std::int64_t in = w.logical_shape[0], out = w.logical_shape[1];
const Tensor dense = ops::dequantize(w); // [out, in], the scheme's float target
std::printf("dense = %s\n", dense.to_string().c_str());
// The packed path and the dense path compute the same values.
const std::shared_ptr packed = QLinearWoQ::make(in, out);
packed->set_weights(w);
const Tensor x = Tensor::full({1, in}, 0.01, DataType::Float32);
const Tensor y = packed->forward(x);
const Tensor yr = ops::linear(x, dense);
std::printf("max |packed - dense| over %lld outputs: %g\n",
static_cast(y.numel()),
ops::amax(ops::abs(ops::sub(y, yr))).item());
std::printf("bytes: dense %zu, packed %zu (%.2fx)\n", dense.nbytes(), w.payload.nbytes(),
static_cast(dense.nbytes()) / static_cast(w.payload.nbytes()));
return 0;
}
```
```python title="check_footprint.py"
def main() -> None:
tensors, _ = crt.io.load_gguf(gguf_path())
q = crt.quantized_view(tensors["blk.0.ffn_up.weight"])
dense = crt.dequantize(q) # [out, in], the scheme's float target
print(f"dense = {dense}")
# The packed payload is the checkpoint's own byte stream; inflating it
# to dense floats shows what staying packed saves. (The packed-vs-dense
# VALUE check runs where the packed matmul lives: the C++ tab.)
dense_bytes = dense.numpy().size * 4
packed_bytes = q.payload.numpy().size
print(f"bytes: dense {dense_bytes}, packed {packed_bytes} "
f"({dense_bytes / packed_bytes:.2f}x)")
if __name__ == "__main__":
main()
```
```text
dense = Tensor(shape=[1536, 576], dtype=Float32, device=CPU, numel=884736, data=[-0.08493, 0.09265, 0.2548, -0.1004, 0.1699, -0.193, ...])
max |packed - dense| over 1536 outputs: 7.18981e-07
bytes: dense 3538944, packed 940032 (3.76x)
```
The packed forward matches the dequantize-then-`ops::linear` reference to float rounding, at 3.76x fewer bytes for Q8_0 (Q4 and Q5 schemes save more).
## When the bytes did not come from GGUF
A packed payload from any other source gets the same treatment through `make_quantized(payload, scheme, logical_shape)`: a UInt8 CPU tensor of packed block rows, the scheme's name (`"GGUF_Q8_0"`, `"MXFP4_E8M0"`), and the element shape it encodes, row-contiguous dimension first. Checkpoints that ship MXFP4 as a split pair, 16 nibble-packed code bytes plus one E8M0 scale byte per 32-element group, go through `make_quantized_mxfp4(blocks, scales, logical_shape)`, which weaves the pair into the packed row layout the runtime consumes. Either way the result is the same `QTensor` the programs above served.
You can now open any GGUF checkpoint, say what every entry is, and serve its weights at their packed size. The bundle's `io` example (chapter `02_gguf`) inspects arbitrary files from the command line, and the [runtime example](../examples.md) wraps modules like `QLinearWoQ` into sessions and batching for serving.
---
# Merge and split
Join ModelGraphs into one graph that runs them all, side by side, with the outputs of one feeding the inputs of another, or across two devices; cut a graph into parts at its cheapest cut or by a function, extract the operators between values, and merge the parts back.
Source: https://docs.clika.io/clikart/how-to/merge-and-split.md
{/* Every block is a program under examples//howto/merge_and_split/: the first
block of each tab is get_the_graphs whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output, and it
reports across_devices as blocked on a host with no accelerator. */}
`merge` joins model graphs into one graph that runs them all. The graphs sit side by side, or the outputs of one feed the inputs of another, on one device or across two. `split` cuts one graph into parts that `merge` joins back, at the cheapest cut through its values or by a function that names each operator's part, and `extract` takes out the operators that compute some values from others. Each call consumes the graphs it is given and returns new ones. The calls are available from C++ and Python, under the same names.
## Graphs to merge and split
A trace returns the graph as built, with every operator as written and nothing optimized or finalized, so it is ready to merge or split. The examples on this page use three models over an input of shape [2, 3]. `rectify` computes `h = Relu(x)`, `twice` computes `y = x + x`, and `pool` negates the sum of each row of `Relu(x)`, which pools the [2, 3] input to [2, 1]. `trace_graph` traces a model under the input and output names a section needs. `label` names a node by its input name, or by its operator's code after the part its name starts with, so a merged graph reads `encoder/Relu`. `labels` lists a graph's nodes, and `io` its input names and then its output names.
This program runs each graph once. A finalized graph runs, and merge and split refuse it, so the other programs merge and split their graphs before `finalize()`.
```cpp title="get_the_graphs.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Error;
using ClikaRT::Tensor;
using ClikaRT::graph::Connection;
using ClikaRT::graph::GraphPart;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::NameKind;
using ClikaRT::graph::Node;
using ClikaRT::graph::NodeKind;
using ClikaRT::graph::Rename;
namespace ops = ClikaRT::ops;
namespace {
// h = Relu(x).
std::vector rectify(const std::vector& inputs) { return {ops::relu(inputs[0])}; }
// y = x + x.
std::vector twice(const std::vector& inputs) { return {ops::add(inputs[0], inputs[0])}; }
// y = -(the sum of each row of Relu(x)): the [2, 3] input pools to [2, 1].
std::vector pool(const std::vector& inputs) {
return {ops::neg(ops::sum(ops::relu(inputs[0]), {1}, true))};
}
// A trace of `model` over one float32 [2, 3] input, named `input`, with its output named `output`.
// A trace returns the graph as built: every operator as written, nothing optimized or finalized.
ModelGraph trace_graph(const ClikaRT::graph::TraceFunction& model, const std::string& input,
const std::string& output) {
const std::vector signature = {{input, DataType::Float32, {2, 3}}};
const std::vector outputs = {output};
return ClikaRT::graph::trace(model, signature, "merge_and_split", outputs);
}
// A node's label: an input's name, or an operator's code after the part its name starts with.
std::string label(const Node& node) {
if (node.kind() == NodeKind::Input) return node.name();
const std::string name = node.name();
const std::size_t slash = name.find('/');
const std::string part = slash == std::string::npos ? std::string() : name.substr(0, slash + 1);
return part + std::string(ClikaRT::graph::op_code_name(node.op_code()));
}
// The labels of a graph's nodes, in order, separated by spaces.
std::string labels(const ModelGraph& graph) {
std::string out;
for (const Node& node : graph.nodes()) out += (out.empty() ? "" : " ") + label(node);
return out;
}
// Names separated by spaces.
std::string names(const std::vector& list) {
std::string out;
for (const std::string& name : list) out += (out.empty() ? "" : " ") + name;
return out;
}
// A graph's input names, then its output names.
std::string io(const ModelGraph& graph) {
return names(graph.input_names()) + " -> " + names(graph.output_names());
}
// What a rename names.
const char* kind_name(NameKind kind) {
switch (kind) {
case NameKind::Node: return "node";
case NameKind::Input: return "input";
case NameKind::Output: return "output";
}
return "node";
}
// The input every program runs, [2, 3].
Tensor input() {
const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F};
return Tensor::from_data(x.data(), {2, 3}, DataType::Float32);
}
// A tensor's values, flattened and separated by spaces.
std::string values(const Tensor& tensor) {
std::string out;
for (const float value : tensor.reshape({-1}).item_as_vec()) {
char text[32];
std::snprintf(text, sizeof(text), "%g", static_cast(value));
out += (out.empty() ? "" : " ") + std::string(text);
}
return out;
}
} // namespace
int main() {
ModelGraph encoder = trace_graph(rectify, "x", "h");
ModelGraph head = trace_graph(twice, "h", "y");
ModelGraph pooled = trace_graph(pool, "x", "y");
for (ModelGraph* graph : {&encoder, &head, &pooled}) {
std::printf("%s | %s\n", labels(*graph).c_str(), io(*graph).c_str());
graph->finalize(); // a finalized graph runs, and merge and split refuse it
std::printf("%s\n", values(graph->run({input()}).front()).c_str());
}
// x Relu | x -> h
// 1 0 3 0 5 0
// h Add | h -> y
// 2 -4 6 -8 10 -12
// x Relu Sum Neg | x -> y
// -4 -5
return 0;
}
```
```python title="get_the_graphs.py"
from collections.abc import Callable
import clika_runtime as crt
from clika_runtime.graph import ModelGraph, NameKind, Node, NodeKind
Model = Callable[[list[crt.Tensor]], list[crt.Tensor]]
def rectify(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
return [crt.relu(inputs[0])] # h = Relu(x)
def twice(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
return [crt.add(inputs[0], inputs[0])] # y = x + x
def pool(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
# y = -(the sum of each row of Relu(x)): the [2, 3] input pools to [2, 1].
return [crt.neg(crt.sum(crt.relu(inputs[0]), [1], True))]
def trace_graph(model: Model, input_name: str, output_name: str) -> ModelGraph:
# A trace of `model` over one float32 [2, 3] input, named `input_name`, with its output named `output_name`.
# A trace returns the graph as built: every operator as written, nothing optimized or finalized.
signature = [crt.TensorSpec(input_name, crt.float32, [2, 3])]
return crt.trace(model, signature, output_names=[output_name]).graph
def label(node: Node) -> str:
# A node's label: an input's name, or an operator's code after the part its name starts with.
if node.kind == NodeKind.Input:
return node.name
part, slash, _ = node.name.partition("/")
return (part + slash if slash else "") + node.op_code.name
def labels(graph: ModelGraph) -> str:
return " ".join(label(node) for node in graph.nodes())
def io(graph: ModelGraph) -> str:
# A graph's input names, then its output names.
return " ".join(graph.input_names()) + " -> " + " ".join(graph.output_names())
def kind_name(kind: NameKind) -> str:
# What a rename names: a node, an input or an output.
return kind.name.lower()
def x() -> crt.Tensor:
# The input every program runs, [2, 3].
return crt.tensor([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], dtype=crt.float32)
def values(tensor: crt.Tensor) -> str:
# A tensor's values, flattened and separated by spaces.
return " ".join(f"{value:g}" for value in crt.to(tensor, "cpu").reshape(-1).tolist())
encoder = trace_graph(rectify, "x", "h")
head = trace_graph(twice, "h", "y")
pooled = trace_graph(pool, "x", "y")
for graph in (encoder, head, pooled):
print(labels(graph), "|", io(graph))
graph.finalize() # a finalized graph runs, and merge and split refuse it
print(values(graph.run([x()])[0]))
# x Relu | x -> h
# 1 0 3 0 5 0
# h Add | h -> y
# 2 -4 6 -8 10 -12
# x Relu Sum Neg | x -> y
# -4 -5
```
Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, the three models and the helpers).
## Merge graphs side by side
`graph::merge(parts)` takes a list of `GraphPart`s, each a graph and the name the merged graph knows it by. Parts that no connection joins sit side by side. The merged graph takes each part's inputs and outputs, parts in the order given, and names every operator of part `p` as `p/`. `MergeOptions::io_names` decides the names of the inputs and outputs. Under `IoNames::PrefixOnCollision`, the default, a name two parts both use becomes `/` and every other name stays as it is. `IoNames::PrefixAlways` prefixes every name, and `IoNames::Keep` keeps every name and refuses a name two parts both use. `MergedGraph::renames` lists each name the merge changed as a `Rename`, which holds its part, whether it names a node, an input or an output, and the name before and after. Here the encoder and the head both read an input named `x`, so the default prefixes the two inputs and keeps the two outputs.
In Python, `graph.merge` takes the parts as a dict from name to graph, or as a list of `(name, graph)` pairs, and returns `(merged, renames)`. `io_names` takes the mode's name, such as `"prefix_always"`, and a `Rename`'s fields are `part`, `kind`, `from_` and `to`.
```cpp title="side_by_side.cpp"
int main() {
// An encoder and a head that both read an input named x.
const auto two_parts = [] {
std::vector parts;
parts.push_back({"encoder", trace_graph(rectify, "x", "h")});
parts.push_back({"head", trace_graph(twice, "x", "y")});
return parts;
};
// IoNames::PrefixOnCollision, the default: a name both parts use takes its part's name as a prefix.
std::vector parts = two_parts();
const ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts);
std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str());
// encoder/x head/x encoder/Relu head/Add | encoder/x head/x -> h y
for (const Rename& rename : merged.renames) { // every name the merge changed
std::printf("%s %s %s -> %s\n", rename.part.c_str(), kind_name(rename.kind), rename.from.c_str(),
rename.to.c_str());
}
// encoder input x -> encoder/x
// head input x -> head/x
// encoder node Relu_1 -> encoder/Relu_1
// head node Add_1 -> head/Add_1
// IoNames::PrefixAlways: every input and output name takes the prefix.
ClikaRT::graph::MergeOptions always;
always.io_names = ClikaRT::graph::IoNames::PrefixAlways;
parts = two_parts();
std::printf("%s\n", io(ClikaRT::graph::merge(parts, {}, always).graph).c_str());
// encoder/x head/x -> encoder/h head/y
// IoNames::Keep keeps every name, so it refuses a name both parts use.
ClikaRT::graph::MergeOptions keep;
keep.io_names = ClikaRT::graph::IoNames::Keep;
parts = two_parts();
try {
ClikaRT::graph::merge(parts, {}, keep);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
}
// INVALID_ARGUMENT | merge: parts 'encoder' and 'head' both have an input named 'x', which IoNames::Keep
// keeps; merge with IoNames::PrefixOnCollision or PrefixAlways
std::printf("%s | %s\n", io(parts[0].graph).c_str(), io(parts[1].graph).c_str());
// x -> h | x -> y: a refused merge leaves every part as it was
return 0;
}
```
```python title="side_by_side.py"
def two_parts() -> list[tuple[str, ModelGraph]]:
# An encoder and a head that both read an input named x.
return [("encoder", trace_graph(rectify, "x", "h")), ("head", trace_graph(twice, "x", "y"))]
# io_names="prefix_on_collision", the default: a name both parts use takes its part's name as a prefix.
merged, renames = crt.graph.merge(two_parts())
print(labels(merged), "|", io(merged))
# encoder/x head/x encoder/Relu head/Add | encoder/x head/x -> h y
for rename in renames: # every name the merge changed
print(rename.part, kind_name(rename.kind), rename.from_, "->", rename.to)
# encoder input x -> encoder/x
# head input x -> head/x
# encoder node Relu_1 -> encoder/Relu_1
# head node Add_1 -> head/Add_1
# io_names="prefix_always": every input and output name takes the prefix.
always, _ = crt.graph.merge(two_parts(), io_names="prefix_always")
print(io(always)) # encoder/x head/x -> encoder/h head/y
# io_names="keep" keeps every name, so it refuses a name both parts use.
parts = two_parts()
try:
crt.graph.merge(parts, io_names="keep")
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | merge: parts 'encoder' and 'head' both have an input named 'x', which IoNames::Keep
# keeps; merge with IoNames::PrefixOnCollision or PrefixAlways
print(io(parts[0][1]), "|", io(parts[1][1])) # x -> h | x -> y: a refused merge leaves every part as it was
```
Merge graphs side by side to run several models as one graph, with one `finalize()` and one `run()`.
## Feed one graph into another
A `Connection` feeds an output of one part into an input of another. `from` names the part and its output, `to` names the part and its input, and a connected output or input is not one of the merged graph's. Here the encoder's `h` feeds the head's `h`, so the merged graph reads `x`, returns `y`, and computes what running the two graphs one after the other computes. The two ends of a connection have one dtype and one rank, and every size both of them fix is the same. In Python, a connection is a pair of strings, `("encoder.h", "head.h")`.
```cpp title="pipeline.cpp"
int main() {
std::vector parts;
parts.push_back({"encoder", trace_graph(rectify, "x", "h")});
parts.push_back({"head", trace_graph(twice, "h", "y")});
// The encoder's output h feeds the head's input h, and neither is the merged graph's any longer.
const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}};
ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections);
std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str());
// x encoder/Relu head/Add | x -> y
merged.graph.optimize(); // optional: the graph optimizer runs over the merged graph
merged.graph.finalize();
// The same input through the two graphs one after the other, for comparison.
ModelGraph encoder = trace_graph(rectify, "x", "h");
ModelGraph head = trace_graph(twice, "h", "y");
encoder.finalize();
head.finalize();
std::printf("%s | %s\n", values(merged.graph.run({input()}).front()).c_str(),
values(head.run(encoder.run({input()})).front()).c_str());
// 2 0 6 0 10 0 | 2 0 6 0 10 0
return 0;
}
```
```python title="pipeline.py"
# The encoder's output h feeds the head's input h, and neither is the merged graph's any longer.
parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")}
merged, _ = crt.graph.merge(parts, [("encoder.h", "head.h")])
print(labels(merged), "|", io(merged)) # x encoder/Relu head/Add | x -> y
merged.optimize() # optional: the graph optimizer runs over the merged graph
merged.finalize()
# The same input through the two graphs one after the other, for comparison.
encoder = trace_graph(rectify, "x", "h")
head = trace_graph(twice, "h", "y")
encoder.finalize()
head.finalize()
print(values(merged.run([x()])[0]), "|", values(head.run(encoder.run([x()]))[0])) # 2 0 6 0 10 0 | 2 0 6 0 10 0
```
Feed one graph into another to serve a model with its preprocessing or its post-processing as one graph.
## Merge a second graph into this one
`ModelGraph::merge(second, io_map)` merges `second` into the graph it is called on. Each `io_map` entry feeds one of this graph's outputs into one of `second`'s inputs. This graph keeps every name, and a name of `second` that this graph already uses takes the first free `_`. The call returns those renames, each under the part `second`, and it consumes `second`. Everything else follows `graph::merge`, and a connection between two devices moves its value. Here both graphs are traces of `rectify`, so the second Relu takes a new name. In Python, `first.merge(second, [("h", "x")])` returns the renames.
```cpp title="merge_into.cpp"
int main() {
ModelGraph first = trace_graph(rectify, "x", "h");
ModelGraph second = trace_graph(rectify, "x", "h");
// first's output h feeds second's input x. first keeps every name, and a name of second that
// first already uses takes the first free _.
const std::vector renames = first.merge(second, {{"h", "x"}});
std::printf("%s | %s\n", labels(first).c_str(), io(first).c_str()); // x Relu Relu | x -> h
for (const Rename& rename : renames) {
std::printf("%s %s %s -> %s\n", rename.part.c_str(), kind_name(rename.kind), rename.from.c_str(),
rename.to.c_str());
}
// second node Relu_1 -> Relu_1_1
std::printf("%zu %zu\n", second.input_names().size(), second.output_names().size()); // 0 0: second is consumed
first.finalize();
std::printf("%s\n", values(first.run({input()}).front()).c_str()); // 1 0 3 0 5 0: Relu(Relu(x))
return 0;
}
```
```python title="merge_into.py"
first = trace_graph(rectify, "x", "h")
second = trace_graph(rectify, "x", "h")
# first's output h feeds second's input x. first keeps every name, and a name of second that
# first already uses takes the first free _.
renames = first.merge(second, [("h", "x")])
print(labels(first), "|", io(first)) # x Relu Relu | x -> h
for rename in renames:
print(rename.part, kind_name(rename.kind), rename.from_, "->", rename.to)
# second node Relu_1 -> Relu_1_1
print(len(second.input_names()), len(second.output_names())) # 0 0: second is consumed
first.finalize()
print(values(first.run([x()])[0])) # 1 0 3 0 5 0: Relu(Relu(x))
```
Merge a second graph into a graph to grow it in place while every name it has stays as it is.
## Connect graphs on two devices
A part's values sit on the device `to()` sent the part to, and otherwise where they are computed. When a connection's two ends sit on two devices, `CrossDevice::Move`, the default, moves the value onto the input's device, and `CrossDevice::Refuse` refuses the merge, naming the connection and both devices. When the parts do not share one device, the merge places each part on its device before it builds, and the merged graph keeps every part where it was placed. Here the head serves on the accelerator that `Device::gpu()` finds, so the merged graph computes the Relu on the CPU, moves `h` to the accelerator, and returns `y` there. On a host with no accelerator, the program prints `BLOCKED: this section needs an accelerator` and exits with status 3. In Python, the option is `cross_device="refuse"`.
```cpp title="across_devices.cpp"
int main() {
const ClikaRT::Device accelerator = ClikaRT::Device::gpu(); // the CPU on a host with no accelerator
if (accelerator.is_cpu()) {
std::printf("BLOCKED: this section needs an accelerator\n");
return 3;
}
// The encoder computes on the CPU, and the head serves on the accelerator.
const auto two_parts = [&accelerator] {
std::vector parts;
parts.push_back({"encoder", trace_graph(rectify, "x", "h")});
parts.push_back({"head", trace_graph(twice, "h", "y")});
parts[1].graph.to(accelerator);
return parts;
};
const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}};
// CrossDevice::Move, the default: the merged graph moves h onto the head's device.
std::vector parts = two_parts();
ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections);
merged.graph.finalize();
const Tensor y = merged.graph.run({input()}).front();
std::printf("%s %s\n", y.device() == accelerator ? "true" : "false",
values(y.to(ClikaRT::Device::cpu())).c_str()); // true 2 0 6 0 10 0: y is on the accelerator
// CrossDevice::Refuse: a connection between two devices refuses the merge.
ClikaRT::graph::MergeOptions refuse;
refuse.cross_device = ClikaRT::graph::CrossDevice::Refuse;
parts = two_parts();
try {
ClikaRT::graph::merge(parts, connections, refuse);
} catch (const Error& error) {
std::printf("%s\n", error.code_name().c_str()); // INVALID_ARGUMENT, naming the connection and its devices
}
std::printf("%s | %s\n", io(parts[0].graph).c_str(), io(parts[1].graph).c_str());
// x -> h | h -> y: a refused merge leaves every part as it was
return 0;
}
```
```python title="across_devices.py"
accelerator = crt.Device.gpu() # the CPU on a host with no accelerator
if accelerator.type == "cpu":
print("BLOCKED: this section needs an accelerator")
raise SystemExit(3)
def two_parts() -> dict[str, ModelGraph]:
# The encoder computes on the CPU, and the head serves on the accelerator.
parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")}
parts["head"].to(accelerator)
return parts
connections = [("encoder.h", "head.h")]
# cross_device="move", the default: the merged graph moves h onto the head's device.
merged, _ = crt.graph.merge(two_parts(), connections)
merged.finalize()
y = merged.run([x()])[0]
print(y.device == accelerator, values(y)) # True 2 0 6 0 10 0: y is on the accelerator
# cross_device="refuse": a connection between two devices refuses the merge.
parts = two_parts()
try:
crt.graph.merge(parts, connections, cross_device="refuse")
except crt.InvalidArgumentError as error:
print(error.code_name) # INVALID_ARGUMENT, naming the connection and its devices
print(io(parts["encoder"]), "|", io(parts["head"]))
# x -> h | h -> y: a refused merge leaves every part as it was
```
Connect graphs on two devices to keep one part on the CPU while the model runs on an accelerator.
## Split a graph at its cheapest cut
`min_value_cut()` finds the cut through a graph's values that hands the fewest bytes from the nodes before it to the nodes after it, as [The cheapest cut](query-a-graph.mdx#the-cheapest-cut) on the Query a graph page shows. `graph::split(graph, cut)` cuts the graph there into two parts. The part `before` holds the operators that produce the cut values and every operator they depend on, and the part `after` holds every other operator. Connections carry the cut values `after` reads and any graph input both parts read, and a value with no name of its own crosses under the name `_`. Here the cheapest values are the [2, 1] sums in `pool`, 8 bytes against 24 for `x`, so `before` computes the sums and `after` negates them. The split consumes the graph. In Python, `graph.split(graph, cut)` returns `(parts, connections)`, the parts as a dict from name to graph and the connections in the form `graph.merge` takes.
```cpp title="split_at_a_cut.cpp"
int main() {
ModelGraph pooled = trace_graph(pool, "x", "y");
// The cheapest cut: the values the nodes before it hand to the nodes after it, and their bytes.
const ClikaRT::graph::ValueCut cut = pooled.min_value_cut();
std::string crossing;
for (const ClikaRT::graph::Value& value : cut.values) {
crossing += (crossing.empty() ? "" : " ") + label(*value.producer());
}
std::printf("%s %llu\n", crossing.c_str(), static_cast(cut.bytes));
// Sum 8: the [2, 1] float32 sums cross, the cheapest value to hand on
ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, cut);
for (const GraphPart& part : split.parts) {
std::printf("%s: %s | %s\n", part.name.c_str(), labels(part.graph).c_str(), io(part.graph).c_str());
}
// before: x Relu Sum | x -> Sum_2_0
// after: Sum_2_0 Neg | Sum_2_0 -> y
for (const Connection& connection : split.connections) {
std::printf("%s.%s -> %s.%s\n", connection.from.part.c_str(), connection.from.name.c_str(),
connection.to.part.c_str(), connection.to.name.c_str());
}
// before.Sum_2_0 -> after.Sum_2_0: a value with no name of its own is named _
std::printf("%zu\n", pooled.input_names().size()); // 0: the split consumed the graph
return 0;
}
```
```python title="split_at_a_cut.py"
pooled = trace_graph(pool, "x", "y")
# The cheapest cut: the values the nodes before it hand to the nodes after it, and their bytes.
cut = pooled.min_value_cut()
print(" ".join(label(value.producer()) for value in cut.values), cut.bytes)
# Sum 8: the [2, 1] float32 sums cross, the cheapest value to hand on
parts, connections = crt.graph.split(pooled, cut)
for name, part in parts.items():
print(f"{name}:", labels(part), "|", io(part))
# before: x Relu Sum | x -> Sum_2_0
# after: Sum_2_0 Neg | Sum_2_0 -> y
for source, target in connections:
print(source, "->", target)
# before.Sum_2_0 -> after.Sum_2_0: a value with no name of its own is named _
print(len(pooled.input_names())) # 0: the split consumed the graph
```
Split a graph at its cheapest cut to run its two halves on two devices, or in two processes, with the least data between them.
## Split a graph by a function
`graph::split(graph, partition)` asks a function you write for each operator's part, a name that is not empty and holds no `/` and no `.`. It asks once per operator, in `nodes()` order, and the parts come in an order every connection runs forward in. A value read in a part other than its producer's becomes an output of the producer's part and an input of each part that reads it, under one name. A graph input goes with the first part that reads it, and every later reader receives it through a connection. Here the function names each operator's part by its code. In Python, the function takes a `Node` and returns a `str`.
```cpp title="split_by_a_function.cpp"
int main() {
ModelGraph pooled = trace_graph(pool, "x", "y");
// Each operator's part, by its code. split asks once per operator, in nodes() order.
const auto by_code = [](const Node& node) {
if (node.op_code() == ClikaRT::graph::OpCode::Relu) return std::string("rectify");
if (node.op_code() == ClikaRT::graph::OpCode::Sum) return std::string("pool");
return std::string("negate");
};
ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, by_code);
for (const GraphPart& part : split.parts) { // in an order every connection runs forward in
std::printf("%s: %s | %s\n", part.name.c_str(), labels(part.graph).c_str(), io(part.graph).c_str());
}
// rectify: x Relu | x -> Relu_1_0
// pool: Relu_1_0 Sum | Relu_1_0 -> Sum_2_0
// negate: Sum_2_0 Neg | Sum_2_0 -> y
for (const Connection& connection : split.connections) {
std::printf("%s.%s -> %s.%s\n", connection.from.part.c_str(), connection.from.name.c_str(),
connection.to.part.c_str(), connection.to.name.c_str());
}
// rectify.Relu_1_0 -> pool.Relu_1_0
// pool.Sum_2_0 -> negate.Sum_2_0
return 0;
}
```
```python title="split_by_a_function.py"
def by_code(node: Node) -> str:
# Each operator's part, by its code. split asks once per operator, in nodes() order.
if node.op_code == crt.graph.OpCode.Relu:
return "rectify"
if node.op_code == crt.graph.OpCode.Sum:
return "pool"
return "negate"
pooled = trace_graph(pool, "x", "y")
parts, connections = crt.graph.split(pooled, by_code)
for name, part in parts.items(): # in an order every connection runs forward in
print(f"{name}:", labels(part), "|", io(part))
# rectify: x Relu | x -> Relu_1_0
# pool: Relu_1_0 Sum | Relu_1_0 -> Sum_2_0
# negate: Sum_2_0 Neg | Sum_2_0 -> y
for source, target in connections:
print(source, "->", target)
# rectify.Relu_1_0 -> pool.Relu_1_0
# pool.Sum_2_0 -> negate.Sum_2_0
```
Split a graph by a function to give the operators you choose, by code, by name or by position, a graph of their own.
## Extract the operators between values
`graph::extract(graph, inputs, outputs)` returns the operators that compute `outputs` from `inputs` as a graph of their own. A graph input keeps its name, and any other input value becomes an input named after it. An output keeps its graph output name when it has one, and is otherwise named after its value. The extract consumes the graph and frees the weights of the operators it leaves out. Weights stored inside an ONNX file, rather than as external data, share one buffer, which the extracted graph keeps in memory as the graph did. In Python, `graph.extract(graph, inputs, outputs)` takes lists of `Value`s.
```cpp title="extract.cpp"
int main() {
ModelGraph pooled = trace_graph(pool, "x", "y");
// The operators between x and the sums: the Relu and the Sum, not the Neg after them.
const Node sum = pooled.find_nodes(ClikaRT::graph::OpCode::Sum).front();
const std::vector inputs = {*pooled.node("x")->output(0)};
const std::vector outputs = {*sum.output(0)};
ModelGraph pooling = ClikaRT::graph::extract(pooled, inputs, outputs);
std::printf("%s | %s\n", labels(pooling).c_str(), io(pooling).c_str());
// x Relu Sum | x -> Sum_2_0: a graph input keeps its name, and an output with none is named after its value
std::printf("%zu\n", pooled.input_names().size()); // 0: the extract consumed the graph
pooling.finalize();
std::printf("%s\n", values(pooling.run({input()}).front()).c_str()); // 4 5: each row's sum of Relu(x)
return 0;
}
```
```python title="extract.py"
pooled = trace_graph(pool, "x", "y")
# The operators between x and the sums: the Relu and the Sum, not the Neg after them.
(sums,) = pooled.find_nodes(crt.graph.OpCode.Sum)
pooling = crt.graph.extract(pooled, [pooled.node("x").output(0)], [sums.output(0)])
print(labels(pooling), "|", io(pooling))
# x Relu Sum | x -> Sum_2_0: a graph input keeps its name, and an output with none is named after its value
print(len(pooled.input_names())) # 0: the extract consumed the graph
pooling.finalize()
print(values(pooling.run([x()])[0])) # 4 5: each row's sum of Relu(x)
```
Extract a block of a model to run it alone, to test it, or to serve it on its own.
## Merge the parts back
`graph::merge(split.parts, split.connections)` joins the parts of a split back into one graph, with the source's inputs and outputs by name and its operators named `/`. The merged graph computes what the source computes. In Python, `graph.merge(*graph.split(graph, cut))` does the same in one call.
```cpp title="round_trip.cpp"
int main() {
ModelGraph pooled = trace_graph(pool, "x", "y");
ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, pooled.min_value_cut());
// The parts merge back into the source's inputs and outputs, by name, with its operators named /.
ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(split.parts, split.connections);
std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str());
// x before/Relu before/Sum after/Neg | x -> y
merged.graph.finalize();
ModelGraph traced = trace_graph(pool, "x", "y"); // the same model traced again, for comparison
traced.finalize();
std::printf("%s | %s\n", values(merged.graph.run({input()}).front()).c_str(),
values(traced.run({input()}).front()).c_str()); // -4 -5 | -4 -5
return 0;
}
```
```python title="round_trip.py"
pooled = trace_graph(pool, "x", "y")
# split returns the parts and connections merge takes, so the parts merge back into the source's inputs and
# outputs, by name, with its operators named /.
merged, _ = crt.graph.merge(*crt.graph.split(pooled, pooled.min_value_cut()))
print(labels(merged), "|", io(merged)) # x before/Relu before/Sum after/Neg | x -> y
merged.finalize()
traced = trace_graph(pool, "x", "y") # the same model traced again, for comparison
traced.finalize()
print(values(merged.run([x()])[0]), "|", values(traced.run([x()])[0])) # -4 -5 | -4 -5
```
Merge the parts back after you place, optimize or edit them one at a time.
## What the calls consume, and what a refusal leaves
A merge consumes every part's graph, and a split or an extract consumes its source. On success those graphs are empty, and a `Node`, `Value` or `Edge` view taken into one of them refuses with the code name `INVALID_ARGUMENT`. Each call checks everything before it changes anything, so a refused call leaves every graph as it was. The one exception is a merge that fixed a dynamic size through a connection before it refused, which leaves that size fixed, the size the merge would give it. The code name is `INVALID_ARGUMENT` for an argument the call cannot take, and `FAILED_PRECONDITION` for a graph it cannot take now, such as a finalized graph. A partition function that throws makes `split` refuse with the function's failure. A thrown `ClikaRT::Error` keeps its status, and any other exception comes back as `INTERNAL`. In Python, `graph.split` raises the function's own exception. [Handle errors by code](handle-errors-by-code.mdx) covers the channels every failure carries.
```cpp title="consumption.cpp"
int main() {
std::vector parts;
parts.push_back({"encoder", trace_graph(rectify, "x", "h")});
parts.push_back({"head", trace_graph(twice, "h", "y")});
const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}};
const Node relu = parts[0].graph.find_nodes(ClikaRT::graph::OpCode::Relu).front(); // a view taken before
// A merge consumes every part's graph, and a view into one refuses from then on.
const ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections);
std::printf("%s | %zu %zu\n", io(merged.graph).c_str(), parts[0].graph.input_names().size(),
parts[1].graph.input_names().size()); // x -> y | 0 0
try {
relu.check();
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
}
// INVALID_ARGUMENT | the node 'Relu_1' is no longer in its graph: that graph no longer exists (a merge,
// a split or an extract consumed it, or it was destroyed or assigned over)
// A refused merge leaves every part as it was.
std::vector again;
again.push_back({"encoder", trace_graph(rectify, "x", "h")});
again.push_back({"head", trace_graph(twice, "h", "y")});
const std::vector wrong = {{{"encoder", "h"}, {"head", "z"}}};
try {
ClikaRT::graph::merge(again, wrong);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
}
// INVALID_ARGUMENT | merge: part 'head' has no input 'z'
again[0].graph.finalize();
try {
ClikaRT::graph::merge(again, connections);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
}
// FAILED_PRECONDITION | merge: part 'encoder' is finalized; merge before finalize()
std::printf("%s | %s\n", io(again[0].graph).c_str(), io(again[1].graph).c_str()); // x -> h | h -> y
// A partition function that throws: split refuses with its failure, and the graph stays as it was.
ModelGraph pooled = trace_graph(pool, "x", "y");
try {
ClikaRT::graph::split(pooled, [](const Node& node) -> std::string {
if (node.op_code() == ClikaRT::graph::OpCode::Neg) {
throw Error(ClikaRT::Status::InvalidArgument, "no part for Neg");
}
return "first";
});
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
}
// INVALID_ARGUMENT | split: the partition function failed on the node 'Neg_3': no part for Neg
std::printf("%s\n", io(pooled).c_str()); // x -> y
return 0;
}
```
```python title="consumption.py"
def refuse_neg(node: Node) -> str:
# A partition function that raises on the Neg.
if node.op_code == crt.graph.OpCode.Neg:
raise ValueError("no part for Neg")
return "first"
parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")}
connections = [("encoder.h", "head.h")]
(relu,) = parts["encoder"].find_nodes(crt.graph.OpCode.Relu) # a view taken before
# A merge consumes every part's graph, and a view into one refuses from then on.
merged, _ = crt.graph.merge(parts, connections)
print(io(merged), "|", len(parts["encoder"].input_names()), len(parts["head"].input_names())) # x -> y | 0 0
try:
relu.check()
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | the node 'Relu_1' is no longer in its graph: that graph no longer exists (a merge,
# a split or an extract consumed it, or it was destroyed or assigned over)
# A refused merge leaves every part as it was.
again = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")}
try:
crt.graph.merge(again, [("encoder.h", "head.z")])
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error) # INVALID_ARGUMENT | merge: part 'head' has no input 'z'
again["encoder"].finalize()
try:
crt.graph.merge(again, connections)
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# FAILED_PRECONDITION | merge: part 'encoder' is finalized; merge before finalize()
print(io(again["encoder"]), "|", io(again["head"])) # x -> h | h -> y
# An exception the partition function raises leaves split as that same exception, and the graph as it was.
pooled = trace_graph(pool, "x", "y")
try:
crt.graph.split(pooled, refuse_neg)
except ValueError as error:
print(type(error).__name__, "|", error) # ValueError | no part for Neg
print(io(pooled)) # x -> y
```
Reuse the graphs a refused call leaves: they are as they were, and a corrected call takes them.
---
# Optimize and finalize
A traced or compiled ModelGraph comes back as built. Run the graph optimizer with optimize() and read its report, choose its transforms and its defaults, place the graph with to(), and finalize() it to run.
Source: https://docs.clika.io/clikart/how-to/optimize-and-finalize.md
{/* Every block is a program under examples//howto/optimize_and_finalize/: the first
block of each tab is get_a_graph whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output. */}
`compile()` and `trace()` return a `ModelGraph` as built, with every operator as the model states it and nothing optimized, placed or finalized. Four calls take it from there. `optimize()` runs the graph optimizer when you want it and reports what it did. `to()` names the device or the stream the graph serves on. `finalize()` places the graph there, packs the weights into their kernel layouts and plans the execution order, and from then on the graph runs. The calls are available from C++ and Python, under the same names.
## A graph as built
The model below computes `y = Relu(Neg(Neg(x)))`. The two Negs cancel each other, which gives the graph optimizer one rewrite to find. Every example on this page traces it. `labels` lists a graph's nodes in order, each by its input name or its operator, and `trace_model` traces the model over an `x` of shape [2, 3]. The input every program runs is `input()` in C++ and `x` in Python, and `values` prints a C++ result on one line.
A graph that is not finalized refuses `run()` with the code name `FAILED_PRECONDITION`, and the message names the call to make.
```cpp title="get_a_graph.cpp"
#include
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Error;
using ClikaRT::Tensor;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::Node;
using ClikaRT::graph::NodeKind;
using ClikaRT::graph::OptimizeOptions;
using ClikaRT::graph::OptimizeReport;
using ClikaRT::graph::Transform;
using ClikaRT::graph::TransformReport;
namespace ops = ClikaRT::ops;
namespace transforms = ClikaRT::graph::transforms;
namespace {
// y = Relu(Neg(Neg(x))): the two Negs cancel, which the graph optimizer finds.
std::vector model(const std::vector& inputs) {
return {ops::relu(ops::neg(ops::neg(inputs[0])))};
}
// A node's label: an input's name, an operator's code.
std::string label(const Node& node) {
return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code()));
}
// The labels of a graph's nodes, in order, separated by spaces.
std::string labels(const ModelGraph& graph) {
std::string out;
for (const Node& node : graph.nodes()) out += (out.empty() ? "" : " ") + label(node);
return out;
}
// A trace returns the graph as built: every operator as written, nothing optimized or finalized.
ModelGraph trace_model() {
const std::vector signature = {{"x", DataType::Float32, {2, 3}}};
const std::vector outputs = {"y"};
return ClikaRT::graph::trace(model, signature, "lifecycle", outputs);
}
// The input every program runs, [2, 3].
Tensor input() {
const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F};
return Tensor::from_data(x.data(), {2, 3}, DataType::Float32);
}
// A tensor's values, flattened and separated by spaces.
std::string values(const Tensor& tensor) {
std::string out;
for (const float value : tensor.reshape({-1}).item_as_vec()) {
char text[32];
std::snprintf(text, sizeof(text), "%g", static_cast(value));
out += (out.empty() ? "" : " ") + std::string(text);
}
return out;
}
} // namespace
int main() {
const ModelGraph graph = trace_model();
std::printf("%s | %s\n", labels(graph).c_str(), graph.is_finalized() ? "true" : "false"); // x Neg Neg Relu | false
// Only a finalized graph runs.
try {
graph.run({input()});
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// FAILED_PRECONDITION | run: the graph is not finalized; call finalize() first
}
return 0;
}
```
```python title="get_a_graph.py"
import numpy as np
import clika_runtime as crt
from clika_runtime.graph import DefaultsOptions, NodeKind, transforms
def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
return [crt.relu(crt.neg(crt.neg(inputs[0])))] # y = Relu(Neg(Neg(x))): the two Negs cancel
def labels(graph: crt.graph.ModelGraph) -> list[str]:
return [node.name if node.kind == NodeKind.Input else node.op_code.name for node in graph.nodes()]
# A trace returns the graph as built: every operator as written, nothing optimized or finalized.
def trace_model() -> crt.graph.ModelGraph:
return crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph
x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32) # the input every program runs (numpy as the data entry)
graph = trace_model()
print(labels(graph), graph.is_finalized()) # ['x', 'Neg', 'Neg', 'Relu'] False
# Only a finalized graph runs.
try:
graph.run([crt.tensor(x)])
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# FAILED_PRECONDITION | run: the graph is not finalized; call finalize() first
```
Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, the helpers, the input and, in Python, the trace).
## Run the graph optimizer
`optimize()` runs the graph's default list of transforms. Each iteration runs the whole list once, and the run stops at the first iteration that changes nothing. It returns an `OptimizeReport`. `iterations` counts the iterations that ran, `converged` says the run stopped on its own, and `changed` says it modified the graph. `transforms` holds one `TransformReport` per list entry, in list order, and each row's `applications` counts the iterations in which its transform changed the graph. Here `remove_double_neg` removes both Negs in the first iteration, and the second iteration changes nothing.
```cpp title="optimize.cpp"
int main() {
ModelGraph graph = trace_model();
const OptimizeReport report = graph.optimize();
std::printf("%s\n", labels(graph).c_str()); // x Relu
std::printf("%lld %s %s\n", static_cast(report.iterations), report.converged ? "true" : "false",
report.changed ? "true" : "false"); // 2 true true
for (const TransformReport& row : report.transforms) {
if (row.applications > 0) {
std::printf("%s %lld\n", row.transform.name().c_str(), static_cast(row.applications));
// remove_double_neg 1
}
}
return 0;
}
```
```python title="optimize.py"
report = graph.optimize()
print(labels(graph)) # ['x', 'Relu']
print(report.iterations, report.converged, report.changed) # 2 True True
print([(row.transform.name, row.applications) for row in report.transforms if row.applications])
# [('remove_double_neg', 1)]
```
A view you took before `optimize()` finds its node again by its key, or refuses once the optimizer removed that node, as [Views across edits](query-a-graph.mdx#views-across-edits) on the Query a graph page shows.
## The default list
`default_transforms()` returns the list `optimize()` runs on this graph, in the order it runs them, as a starting point to edit and pass back. `transforms::defaults()` (in Python, `transforms.defaults()`) returns the default list for a set of choices without a graph. A graph's own list follows how the graph was built, so the two can differ. For a traced graph with no choices set, they are the same list.
Each entry is a `Transform` handle named as its factory is, so `transforms::remove_double_neg()` names `remove_double_neg` (`name()` in C++, the `name` property in Python). Compare and store transforms by name, since the values behind `kind()` move when the runtime's transform list changes.
```cpp title="default_list.cpp"
int main() {
const ModelGraph graph = trace_model();
const std::vector listed = graph.default_transforms();
const auto holds = [&listed](const Transform& transform) {
return std::find(listed.begin(), listed.end(), transform) != listed.end() ? "true" : "false";
};
std::printf("%s\n", listed.front().name().c_str()); // cleanup_graph
std::printf("%s %s\n", holds(transforms::remove_double_neg()),
holds(transforms::fixate_dyn_shape_queues())); // true false
std::printf("%s\n", listed == transforms::defaults() ? "true" : "false"); // true
return 0;
}
```
```python title="default_list.py"
listed = graph.default_transforms()
print(listed[0].name) # cleanup_graph
print(transforms.remove_double_neg() in listed,
transforms.fixate_dyn_shape_queues() in listed) # True False
print(listed == transforms.defaults()) # True
```
## Run a list of your own
`OptimizeOptions::transforms` runs the list you give it instead, in order, each iteration. In Python, pass the list as the first argument of `optimize()`. A list may hold a transform more than once. An empty list runs none of them, though the run still drops the nodes nothing reads.
The list is checked before anything runs. A transform placed ahead of one it must follow is refused with the code name `INVALID_ARGUMENT`, and the message names both entries and the order to use. The refused call leaves the graph as it was.
```cpp title="own_list.cpp"
int main() {
ModelGraph graph = trace_model();
OptimizeOptions options;
options.transforms = std::vector{transforms::remove_double_neg()};
const OptimizeReport report = graph.optimize(options);
const TransformReport& row = report.transforms.front();
std::printf("%s | %zu %s %lld\n", labels(graph).c_str(), report.transforms.size(), row.transform.name().c_str(),
static_cast(row.applications)); // x Relu | 1 remove_double_neg 1
// A list is checked before anything runs: matmul_absorb_transpose must run after remove_double_permute.
ModelGraph other = trace_model();
options.transforms = std::vector{transforms::matmul_absorb_transpose(), transforms::remove_double_permute()};
try {
other.optimize(options);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// INVALID_ARGUMENT | optimize: transforms[0] (matmul_absorb_transpose) must run after remove_double_permute, which the list places later, at transforms[1]; list remove_double_permute ahead of it
}
std::printf("%s\n", labels(other).c_str()); // x Neg Neg Relu
return 0;
}
```
```python title="own_list.py"
report = graph.optimize([transforms.remove_double_neg()])
print(labels(graph), [(row.transform.name, row.applications) for row in report.transforms])
# ['x', 'Relu'] [('remove_double_neg', 1)]
# A list is checked before anything runs: matmul_absorb_transpose must run after remove_double_permute.
other = trace_model()
try:
other.optimize([transforms.matmul_absorb_transpose(), transforms.remove_double_permute()])
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | optimize: transforms[0] (matmul_absorb_transpose) must run after remove_double_permute, which the list places later, at transforms[1]; list remove_double_permute ahead of it
print(labels(other)) # ['x', 'Neg', 'Neg', 'Relu']
```
A list can also hold a transform you write, a function that edits the graph and runs beside the runtime's transforms (`Transform::from_function` in C++, `Transform(name, fn)` in Python). [Write a transform](write-a-transform.mdx) covers them.
## Cap the iterations
`max_num_iterations` caps one run at that many iterations. It is 100 unless you set it, and a value outside 1 to 2147483647 is refused with `INVALID_ARGUMENT`. A run that reaches the cap still leaves a correct graph. Its report reads `converged` as false, `firing_at_cap` marks the transforms that changed the graph in the last iteration, and the runtime logs a warning that names them. Here the single iteration applies `remove_double_neg`, and the cap ends the run before a second iteration can show that nothing else changes.
```cpp title="iteration_cap.cpp"
int main() {
ModelGraph graph = trace_model();
OptimizeOptions options;
options.max_num_iterations = 1; // one pass over the list
const OptimizeReport report = graph.optimize(options);
std::printf("%lld %s %s\n", static_cast(report.iterations), report.converged ? "true" : "false",
report.changed ? "true" : "false"); // 1 false true
for (const TransformReport& row : report.transforms) {
if (row.firing_at_cap) std::printf("%s\n", row.transform.name().c_str()); // remove_double_neg
}
std::printf("%s\n", labels(graph).c_str()); // x Relu
options.max_num_iterations = 0;
try {
graph.optimize(options);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// INVALID_ARGUMENT | optimize: max_num_iterations must be between 1 and 2147483647, got 0
}
return 0;
}
```
```python title="iteration_cap.py"
report = graph.optimize(max_num_iterations=1) # one pass over the list
print(report.iterations, report.converged, report.changed) # 1 False True
print([row.transform.name for row in report.transforms if row.firing_at_cap]) # ['remove_double_neg']
print(labels(graph)) # ['x', 'Relu']
try:
graph.optimize(max_num_iterations=0)
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | optimize: max_num_iterations must be between 1 and 2147483647, got 0
```
A transform of the runtime's that fails is undone and skipped for the rest of the run, which goes on. Its row reads `failed`, with the status name in `failure_code` and the message in `failure_message`. The graph stays correct, only less optimized.
## Fix the shapes
`DefaultsOptions` holds the choices the transforms read, each off by default. With every choice off, the graph keeps its input and output contract and serves every shape it admits. Two choices fix the graph to its shapes instead, collapsing shape arithmetic such as shape cones, masks and reshape templates into constants. Each adds the shape-fixation transforms, `fixate_dyn_shape_queues` among them, to the default list.
A traced graph follows `shape_fixation`, which fixes it to the shapes it was traced with. A graph compiled from ONNX follows `specialize`, which fixes it to the input shapes the compile pinned (`CompileOptions::specs` or `sample_inputs`). Without pinned shapes, the arithmetic the model's own static shapes fix collapses either way. The graph-free list takes the choices as stated.
Pass the choices to `optimize()` in `OptimizeOptions::defaults` (in Python, the `defaults=` keyword), and to `default_transforms()` to see the list they give. A list of your own that holds a transform these choices switch off is refused, and the message names the choice to set.
```cpp title="defaults.cpp"
int main() {
const ModelGraph graph = trace_model();
const Transform fixate = transforms::fixate_dyn_shape_queues();
const auto holds = [&fixate](const std::vector& listed) {
return std::find(listed.begin(), listed.end(), fixate) != listed.end() ? "true" : "false";
};
// A traced graph follows shape_fixation: its list gains the transforms that fix it to the traced shapes.
ClikaRT::graph::DefaultsOptions fixation;
fixation.shape_fixation = true;
std::printf("%s %s\n", holds(graph.default_transforms()), holds(graph.default_transforms(fixation))); // false true
// A graph compiled from ONNX follows specialize; the graph-free list takes the choice as stated.
ClikaRT::graph::DefaultsOptions specialize;
specialize.specialize = true;
std::printf("%s %s\n", holds(transforms::defaults()), holds(transforms::defaults(specialize))); // false true
return 0;
}
```
```python title="defaults.py"
fixate = transforms.fixate_dyn_shape_queues()
# A traced graph follows shape_fixation: its list gains the transforms that fix it to the traced shapes.
print(fixate in graph.default_transforms(),
fixate in graph.default_transforms(DefaultsOptions(shape_fixation=True))) # False True
# A graph compiled from ONNX follows specialize; the graph-free list takes the choice as stated.
print(fixate in transforms.defaults(),
fixate in transforms.defaults(DefaultsOptions(specialize=True))) # False True
```
## Channels-last inputs and outputs
A model exported channels-first reaches the runtime's channels-last operators through a conversion at each input a spatial operator, such as a convolution, reads. `channels_last_inputs` omits that conversion, so the input is supplied channels-last (`[N, spatial..., C]`) from then on. `channels_last_outputs` does the same for an output that leaves through a conversion. Both change the graph's input and output contract, and `inputs()` and `outputs()` report the layout to supply and the layout delivered.
The program builds its ONNX model with a helper outside the block, an `X` of shape [1, 2, 4, 4] that a 1x1 Conv reads, followed by a Relu, and `layout` prints an input's name and dims. `drop_boundary_permutes`, the transform that omits the conversions, runs only under one of the two choices.
```cpp title="channels_last.cpp"
int main() {
ModelGraph graph = compile_conv_model();
std::printf("%s\n", layout(graph.inputs().front()).c_str()); // X [1, 2, 4, 4]
// drop_boundary_permutes runs only under a channels-last choice.
OptimizeOptions listed;
listed.transforms = std::vector{transforms::drop_boundary_permutes()};
try {
graph.optimize(listed);
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// INVALID_ARGUMENT | optimize: transforms[0] (drop_boundary_permutes) changes the layout of the graph's inputs and outputs; set DefaultsOptions::channels_last_inputs or channels_last_outputs to run it
}
OptimizeOptions options;
options.defaults.channels_last_inputs = true; // X is supplied channels-last from here on
graph.optimize(options);
std::printf("%s\n", layout(graph.inputs().front()).c_str()); // X [1, 4, 4, 2]
return 0;
}
```
```python title="channels_last.py"
graph = compile_conv_model()
print(layout(graph)) # X [1, 2, 4, 4]
# drop_boundary_permutes runs only under a channels-last choice.
try:
graph.optimize([transforms.drop_boundary_permutes()])
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | optimize: transforms[0] (drop_boundary_permutes) changes the layout of the graph's inputs and outputs; set DefaultsOptions::channels_last_inputs or channels_last_outputs to run it
graph.optimize(defaults=DefaultsOptions(channels_last_inputs=True)) # X is supplied channels-last from here on
print(layout(graph)) # X [1, 4, 4, 2]
```
## Place the graph with to()
`to()` takes a device or a stream. Before `finalize()`, it records where finalizing places the graph and moves nothing, and `device()` reports the new device at once. After `finalize()`, it moves the finalized graph. An operator that changes device re-packs its weights there once, and no copy of a moved weight stays behind. A graph no `to()` moved serves where its operators sit, the CPU for this trace, and a graph compiled from ONNX serves on the device `CompileOptions::weights` names until `to()` names another. In Python, `device` is a property.
To compute on a stream of your own, create it and pass it to `to()` once, and every run computes there. On a host with an accelerator, `to(Device::cuda(0))` places the operators and their weights on it; the program stays on the CPU, so it runs on any host. `to()` refuses while a run is in flight, while a KV cache is attached (call it before `attach_kv_cache()`), and while `optimize()`, `finalize()` or an edit runs on the graph.
```cpp title="placement.cpp"
int main() {
ModelGraph graph = trace_model();
std::printf("%s\n", ClikaRT::device::compute_api_name(graph.device().api)); // CPU: where its operators sit
// Before finalize(), to() records where finalize() places the graph; nothing moves yet.
graph.to(ClikaRT::Stream::create(ClikaRT::Device::cpu()));
std::printf("%s %s\n", ClikaRT::device::compute_api_name(graph.device().api),
graph.is_finalized() ? "true" : "false"); // CPU false
graph.optimize();
graph.finalize(); // places the operators on that stream and packs their weights there
std::printf("%s\n", values(graph.run({input()}).front()).c_str()); // 1 0 3 0 5 0
// After finalize(), to() moves the finalized graph.
graph.to(ClikaRT::Device::cpu());
std::printf("%s\n", values(graph.run({input()}).front()).c_str()); // 1 0 3 0 5 0
return 0;
}
```
```python title="placement.py"
print(graph.device) # cpu: where its operators sit
# Before finalize(), to() records where finalize() places the graph; nothing moves yet.
graph.to(crt.Stream.create(crt.Device.cpu()))
print(graph.device, graph.is_finalized()) # cpu False
graph.optimize()
graph.finalize() # places the operators on that stream and packs their weights there
(y,) = graph.run([crt.tensor(x)])
print(y.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]]
# After finalize(), to() moves the finalized graph.
graph.to(crt.Device.cpu())
(again,) = graph.run([crt.tensor(x)])
print(again.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]]
```
## Finalize and run
`finalize()` places the graph on its device, packs the weights into their kernel layouts, absorbs the constants the operators hold and plans the execution order. A second call succeeds and changes nothing. A finalized graph runs, benches and takes a KV cache. It refuses `optimize()` and every edit with the code name `FAILED_PRECONDITION`, and the message says to compile or trace the model again, while `run()` and `to()` keep working. Each program checks the run against a reference it computes itself, the Python one with numpy.
```cpp title="finalize.cpp"
int main() {
ModelGraph graph = trace_model();
graph.optimize();
graph.finalize();
graph.finalize(); // a second call succeeds and changes nothing
std::printf("%s\n", graph.is_finalized() ? "true" : "false"); // true
const Tensor y = graph.run({input()}).front();
const std::vector got = y.reshape({-1}).item_as_vec();
const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F};
bool matches = got.size() == x.size();
for (std::size_t i = 0; matches && i < x.size(); ++i) {
matches = got[i] == std::max(x[i], 0.0F); // the reference, Relu(x), by hand
}
std::printf("%s\n", values(y).c_str()); // 1 0 3 0 5 0
std::printf("%s\n", matches ? "true" : "false"); // true
// A finalized graph refuses optimize() and every edit; run() and to() keep working.
try {
graph.optimize();
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// FAILED_PRECONDITION | optimize: the graph is finalized; optimize() runs before finalize(), so compile or trace the model again to optimize it
}
try {
graph.rename_output("y", "out");
} catch (const Error& error) {
std::printf("%s\n", error.code_name().c_str()); // FAILED_PRECONDITION
}
return 0;
}
```
```python title="finalize.py"
graph.optimize()
graph.finalize()
graph.finalize() # a second call succeeds and changes nothing
print(graph.is_finalized()) # True
(y,) = graph.run([crt.tensor(x)])
print(y.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]]
print(np.array_equal(y.numpy(), np.maximum(x, 0))) # True (numpy states the reference)
# A finalized graph refuses optimize() and every edit; run() and to() keep working.
try:
graph.optimize()
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# FAILED_PRECONDITION | optimize: the graph is finalized; optimize() runs before finalize(), so compile or trace the model again to optimize it
try:
graph.rename_output("y", "out")
except crt.InvalidArgumentError as error:
print(error.code_name) # FAILED_PRECONDITION
```
Optimize and edit before you finalize. [Edit a graph](edit-a-graph.mdx) covers the edits, and [Query a graph](query-a-graph.mdx) reads the graph at any point in its lifecycle.
---
# Package ClikaRT in an Android app
What an Android app ships to run the runtime and the model library on the phone: the one artifact and the libraries inside it, the build requirements, the license credential, the thread count, the cache root, and a server that outlives the screen.
Source: https://docs.clika.io/clikart/how-to/package-clikart-in-an-android-app.md
{/* CERTIFICATION: every fact on this page is read from the Kotlin binding's README and gradle files of the
pinned release (compileSdk, minSdk, the plugin versions, the AAR contents, the load entries) and from the
release's Kotlin artifact itself (its Maven repository directory); no app on this page is compiled by the
docs build. */}
The [tutorial's last part](../getting-started/first-program/06-deploy-to-mobile.mdx) pushes a
binary to a phone over `adb`. An app does the same thing from inside its own process: it ships
the runtime library, loads it through the Kotlin binding, and places the license credential
before the first call. This page is the inventory of what the app carries and the few facts
that are easy to get wrong the first time.
## The one artifact
One Maven artifact, `io.clika:clika-runtime`, carries the runtime and the model library's Kotlin API
together. It is published as a Maven repository directory inside `clika-runtime-maven-.zip`,
a download of the platform beside the release archives ([Get ClikaRT](../getting-started/get-clikart.mdx));
extract it anywhere and name the directory as a repository, then one dependency line serves both
products:
```kotlin
repositories {
maven { url = uri("/path/to/clika-runtime-maven-") }
google()
mavenCentral()
}
dependencies {
implementation("io.clika:clika-runtime:")
}
```
The artifact is a Kotlin Multiplatform library with two variants; gradle reads its module metadata and
picks the one the project builds for:
| Variant | Form | Carries |
| --- | --- | --- |
| Android (`arm64-v8a`) | an AAR | the Kotlin classes of both packages (`io.clika.runtime`, `io.clika.modelverse`), the two JNI bridges, every library of the Android release archive (`libClikaRT.so`, its backend images, `libClikaRT_modelverse.so`) under `jni/arm64-v8a/`, the consumer keep rules, and the archive's third-party notices under `META-INF/licenses/` |
| desktop JVM (Linux, Windows, macOS) | a jar | the Kotlin classes of both packages, the two JNI bridges per desktop platform, and the notices; the runtime and model libraries come from the release archive's `lib/` on `java.library.path` |
On Android the AAR carries everything: the app copies nothing into `jniLibs`. At run time the
binding loads `libClikaRT.so` first, then the runtime bridge; `Modelverse.load` loads
`libClikaRT_modelverse.so` and the model library's bridge after it. The binding is built against one
release and refuses to load another, naming both versions (`BINDING_VERSION_MISMATCH`).
## The build requirements
| Requirement | Value |
| --- | --- |
| `compileSdk` | 36, the level the binding's modules compile against |
| `minSdk` | 28: a device below it installs the AAR but cannot load the runtime library |
| Android Gradle plugin | 9.3.2, which needs gradle 9.5 or newer |
| Kotlin | 2.3 or newer in the app: the binding's classes are compiled by Kotlin 2.3, and an older compiler refuses their metadata |
| JDK | 17 |
| NDK | none: an app that takes the published artifact builds nothing native |
| ABI | `arm64-v8a` alone |
The binding reaches nothing by reflection, so an app's R8 pass needs no rule of its own: the AAR
carries the consumer rules for the classes its bridges bind by name.
## Load, and place the credential first
An app has no shell to export a variable from and no per-user license file, so the credential is
an argument of the load. `ClikaRtAndroid.load(context, license = "CLIKA1-...")` places it for
the process and loads the runtime; an app that uses the model library calls
`Modelverse.load(context, license)` instead, which runs the whole sequence, the runtime first.
A credential given there replaces one already in the process environment; the default, `null`,
leaves the environment as it is. The credential's text is never logged. Ship it the way you ship
any other secret the app needs at start-up: a build-time field read from a file outside the
source tree, or a value the app asks for once and keeps in its own secret store. Without a valid
credential the first operator throws a `ClikaRtException` (a `ModelverseException` from the
model library) whose `codeName` is `LICENSE_FAILED`, or `LICENSE_EXPIRED` past the end date.
Two more facts are placed before the first call. The CPU worker count, `CLIKA_RT_NUM_THREADS`,
is read once before the first compute; an app sets it in its own process with
`Os.setenv("CLIKA_RT_NUM_THREADS", "", true)` before the load. A phone with a few large
cores and several small ones runs best at the large cores' count; measure on the device. The cache
root is the app's own cache directory: `ClikaRtAndroid.load` sets `XDG_CACHE_HOME` to
`Context.getCacheDir()` unless the app set it itself, and the hub cache the model library
downloads into sits under it in the Hugging Face layout, so a snapshot downloaded on another
machine and copied there serves with no network.
Every call into the runtime runs on a thread of the app's own, never on the main thread: a model
load blocks for the load, and `generate` blocks for the reply. After the load,
`ClikaRtAndroid.trimOnMemoryPressure(context)` lets the runtime give back its cached memory when
the platform reports pressure, and `ClikaRtAndroid.bridgeLogs()` routes the runtime's log lines
to logcat.
## A model on the phone
`AutoModelForCausalLM.fromPretrained(source, LoadOptions(contextLength = 4096, cacheDir = ...))`
takes a hub repository id (the snapshot downloads inside the load), a snapshot directory, or a
`.gguf` file. Set the context length on a phone: the key-value cache is sized for the window, and
a checkpoint's own window is often tens of thousands of tokens. `chat(messages, config, listener)`
returns at once and calls the listener on the model's worker thread, `onToken` per visible piece
and `onDone` with the report; the handle's `cancel()` stops the decode at its next token.
`generate(prompt = ...)` is the blocking form. `Modelverse.registry()` lists the families the
library serves with the keys a checkpoint is matched by, so an app checks a model before it
downloads one. The whole surface, kind by kind, is the binding's README.
## A server that outlives the screen
`Modelverse.serve(model, ServeOptions(port = 8129))` hosts the library's chat API on the phone
(the OpenAI chat route, the model list and the chat page when enabled), on the server's own
threads; `start()` returns at once and `baseUrl` is the address another device on the network
points its client at. A server started
from an `adb shell` dies with the shell and with the phone's next reboot; an app keeps it alive
by hosting it in a foreground service with an ongoing notification (the connected-device service
type, which Android 14 and newer admits with a network-class permission beside it), so the
process survives the activity leaving the screen and starts again with the app.
## Where the runnable proof is
The examples archive ([Additional examples](../examples.md)) carries the Android programs. The Kotlin
chapters under its `kotlin/clika_rt/` directory are instrumented tests that run the same calls on a
connected device (`gradle connectedAndroidTest -PclikaRtMavenRepo=`): the operator surface
through `Ops`, and the device and backend facts through `Backends`. Its `android/` directory holds
three sample apps over the same artifact, one screen each: `hello` loads the runtime and prints the
version, the backends with their devices and one operator's result; `chat` loads a chat model and
streams its replies; `serve` hosts a loaded model for the other devices on the network from a
foreground service. Every one of them takes the artifact through the `clikaRtMavenRepo` property and
copies nothing into `jniLibs`.
---
# Preprocess inputs with processors
Model-specific image and audio preprocessing: build a resize/rescale/normalize pipeline or a log-mel front end from knobs, or load it from the model's own config.
Source: https://docs.clika.io/clikart/how-to/preprocess-with-processors.md
Decoding a file into pixels or samples is half the input story; the other half is what the MODEL was trained to receive: resized, rescaled, normalized pixels, or a log-mel spectrogram at a fixed rate. `processor::ImageProcessor` and `processor::AudioProcessor` carry that half. Build one from explicit knobs when you know the recipe, or load it straight from the model's own `preprocessor_config.json` so the recipe cannot drift from the checkpoint.
[Load images and audio for inference](load-images-and-audio.mdx) covers the decode step this page starts after; the processors consume the same tensors it produces.
## Build an image pipeline from knobs
`from_args` covers the common recipe with no config file: resize, center-crop, rescale, normalize. The example uses a constant image so the result is checkable by hand: every input pixel 127 becomes `(127/255 - 0.5) / 0.5`.
```cpp title="image_knobs.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
using ClikaRT::processor::ImageProcessor;
int main() {
const Tensor image = Tensor::full({8, 8, 3}, 127, DataType::UInt8);
const float mean[] = {0.5F, 0.5F, 0.5F};
const float std_[] = {0.5F, 0.5F, 0.5F};
ImageProcessor proc = ImageProcessor::from_args(
/*resize=*/{{4, 4}}, /*crop=*/{}, /*rescale=*/{}, mean, std_);
const Tensor out = proc.process(image);
std::printf("features = %s\n", out.to_string().c_str());
return 0;
}
```
```python title="image_knobs.py"
import numpy as np
import clika_runtime as crt
image = crt.tensor(np.full((8, 8, 3), 127, dtype=np.uint8))
proc = crt.processor.ImageProcessor.from_args(
resize=(4, 4), mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]
)
out = proc.process(image)
print(out.dtype, out.shape) # clika_runtime.float32 clika_runtime.Size([4, 4, 3])
print(out.numpy()[0, 0, 0]) # (127/255 - 0.5) / 0.5
```
The Kotlin binding carries the processor family (`ImageProcessor`, `AudioProcessor`, `VideoProcessor` and the `Processor` composite); the worked programs on this page are the C++ and Python ones, and [Language bindings](../bindings.md) says what Kotlin carries.
```text
features = Tensor(shape=[4, 4, 3], dtype=Float32, device=CPU, numel=48, data=[-0.003922, -0.003922, -0.003922, -0.003922, -0.003922, -0.003922, ...])
```
Knob semantics worth knowing: `rescale` defaults to 1/255; `mean` and `std` must arrive together (both present turns normalize on, both absent leaves it off); and `process` also takes a file path or an in-memory byte span, folding the decode step in when you have not done it yourself.
## Or load the model's own recipe
A checkpoint that ships `preprocessor_config.json` names its exact pipeline. `from_huggingface` reads that file (or the model directory holding it), so preprocessing follows the checkpoint instead of a hand-copied recipe. Unrecognized keys are ignored with a logged warning naming them.
```python
proc = crt.processor.ImageProcessor.from_huggingface("SmolVLM-Instruct")
features = proc.process("photo.jpg") # decode + the model's own recipe
```
## The audio front end
`AudioProcessor.from_args()` is the standard log-mel front end, 16 kHz and 80 mel bins by default; feed it a waveform and the sample rate it actually has.
```python title="audio_frontend.py"
import math
import clika_runtime as crt
# 1 s of a 440 Hz tone, composed on the runtime: t = n / rate, s = 0.5 sin(2 pi f t).
t = crt.arange(0.0, 16000.0, 1.0, dtype=crt.float32) / 16000.0
waveform = 0.5 * crt.sin(2.0 * math.pi * 440.0 * t)
audio = crt.processor.AudioProcessor.from_args()
feats = audio.process(waveform, sample_rate=16000)
print(feats.dtype, list(feats.shape)) # float32, an 80-mel-bin axis inside
```
The frame axis follows the hop length, so its extent depends on the clip; the 80-bin axis is the front end's signature. The same `from_huggingface` path exists here, reading the audio half of a checkpoint's preprocessor config.
Where this fits: [decode](load-images-and-audio.mdx) turns files into tensors; processors turn tensors into MODEL inputs; and the tensor a processor returns feeds a forward or a [compiled graph](run-an-onnx-model.mdx) directly.
---
# Structure inputs and outputs as pytrees
Flatten and rebuild nested containers of tensors, name every leaf by its key path, register your own classes and dataclasses as nodes, serialize a structure, and pass whole trees through compile, trace, eval and save.
Source: https://docs.clika.io/clikart/how-to/pytrees.md
A pytree is a nested arrangement of Python containers whose ends are the values you care about: a dict of lists of tensors, a tuple of dicts, a dataclass holding both. `clika_runtime.pytree` takes such a tree apart into a flat list of leaves plus a `TreeSpec` that remembers the shape, and puts the two halves back together. Every boundary of the Python surface that accepts several tensors at once (`compile`, `trace`, `eval`, `save` and `load`) accepts a pytree, so a model's inputs and outputs keep the structure the code was written in.
The functions carry PyTorch's spellings (`tree_flatten`, `tree_map`, `register_pytree_node`), so code written against `torch.utils._pytree` reads the same here. Two rules differ from a plain container walk and are worth reading first.
## Leaves, and why None is one of them
Lists, tuples, dicts, `OrderedDict`, `defaultdict`, `deque`, namedtuples, and every class registered with the node registry are nodes; everything else (numbers, strings, bytes, tensors, objects the registry does not know) is a leaf. Leaves come out depth first, left to right, and a flatten followed by an unflatten reproduces the same container types.
```python title="flatten.py"
from clika_runtime import pytree
tree = {"a": [1, (2, None)], "b": 3, "c": {"d": (4,), "e": []}}
leaves, spec = pytree.tree_flatten(tree)
print(leaves) # [1, 2, None, 3, 4]
print(spec.num_leaves) # 5
print(spec) # TreeSpec({'a': [*, (*, *)], 'b': *, 'c': {'d': (*,), 'e': []}})
rebuilt = pytree.tree_unflatten(spec, leaves)
assert rebuilt == tree and type(rebuilt) is dict
```
`None` is a leaf. A model that takes an optional input (a KV cache that is empty on the first step) keeps the same flat index for every other input whether or not the optional one is present, so a compiled graph's argument slots stay stable. `none_is_leaf=False` turns `None` into a childless node instead, and `is_leaf=lambda x: isinstance(x, list)` stops the descent at every list, which then counts as one leaf.
```python title="leaf_rules.py"
from clika_runtime import pytree
with_cache = {"input_ids": 1, "cache": (2, 3)}
without_cache = {"input_ids": 1, "cache": None}
leaves_a, _ = pytree.tree_flatten(with_cache)
leaves_b, spec_b = pytree.tree_flatten(without_cache)
print(leaves_a[0], leaves_b[0]) # 1 1
print(leaves_b) # [1, None]
print(pytree.tree_unflatten(spec_b, [10, None])) # {'input_ids': 10, 'cache': None}
leaves, spec = pytree.tree_flatten({"x": 1, "cache": None}, none_is_leaf=False)
print(leaves) # [1]
print(spec) # TreeSpec({'x': *, 'cache': None}, none_is_leaf=False)
```
## Map and reduce
`tree_map` applies a function to every leaf and rebuilds the same structure; extra trees pair their leaves by position and must share the first tree's structure (a mismatch raises `ValueError`). `tree_map_` runs the function for its side effect and returns the tree it was given; `tree_map_only` filters by type or by predicate. The reductions (`tree_all`, `tree_any`, `tree_reduce`, `tree_sum`, `tree_max`, `tree_min`) read the leaves without rebuilding anything, `tree_leaves` returns the flat list and `tree_iter` yields it lazily.
```python title="map.py"
from clika_runtime import pytree
print(pytree.tree_map(lambda x: x is None, {"x": 1, "y": None})) # {'x': False, 'y': True}
print(pytree.tree_map(lambda x, y: x + y, {"a": 1, "b": 2}, {"a": 10, "b": 20})) # {'a': 11, 'b': 22}
mixed = {"a": 1, "b": "text", "c": [2, None, 2.5]}
print(pytree.tree_map_only(int, lambda x: x + 1, mixed)) # {'a': 2, 'b': 'text', 'c': [3, None, 2.5]}
```
## Name every leaf by its key path
`tree_flatten_with_path` returns `(path, leaf)` pairs; a path is a tuple of key entries (`MappingKey`, `SequenceKey`, `GetAttrKey`, `DataclassKey`), `keystr` renders it as the indexing expression that reaches the leaf, and `key_get` follows it. Key paths are how the file format names the leaves of a saved tree (below) and how a diagnostic points at one tensor inside a large state.
```python title="paths.py"
from typing import NamedTuple
from clika_runtime import pytree
class Point(NamedTuple):
x: object
y: object
tree = {"a": [1, 2], "p": Point(3, 4)}
pairs = pytree.tree_flatten_with_path(tree)[0]
print([leaf for _, leaf in pairs]) # [1, 2, 3, 4]
print([pytree.keystr(path) for path, _ in pairs]) # ["['a'][0]", "['a'][1]", "['p'].x", "['p'].y"]
path = (pytree.MappingKey("a"), pytree.SequenceKey(1))
print(pytree.key_get(tree, path)) # 2
```
## Register your own containers
A class the registry does not know is a leaf. `register_pytree_node(cls, flatten_fn, unflatten_fn)` opens it: `flatten_fn` returns `(children, context)` or `(children, context, entries)` (the entries name the children for key paths), and `unflatten_fn(children, context)` rebuilds an instance. The argument order is PyTorch's. After registration, every tree function descends into the class and every rebuild produces an instance of it.
```python title="register.py"
from clika_runtime import pytree
class Point:
def __init__(self, x, y):
self.x, self.y = x, y
pytree.register_pytree_node(
Point,
lambda p: ((p.x, p.y), None, ("x", "y")),
lambda children, _ctx: Point(*children),
)
print(pytree.tree_leaves({"p": Point(1, 2), "n": 3})) # [1, 2, 3]
mapped = pytree.tree_map(lambda v: v * 10, Point(1, 2))
print(type(mapped).__name__, mapped.x, mapped.y) # Point 10 20
spec = pytree.tree_structure(Point(1, 2))
print(spec.entries(), spec.paths()) # ['x', 'y'] [('x',), ('y',)]
```
A class can carry its own flattening as a method pair and register with the decorator. Note the method order: `__tree_unflatten__(cls, context, children)` takes the context first, while the function form above takes `(children, context)`.
```python title="register_class.py"
from clika_runtime import pytree
@pytree.register_pytree_node_class
class Pair:
def __init__(self, a, b):
self.a, self.b = a, b
def __tree_flatten__(self):
return (self.a, self.b), "pair", ("a", "b")
@classmethod
def __tree_unflatten__(cls, context, children):
return cls(*children)
leaves, spec = pytree.tree_flatten(Pair(1, [2, 3]))
print(leaves, spec.context) # [1, 2, 3] pair
rebuilt = pytree.tree_unflatten(spec, [10, 20, 30])
print(rebuilt.a, rebuilt.b) # 10 [20, 30]
```
Dataclasses register by field: the data fields become children, the meta fields ride the context and come back unchanged.
```python title="register_dataclass.py"
import dataclasses
from clika_runtime import pytree
@dataclasses.dataclass
class Sample:
weight: object
bias: object
name: str
pytree.register_dataclass(Sample, data_fields=["weight", "bias"], meta_fields=["name"])
leaves, spec = pytree.tree_flatten(Sample(1, 2, "s"))
print(leaves) # [1, 2]
print(pytree.tree_unflatten(spec, [10, 20])) # Sample(weight=10, bias=20, name='s')
print([pytree.keystr(p) for p, _ in pytree.tree_flatten_with_path(Sample(1, 2, "s"))[0]]) # ['.weight', '.bias']
```
Registering a built-in container, an instance instead of a class, or the same class twice raises; `unregister_pytree_node(cls)` makes instances leaves again. A registration meant for one library goes into a namespace so it never changes what other code sees: `register_pytree_node(..., namespace="mylib")` opens the class only for calls that pass `namespace="mylib"` (`tree_leaves(tree, namespace="mylib")`), `is_registered(cls, namespace="mylib")` reports it, and a namespace with no registration of its own falls back to the global one.
## Dictionary order
Dicts flatten in insertion order, and the order is part of the structure: `{"b": 1, "a": 2}` and `{"a": 2, "b": 1}` have different specs. Code that wants two dicts with the same keys to share one structure regardless of insertion order wraps the calls in `dict_insertion_ordered(False)`, which flattens by sorted key inside the block; the rebuilt dict still comes back in the original order.
```python title="dict_order.py"
from clika_runtime import pytree
tree = {"b": 1, "a": 2, "c": 3}
leaves, spec = pytree.tree_flatten(tree)
print(leaves) # [1, 2, 3]
print(list(pytree.tree_unflatten(spec, leaves))) # ['b', 'a', 'c']
with pytree.dict_insertion_ordered(False):
print(pytree.tree_leaves({"b": 1, "a": 2})) # [2, 1]
print(pytree.tree_structure({"b": 1, "a": 2}) == pytree.tree_structure({"a": 0, "b": 0})) # True
```
## Treespecs
A `TreeSpec` is a value: it compares and hashes by structure, prints as the container shape with a `*` per leaf, and rebuilds a tree from any sequence of the right length (`spec.unflatten(leaves)`; a wrong count raises `ValueError` naming the expected number). `is_prefix` asks whether one structure is the top of another, `compose` grafts an inner structure onto every leaf of an outer one, and `flatten_up_to` flattens a tree only as deep as the spec goes.
```python title="treespec.py"
from clika_runtime import pytree
spec = pytree.tree_structure({"a": [1, (2, None)], "b": 3})
print(spec) # TreeSpec({'a': [*, (*, *)], 'b': *})
print(spec.num_leaves, spec.num_nodes, spec.num_children) # 4 7 2
print(spec.paths()) # [('a', 0), ('a', 1, 0), ('a', 1, 1), ('b',)]
short = pytree.tree_structure([1, 2])
deep = pytree.tree_structure([1, (2, 3)])
print(short.is_prefix(deep), deep.is_prefix(short)) # True False
outer = pytree.tree_structure([1, 2])
inner = pytree.tree_structure((1, 2))
print(outer.compose(inner)) # TreeSpec([(*, *), (*, *)])
```
`treespec_dumps` writes a spec as JSON and `treespec_loads` reads it back, so a structure can travel beside a file or a request. Built-in containers and namedtuples serialize as they are; a registered class needs a `serialized_type_name`, and a context that is not JSON needs the `to_dumpable_context` / `from_dumpable_context` pair at registration. An unnamed custom node or an unknown type name in the document raises `ValueError`.
```python title="serialize.py"
import json
from clika_runtime import pytree
spec = pytree.tree_structure({"a": [1, (2, None)], "b": 3})
text = pytree.treespec_dumps(spec)
print(json.loads(text)["version"] == pytree.SERIALIZATION_PROTOCOL) # True
loaded = pytree.treespec_loads(text)
print(loaded == spec) # True
print(loaded.unflatten([1, 2, None, 3])) # {'a': [1, (2, None)], 'b': 3}
```
## Pytrees at the boundaries
`crt.compile` takes a callable whose arguments and return value are pytrees of tensors: it flattens the call's inputs, treats the tensor leaves as graph inputs and every other leaf as a static value that is part of the guard, and rebuilds the output structure on the way out; `dynamic=True` keeps one graph across batch sizes.
```python title="compile_tree.py"
import numpy as np
import clika_runtime as crt
import clika_runtime.nn as nn
BATCH, D_IN, D_OUT = 5, 131, 64
rng = np.random.default_rng(0)
w = rng.standard_normal((D_IN, D_OUT)).astype(np.float32) * 0.1
b = rng.standard_normal(D_OUT).astype(np.float32)
x = rng.standard_normal((BATCH, D_IN)).astype(np.float32)
class Head(nn.Module):
def __init__(self, w: np.ndarray, b: np.ndarray) -> None:
super().__init__()
self.weight = nn.Parameter(crt.tensor(w))
self.bias = nn.Parameter(crt.tensor(b))
def forward(self, x: crt.Tensor) -> crt.Tensor:
return crt.softmax(crt.relu(crt.add(crt.matmul(x, self.weight), self.bias)), dim=-1)
class Wrapped(nn.Module):
def __init__(self) -> None:
super().__init__()
self.head = Head(w, b)
def forward(self, batch: dict[str, crt.Tensor]) -> dict[str, crt.Tensor]:
p = self.head(batch["x"])
return {"probs": p, "argmax": crt.argmax(p, dims=[1])}
compiled = crt.compile(Wrapped(), dynamic=True)
out = compiled({"x": crt.tensor(x)})
print(sorted(out)) # ['argmax', 'probs']
again = compiled({"x": crt.tensor(np.tile(x, (2, 1)))})
print(tuple(again["probs"].shape)) # (10, 64)
```
`crt.trace` records a function once over stand-ins and returns a graph that runs by position or by name; the traced function takes its tensor inputs as one list and returns a list, and [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) walks through it.
Operators return before their work runs. `crt.eval(*trees)` flattens whatever trees it is given, skips the leaves that are not tensors, and returns once every tensor leaf holds its value; a later read then costs no wait.
```python title="eval_tree.py"
import numpy as np
import clika_runtime as crt
x = np.random.default_rng(0).standard_normal((5, 131)).astype(np.float32)
t = crt.tensor(x)
tree = {"a": crt.exp(t), "b": [crt.sum(t), None, "text"], "c": (crt.relu(t),)}
crt.eval(tree, crt.abs(t))
print(np.allclose(tree["a"].numpy(), np.exp(x.astype(np.float64)), rtol=1e-5, atol=1e-6)) # True
```
`crt.save` writes a tensor, a state dict, or any pytree of tensors as a safetensors file, and `crt.load` reads it back. A state dict (`crt.save(state, path)`, `crt.load(path, device="cpu")`) saves under its own keys in insertion order, so another safetensors reader sees it as it is; a deeper tree saves one entry per leaf named by `keystr` of its key path, with the structure carried in the file's metadata, and `crt.load` rebuilds the same tree.
```python title="save_tree.py"
from pathlib import Path
import numpy as np
import clika_runtime as crt
from clika_runtime import pytree
rng = np.random.default_rng(0)
tree = {
"layers": [
{"w": crt.tensor(rng.standard_normal((32, 131)).astype(np.float32)), "b": crt.tensor(np.zeros(32, np.float32))}
for _ in range(2)
],
"step": crt.tensor(np.array([7], dtype=np.int64)),
}
crt.save(tree, Path("tree.safetensors"))
back = crt.load(Path("tree.safetensors"))
print(pytree.tree_structure(back) == pytree.tree_structure(tree)) # True
print(sorted(pytree.keystr(kp) for kp, _ in pytree.tree_flatten_with_path(tree)[0]))
# ["['layers'][0]['b']", "['layers'][0]['w']", "['layers'][1]['b']", "['layers'][1]['w']", "['step']"]
```
The same structure rule holds for a module's state: `crt.save(module.state_dict(), path)` followed by `other.load_state_dict(crt.load(path))` reproduces the forward, and the file is plain safetensors any other tool reads. [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module surface the leaves belong to.
---
# Query a graph
Read a ModelGraph as data: its nodes, values and edges, search by name or predicate, walks and topological orders, paths and the critical path, dominance, regions, cones and the cheapest cut, views across edits, and networkx.
Source: https://docs.clika.io/clikart/how-to/query-a-graph.md
{/* Every block is a program under examples//howto/query_a_graph/: the first
block of each tab is get_a_graph whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output. */}
A `ModelGraph`, traced from your own code or compiled from a model file, answers questions about its own structure: which operators it holds, what feeds what, which orders it can run in, which paths join two nodes, and where it can be split. The answers are views (`Node`, `Value` and `Edge`) that you read, compare and use as keys, and every query reads the graph as it stands. The graph queries are available from C++ and Python.
## Get a graph
A trace records a function once over stand-in tensors and returns the graph as built: every operator stays as written, nothing optimized or finalized, which makes the graph easy to read. The model below has two branches that meet, one tensor read twice and two outputs. Every example on this page uses it, and `label` names a node by its input name or its operator.
```cpp title="get_a_graph.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::Node;
using ClikaRT::graph::NodeKind;
using ClikaRT::graph::OpCode;
namespace ops = ClikaRT::ops;
namespace {
// Two branches that meet, one tensor read twice, two outputs.
std::vector model(const std::vector& inputs) {
const Tensor a = ops::relu(inputs[0]);
const Tensor b = ops::sigmoid(inputs[0]);
const Tensor c = ops::add(a, b);
const Tensor d = ops::mul(c, c); // c feeds d twice: two edges, on input ports 0 and 1
const Tensor e = ops::tanh(a);
return {ops::sub(d, e), b};
}
// A node's label: an input's name, an operator's code.
std::string label(const Node& node) {
return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code()));
}
// The labels of `nodes`, separated by spaces.
std::string labels(const std::vector& nodes) {
std::string out;
for (const Node& node : nodes) out += (out.empty() ? "" : " ") + label(node);
return out;
}
// The trace returns the graph as recorded: every operator stays as written,
// nothing optimized or finalized.
ModelGraph trace_model() {
const std::vector signature = {{"x", DataType::Float32, {2, 3}}};
return ClikaRT::graph::trace(model, signature, "query");
}
} // namespace
int main() {
const ModelGraph graph = trace_model();
std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu Sigmoid Add Mul Tanh Sub
std::printf("%zu\n", graph.find_nodes(OpCode::Mul).front().in_degree()); // 2
return 0;
}
```
```python title="get_a_graph.py"
import clika_runtime as crt
from clika_runtime.graph import NodeKind, OpCode
def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
x = inputs[0]
a, b = crt.relu(x), crt.sigmoid(x)
c = a + b
d = c * c # c feeds d twice: two edges, on input ports 0 and 1
e = crt.tanh(a)
return [d - e, b]
def label(node: crt.graph.Node) -> str:
return node.name if node.kind == NodeKind.Input else node.op_code.name
# The trace returns the graph as recorded: every operator stays as written,
# nothing optimized or finalized.
graph = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])]).graph
print([label(node) for node in graph.nodes()]) # ['x', 'Relu', 'Sigmoid', 'Add', 'Mul', 'Tanh', 'Sub']
print(graph.find_nodes(OpCode.Mul)[0].in_degree()) # 2
```
A compiled model answers the same queries: `compile()` returns the same `ModelGraph` type. The program builds a one-operator ONNX file first; any `.onnx` file works in its place.
```cpp title="compile_a_model.cpp"
int main() {
const ModelGraph graph = OnnxModel::open(relu_model_path()).compile();
std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu
return 0;
}
```
```python title="compile_a_model.py"
graph = crt.io.OnnxModel.open(path).compile()
print([label(node) for node in graph.nodes()]) # ['x', 'Relu']
```
Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, `label` and the trace).
## Nodes, values and edges
A `Node` is an operator or a graph input. A `Value` is a tensor that a node produces on one output port, and an `Edge` is one read of a value, from the producer's output port to the reader's input port. The graph is a multigraph keyed by those ports: `d = c * c` reads one value over two edges. A node answers its inputs and outputs, `input(port)`, `output(port)`, `in_edges()` and `out_edges()`; a value answers `producer()`, `consumers()` and `uses()`.
```cpp title="nodes_and_edges.cpp"
int main() {
const ModelGraph graph = trace_model();
const Node mul = graph.find_nodes(OpCode::Mul).front();
const std::vector reads = mul.in_edges(); // one edge per input port, each keyed by its two ports
for (const Edge& edge : reads) {
std::printf("%s %d -> %s %d\n", label(edge.src()).c_str(), edge.src_port(), label(edge.dst()).c_str(),
edge.dst_port());
}
// Add 0 -> Mul 0
// Add 0 -> Mul 1
const Value value = *mul.input(0); // the tensor both edges carry
const Node add = *value.producer();
std::printf("%s %s %s\n", value == *mul.input(1) ? "true" : "false", label(add).c_str(),
labels(value.consumers()).c_str()); // true Add Mul
std::printf("%zu %zu\n", graph.edges().size(), graph.edges_between(add, mul).size()); // 9 2
return 0;
}
```
```python title="nodes_and_edges.py"
(mul,) = graph.find_nodes(OpCode.Mul)
for edge in mul.in_edges(): # one edge per input port, each keyed by its two ports
print(label(edge.src()), edge.src_port(), "->", label(edge.dst()), edge.dst_port())
# Add 0 -> Mul 0
# Add 0 -> Mul 1
value = mul.input(0) # the tensor both edges carry
print(value == mul.input(1), label(value.producer()), [label(node) for node in value.consumers()]) # True Add ['Mul']
print(len(graph.edges()), len(graph.edges_between(value.producer(), mul))) # 9 2
```
Read the views when your own code walks the structure, as an exporter or a quantizer does.
## Search by name and by predicate
`find_nodes_by_name` matches a regular expression against a node's whole name, or any part of it with the `Anywhere` name match. `find_nodes` takes an `OpCode` or a predicate over a `Node`, and `node(name)` looks up one node. Each list comes back in name order. Operator names carry the operator and a number, so match them with a pattern rather than a literal.
```cpp title="search.cpp"
int main() {
const ModelGraph graph = trace_model();
const Regex relu = Regex::compile("Relu_[0-9]+");
const Regex sig = Regex::compile("Sig");
std::printf("%s\n", labels(graph.find_nodes_by_name(relu)).c_str()); // Relu
std::printf("%s\n", labels(graph.find_nodes_by_name(sig, NameMatch::Anywhere)).c_str()); // Sigmoid
std::printf("%s\n", labels(graph.find_nodes([](const Node& node) { return node.in_degree() == 2; })).c_str());
// Add Mul Sub
std::printf("%s %s\n", label(*graph.node("x")).c_str(), labels(graph.find_nodes(OpCode::Tanh)).c_str()); // x Tanh
return 0;
}
```
```python title="search.py"
print([label(node) for node in graph.find_nodes_by_name("Relu_[0-9]+")]) # ['Relu']
print([label(node) for node in graph.find_nodes_by_name("Sig", crt.graph.NameMatch.Anywhere)]) # ['Sigmoid']
print([label(node) for node in graph.find_nodes(lambda node: node.in_degree() == 2)]) # ['Add', 'Mul', 'Sub']
print(label(graph.node("x")), [label(node) for node in graph.find_nodes(OpCode.Tanh)]) # x ['Tanh']
```
## Walks, visits, ancestors and descendants
`walk(start)` lists the nodes a breadth-first walk reaches, the start first. The options `direction`, `max_depth`, `max_nodes`, `stop_at` and `edge_filter` (the fields of `WalkOptions` in C++, keyword arguments in Python) bound any walk. `descendants(node)` and `ancestors(node)` leave the node itself out and list the nearest first. `visit(start, visitor)` shows a callable each node, and the callable answers the `VisitAction` `Continue`, `Prune` (do not go past this node) or `Stop`.
```cpp title="walks.cpp"
int main() {
const ModelGraph graph = trace_model();
const Node x = *graph.node("x");
const Node relu = graph.find_nodes(OpCode::Relu).front();
const Node sub = graph.find_nodes(OpCode::Sub).front();
WalkOptions one_step;
one_step.max_depth = 1;
std::printf("%s\n", labels(graph.walk(x, one_step)).c_str()); // x Relu Sigmoid
std::printf("%s\n", labels(graph.descendants(relu)).c_str()); // Add Tanh Mul Sub
std::printf("%s\n", labels(graph.ancestors(sub)).c_str()); // Mul Tanh Add Relu Sigmoid x
std::vector seen;
graph.visit(x, [&seen](const Node& node) {
seen.push_back(node);
return node.op_code() == OpCode::Add ? VisitAction::Prune : VisitAction::Continue;
}); // the walk does not go past the Add, so the Mul is never shown
std::printf("%s\n", labels(seen).c_str()); // x Relu Sigmoid Add Tanh Sub
return 0;
}
```
```python title="walks.py"
x = graph.node("x")
(relu,) = graph.find_nodes(OpCode.Relu)
(sub,) = graph.find_nodes(OpCode.Sub)
print([label(node) for node in graph.walk(x, max_depth=1)]) # ['x', 'Relu', 'Sigmoid']
print([label(node) for node in graph.descendants(relu)]) # ['Add', 'Tanh', 'Mul', 'Sub']
print([label(node) for node in graph.ancestors(sub)]) # ['Mul', 'Tanh', 'Add', 'Relu', 'Sigmoid', 'x']
seen: list[str] = []
def look(node: crt.graph.Node) -> crt.graph.VisitAction:
seen.append(label(node))
return crt.graph.VisitAction.Prune if node.op_code == OpCode.Add else crt.graph.VisitAction.Continue
graph.visit(x, look) # the walk does not go past the Add, so the Mul is never shown
print(seen) # ['x', 'Relu', 'Sigmoid', 'Add', 'Tanh', 'Sub']
```
Use a walk to collect everything upstream or downstream of a node, and a visit when the decision to go on depends on the node you are at.
## Topological orders, generations and depths
`topological_order()` lists every node after the nodes it reads, in the `Deterministic` order by default (`Structural` and `MinMemory` are the other two `TopologicalOrder` values). `topological_generations()` groups the nodes by the longest path that ends at them, and `depths()` gives that length per node. `random_topological_order(seed)` draws one of the valid orders; the programs check the order they draw instead of printing it, since the order depends on the seed.
```cpp title="orders.cpp"
int main() {
const ModelGraph graph = trace_model();
std::printf("%s\n", labels(graph.topological_order()).c_str()); // x Relu Sigmoid Add Mul Tanh Sub
std::string layers;
for (const std::vector& layer : graph.topological_generations()) layers += "[" + labels(layer) + "]";
std::printf("%s\n", layers.c_str()); // [x][Relu Sigmoid][Add Tanh][Mul][Sub]
std::string depths;
for (const auto& [node, depth] : graph.depths()) {
depths += (depths.empty() ? "" : " ") + label(node) + "=" + std::to_string(depth);
}
std::printf("%s\n", depths.c_str()); // x=0 Relu=1 Sigmoid=1 Add=2 Mul=3 Tanh=2 Sub=4
// Any valid order, drawn from the seed: check it rather than print it.
std::printf("%s\n", graph.is_topological_order(graph.random_topological_order(7)) ? "true" : "false"); // true
return 0;
}
```
```python title="orders.py"
print([label(node) for node in graph.topological_order()])
# ['x', 'Relu', 'Sigmoid', 'Add', 'Mul', 'Tanh', 'Sub']
print([[label(node) for node in layer] for layer in graph.topological_generations()])
# [['x'], ['Relu', 'Sigmoid'], ['Add', 'Tanh'], ['Mul'], ['Sub']]
print({label(node): depth for node, depth in graph.depths().items()})
# {'x': 0, 'Relu': 1, 'Sigmoid': 1, 'Add': 2, 'Mul': 3, 'Tanh': 2, 'Sub': 4}
print(graph.is_topological_order(graph.random_topological_order(7))) # True: any valid order, drawn from the seed
```
A generation's nodes read only earlier generations, so they are the nodes that can run at the same time.
## Paths and path counts
A path never repeats a node. `all_simple_paths(src, dst)` lists the paths as nodes, where parallel edges give one path, and `all_simple_edge_paths` lists them as edges, where each parallel edge gives its own. Both search at the call, and Python hands the paths back through an iterator. Without a `limit`, a search that finds more than `kMaxPathsWithoutLimit` paths (`MAX_PATHS_WITHOUT_LIMIT` in Python) fails with an invalid-argument error instead of returning part of the list; `count_paths` counts without listing and has no such cap. `shortest_path` gives the path with the fewest edges and `has_path` answers reachability. The options `cutoff`, `avoid` and `edge_filter` (the fields of `PathOptions` in C++, keyword arguments in Python) bound every path query.
```cpp title="paths.cpp"
int main() {
const ModelGraph graph = trace_model();
const Node x = *graph.node("x");
const Node sub = graph.find_nodes(OpCode::Sub).front();
// The two edges into the Mul give one path.
const std::vector> paths = graph.all_simple_paths(x, sub);
for (const std::vector& path : paths) std::printf("%s\n", labels(path).c_str());
// x Relu Add Mul Sub
// x Relu Tanh Sub
// x Sigmoid Add Mul Sub
std::printf("%llu %zu\n", static_cast(graph.count_paths(x, sub)),
graph.all_simple_edge_paths(x, sub).size()); // 3 5
std::printf("%s %s\n", labels(graph.shortest_path(x, sub)).c_str(),
graph.has_path(sub, x) ? "true" : "false"); // x Relu Tanh Sub false
PathOptions first;
first.limit = 1;
std::printf("%zu %zu\n", graph.all_simple_paths(x, sub, first).size(),
ClikaRT::graph::kMaxPathsWithoutLimit); // 1 10000
return 0;
}
```
```python title="paths.py"
x = graph.node("x")
(sub,) = graph.find_nodes(OpCode.Sub)
for path in graph.all_simple_paths(x, sub): # the two edges into the Mul give one path
print([label(node) for node in path])
# ['x', 'Relu', 'Add', 'Mul', 'Sub']
# ['x', 'Relu', 'Tanh', 'Sub']
# ['x', 'Sigmoid', 'Add', 'Mul', 'Sub']
print(graph.count_paths(x, sub), len(list(graph.all_simple_edge_paths(x, sub)))) # 3 5
print([label(node) for node in graph.shortest_path(x, sub)], graph.has_path(sub, x)) # ['x', 'Relu', 'Tanh', 'Sub'] False
print(len(list(graph.all_simple_paths(x, sub, limit=1))), crt.graph.MAX_PATHS_WITHOUT_LIMIT) # 1 10000
```
## The critical path
`critical_path(node_cost, edge_cost)` returns the costliest path as a `WeightedPath` carrying its nodes, its edges and its summed cost: the longest chain of work when each node costs its run time. Given an `edge_cost`, `shortest_path` returns the cheapest path in the same form.
```cpp title="critical_path.cpp"
int main() {
const ModelGraph graph = trace_model();
// A cost per node (and optionally per edge): here the Sigmoid costs 2, every other node 1.
const WeightedPath path =
graph.critical_path([](const Node& node) { return node.op_code() == OpCode::Sigmoid ? 2.0 : 1.0; });
std::printf("%s %g\n", labels(path.nodes).c_str(), path.cost); // x Sigmoid Add Mul Sub 6
return 0;
}
```
```python title="critical_path.py"
# A cost per node (and optionally per edge): here the Sigmoid costs 2, every other node 1.
path = graph.critical_path(lambda node: 2.0 if node.op_code == OpCode.Sigmoid else 1.0)
print([label(node) for node in path.nodes], path.cost) # ['x', 'Sigmoid', 'Add', 'Mul', 'Sub'] 6.0
```
## Structure: cycles, dominance, regions and cones
`find_cycle()` is empty for every traced or compiled graph, and `would_create_cycle(src, dst)` checks an edge before an edit adds it. `weakly_connected_components()` groups the nodes that edges join in either direction. `dominators()` maps each node to the last node that every path from the inputs to it crosses, `split_points()` lists the nodes that every path from the inputs to the outputs crosses, and `merge_point(nodes)` finds where the paths from several nodes meet. `region(entry, exit)` is the single-entry, single-exit block between two nodes, `producer_cone` and `consumer_cone` are everything a set of nodes reads or feeds, and `is_convex` tells whether a set can be cut out and replaced as one piece.
```cpp title="structure.cpp"
int main() {
const ModelGraph graph = trace_model();
std::map node;
for (const Node& n : graph.nodes()) node.emplace(label(n), n);
std::printf("%zu %s\n", graph.find_cycle().size(),
graph.would_create_cycle(node.at("Sub"), node.at("x")) ? "true" : "false"); // 0 true
std::printf("%zu\n", graph.weakly_connected_components().size()); // 1
std::string dominators;
for (const auto& [n, dominator] : graph.dominators()) {
dominators += (dominators.empty() ? "" : " ") + label(n) + ":" + label(dominator);
}
std::printf("%s\n", dominators.c_str()); // x:x Relu:x Sigmoid:x Add:x Mul:Add Tanh:Relu Sub:x
std::printf("%s %s\n", labels(graph.split_points()).c_str(),
label(*graph.merge_point({node.at("Add"), node.at("Tanh")})).c_str()); // x Sub
std::printf("%s\n", labels(graph.region(node.at("Add"), node.at("Mul"))).c_str()); // Add Mul
std::printf("%s\n", labels(graph.producer_cone({node.at("Mul")})).c_str()); // x Relu Sigmoid Add
std::printf("%s %s\n", graph.is_convex({node.at("Add"), node.at("Mul")}) ? "true" : "false",
graph.is_convex({node.at("Relu"), node.at("Mul")}) ? "true" : "false"); // true false
return 0;
}
```
```python title="structure.py"
node = {label(n): n for n in graph.nodes()}
print(graph.find_cycle(), graph.would_create_cycle(node["Sub"], node["x"])) # [] True
print(len(graph.weakly_connected_components())) # 1
print({label(n): label(dominator) for n, dominator in graph.dominators().items()})
# {'x': 'x', 'Relu': 'x', 'Sigmoid': 'x', 'Add': 'x', 'Mul': 'Add', 'Tanh': 'Relu', 'Sub': 'x'}
print([label(n) for n in graph.split_points()], label(graph.merge_point([node["Add"], node["Tanh"]]))) # ['x'] Sub
print([label(n) for n in graph.region(node["Add"], node["Mul"])]) # ['Add', 'Mul']
print([label(n) for n in graph.producer_cone([node["Mul"]])]) # ['x', 'Relu', 'Sigmoid', 'Add']
print(graph.is_convex([node["Add"], node["Mul"]]), graph.is_convex([node["Relu"], node["Mul"]])) # True False
```
## The cheapest cut
`min_value_cut(before, after)` finds the split of the nodes into two parts that hands the fewest bytes from the first part to the second, and returns a `ValueCut` of the values that cross and their bytes. `before` and `after` pin nodes to either part. Of the cheapest cuts it returns the one nearest the inputs; networkx's `minimum_cut` returns the one nearest the sink instead. Use it to decide where to split a model across two devices or two processes.
```cpp title="value_cut.cpp"
int main() {
const ModelGraph graph = trace_model();
std::map node;
for (const Node& n : graph.nodes()) node.emplace(label(n), n);
// The producers of the values that cross the cut, and their bytes.
const auto show = [](const ValueCut& cut) {
std::string crossing;
for (const Value& value : cut.values) crossing += (crossing.empty() ? "" : " ") + label(*value.producer());
std::printf("%s %llu\n", crossing.c_str(), static_cast(cut.bytes));
};
show(graph.min_value_cut()); // x 24: every float32 [2, 3] value holds 24 bytes
show(graph.min_value_cut({node.at("Add")}, {node.at("Sub")})); // Relu Sigmoid Add 72
return 0;
}
```
```python title="value_cut.py"
node = {label(n): n for n in graph.nodes()}
cut = graph.min_value_cut() # every float32 [2, 3] value holds 24 bytes
print([label(value.producer()) for value in cut.values], cut.bytes) # ['x'] 24
pinned = graph.min_value_cut(before=[node["Add"]], after=[node["Sub"]])
print([label(value.producer()) for value in pinned.values], pinned.bytes) # ['Relu', 'Sigmoid', 'Add'] 72
```
## Views across edits
An edit such as `optimize()` changes the graph under the views you hold. A view finds its node, value or edge again by its key (a node's name, a value's producer and port, an edge's two ends and ports) and keeps answering. Once the key is gone, `check()` refuses, naming the key, and so does a query given the view. In C++ the view's other reads answer empty; in Python they raise the same `InvalidArgumentError` as `check()`. Equality and hashing keep working in both, so a map or dictionary keyed by views still finds its entries.
```cpp title="edits.cpp"
int main() {
const std::vector signature = {{"x", DataType::Float32, {2, 3}}};
ModelGraph twice = ClikaRT::graph::trace(
[](const std::vector& inputs) -> std::vector { return {ops::relu(ops::relu(inputs[0]))}; },
signature, "twice");
const Node inner = twice.node("x")->successors().front();
const Node outer = inner.successors().front();
// Views key a map across edits.
const std::unordered_map index = {{inner, "inner"}, {outer, "outer"}};
twice.optimize(); // Relu(Relu(x)) is Relu(x): the optimizer removes the outer Relu
inner.check(); // the kept Relu still answers
std::printf("%s %s\n", label(inner).c_str(), inner.output(0)->is_graph_output() ? "true" : "false"); // Relu true
try {
outer.check(); // refuses, naming the node, and so does a query given the view
} catch (const Error& error) {
std::printf("%s\n", error.code_name().c_str()); // INVALID_ARGUMENT
}
// Every other read of the gone view answers empty.
std::printf("%s '%s'\n", index.at(outer).c_str(), outer.name().c_str()); // outer ''
return 0;
}
```
```python title="edits.py"
twice = crt.trace(lambda inputs: [crt.relu(crt.relu(inputs[0]))],
[crt.TensorSpec("x", crt.float32, [2, 3])]).graph
(inner,) = twice.node("x").successors()
(outer,) = inner.successors()
index = {inner: "inner", outer: "outer"} # views key a dict across edits
twice.optimize() # Relu(Relu(x)) is Relu(x): the optimizer removes the outer Relu
print(inner.check(), label(inner), inner.output(0).is_graph_output()) # None Relu True
try:
outer.check() # so does every other read of the view, and any query given it
except crt.InvalidArgumentError:
print("the outer Relu is gone") # the outer Relu is gone
print(index[outer], repr(outer).startswith("Node(gone: ")) # outer True
```
## Hand the graph to networkx
`to_networkx()` returns the graph as a `networkx.MultiDiGraph`, for the algorithms networkx provides. Its nodes are the `Node` views, which keep the graph alive, and each edge is keyed `(src_port, dst_port)`. It needs the `networkx` package, which the program imports as `nx`.
networkx is a Python library, so this section has a Python program only. From C++, `edges()` and the node views carry the same structure to any graph library.
```python title="to_networkx.py"
nx_graph = graph.to_networkx() # a networkx.MultiDiGraph whose nodes are the Node views
print(nx_graph.number_of_nodes(), nx_graph.number_of_edges()) # 7 9
print(nx.dag_longest_path_length(nx_graph)) # 4
print(sorted(label(node) for node in nx_graph.successors(graph.node("x")))) # ['Relu', 'Sigmoid']
```
---
# Run an ONNX model
Open an ONNX file (or build one from scratch), compile it into a ModelGraph, optimize and finalize it, and execute it by position or by name.
Source: https://docs.clika.io/clikart/how-to/run-an-onnx-model.md
{/* CERTIFICATION: the python and C++ arms are verified against the pinned
release (tools/tutorial_check.py over the programs the page embeds:
build_onnx, inspect_onnx and run_onnx all match their recordings in both
languages, on the release's cp313 wheel and its payload). Given no path,
the C++ and Python IO-contract programs build the tiny MLP into a
temporary file, so they run with no file on disk. */}
You have a model as an `.onnx` file and want ClikaRT to run it. Three types carry the whole story. `io::OnnxModel` is the format-level object, the ONNX graph as data: open it, inspect it, edit it, save it. `compile()` is the one crossing into the executable world, where every node parses into its runtime operator, weights bind, and shapes resolve. The result is a `ModelGraph` as built. `optimize()` runs the graph optimizer when you want it, `finalize()` places the graph and packs its weights, and the finalized graph has one `run`.
The programs below first open an existing file and read its contract, then build a tiny model from scratch so the page runs with no file on disk, then compile and run it. Any `.onnx` file works in the first section; nothing depends on the architecture.
## Open a model and read its IO contract
Opening is cheap and does not compile anything: you get the graph as data. The IO contract (names, dtypes, dims, with dynamic dims reading as named placeholders or -1) is what you need to prepare feeds; print it before anything else when a model is new to you.
```cpp title="inspect_onnx.cpp (the inspection)"
int main(int argc, char** argv) {
const std::string path = argc > 1 ? argv[1] : tiny_model_path();
OnnxModel model = OnnxModel::open(path);
std::printf("nodes: %zu initializers: %zu opset: %lld\n",
model.num_nodes(), model.num_initializers(),
static_cast(model.opset()));
// Hold the returned vectors: a range-for over a temporary would iterate
// storage that is gone before the first step.
const std::vector inputs = model.inputs();
const std::vector outputs = model.outputs();
for (const auto& s : inputs)
std::printf(" input %s\n", s.name.c_str());
for (const auto& s : outputs)
std::printf(" output %s\n", s.name.c_str());
return 0;
}
```
```python title="inspect_onnx.py (the inspection)"
model = crt.io.OnnxModel.open(sys.argv[1] if len(sys.argv) > 1 else tiny_model_path())
print(f"nodes: {model.num_nodes} initializers: {model.num_initializers} "
f"opset: {model.opset}")
for name, dtype, dims in model.inputs():
print(f" input {name} {list(dims)}") # a dynamic dim reads as -1
for name, dtype, dims in model.outputs():
print(f" output {name} {list(dims)}")
```
The program opens the path on its command line; with none, `tiny_model_path()` builds the MLP of the next section into a temporary file, so the page runs with no file on disk.
## Build a graph from scratch
When there is no file yet (a test, a fixture, a tool that emits ONNX), the same object builds a graph node by node: declare inputs, add initializers (weights enter as ordinary tensors), add nodes by operator type, name the outputs. The tiny MLP here is `y = relu(X W + B)` with a dynamic batch dimension.
```cpp title="build_onnx.cpp (excerpt of the build stage)"
OnnxModel build_tiny_mlp() {
Dim batch; // one Dim object: every use is the SAME dynamic dimension
OnnxModel model = OnnxModel::create("tiny_mlp", /*opset_version=*/21);
model.add_input("X", DataType::Float32, {batch, 4});
const std::string w = model.add_initializer(
Tensor::full({4, 3}, 0.5, DataType::Float32));
const std::string b = model.add_initializer(
Tensor::full({3}, 0.25, DataType::Float32));
const std::vector mm = model.add_node("MatMul", {"X", w}, 1);
const std::vector sum = model.add_node("Add", {mm[0], b}, 1);
const std::vector y = model.add_node("Relu", {sum[0]}, 1);
model.add_output(y[0], DataType::Float32, {batch, 3});
return model;
}
```
```python title="build_onnx.py (the build stage)"
def build_tiny_mlp() -> "crt.io.OnnxModel":
model = crt.io.OnnxModel.create("tiny_mlp", opset_version=21)
model.add_input("X", crt.float32, [-1, 4]) # any value <= 0 is dynamic
w = model.add_initializer(crt.tensor(np.full((4, 3), 0.5, dtype=np.float32)))
b = model.add_initializer(crt.tensor(np.full((3,), 0.25, dtype=np.float32)))
(mm,) = model.add_node("MatMul", ["X", w])
(summed,) = model.add_node("Add", [mm, b])
(out,) = model.add_node("Relu", [summed])
model.add_output(out, crt.float32, [-1, 3])
return model
```
`save(path)` writes the graph; `open(path)` round-trips it structure-intact, and `optimize()` runs the rewrite pipeline to a fixed point (a minimal graph survives unchanged).
## Compile and run
`compile()` can fail like any load of real weights and shapes, so it returns through the error contract; branch on the code name. It returns the graph as built. `optimize()` runs the graph optimizer when you want it, and `finalize()` readies the graph to run ([Optimize and finalize](optimize-and-finalize.mdx) covers both). The finalized `ModelGraph` runs positionally (one tensor per `input_names()` entry, in that order) and, in Python, also by name.
```cpp title="run_onnx.cpp (compile and run)"
int run(ClikaRT::io::OnnxModel& model) {
ClikaRT::Result compiled = CLIKART_TRY(model.compile());
if (!compiled.ok()) {
std::printf("compile: FAILED [%s]\n", compiled.code_name().c_str());
return 1;
}
ModelGraph graph = std::move(compiled.value());
graph.optimize(); // the graph optimizer, when wanted
graph.finalize(); // only a finalized graph runs
const std::vector x = {1.0F, 2.0F, 3.0F, 4.0F, -1.0F, 0.5F, 2.0F, -2.0F};
std::vector outputs = graph.run(
{Tensor::from_data(x.data(), {2, 4}, DataType::Float32)});
std::printf("y = %s\n", outputs[0].to_string().c_str());
return 0;
}
```
```python title="run_onnx.py (compile and run)"
graph = model.compile()
graph.optimize() # the graph optimizer, when wanted
graph.finalize() # only a finalized graph runs
print(f"inputs {graph.input_names()} -> outputs {graph.output_names()}")
x = crt.tensor(np.array([[1, 2, 3, 4], [-1, 0.5, 2, -2]], dtype=np.float32))
(y,) = graph.run([x]) # positional: input_names() order
(y2,) = graph.run({"X": x}) # named: the same result
print(y.numpy())
codes = [n.op_code for n in graph.nodes()] # the compiled graph as data
assert crt.graph.OpCode.MatMul in codes
# The one-call loader: open, compile, optimize and finalize, then call the model with named inputs.
model.save("tiny_mlp.onnx")
loaded = crt.onnx.load("tiny_mlp.onnx", dynamic_axes={"X": {0: "batch"}}, device="cpu")
out = loaded(X=x) # a dict keyed by output name
assert list(out) == loaded.output_names
```
`crt.onnx.load` opens, compiles, optimizes and finalizes in one call: `dynamic_axes` names the axes that vary between calls, `device` places the weights, and the returned model takes its inputs by keyword and returns a dict keyed by output name. `optimize=False` returns the graph as built, which runs after `model.graph.finalize()`. `graph.nodes()` reads the compiled graph as data, one `crt.graph.OpCode` per node; [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) walks that surface.
With the fixed all-half weights and the 0.25 bias above, the math fits in your head; the C++ program's run over the two-row input prints:
```text
y = relu(X W + B):
[5.25, 5.25, 5.25]
[0, 0, 0]
```
The first row is `(1+2+3+4) * 0.5 + 0.25 = 5.25` per output; the second row's pre-activation is negative in every column, so relu zeroes it.
Two pointers from here. A compiled `ModelGraph` is the same type that [tracing eager code](trace-eager-code-to-graphs.mdx) produces, so everything downstream of `compile()` is shared, [optimizing and finalizing](optimize-and-finalize.mdx) included. And a model too big to build by hand arrives as a file: the IO-contract section works unchanged on a checkpoint you downloaded.
---
# Serve a model over HTTP
Put compute behind HTTP endpoints with the built-in server, answer JSON requests with tensor results, and stream tokens with server-sent events.
Source: https://docs.clika.io/clikart/how-to/serve-over-http.md
You have working compute and want it behind an HTTP endpoint. `ClikaRT::http` ships a server in the same library: declare routes with lambdas, parse and build bodies with `ClikaRT::Json`, and stream with server-sent events. No web framework enters the ship path.
Each program below starts a server, drives it with the built-in HTTP client in the same process, and prints the exchange, so it runs self-contained. To poke a server from outside instead, replace `bind_to_any_port` with `listen_async("0.0.0.0", 8080)` and use curl.
A served process runs compute like any other, so give it a license credential where you set the rest of its environment (`CLIKA_RT_LICENSE`, or the per-user file `clikart-license-init` writes; [Get ClikaRT](../getting-started/get-clikart.mdx#license-credential)). Without one, a handler the runtime refuses for licensing reports the code name `LICENSE_FAILED`.
## Start a server and add routes
Routes are declared before the server starts: `get`/`post` for fixed paths, `route` with a regex for path parameters (`req.param(0)` is the first capture). Handlers return a `ServerResponse` and run concurrently on the server's I/O pool, so anything they share needs a lock.
```cpp title="routes.cpp"
#include
#include
#include
namespace http = ClikaRT::http;
using ClikaRT::json::Json;
using ClikaRT::Result;
int main() {
http::HttpServer server = http::HttpServer::create();
server.get("/health", [](http::ServerRequest&) -> Result {
Json o = Json::object();
o["status"] = "ok";
o["runtime"] = ClikaRT::GetVersionInfo();
return http::ServerResponse::json(o.dump());
});
server.route(http::Method::Get, R"(/models/(\w+))",
[](http::ServerRequest& req) -> Result {
Json o = Json::object();
o["model"] = req.param(0);
o["loaded"] = false;
return http::ServerResponse::json(o.dump());
});
const int port = server.bind_to_any_port("127.0.0.1");
server.wait_until_ready();
const std::string base = "http://127.0.0.1:" + std::to_string(port);
std::printf("GET /health -> %s\n", http::get_text(base + "/health").c_str());
std::printf("GET /models/smol -> %s\n", http::get_text(base + "/models/smol").c_str());
server.stop();
return 0;
}
```
The Python server lands with the HTTP bindings of the `clika_runtime.http` module; the samples on this page run once the wheel carries it. Routes are decorators or calls: `get` / `post` for fixed paths, `route(method, pattern, pattern=True)` for a regular expression with named captures, read back with `request.param(name)`. A handler returns an `http.Response`, a `str` (sent as `text/plain`) or a `dict` / `list` (sent as JSON). Handlers run on the server's own threads and hold the interpreter lock only while they run Python code.
```python title="routes.py"
import http.client
import json
from clika_runtime import http as crt_http
server = crt_http.HttpServer()
@server.get("/health")
def health(request: crt_http.Request) -> dict[str, str]:
return {"status": "ok"} # a dict is sent as JSON
@server.route("GET", r"/models/(?P[^/]+)", pattern=True)
def model(request: crt_http.Request) -> str:
return f"model {request.param('name')}" # a str is sent as text/plain
@server.post("/echo")
def echo(request: crt_http.Request) -> crt_http.Response:
return crt_http.Response.bytes(request.body, "application/octet-stream")
with server: # the block ends with server.stop()
port = server.bind_to_any_port("127.0.0.1")
client = http.client.HTTPConnection("127.0.0.1", port)
client.request("GET", "/health")
print(json.loads(client.getresponse().read())) # {'status': 'ok'}
client.request("GET", "/models/tiny")
print(client.getresponse().read().decode()) # model tiny
client.request("POST", "/echo", body=b"ping")
print(client.getresponse().read()) # b'ping'
```
`server.use(middleware)` wraps every route with `middleware(request, next)`; `server.group(prefix)` scopes routes under a path prefix; `listen(host, port)` blocks the calling thread until `stop()` and `listen_async` serves on the server's own thread. Every one of them releases the interpreter lock while it waits.
The Kotlin binding does not carry the serving runtime; the C++ arm is the serving story today. A served model is consumed from Kotlin with any JVM HTTP client, and [Modelverse's serving guide](/modelverse/how-to/serve-openai-compatible) shows the OpenAI-compatible route.
```text
GET /health -> {"status":"ok","runtime":"0.6.4"}
GET /models/smol -> {"model":"smol","loaded":false}
```
## A JSON inference endpoint
The serving shape every model endpoint repeats: parse the body, validate, build a tensor from the request, compute, and put the result back into JSON. Bad input gets a clean 400 with a reason, not an exception. The compute here is one dense layer with fixed weights, standing in for a loaded model ([the GGUF guide](load-quantized-weights.mdx) is where real weights come from). The Python tab carries the compute half for real: the same scoring model compiled once to a `ModelGraph` and run per request; only the HTTP transport stays C++.
```cpp title="score_endpoint.cpp"
#include
#include
#include
#include
#include
namespace http = ClikaRT::http;
namespace ops = ClikaRT::ops;
using ClikaRT::DataType;
using ClikaRT::json::Json;
using ClikaRT::Result;
using ClikaRT::Tensor;
constexpr std::int64_t kFeatures = 4;
int main() {
// The "model": y = x * W^T + b, weights fixed for a reproducible page.
const Tensor w = Tensor::full({2, kFeatures}, 0.5, DataType::Float32);
const Tensor b = Tensor::full({2}, 0.25, DataType::Float32);
http::HttpServer server = http::HttpServer::create();
server.post("/score", [&](http::ServerRequest& req) -> Result {
Result body = CLIKART_TRY(Json::parse(req.body()));
if (!body.ok())
return http::ServerResponse::json(R"({"error":"invalid json"})", 400);
Result feats = CLIKART_TRY(body.value().at("features"));
if (!feats.ok() || feats.value().size() != kFeatures) {
Json err = Json::object();
err["error"] = Json("'features' must hold " + std::to_string(kFeatures) + " numbers");
return http::ServerResponse::json(err.dump(), 400);
}
std::vector x(kFeatures);
for (std::size_t i = 0; i < kFeatures; ++i)
x[i] = static_cast(feats.value().at(i).as_double());
const Tensor input = Tensor::from_data(x.data(), {1, kFeatures}, DataType::Float32);
const std::vector scores =
ops::linear(input, w, b).reshape({-1}).item_as_vec();
Json out = Json::object();
out["scores"] = Json::array();
for (float s : scores) out["scores"].push_back(Json(static_cast(s)));
return http::ServerResponse::json(out.dump());
});
const int port = server.bind_to_any_port("127.0.0.1");
server.wait_until_ready();
const std::string base = "http://127.0.0.1:" + std::to_string(port);
std::printf("POST /score [1,2,3,4] -> %s\n",
http::post_text(base + "/score", R"({"features":[1, 2, 3, 4]})").c_str());
// Client helpers raise ClikaRT::Error on any status >= 400; CLIKART_TRY
// captures that as a Result when a failure is an expected outcome.
Result bad = CLIKART_TRY(http::post_text(base + "/score", R"({"features":[1]})"));
std::printf("POST /score [1] -> %s\n",
bad.ok() ? bad.value().c_str() : bad.message().c_str());
server.stop();
return 0;
}
```
```python title="score_compute.py"
# The model is compiled to a ModelGraph once at startup and run per request;
# the route below is the served endpoint (the ONNX chapters of the python
# examples build bigger graphs the same way).
import json
import numpy as np
import clika_runtime as crt
from clika_runtime import http as crt_http
FEATURES = 4
# The "model": y = x * W^T + b, as a compiled graph with fixed weights.
model = crt.io.OnnxModel.create("score", opset_version=21)
model.add_input("X", crt.float32, [1, FEATURES])
w = model.add_initializer(crt.tensor(np.full((FEATURES, 2), 0.5, dtype=np.float32)))
b = model.add_initializer(crt.tensor(np.full(2, 0.25, dtype=np.float32)))
(mm,) = model.add_node("MatMul", ["X", w])
(scores,) = model.add_node("Add", [mm, b])
model.add_output(scores, crt.float32, [1, 2])
graph = model.compile() # every node parses, weights bind, shapes resolve
graph.optimize() # the graph optimizer
graph.finalize() # only a finalized graph runs
def score(body: str) -> tuple[int, str]:
"""The handler shape: parse, validate, run the graph, answer JSON."""
try:
feats = json.loads(body).get("features")
except ValueError:
return 400, json.dumps({"error": "invalid json"})
if not isinstance(feats, list) or len(feats) != FEATURES:
return 400, json.dumps({"error": "'features' must hold 4 numbers"})
x = crt.tensor(np.asarray([feats], dtype=np.float32))
(y,) = graph.run({"X": x})
return 200, json.dumps({"scores": y.numpy().reshape(-1).tolist()})
server = crt_http.HttpServer()
@server.post("/score")
def score_route(request: crt_http.Request) -> crt_http.Response:
status, body = score(request.body.decode())
return crt_http.Response.json(json.loads(body), status=status)
print("POST /score [1,2,3,4] ->", *score('{"features":[1, 2, 3, 4]}')) # 200 {"scores": [5.25, 5.25]}
print("POST /score [1] ->", *score('{"features":[1]}')) # 400 {"error": "'features' must hold 4 numbers"}
```
A bad body answers a 400 with a reason, a good one a 200 with the scores; the route runs the same `score` the two prints call.
The Kotlin binding carries the serving runtime (`FunctionModel`, `Executor`, `Pipeline`) and the HTTP server (`HttpServer`); the worked program on this page is the C++ one, and [Language bindings](../bindings.md) says what Kotlin carries.
```text
POST /score [1,2,3,4] -> {"scores":[5.25,5.25]}
POST /score [1] -> 400 Bad Request on POST /score
```
The handler captures the weight tensors by reference; they outlive the server. A real model swaps the `ops::linear` line for its forward and nothing else changes shape. On the failure, the in-process client surfaces the status as the `Result`'s message; an external client (curl, a browser) reads the JSON error body the handler wrote.
## Stream results with server-sent events
Token-by-token streaming (the transport behind LLM chat responses) is one factory away: `ServerResponse::sse` takes a `next` callback, and the server pulls it until it returns an empty optional, one SSE frame per event. The client here buffers the finite stream and prints the raw frames; a browser or an SSE-aware client consumes them incrementally.
```cpp title="stream_tokens.cpp"
#include
#include
#include
#include
#include
namespace http = ClikaRT::http;
using ClikaRT::Result;
namespace {
constexpr const char* kTokens[] = {"Tensors ", "stream ", "one ", "by ", "one."};
constexpr int kTokenCount = static_cast(sizeof kTokens / sizeof kTokens[0]);
} // namespace
int main() {
http::HttpServer server = http::HttpServer::create();
server.get("/generate", [](http::ServerRequest&) -> Result {
auto sent = std::make_shared(0); // per-connection cursor
return http::ServerResponse::sse(
[sent]() -> Result> {
if (*sent >= kTokenCount) return std::optional{}; // close
http::ServerSentEvent ev;
ev.event = "token";
ev.data = kTokens[(*sent)++];
return std::optional{ev};
});
});
const int port = server.bind_to_any_port("127.0.0.1");
server.wait_until_ready();
const std::string raw =
http::get_text("http://127.0.0.1:" + std::to_string(port) + "/generate");
std::printf("raw SSE body:\n%s", raw.c_str());
server.stop();
return 0;
}
```
A streamed body pulls its frames from a generator: `crt_http.Response.sse(events)` sends one `event:` / `data:` frame per item, and the server pulls the next item on its own thread as the client reads. The producer is the generation loop; [the tokenizer guide's streaming-decode section](tokenize-and-chat-templates.mdx) turns its ids into exactly the text pieces a stream sends, one non-empty piece per `token` event.
```python title="sse.py"
from clika_runtime import http as crt_http
server = crt_http.HttpServer()
@server.get("/events")
def events(request: crt_http.Request) -> crt_http.Response:
def frames():
for piece in ("Tensors ", "stream ", "one ", "by ", "one."):
yield {"event": "token", "data": piece} # one frame per generated piece
return crt_http.Response.sse(frames())
```
The same route with a decoder behind it yields each non-empty `push` result of the streaming decoder as a frame's `data`.
The Kotlin binding carries the serving runtime, the streaming decoder and the HTTP server; the worked program on this page is the C++ one, and [Language bindings](../bindings.md) says what Kotlin carries.
```text
raw SSE body:
event: token
data: Tensors
event: token
data: stream
event: token
data: one
event: token
data: by
event: token
data: one.
```
Each frame is an `event:` line, a `data:` line, and a blank line; a generation loop replaces the fixed token array with reads from its decoder, and [the tokenizer guide's streaming-decode section](tokenize-and-chat-templates.mdx) is that decoder: each non-empty `push` result is one frame's `data`. The server pulls `next` on an I/O worker, so a slow producer stalls only its own connection.
Middleware (logging, auth), static mounts, and an image-upload endpoint that runs compute are in the bundle's `http_server` example, chapters `03` to `06`. Wiring a served endpoint into sessions and continuous batching is the `runtime` example's ground. The finished version of this page's story ships in Modelverse: [an OpenAI-compatible endpoint](/modelverse/how-to/serve-openai-compatible) over these same server pieces, with its `01_serve` example as the smallest complete server ([Additional examples](/modelverse/examples) has it).
---
# Tokenize text and apply a chat template
Turn text into token ids and back, align tokens to source bytes, render a conversation with the model's own chat template, and batch for a model.
Source: https://docs.clika.io/clikart/how-to/tokenize-and-chat-templates.md
Your model consumes token ids, and a chat model expects its prompt formatted exactly the way it was trained. `ClikaRT::Tokenizer` covers both: one loader reads a HuggingFace model directory, `encode`/`decode` convert text to ids and back, and the model's own chat template renders conversations. No Python and no external tokenizer library are involved.
The programs below use the tokenizer of [SmolLM2-135M-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct), the model from [the GGUF guide](load-quantized-weights.mdx). Two small files are all a tokenizer needs:
```bash
mkdir -p SmolLM2-135M-Instruct && cd SmolLM2-135M-Instruct
curl -LO "https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/resolve/main/tokenizer.json"
curl -LO "https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/resolve/main/tokenizer_config.json"
cd ..
```
The bundle also ships a self-contained tokenizer under `examples/src/tokenizer/data/hf_model`, if you would rather not download anything.
## Load a tokenizer and round-trip some text
`Tokenizer::from_huggingface` takes the model directory, detects the artifact inside it (`tokenizer.json`, `tokenizer.model`, `tekken.json`, or `vocab.json` plus `merges.txt`), and overlays `tokenizer_config.json` for the special-token ids and the chat template. `Tokenizer::from_file` loads one tokenizer file (or a directory holding one) directly, detecting its format the same way.
```cpp title="roundtrip.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
std::printf("vocab %lld, bos %lld, eos %lld, chat template: %s\n",
static_cast(tok.vocab_size()), static_cast(tok.bos_id()),
static_cast(tok.eos_id()), tok.has_chat_template() ? "yes" : "no");
const std::string text = "ClikaRT runs the same code on every backend.";
const std::vector ids = tok.encode(text);
std::printf("encoded %zu tokens:", ids.size());
for (std::int32_t id : ids) std::printf(" %d", id);
std::printf("\ndecoded: %s\n", tok.decode(ids).c_str());
return 0;
}
```
```python title="roundtrip.py"
def main() -> None:
tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir())
print(f"vocab {tok.vocab_size}, bos {tok.bos_id}, eos {tok.eos_id}, "
f"chat template: {'yes' if tok.has_chat_template else 'no'}")
text = "ClikaRT runs the same code on every backend."
ids = tok.encode(text)
print(f"encoded {len(ids)} tokens:", *ids)
print(f"decoded: {tok.decode(ids)}")
if __name__ == "__main__":
main()
```
```text
vocab 49152, bos 1, eos 2, chat template: yes
encoded 12 tokens: 51 1418 6335 16895 7313 260 1142 2909 335 897 25817 30
decoded: ClikaRT runs the same code on every backend.
```
## See where each token came from
`tokenize` returns one `Token` per piece: the id, the surface string, and the byte span `[begin, end)` in the original text. Slice the original by that span when you need alignment (highlighting, span labeling, streaming cursors); the spans line up exactly, dropped spaces included.
```cpp title="offsets.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
const std::string text = "Quantized weights stay packed.";
const std::vector tokens = tok.tokenize(text, /*add_special_tokens=*/false);
std::printf(" id [begin,end) source span\n");
for (const Token& t : tokens) {
const std::string_view span(text.data() + t.begin, t.end - t.begin);
std::printf(" %-6d [%2zu,%2zu) \"%.*s\"\n",
t.id, t.begin, t.end, static_cast(span.size()), span.data());
}
return 0;
}
```
```python title="offsets.py"
def main() -> None:
tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir())
text = "Quantized weights stay packed."
ids = tok.encode(text, add_special_tokens=False)
# The python binding returns ids; id_to_token shows each piece. A
# leading 'G-with-breve' marks a token that starts with a space; the
# byte-span view is the C++ tokenize surface.
print(" id token")
for i in ids:
print(f" {i:<6} {tok.id_to_token(i)!r}")
if __name__ == "__main__":
main()
```
```text
id [begin,end) source span
24696 [ 0, 5) "Quant"
1005 [ 5, 9) "ized"
10379 [ 9,17) " weights"
2951 [17,22) " stay"
13448 [22,29) " packed"
30 [29,30) "."
```
## Render a conversation with the model's chat template
A chat model's prompt format (its role markers, turn separators, generation priming) ships with the model as a Jinja2 template in `tokenizer_config.json`, and the loader attached it above. `apply_chat_template` renders a `messages` array the OpenAI-API shape into the exact prompt string; `encode_chat` goes straight to ids. Never hand-build these markers: the template is the model's contract.
```cpp title="chat_template.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
Json messages = Json::array();
Json system = Json::object();
system["role"] = "system";
system["content"] = "You are a concise assistant.";
messages.push_back(std::move(system));
Json user = Json::object();
user["role"] = "user";
user["content"] = "What does a tokenizer do?";
messages.push_back(std::move(user));
const std::string prompt = tok.apply_chat_template(messages);
std::printf("=== rendered prompt ===\n%s\n=======================\n", prompt.c_str());
// The template writes the prompt's own bos/eos framing, so encode_chat adds
// no special tokens on top of it.
const std::vector ids = tok.encode_chat(messages, /*add_generation_prompt=*/true);
std::printf("encode_chat produced %zu tokens\n", ids.size());
return 0;
}
```
```python title="chat_template.py"
def main() -> None:
tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir())
# The messages array travels as JSON text at this surface.
messages = json.dumps([
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What does a tokenizer do?"},
])
prompt = tok.apply_chat_template(messages)
print(f"=== rendered prompt ===\n{prompt}\n=======================")
# The template writes the prompt's own bos/eos framing, so encode_chat adds
# no special tokens on top of it.
ids = tok.encode_chat(messages, add_generation_prompt=True)
print(f"encode_chat produced {len(ids)} tokens")
if __name__ == "__main__":
main()
```
```text
=== rendered prompt ===
<|im_start|>system
You are a concise assistant.<|im_end|>
<|im_start|>user
What does a tokenizer do?<|im_end|>
<|im_start|>assistant
=======================
encode_chat produced 107 tokens
```
The rendered prompt ends with the assistant-turn priming (`add_generation_prompt` defaults to true), so the model continues as the assistant. For tool calling, extra template variables, or a reproducible clock, pass a `ChatTemplateInputs` instead of the bare messages array; the two-argument form above covers plain conversations.
## Batch for a model
Feeding a model takes tensors, not vectors. `encode_batch` with `return_tensors` produces the standard quartet: padded `input_ids` `[B, S]`, an `attention_mask`, per-sequence lengths, and the `cu_seqlens` prefix-sum table. `varlen = true` skips padding entirely and lays the ids out flat, the shape variable-length attention consumes.
```cpp title="batch.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
const std::vector texts = {
"Short prompt.",
"A somewhat longer prompt that pads the short one.",
};
EncodeOptions opts;
opts.return_tensors = true;
// Decoder-only checkpoints often ship no pad token; designate one (eos is
// the usual choice) or the padded encode raises ClikaRT::Error.
opts.pad_id = static_cast(tok.eos_id());
const ClikaRT::tokenizer::Encoded batch = tok.encode_batch(texts, opts);
std::printf("input_ids %s\n", batch.input_ids->to_string().c_str());
std::printf("attention_mask %s\n", batch.attention_mask->to_string().c_str());
std::printf("seq_lengths %s\n", batch.seq_lengths->to_string().c_str());
opts.varlen = true;
const ClikaRT::tokenizer::Encoded flat = tok.encode_batch(texts, opts);
std::printf("varlen ids %s\n", flat.input_ids->to_string().c_str());
std::printf("cu_seqlens %s\n", flat.cu_seqlens->to_string().c_str());
return 0;
}
```
```python title="batch.py"
def main() -> None:
tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir())
texts = [
"Short prompt.",
"A somewhat longer prompt that pads the short one.",
]
# The python binding returns the ragged ids, one list per text; pad on
# the tensor side with the lengths below. The padded quartet (input_ids,
# attention_mask, seq_lengths, cu_seqlens) is the C++ and C surface.
batch = tok.encode_batch(texts)
for row, ids in enumerate(batch.ids):
print(f"text {row}: {len(ids):2} ids {ids}")
if __name__ == "__main__":
main()
```
```text
input_ids Tensor(shape=[2, 10], dtype=Int32, device=CPU, numel=20, data=[20355, 6011, 30, 2, 2, 2, ...])
attention_mask Tensor(shape=[2, 10], dtype=Int32, device=CPU, numel=20, data=[1, 1, 1, 0, 0, 0, ...])
seq_lengths Tensor(shape=[2], dtype=Int32, device=CPU, numel=2, data=[3, 10])
varlen ids Tensor(shape=[13], dtype=Int32, device=CPU, numel=13, data=[20355, 6011, 30, 49, 7932, 2848, ...])
cu_seqlens Tensor(shape=[3], dtype=Int32, device=CPU, numel=3, data=[0, 3, 13])
```
`EncodeOptions` also carries truncation (`max_length`, `truncation_side`), the padding side (`Left` suits decoder-only batch generation), and a target `device` so the tensors land where the model computes. The bundle's `tokenizer` example walks each of these one chapter at a time, and the `templating` example covers the Jinja2-compatible engine behind `apply_chat_template` on its own.
## Stream the decode of a generation loop
A generation loop produces ids one at a time, and `decode(ids)` over the growing list re-decodes everything on every step. `Tokenizer::streaming_decoder` is the incremental form: `push(id)` returns exactly the newly-stable text, and the pieces concatenate to what `decode` would have produced. The catch it handles for you is the UTF-8 boundary: one code point can span tokens, so `push` holds bytes back until they are displayable and returns an empty string meanwhile; you never emit half a character. `finish()` flushes whatever the tail held (a trailing incomplete sequence as-is) and resets the decoder for a fresh stream; a decoder serves one stream at a time.
```cpp title="stream_decode.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
// Stand-in for a generation loop: the ids a real decoder would emit one
// at a time (the roundtrip section's sentence, so the ids match).
const std::vector ids =
tok.encode("ClikaRT runs the same code on every backend.");
StreamingDecoder stream = tok.streaming_decoder();
std::string assembled;
int emitted = 0;
for (std::int32_t id : ids) {
// push returns exactly the newly-stable text: empty while a
// multi-byte code point is still incomplete, never a torn character.
const std::string piece = stream.push(id);
if (!piece.empty()) ++emitted;
assembled += piece;
}
assembled += stream.finish(); // flush the tail; the decoder resets
std::printf("%zu ids -> %d incremental pieces\n", ids.size(), emitted);
std::printf("assembled: %s\n", assembled.c_str());
std::printf("assembled == decode(ids): %s\n",
assembled == tok.decode(ids) ? "yes" : "no");
return 0;
}
```
```python title="stream_decode.py"
def main() -> None:
tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir())
# Stand-in for a generation loop: the ids a real decoder would emit one
# at a time (the roundtrip section's sentence, so the ids match).
ids = tok.encode("ClikaRT runs the same code on every backend.")
stream = tok.streaming_decoder()
pieces = [stream.push(i) for i in ids] # "" while a code point is incomplete
assembled = "".join(pieces) + stream.finish() # flush; the decoder resets
print(f"{len(ids)} ids -> {sum(1 for p in pieces if p)} incremental pieces")
print(f"assembled: {assembled}")
print(f"assembled == decode(ids): "
f"{'yes' if assembled == tok.decode(ids) else 'no'}")
if __name__ == "__main__":
main()
```
{/* CERTIFICATION: the run below follows the API contract (the pieces concatenate to decode(ids));
the piece count is not verified against a built bundle. */}
```text
32 ids -> 32 incremental pieces
assembled: ClikaRT runs the same code on every backend.
assembled == decode(ids): yes
```
Every push emitted text here because the sentence is plain ASCII; text with accents, CJK, or emoji is where the empty returns appear, and exactly why the boundary handling exists. The decoder skips special tokens by default (`streaming_decoder(false)` keeps them), and the handle stays valid even after the `Tokenizer` that made it is gone, so a generation worker can own just the decoder.
This is the producer half of token streaming: each non-empty piece is one frame for the transport. [The serving guide](serve-over-http.mdx) sends exactly these pieces as `token` events over server-sent events.
---
# Trace eager code to graphs
Capture an eager function or module as a ModelGraph with trace, inspect it, optimize, finalize and run it, and wrap hot paths in compile for capture-and-replay.
Source: https://docs.clika.io/clikart/how-to/trace-eager-code-to-graphs.md
Eager code runs op by op. Tracing runs your function ONCE over data-free stand-ins (shapes and dtypes matter, values are never read) and captures the operator graph as a `ModelGraph`, the same type the ONNX loader compiles to. The graph comes back as recorded: `optimize()` runs the graph optimizer when you want it, and `finalize()` readies it for the same `run`.
Two rules make a function traceable:
- **List in, list out.** A traced callable receives its input tensors as one list and returns its outputs as a list. A bare tensor return does not auto-wrap; return `[y]`.
- **No value reads.** The stand-ins carry no data, so reading a value during tracing raises. Shape-driven math is fine.
In Python the two rules relax to pytrees: a traced function may take a list, a dict or any nested container of tensors and return one, and the graph keeps the names. No value is read during the capture in either form.
```python title="trace_function.py"
import numpy as np
import clika_runtime as crt
w = crt.tensor(np.full((4, 3), 0.1, dtype=np.float32))
b = crt.tensor(np.zeros(3, dtype=np.float32))
def fn(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
return [crt.softmax(crt.relu(crt.add(crt.matmul(inputs[0], w), b)), dim=-1)]
x = crt.tensor(np.ones((2, 4), dtype=np.float32))
# Capture: fn runs once over stand-ins; the graph names its inputs and outputs.
graph = crt.trace(fn, example_inputs=[x], input_names=["x"], output_names=["p"])
print(graph.input_names(), graph.output_names()) # ['x'] ['p']
# The graph comes back as recorded: optimize it, then finalize it to run.
graph.optimize()
graph.finalize()
# The graph runs like the function did, by position or by name.
(positional,) = graph.run([x])
(named,) = graph.run({"x": x})
print(positional.numpy()[0]) # [0.33333334 0.33333334 0.33333334]
```
A module traces the same way, and the traced graph is inspectable node by node: each node carries an `op_code` from `crt.graph.OpCode` and a `kind`.
```python title="trace_module.py"
import numpy as np
import clika_runtime as crt
import clika_runtime.nn as nn
class Head(nn.Module):
def __init__(self, w: np.ndarray, b: np.ndarray) -> None:
super().__init__()
self.weight = nn.Parameter(crt.tensor(w))
self.bias = nn.Parameter(crt.tensor(b))
def forward(self, x: crt.Tensor) -> crt.Tensor:
return crt.softmax(crt.relu(crt.add(crt.matmul(x, self.weight), self.bias)), dim=-1)
rng = np.random.default_rng(0)
head = Head(rng.standard_normal((4, 3)).astype(np.float32), np.zeros(3, dtype=np.float32))
x = crt.tensor(np.ones((2, 4), dtype=np.float32))
traced = crt.trace(head, example_inputs=(x,))
graph = traced.graph # the ModelGraph, as recorded
codes = {node.op_code for node in graph.nodes()}
print(crt.graph.OpCode.MatMul in codes, crt.graph.OpCode.Softmax in codes) # True True
print(sum(node.kind == crt.graph.NodeKind.Input for node in graph.nodes())) # 1
```
As recorded, the graph holds one node per operator the forward calls, the `MatMul`, the `Add`, the `Relu` and the `Softmax`, after its one input node. `optimize()` folds the bias add and the relu into the `MatMul`, which leaves it and the `Softmax`.
The same capture surface in C++: `ClikaRT::graph::trace` takes the callable and the example inputs and returns the `ModelGraph` as recorded, and `ClikaRT::compile` wraps a callable in capture-and-replay exactly as below. [Optimize and finalize](optimize-and-finalize.mdx) walks `optimize()`, `to()` and `finalize()` in C++, and the shipped examples include a full tracing walkthrough.
## The compile wrapper
`compile` wraps a callable in capture-and-replay: the first call runs eagerly AND captures; later calls replay the graph. A shape change recaptures, invisible to values and visible on the counter.
```python title="compile_function.py"
step = crt.compile(fn, [crt.TensorSpec("x", crt.float32, [2, 4])])
print(step.state) # State.Pending (State.Pending: nothing captured yet)
(first,) = step([x]) # runs eagerly and records the graph
print(step.state) # State.Compiled (State.Compiled: later calls replay)
(again,) = step([x]) # replays the captured graph; Python is not called
print(step.recapture_count) # 0 (0)
graph = step.take_graph() # the captured ModelGraph, for standalone use
(replayed,) = graph.run([x])
```
The pytree form takes a module or a function over dicts and returns dicts; `dynamic=True` keeps one graph across batch sizes instead of recapturing on a shape change.
```python title="compile_module.py"
class Wrapped(nn.Module):
def __init__(self) -> None:
super().__init__()
self.head = head
def forward(self, batch: dict[str, crt.Tensor]) -> dict[str, crt.Tensor]:
p = self.head(batch["x"])
return {"probs": p, "argmax": crt.argmax(p, dims=[1])}
compiled = crt.compile(Wrapped(), dynamic=True)
out = compiled({"x": x})
print(sorted(out)) # ['argmax', 'probs'] (['argmax', 'probs'])
again = compiled({"x": crt.tensor(np.ones((4, 4), dtype=np.float32))})
print(again["probs"].shape) # clika_runtime.Size([4, 3])
```
`fullgraph=True` turns a fallback into an error: a Python branch on a tensor value cannot be recorded, and `crt.compile(fn, fullgraph=True)` raises `crt.ClikaRTError` at the call instead of serving it eagerly. `step.reset()` re-arms the wrapper, and `step.state` reads `Pending`, `Compiled` or `Fallback`.
---
# Use ClikaRT from Python
The clika-runtime wheel: NumPy in and out with explicit copy semantics, math that reads as math, modes as strings, models as nn.Module, and errors typed by class.
Source: https://docs.clika.io/clikart/how-to/use-clikart-from-python.md
The `clika-runtime` wheel puts the runtime behind one import: `import clika_runtime as crt` loads `libClikaRT.so` and its backends from inside the wheel, with no library paths to set. The Python surface follows PyTorch's shapes: dtype objects such as `crt.float32`, a string spelling for every mode argument, `nn.Module` for models, and one exception class per failure kind. NumPy plays three roles: data entry, data exit, and the independent oracle you check results against. Compute runs in the runtime.
This guide assumes the wheel is installed; [First steps](../getting-started/index.md) covers getting it. Everything below is one script's worth of ground: the license credential, the NumPy boundary, operator chains, modes as strings, device placement, a model as `nn.Module`, and the error contract.
## The credential goes in before the import
The import is what loads the runtime, so `CLIKA_RT_LICENSE` has to hold the credential before `import clika_runtime` runs. Export it in the shell, or assign it above the import:
```python title="license.py"
import os
os.environ["CLIKA_RT_LICENSE"] = "CLIKA1-..." # the credential text, or the path of a file holding it
import clika_runtime as crt
```
The alternative drops the variable: `clikart-license-init `, a console script the wheel installs, stores the credential once under your user account. `clika_runtime.torch` and `clika_runtime.modelverse` follow the same rule, because all three are the one runtime. [License the runtime](license-the-runtime.mdx) is the whole contract; without a valid credential a call raises `crt.ClikaRTError` with `code_name` `LICENSE_FAILED`.
## NumPy in, NumPy out
`crt.tensor(array)` copies the array in, `crt.from_numpy(array)` borrows its memory, and `t.numpy()` is the exit. Lists and scalars enter too, at the dtype NumPy would pick for them.
```python title="boundary.py"
import numpy as np
import clika_runtime as crt
a = np.ones((2, 5), dtype=np.float32)
t = crt.tensor(a)
print(t.shape, t.dtype, t.device) # clika_runtime.Size([2, 5]) clika_runtime.float32 cpu
assert t.dtype == crt.float32 # one dtype object per storable dtype
assert isinstance(crt.float32, crt.dtype)
assert t.device == crt.Device("cpu")
a[0, 0] = 999.0 # entry copied: the tensor is unmoved
assert t.numpy()[0, 0] == 1.0
borrowed = crt.from_numpy(a) # from_numpy shares the array's memory
assert borrowed.numpy()[0, 0] == 999.0
assert crt.tensor([1, 2, 3]).dtype == crt.int64
half = t.to(crt.float16) # narrowing is explicit
assert half.dtype == crt.float16
print(half)
# tensor([[1., 1., 1., 1., 1.],
# [1., 1., 1., 1., 1.]], dtype=clika_runtime.float16)
```
Every NumPy-native dtype enters as itself (the float family, the signed ints, `uint8`, `bool`); a dtype with no tensor twin, `complex64` for example, is refused with a `TypeError` that names the routes out. Payload dtypes NumPy cannot spell, bfloat16 among them, cross through `bytes()` and `Tensor.from_bytes()` instead of the array bridge. A tensor prints as `tensor([...])`: the dtype is named when it is not `float32`, the device when it is not the CPU, and a tensor above a thousand elements is abbreviated to its edge items.
A tensor prints as its values, wrapped in `tensor(...)`, with no shape or device header:
```python
print(2 * crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3)) + 3)
```
```text
tensor([[ 3., 5., 7.],
[ 9., 11., 13.]])
```
## Math that reads as math
Operators compose the way the expression reads: Python numbers broadcast, `@` is matmul, and method chains mirror the functional forms. Check anything against NumPy; that is what the oracle role means.
```python title="tensor_math.py"
import numpy as np
import clika_runtime as crt
x = crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3))
y = 2.0 * x + 3.0 # scalars broadcast
z = (y - 3.0).abs().amax().item() # a method chain down to one float
print(z) # 10.0
assert np.allclose((x ** 2).numpy(), x.numpy() ** 2)
a = crt.tensor(np.ones((2, 3), dtype=np.float32))
b = crt.tensor(np.ones((3, 2), dtype=np.float32))
print((a @ b).numpy())
# [[3. 3.]
# [3. 3.]]
```
## Modes are strings
A mode argument takes its spelling as a string: `approximate="tanh"`, `mode="reflect"`, `rounding_mode="floor"`, `activation="relu"`. An unknown spelling raises a `ValueError` that lists the accepted ones.
```python title="modes.py"
import numpy as np
import clika_runtime as crt
x = crt.tensor(np.linspace(-3.0, 3.0, 7, dtype=np.float32))
tanh_form = crt.gelu(x, approximate="tanh")
assert not np.array_equal(tanh_form.numpy(), crt.gelu(x).numpy())
assert np.array_equal(crt.gelu(x).numpy(), crt.gelu(x, approximate="none").numpy())
padded = crt.pad(x, [2, 2], mode="reflect")
assert np.array_equal(padded.numpy(), np.pad(x.numpy(), (2, 2), mode="reflect"))
try:
crt.div(x, 2.0, rounding_mode="ceil")
except ValueError as e:
print(e) # rounding_mode: 'ceil' is not a RoundingMode; choose one of 'none', 'trunc', 'floor'
```
## Placement
Placement is a constructor argument or a move: `crt.tensor(arr, device=...)` lands data where you say, `.to("cpu")` moves it, and `crt.Device.gpu()` names the machine's accelerator, or the CPU when it has none, so the same script runs everywhere. Each backend has a namespace: `crt.cuda`, `crt.vulkan` and `crt.metal` mirror `crt.accelerator`, with `is_available()`, `device(index)` and `synchronize()`.
```python title="placement.py"
import numpy as np
import clika_runtime as crt
gpu = crt.Device.gpu()
t = crt.zeros(2, device=gpu)
assert t.device == gpu
assert np.array_equal(crt.to(t, "cpu").numpy(), np.zeros(2, dtype=np.float32))
if crt.cuda.is_available():
x = crt.ones(2, 3, device=crt.cuda.device(0))
crt.cuda.synchronize()
print(crt.Device("cpu")) # cpu
```
## A model is an nn.Module
Assigning a layer in `__init__` registers it, as in PyTorch: `load_state_dict` binds dotted names, `named_parameters()` enumerates them, `state_dict()` exports the same names back out, and the instance is callable. A `Parameter` is a `Tensor`.
```python title="model.py"
import numpy as np
import clika_runtime as crt
import clika_runtime.nn as nn
class TinyMlp(nn.Module):
def __init__(self, d_in: int, d_hidden: int, d_out: int) -> None:
super().__init__()
# The first layer fuses its activation as an epilogue.
self.up = nn.Linear(d_in, d_hidden, activation="relu")
self.down = nn.Linear(d_hidden, d_out, bias=False)
def forward(self, x: crt.Tensor) -> crt.Tensor:
return self.down(self.up(x))
model = TinyMlp(4, 8, 2)
assert isinstance(model.up.weight, nn.Parameter) and isinstance(model.up.weight, crt.Tensor)
result = model.load_state_dict({
"up.weight": crt.tensor(np.full((8, 4), 0.1, dtype=np.float32)),
"up.bias": crt.tensor(np.zeros(8, dtype=np.float32)),
"down.weight": crt.tensor(np.full((2, 8), 0.1, dtype=np.float32)),
})
assert result.missing_keys == [] and result.unexpected_keys == []
y = model(crt.tensor(np.ones((3, 4), dtype=np.float32)))
print(y.shape) # clika_runtime.Size([3, 2])
exported = model.state_dict() # the same dotted names back out
print(sorted(exported)) # ['down.weight', 'up.bias', 'up.weight']
```
`load_state_dict` is strict by default: a missing or unexpected key raises a `RuntimeError` naming both sets; `strict=False` returns the report instead, with `missing_keys` and `unexpected_keys`. The `state_dict()` -> fresh `load_state_dict()` round trip reproduces the forward, which is the portable way to hand weights between processes. [Author a model in Python](author-a-model-in-python.mdx) builds a full decoder this way.
## When it fails
A runtime failure raises an exception typed by its kind: `crt.InvalidArgumentError` (also a `ValueError`) for a shape, dtype, device or option the call cannot accept, `crt.NotFoundError` (also a `FileNotFoundError`) for a missing file or entry, `crt.UnsupportedError`, `crt.OutOfMemoryError` and `crt.UnavailableError` for their kinds, all under `crt.ClikaRTError`. Every instance carries `.code_name` (the fine code, stable across builds) and `.status` (the coarse class); `str(err)` is the message alone. Branch on the class or the code name, never on the message text: an argument mistake reads as a sentence naming the operation and the values, while an `E` message is an internal fault code specific to the build that produced it (report it verbatim with the runtime version).
```python title="errors.py"
import numpy as np
import clika_runtime as crt
try:
a = crt.tensor(np.ones((2, 3), dtype=np.float32))
b = crt.tensor(np.ones((4, 5), dtype=np.float32))
_ = a @ b # shape mismatch
except crt.InvalidArgumentError as e:
assert isinstance(e, ValueError)
print(f"failed with code {e.code_name!r}") # failed with code 'INVALID_ARGUMENT'
try:
crt.load("absent.safetensors")
except crt.NotFoundError as e:
assert isinstance(e, FileNotFoundError)
assert e.code_name == "NO_SUCHFILE"
```
From here, the rest of the Python surface follows the same grammar: [GGUF and quantized weights](load-quantized-weights.mdx) and [tokenizers](tokenize-and-chat-templates.mdx) have Python arms on their pages, [ONNX models](run-an-onnx-model.mdx) compile and run, [tracing](trace-eager-code-to-graphs.mdx) turns eager functions into graphs, [pytrees](pytrees.mdx) carry structured inputs and outputs across those boundaries, and [PyTorch interop](use-clikart-with-pytorch.mdx) covers the torch backend and tensor exchange. The wheel's own example programs double as a smoke suite for an installed wheel.
---
# Use ClikaRT with PyTorch
Run a torch.compile model on the runtime with backend=\"clika\", move tensors across the two libraries over DLPack without a copy, and convert an eager torch.nn.Module tree into runtime layers.
Source: https://docs.clika.io/clikart/how-to/use-clikart-with-pytorch.md
`clika_runtime.torch` is the PyTorch side of the wheel: a `torch.compile` backend that lowers the captured graph onto the runtime's operators, a tensor exchange over DLPack that shares memory instead of copying, and a converter from `torch.nn.Module` trees to `clika_runtime.nn` layers. It ships with the `torch` extra:
```bash
pip install "clika-runtime[torch]"
```
Importing `clika_runtime` on its own never imports torch; only `clika_runtime.torch` reaches for it, at the first use, and it raises an `ImportError` naming the extra when torch is absent. A process that never touches the interop package pays nothing for it.
## Compile a model onto the runtime
Registration puts the backend under the name `"clika"`; from then on `torch.compile(model, backend="clika")` sends every captured graph to the runtime, and the compiled module returns torch tensors like any other backend. Registering twice is a no-op; registering under a name torch already owns (`"inductor"`) raises.
```python title="compile.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime.torch as crt_torch
crt_torch.register_torch_backend()
print("clika" in torch._dynamo.list_backends()) # True
torch.manual_seed(2)
model = torch.nn.Sequential(torch.nn.Linear(131, 64), torch.nn.ReLU(), torch.nn.Linear(64, 131)).eval()
x = torch.randn(5, 131)
with torch.no_grad():
expected = model(x)
got = torch.compile(model, backend="clika")(x)
print(got.shape) # torch.Size([5, 131])
print(torch.allclose(got, expected, rtol=1e-3, atol=1e-3)) # True
```
The wheel also declares the backend as a `torch_dynamo_backends` entry point, so an installed distribution serves `torch.compile(model, backend="clika")` with no `clika_runtime.torch` import in the calling code: torch resolves the name through `importlib.metadata.entry_points(group="torch_dynamo_backends")`, where `clika` loads `clika_runtime.torch.clika_backend`. Passing the function itself works too: `torch.compile(model, backend=crt_torch.clika_backend)`.
`torch_backend(name)` registers under a name of your choosing for the length of a `with` block and restores torch's backend registry on exit, byte for byte:
```python title="scoped_backend.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime.torch as crt_torch
torch._dynamo.list_backends(None) # torch imports its own backends lazily
with crt_torch.torch_backend("clika-scoped") as backend:
print("clika-scoped" in torch._dynamo.list_backends()) # True
print("clika-scoped" in torch._dynamo.list_backends()) # False
```
### What the backend does with a graph
`torch.compile` hands the backend an FX graph: placeholders for the inputs, one node per operation, an output node. The backend lowers every node onto a runtime operator once and records the result; every later call converts the inputs, replays the recorded runtime graph, and converts the outputs back. Static shapes specialize the graph, as they do for any dynamo backend.
A graph break in the function (a `print`, a Python branch on a value) splits the capture into several graphs, each lowered on its own; `captured_graphs()` lists what the backend has lowered in this process and `clear_captured_graphs()` empties the list. `clika_runtime.torch.compile(fn, **torch_compile_kwargs)` is `torch.compile(fn, backend=clika_backend, ...)` with one addition: an operator the runtime cannot lower raises `LoweringError` out of the compiled call, and one error names every unsupported operator in the graph rather than the first.
```python title="graph_breaks.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime.torch as crt_torch
def fn(a: torch.Tensor, b: torch.Tensor) -> torch.Tensor:
h = torch.relu(a @ b)
print("break") # a graph break: two graphs; prints twice (the recording run, then the replay)
return torch.tanh(h).sum(dim=-1)
a, b = torch.randn(7, 131), torch.randn(131, 29)
got = crt_torch.compile(fn)(a, b)
print(torch.allclose(got, fn(a, b), rtol=1e-3, atol=1e-3)) # True
print(len(crt_torch.captured_graphs())) # 2
def unsupported(a: torch.Tensor) -> torch.Tensor:
return torch.special.zeta(a, a) + torch.special.bessel_j0(a)
try:
crt_torch.compile(unsupported, fullgraph=True)(torch.rand(3, 131) + 2.0)
except crt_torch.LoweringError as error:
print(error)
# 2 unsupported torch operation(s) in the graph:
# torch._C._special.special_zeta (node: special_zeta)
# torch._C._special.special_bessel_j0 (node: special_bessel_j0)
```
The compiled path follows torch's dtype: a float32 model compares with its eager result to within 1e-3 relative and absolute, the tolerance the interop test suite holds every compiled model to.
## Tensors across the two libraries
`from_torch` and `to_torch` exchange tensors through the DLPack protocol. The result shares the source's memory (a write through either side is visible from the other), the source buffer stays alive for the result's lifetime, and nothing is copied: on the CPU, and on a CUDA device both libraries address. The dtypes that cross are the ones both sides spell in the protocol: bool, the four unsigned and four signed integer widths, float16, bfloat16, float32 and float64; a torch float8 tensor or a runtime sub-byte tensor is refused with a `TypeError` naming the dtype and the cast that gets it across.
```python title="exchange.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime as crt
import clika_runtime.torch as crt_torch
source = torch.randn(5, 131)
runtime_tensor = crt_torch.from_torch(source)
print(runtime_tensor.dtype == crt.float32, tuple(runtime_tensor.shape)) # True (5, 131)
print(runtime_tensor.numpy().ctypes.data == source.data_ptr()) # True: one buffer
source[0, 0] = 7.0
print(float(runtime_tensor.numpy()[0, 0])) # 7.0
view = crt_torch.to_torch(runtime_tensor)
print(view.data_ptr() == source.data_ptr()) # True: the round trip never copies
view[1, 1] = -3.0
print(float(source[1, 1])) # -3.0
print(crt_torch.to_clika_dtype(torch.bfloat16) == crt.bfloat16) # True
print(crt_torch.to_torch_dtype(crt.bfloat16) == torch.bfloat16) # True
```
Two things the exchange does on purpose. A torch tensor that is not contiguous is made contiguous first (one copy on the torch side), so the runtime tensor shares that dense buffer and later writes through the strided source do not reach it. And the exchange never moves a tensor: `device=` on either function names where the result must live, and a device other than the tensor's own raises `ValueError` naming both; move the tensor first (`tensor.to("cuda:0")`, `clika_runtime.to(tensor, "cpu")`) and convert the result.
```python title="refusals.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime as crt
import clika_runtime.torch as crt_torch
try:
crt_torch.from_torch(torch.ones(4, 131).to(torch.float8_e4m3fn))
except TypeError as error:
print("float8_e4m3fn" in str(error)) # True
try:
crt_torch.from_torch(torch.ones(3), device="cuda:0")
except ValueError as error:
print("cpu" in str(error) and "cuda:0" in str(error)) # True
strided = torch.randn(5, 262)[:, ::2]
dense = crt_torch.from_torch(strided)
print(dense.is_contiguous(), tuple(dense.shape)) # True (5, 131)
```
On a machine where torch and the runtime both see a CUDA device (`torch.cuda.is_available()` and `crt.is_available("cuda")`), `from_torch(torch.randn(5, 131, device="cuda"))` lands on the runtime's `cuda:0` without a copy, and `to_torch` of that tensor reads the same device address.
`from_torch_state_dict(model.state_dict())` converts a checkpoint one tensor at a time and keeps tied weights tied: two entries over the same storage, offset, shape and strides come back as one runtime tensor under both names.
## Convert a module tree
`from_torch_module(module)` builds the `clika_runtime.nn` twin of a `torch.nn.Module` and loads its weights. Leaves convert by class (`Linear`, `Conv1d/2d/3d`, `ConvTranspose1d/2d/3d`, `Embedding`, `LayerNorm`, `RMSNorm`, the stateless activations, `Dropout`, `Identity`, `Flatten`) and the containers (`Sequential`, `ModuleList`, `ModuleDict`) keep their shape. The runtime layers are declared without storage first, then the weights bind through `load_state_dict(assign=True)` one tensor at a time, so each weight exists once: the converted module's parameters carry the same names, in the same order, at the same byte count as the torch module's.
```python title="convert.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime as crt
import clika_runtime.torch as crt_torch
torch.manual_seed(1)
model = torch.nn.Sequential(
torch.nn.Linear(131, 64), torch.nn.GELU(), torch.nn.LayerNorm(64), torch.nn.Linear(64, 16)
).eval()
converted = crt_torch.from_torch_module(model)
print(isinstance(converted, crt.nn.Module)) # True
print([name for name, _ in converted.named_parameters()] == [name for name, _ in model.named_parameters()]) # True
torch_bytes = sum(p.numel() * p.element_size() for p in model.parameters())
runtime_bytes = sum(p.nbytes for p in converted.parameters())
print(runtime_bytes == torch_bytes) # True: one copy of every weight, at its dtype
x = torch.randn(5, 131)
with torch.no_grad():
expected = model(x)
got = crt_torch.to_torch(converted(crt_torch.from_torch(x)))
print(torch.allclose(got, expected, rtol=1e-3, atol=1e-3)) # True
```
Tied weights stay one resident copy. A head whose `weight` is the embedding table binds the same runtime tensor into both slots, the converted `state_dict()` lists both keys, and `parameters()` counts the table once, as torch does:
```python title="tied.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime.torch as crt_torch
torch.manual_seed(3)
embed = torch.nn.Embedding(50, 64)
head = torch.nn.Linear(64, 50, bias=False)
head.weight = embed.weight
model = torch.nn.Sequential(embed, torch.nn.LayerNorm(64), head).eval()
converted = crt_torch.from_torch_module(model)
print(list(converted.state_dict()) == list(model.state_dict())) # True
distinct_torch = sum(p.numel() * p.element_size() for p in {id(p): p for p in model.parameters()}.values())
distinct_runtime = sum(p.nbytes for p in {id(p): p for p in converted.parameters()}.values())
print(distinct_runtime == distinct_torch) # True: the tied table is one resident copy
```
Convolution weights are permuted on the way in, from torch's channels-first layout (`OIHW`; `IOHW` for a transposed convolution) to the runtime's channels-last one (`OHWI`; `IHWO`), so a converted convolution consumes the runtime's channels-last activations directly. `device=` lands every converted weight on that device (each tensor crosses on the CPU and moves once on the runtime side; the torch module itself is never moved), and `dtype=` casts the floating-point weights after the move.
A leaf class the converter does not know raises `NotImplementedError` naming it; `strict=False` keeps such a module as a structure-only container whose children, parameters and buffers are converted under their names and whose own forward is left to the compiled path. `register_module_converter(name)` adds a converter for a class of your own.
```python title="strict.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime as crt
import clika_runtime.torch as crt_torch
class Odd(torch.nn.Module):
def forward(self, x: torch.Tensor) -> torch.Tensor:
return x.flip(0)
try:
crt_torch.from_torch_module(torch.nn.Sequential(torch.nn.Linear(131, 8), Odd()))
except NotImplementedError as error:
print("Odd" in str(error)) # True
lenient = crt_torch.from_torch_module(torch.nn.Sequential(torch.nn.Linear(131, 8), Odd()), strict=False)
print(isinstance(lenient, crt.nn.Module)) # True
```
## torch containers as pytrees
Importing `clika_runtime.torch` registers `torch.Size` (and the FX immutable list and dict) with the pytree registry, so a shape inside a nested input flattens to its integers and rebuilds as a `torch.Size`; `register_torch_pytree_nodes()` performs the same registration explicitly. With `transformers` installed, `register_hf_pytree_nodes()` registers its `ModelOutput` classes and caches, so a model output flattens to the fields that are set.
```python title="size_pytree.py"
try:
import torch
except ImportError:
print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")")
raise SystemExit(3)
import clika_runtime.torch as crt_torch
from clika_runtime import pytree
crt_torch.register_torch_pytree_nodes()
size = torch.Size([2, 3, 131])
leaves, spec = pytree.tree_flatten({"shape": size, "n": 1})
print(leaves) # [2, 3, 131, 1]
rebuilt = pytree.tree_unflatten(spec, leaves)
print(isinstance(rebuilt["shape"], torch.Size), rebuilt["shape"] == size) # True True
```
[Structure inputs and outputs as pytrees](pytrees.mdx) covers the registry these nodes join, and [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module surface the converted model lands on.
---
# Wrap existing memory without copying
Tensor::from_blob over buffers your application already owns: the borrow and adopt contracts, strided views, and what happens on a device move.
Source: https://docs.clika.io/clikart/how-to/wrap-existing-memory.md
Your data already sits in memory that some other part of the program owns: a decoder's output buffer, an arena, a mapped file, another library's array. `Tensor::from_blob` wraps that memory as a tensor with no copy; the tensor's data pointer is your pointer. What needs deciding is ownership, and the API makes the two contracts explicit:
- **No deleter passed: borrowed.** You keep ownership. The buffer must stay alive, and its layout unchanged, for as long as the tensor or any view of it is in use. Writes through the buffer are visible through the tensor and the other way around; it is the same memory.
- **Deleter passed: adopted.** The tensor takes ownership and calls `deleter(data)` once, when the last reference drops.
The Python variant borrows a numpy array's memory.
## Borrow a buffer and compute on it
```cpp title="borrow.cpp"
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
int main() {
std::vector buf(8, 1.0F); // memory the application owns
const Tensor view = Tensor::from_blob(buf.data(), {8}, DataType::Float32);
std::printf("shares memory: %s\n",
view.const_data_ptr() == buf.data() ? "yes (no copy)" : "no");
std::printf("sum = %g\n", ops::sum(view).item());
buf[0] = 100.0F; // write through the buffer...
std::printf("sum after buf[0] = 100: %g\n", ops::sum(view).item());
return 0;
}
```
```python title="main.py"
import numpy as np
import clika_runtime as crt
def main() -> None:
# Borrow: from_numpy wraps the array's own memory, no copy. The array
# and the tensor see the same bytes, mutation aliases both ways, and
# the tensor keeps the array alive. The array must be writable and
# contiguous; the refusals name the fix.
buf = np.zeros(4, dtype=np.float32)
view = crt.from_numpy(buf)
buf[0] = 100.0
print(f"sum after buf[0] = 100: {view.sum().item():g}")
# The other direction: numpy() on a CPU tensor is a zero-copy view of
# the tensor's memory; a device tensor asks you to move it first
# (t.to('cpu').numpy()).
t = crt.ones(2, 2)
arr = t.numpy()
print(f"shared bytes: {arr.sum():g}")
# A non-contiguous array does not borrow; the error names the fix.
try:
crt.from_numpy(np.zeros((4, 4), dtype=np.float32)[:, ::2])
except TypeError as e:
print(f"non-contiguous refused: {e}")
if __name__ == "__main__":
main()
```
`crt.tensor(array)` stays the copying entry when an independent tensor is
wanted.
```text
shares memory: yes (no copy)
sum = 8
sum after buf[0] = 100: 107
```
The borrow contract in one sentence: the runtime never reuses or overwrites borrowed memory, and in exchange you guarantee it outlives every tensor that sees it. A `vector` that reallocates (or a stack buffer that goes out of scope) under a live view is the bug this contract exists to name.
## Hand ownership over with a deleter
When the producer wants to fire and forget, pass a deleter. The tensor (and every tensor computed from it) keeps the buffer alive; the deleter runs exactly once, when the last reference drops, and it runs in your runtime: an exception it throws never crosses the library boundary. It also frees the runtime to reuse the buffer as scratch, which the borrow contract forbids.
```cpp title="adopt.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
namespace {
constexpr std::int64_t kCount = 8; // elements in the caller-owned buffer
} // namespace
int main() {
float* buf = static_cast(std::malloc(kCount * sizeof(float)));
for (std::int64_t i = 0; i < kCount; ++i) buf[i] = static_cast(i);
{
const Tensor adopted = Tensor::from_blob(
buf, {kCount}, DataType::Float32, {},
[](void* p) { std::printf("deleter: buffer released\n"); std::free(p); });
std::printf("mean = %g\n", ops::mean(adopted).item());
std::printf("leaving the tensor's scope...\n");
}
std::printf("scope closed\n");
return 0;
}
```
Handing ownership over with a deleter is part of the C++ API today; the
C++ tab shows it. The python borrow keeps the ARRAY as the owner:
`crt.from_numpy(array)` holds the array alive for the tensor's lifetime,
so no deleter changes hands.
```text
mean = 3.5
leaving the tensor's scope...
deleter: buffer released
scope closed
```
The deleter must not throw (a throw is swallowed). Adoption is the right contract at module boundaries: the producer allocates, the consumer wraps and forgets the allocation ever existed.
## Wrap non-contiguous memory with strides
The strided overload views memory that is not laid out contiguously, without rearranging a byte. Strides are in elements, one per dimension. A worked case: cropping a region of interest out of a pitched image buffer, the layout every camera API and GPU readback hands you (rows padded to a pitch wider than the image).
```cpp title="strided_roi.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::Tensor;
namespace ops = ClikaRT::ops;
namespace {
constexpr std::int64_t kRows = 4, kPitch = 6; // the whole image, row-major
constexpr std::int64_t kRoiRow = 1, kRoiCol = 2; // where the region starts
constexpr std::int64_t kRoiRows = 2, kRoiCols = 3; // its extent
} // namespace
int main() {
// A 4x6 single-channel image, row-major. buf[r][c] = r*10 + c.
std::vector buf(kRows * kPitch);
for (std::int64_t r = 0; r < kRows; ++r)
for (std::int64_t c = 0; c < kPitch; ++c) buf[r * kPitch + c] = static_cast(r * 10 + c);
// The 2x3 region starting at row 1, column 2: shape {2, 3}, and the
// ORIGINAL row pitch as the row stride. No pixel is copied.
const Tensor roi = Tensor::from_blob(buf.data() + kRoiRow * kPitch + kRoiCol,
{kRoiRows, kRoiCols}, {kPitch, 1}, DataType::Float32);
std::printf("roi = %s\n", roi.to_string().c_str());
std::printf("sum = %g (12+13+14+22+23+24 = 108)\n", ops::sum(roi).item());
return 0;
}
```
Wrapping strided memory with explicit strides is part of the C++ API
today; the C++ tab shows it. `crt.from_numpy` takes a C-contiguous
array; a strided source enters through `np.ascontiguousarray` (the
refusal names it).
```text
roi = Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[12, 13, 14, 22, 23, 24])
sum = 108 (12+13+14+22+23+24 = 108)
```
The same shape covers any pitched or tiled layout: a submatrix of a row-major matrix, a plane in a planar image, a batch entry inside a larger allocation. `ops::contiguous` materializes an owned compact copy when a consumer needs one.
## Device moves and pinned memory
`.to(device)` on a wrapped tensor behaves like on any other: on a discrete accelerator the move is a real transfer to device memory (the wrap saved the host-side copy, not the transfer), while unified-memory hardware moves for free. Two related notes. `from_data` is the copying cousin: it copies your bytes into an owned tensor so the source's lifetime stops mattering; take it when the buffer is short-lived and the tensor is not. And `from_blob`'s `pinned_for` parameter is tag-only: it asserts pages you already page-locked for a device, letting transfers take the pinned path; it cannot pin memory for you.
## Give the pool's idle reserve back
The runtime side of the memory story: the pool keeps memory it handed out and got back (`MemoryStats::cached_bytes`), so the next allocation is cheap. After a model unloads, or before a second model must fit beside the first, that idle reserve is memory the device (on a shared-memory part, the host) cannot use for anything else. `device::release_cached_memory(device)` returns it to the driver or the OS: deferred reservations drain, parked buffers retire and empty blocks release, the same three steps the runtime takes before it reports out-of-memory. Memory still in use, or whose last use has not completed on the device, is never touched; a later call can release more once that work retires. Automatic placement calls it on a failed accelerator before the CPU fallback loads.
```cpp title="release_cached.cpp"
#include
#include
using ClikaRT::DataType;
using ClikaRT::Device;
using ClikaRT::Tensor;
static void report(const char* when) {
const ClikaRT::device::MemoryStats s =
ClikaRT::device::memory_stats(Device::cpu());
std::printf("%-14s active %8.1f MiB, cached %8.1f MiB\n", when,
s.active_bytes / 1048576.0, s.cached_bytes / 1048576.0);
}
int main() {
{
// 256 MiB of Float32 work: the pool reserves real memory for it.
Tensor big = Tensor::zeros({64, 1024, 1024}, DataType::Float32);
big.synchronize();
report("in use:");
}
// The tensor is gone, but the pool keeps its bytes idle for reuse.
report("dropped:");
// Hand the idle reserve back to the system. Memory still in use, or
// whose last use has not completed, is never touched.
ClikaRT::device::release_cached_memory(Device::cpu());
report("released:");
return 0;
}
```
```text
in use: active 256.0 MiB, cached 0.0 MiB
dropped: active 0.0 MiB, cached 256.0 MiB
released: active 0.0 MiB, cached 0.0 MiB
```
The bundle's `compute` example covers device movement and this wrap in its `03_data_movement` and `06_zero_copy` chapters; [the custom-operator guide](write-a-custom-operator.mdx) uses `from_blob` to hand a hand-written kernel's output back to the runtime.
---
# Write a custom operator
Subclass nn::Module two ways: compose built-ins into a reusable block, or run your own kernel through the runtime with output_shapes and compute.
Source: https://docs.clika.io/clikart/how-to/write-a-custom-operator.md
The built-in `ops::` library does not have to be the end of the line. `nn::Module` is both the parameter container (PyTorch-style dotted names, so HuggingFace checkpoints bind by name) and the custom-op entry point: a leaf module that overrides `output_shapes()` and `compute()` runs its own kernel through the runtime, eager and tracing alike. This guide builds one of each.
## Compose built-ins into a module
A composite module subclasses `nn::Module`, registers its children in the constructor, and defines its own `forward` composing `ops::` and child modules. Registration is what buys the dotted names: the block below enumerates `up.weight` and `down.weight`, which is exactly how a checkpoint refers to them.
```cpp title="mlp_block.cpp"
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::nn::Linear;
using ClikaRT::NamedTensors;
using ClikaRT::Tensor;
namespace nn = ClikaRT::nn;
namespace ops = ClikaRT::ops;
// A residual MLP block: y = down(relu(up(x))) + x.
class MlpBlock final : public nn::Module {
public:
static std::shared_ptr make(std::int64_t dim, std::int64_t hidden) {
std::shared_ptr m(new MlpBlock());
m->up_ = Linear::make(dim, hidden);
m->down_ = Linear::make(hidden, dim);
m->register_module("up", m->up_);
m->register_module("down", m->down_);
return m;
}
Tensor forward(const Tensor& x) const {
Tensor h = this->up_->forward(x);
ops::relu_(h); // fresh from the projection: nothing else sees it
h = this->down_->forward(h);
ops::add_(h, x); // the residual add writes the fresh down output; x stays a read
return h;
}
private:
MlpBlock() = default;
std::shared_ptr up_;
std::shared_ptr down_;
};
int main() {
const std::shared_ptr block = MlpBlock::make(4, 8);
std::printf("parameters (dotted, checkpoint-shaped):\n");
for (const auto& [name, t] : block->named_parameters())
std::printf(" %-12s fake=%d\n", name.c_str(), static_cast(t.is_fake()));
// Bind weights by name, the way a checkpoint would.
NamedTensors ckpt;
ckpt.set("up.weight", Tensor::full({8, 4}, 0.1, DataType::Float32));
ckpt.set("down.weight", Tensor::full({4, 8}, 0.1, DataType::Float32));
block->load_state_dict(ckpt);
const Tensor x = Tensor::ones({1, 4}, DataType::Float32);
std::printf("y = %s\n", block->forward(x).to_string().c_str());
return 0;
}
```
```python title="mlp_block.py"
import clika_runtime as crt
import clika_runtime.nn as nn
class MlpBlock(nn.Module):
"""A residual MLP block: y = down(relu(up(x))) + x. A model is a Module
subclass: assigning a module REGISTERS it: load_state_dict,
named_parameters and repr all see `up` and `down` automatically. The
first linear fuses its ReLU as an activation epilogue, named by its
string spelling; without the epilogue the same block writes
F.relu(self.up(x)) with `import clika_runtime.nn.functional as F`.
Bias defaults to on, so these layers opt out explicitly."""
def __init__(self, dim: int, hidden: int) -> None:
super().__init__()
self.up = nn.Linear(dim, hidden, bias=False, activation="relu")
self.down = nn.Linear(hidden, dim, bias=False)
def forward(self, x: crt.Tensor) -> crt.Tensor:
return self.down(self.up(x)) + x
def main() -> None:
block = MlpBlock(4, 8)
print(block) # the module tree, registered names included
print("parameters (dotted, checkpoint-shaped):")
for name, t in block.named_parameters():
print(f" {name}")
# Weights arrive as a plain dict; dotted keys route to the registered
# submodules.
block.load_state_dict({
"up.weight": crt.full((8, 4), 0.1),
"down.weight": crt.full((4, 8), 0.1),
})
y = block(crt.ones(1, 4)) # calling the module runs forward
print(f"y = {y}")
if __name__ == "__main__":
main()
```
```text
parameters (dotted, checkpoint-shaped):
up.weight fake=1
down.weight fake=1
y = Tensor(shape=[1, 4], dtype=Float32, device=CPU, numel=4, data=[1.32, 1.32, 1.32, 1.32])
```
The math checks out by hand: `up` maps ones to 0.4 per unit, relu passes it, `down` sums 8 of them times 0.1 to 0.32, and the residual adds the input back, 1.32. Python prints a tensor as its values alone, so its last line reads `y = tensor([[1.3200, 1.3200, 1.3200, 1.3200]])`. `Linear::make` declares storage-free slots, so the parameters read as fake until `load_state_dict` binds them; `block->to(device)` moves the whole tree.
## Run your own kernel through the runtime
A leaf custom op overrides two virtuals. `output_shapes()` is the shape rule: it reads the inputs' metadata (`FakeTensor`: symbolic dims, dtype, placement) and returns the outputs' metadata. `compute()` is the kernel, with one accessor pair per residence: for a plain host loop over CPU-resident tensors, inputs read through `const_data_ptr()` and the pre-allocated outputs write through `mutable_data_ptr()` (both wait for the data and refuse a device tensor); a kernel on the op's device reads `device_const_data_ptr()` and writes `device_mutable_data_ptr()`, with no host wait and the pointer ordered on the op's stream. An exception thrown inside `compute()` surfaces as a typed error in the caller's runtime instead of crossing the library boundary. `forward` hands both to `dispatch()`, which runs the op through the runtime: eager mode calls `compute()` now, a tracing scope records the op from the shape rule alone. A leaf that dispatches must be owned by a `shared_ptr`.
```cpp title="softclip_op.cpp"
#include
#include
#include
#include
#include
#include
using ClikaRT::DataType;
using ClikaRT::FakeTensor;
using ClikaRT::Tensor;
namespace nn = ClikaRT::nn;
// Elementwise soft clip: y = x / (1 + |x|).
class SoftClip final : public nn::Module {
public:
static std::shared_ptr make() {
return std::shared_ptr(new SoftClip());
}
Tensor forward(const Tensor& x) const { return dispatch({&x, 1})[0]; }
// Shape rule: one output, same shape and dtype as the input.
std::vector output_shapes(
ClikaRT::Span inputs) const override {
return {FakeTensor(inputs[0].shape(), inputs[0].dtype(), inputs[0].stream())};
}
// Kernel: runs on the op's stream; here the CPU path, a plain loop over
// host pointers. A device kernel (a CUDA arm, say) reads
// `device_const_data_ptr()`, writes `device_mutable_data_ptr()`, and
// launches on `this->stream().native_handle()`; the host accessors below
// wait for the data and refuse a device tensor.
void compute(ClikaRT::Span inputs,
ClikaRT::Span outputs) const override {
const Tensor& in = inputs[0];
const float* x = static_cast(in.const_data_ptr());
float* y = static_cast(outputs[0].mutable_data_ptr());
const std::int64_t n = in.numel();
for (std::int64_t i = 0; i < n; ++i)
y[i] = x[i] / (1.0F + std::fabs(x[i]));
}
private:
SoftClip() = default;
};
int main() {
const std::shared_ptr clip = SoftClip::make();
std::vector v = {-9.0F, -1.0F, 0.0F, 1.0F, 9.0F};
const Tensor x = Tensor::from_data(v.data(), {5}, DataType::Float32);
std::printf("y = %s\n", clip->forward(x).to_string().c_str());
return 0;
}
```
```python title="softclip_op.py"
import clika_runtime as crt
import clika_runtime.nn as nn
class SoftClip(nn.Module):
"""Elementwise soft clip: y = x / (1 + |x|), composed from the built-in
operations; a custom module needs nothing beyond a forward. Writing a
custom KERNEL (your own shape rule and compute) is done through the C++
API; the C++ tab walks it."""
def forward(self, x: crt.Tensor) -> crt.Tensor:
return x / (1 + x.abs())
def main() -> None:
clip = SoftClip()
x = crt.tensor([-9.0, -1.0, 0.0, 1.0, 9.0])
print(f"y = {clip(x)}")
if __name__ == "__main__":
main()
```
```text
y = Tensor(shape=[5], dtype=Float32, device=CPU, numel=5, data=[-0.9, -0.5, 0, 0.5, 0.9])
```
The runtime allocated the output, placed it beside the input, and ran the kernel; the op composes with everything else (`ops::` calls before and after, `on_complete`, the scopes) because it went through `dispatch` like a built-in. Python's line reads `y = tensor([-0.9000, -0.5000, 0.0000, 0.5000, 0.9000])`: the same values in its own tensor form.
## Launching on an accelerator
`compute()` runs on whatever device its tensors live on. Inside it, `this->stream()` is the op's stream: `stream().device()` says where you are, and `stream().native_handle()` is the backend's queue as an opaque pointer, `cudaStream_t` on CUDA (`nullptr` on the CPU backend, which has no device queue). The launch pattern, from the `nn/module.h` contract:
```text
Stream s = this->stream(); // the op's stream
auto cu = static_cast(s.native_handle()); // the CUDA queue
const float* q = static_cast(inputs[0].device_const_data_ptr());
float* o = static_cast(outputs[0].device_mutable_data_ptr());
my_kernel<<>>(q, ..., o); // your kernel
```
The device accessors return the device address with no host wait, ordered on the op's stream, so the launch above is valid as written; the host accessors (`const_data_ptr()`, `mutable_data_ptr()`) are for CPU-resident tensors and would refuse here.
Size the launch from `get_device_properties(stream().device())`, and take stream-ordered scratch from `stream().allocate(shape, dtype)`, which recycles safely on that stream only. The handle is owned by the runtime: never destroy it, and keep the `Stream` alive while using it. The bundle's `flash_attention` example is the complete worked case, a hand-written CUDA attention kernel dispatched through this exact interface.
A rule of thumb for choosing the shape: compose built-ins when the math decomposes into `ops::` (the runtime already fuses and places them); write a leaf when you have a kernel the library does not, and keep its `output_shapes` honest, because tracing trusts it without running `compute`.
---
# Write a transform
Write graph transforms of your own: functions over the ModelGraph edits that optimize() runs beside the runtime's transforms, operators built with add_node, insert and replace, and rewrite rules that replace each occurrence of a Pattern.
Source: https://docs.clika.io/clikart/how-to/write-a-transform.md
{/* Every block is a program under examples//howto/write_a_transform/: the first
block of each tab is get_a_graph whole, and every later block is its program's docs
region. tools/tutorial_check.py runs each program against its recorded output. */}
A transform rewrites a `ModelGraph` inside `optimize()`. The runtime ships its own transforms, and you can write more. A transform you write is a function over the graph edits that [Edit a graph](edit-a-graph.mdx) covers, and `optimize()` runs it beside the runtime's transforms in one list, at its place, once per iteration. Three more calls build operators inside a transform: `add_node` adds one operator, and `insert` and `replace` add the operators that a function over the public ops records. A rewrite rule is a transform that finds each occurrence of a `Pattern` and replaces it. The calls are available from C++ and Python, under the same names.
## A graph to transform
The model below computes `y = Relu(Relu(x)) + Neg(x)`, a Relu over a Relu and a Neg that the Add reads, and the examples on this page rewrite both. `trace_model` traces a function over an `x` of shape [2, 3], the model unless it is given another, and returns the graph as built. Each program prints a graph as the formula it returns. `formula` names a value by its operator and the values that operator reads, a graph input by its name and a constant as `c`, and `formula_of` gives the formula of the graph's output, which the node names and their order do not change. `rows` lists a report's rows as each transform's name and applications, and in C++ `optimize_with` runs a list through `OptimizeOptions::transforms`. `drop_repeated_relu` is the first transform on this page. It bypasses each Relu that reads a Relu with `bypass_node`, so that Relu's readers read the one before it.
The program runs `remove_redundant_relu`, one of the runtime's transforms, in a list of its own, as [Optimize and finalize](optimize-and-finalize.mdx#run-a-list-of-your-own) shows. It leaves one Relu where the model has two, and its row counts one application.
```cpp title="get_a_graph.cpp"
#include
#include
#include
#include
#include
#include
#include
#include
#include
using ClikaRT::Error;
using ClikaRT::Result;
using ClikaRT::Tensor;
using ClikaRT::graph::ModelGraph;
using ClikaRT::graph::Node;
using ClikaRT::graph::NodeKind;
using ClikaRT::graph::OpCode;
using ClikaRT::graph::OptimizeOptions;
using ClikaRT::graph::OptimizeReport;
using ClikaRT::graph::Pattern;
using ClikaRT::graph::RewriteMatch;
using ClikaRT::graph::Transform;
using ClikaRT::graph::TransformReport;
using ClikaRT::graph::Value;
namespace ops = ClikaRT::ops;
namespace transforms = ClikaRT::graph::transforms;
namespace {
// y = Relu(Relu(x)) + Neg(x): a Relu over a Relu, and a Neg that the Add reads.
std::vector model(const std::vector& inputs) {
const Tensor rectified = ops::relu(ops::relu(inputs[0]));
const Tensor negated = ops::neg(inputs[0]);
return {ops::add(rectified, negated)};
}
// The formula a value computes: its operator over the values it reads, a graph input by its name, and a
// constant as c.
std::string formula(const Value& value) {
const std::optional producer = value.producer();
if (!producer.has_value()) return "c";
if (producer->kind() == NodeKind::Input) return producer->name();
std::string reads;
for (const Value& read : producer->inputs()) reads += (reads.empty() ? "" : ", ") + formula(read);
return std::string(ClikaRT::graph::op_code_name(producer->op_code())) + "(" + reads + ")";
}
// The formula of the value the graph returns.
std::string formula_of(const ModelGraph& graph) {
for (const Node& node : graph.nodes()) {
for (const Value& value : node.outputs()) {
if (value.is_graph_output()) return formula(value);
}
}
return "";
}
// A trace returns the graph as built: every operator as written, nothing optimized or finalized.
ModelGraph trace_model(const ClikaRT::graph::TraceFunction& fn = model) {
const std::vector signature = {{"x", ClikaRT::DataType::Float32, {2, 3}}};
const std::vector outputs = {"y"};
return ClikaRT::graph::trace(fn, signature, "transform", outputs);
}
// optimize() over `list`, in order.
OptimizeReport optimize_with(ModelGraph& graph, std::vector list) {
OptimizeOptions options;
options.transforms = std::move(list);
return graph.optimize(options);
}
// Each row of an optimize() report, as "name applications", separated by commas.
std::string rows(const OptimizeReport& report) {
std::string out;
for (const TransformReport& row : report.transforms) {
out += (out.empty() ? "" : ", ") + row.transform.name() + " " + std::to_string(row.applications);
}
return out;
}
// A Relu that reads a Relu is bypassed: its readers read the first Relu instead.
Result drop_repeated_relu(ModelGraph& graph) {
for (const Node& relu : graph.find_nodes(OpCode::Relu)) {
const std::optional producer = relu.input(0)->producer();
if (producer.has_value() && producer->op_code() == OpCode::Relu) graph.bypass_node(relu);
}
return {};
}
} // namespace
int main() {
ModelGraph graph = trace_model();
std::printf("%s\n", formula_of(graph).c_str()); // Add(Relu(Relu(x)), Neg(x))
// optimize() runs a list of transforms. This one of the runtime's removes the repeated Relu.
const OptimizeReport report = optimize_with(graph, {transforms::remove_redundant_relu()});
std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str());
// Add(Relu(x), Neg(x)) | remove_redundant_relu 1
return 0;
}
```
```python title="get_a_graph.py"
from collections.abc import Callable
import clika_runtime as crt
from clika_runtime.graph import NodeKind, OpCode, Pattern, RewriteMatch, Transform, transforms
def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
x = inputs[0]
return [crt.relu(crt.relu(x)) + crt.neg(x)] # y = Relu(Relu(x)) + Neg(x)
# The formula a value computes: its operator over the values it reads, a graph input by its name, and a
# constant as c.
def formula(value: crt.graph.Value) -> str:
producer = value.producer()
if producer is None:
return "c"
if producer.kind == NodeKind.Input:
return producer.name
return f"{producer.op_code.name}({', '.join(formula(read) for read in producer.inputs)})"
# The formula of the value the graph returns.
def formula_of(graph: crt.graph.ModelGraph) -> str:
(returned,) = [value for node in graph.nodes() for value in node.outputs if value.is_graph_output()]
return formula(returned)
# A trace returns the graph as built: every operator as written, nothing optimized or finalized.
def trace_model(fn: Callable[[list[crt.Tensor]], list[crt.Tensor]] = model) -> crt.graph.ModelGraph:
return crt.trace(fn, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph
# Each row of an optimize() report, as (name, applications).
def rows(report: crt.graph.OptimizeReport) -> list[tuple[str, int]]:
return [(row.transform.name, row.applications) for row in report.transforms]
# A Relu that reads a Relu is bypassed: its readers read the first Relu instead.
def drop_repeated_relu(graph: crt.graph.ModelGraph) -> None:
for relu in graph.find_nodes(OpCode.Relu):
producer = relu.input(0).producer()
if producer is not None and producer.op_code == OpCode.Relu:
graph.bypass_node(relu)
graph = trace_model()
print(formula_of(graph)) # Add(Relu(Relu(x)), Neg(x))
# optimize() runs a list of transforms. This one of the runtime's removes the repeated Relu.
report = graph.optimize([transforms.remove_redundant_relu()])
print(formula_of(graph), rows(report)) # Add(Relu(x), Neg(x)) [('remove_redundant_relu', 1)]
```
Write a transform when a rewrite should run inside `optimize()`, beside the runtime's transforms and at every iteration, so it sees what their rewrites expose and they see what it changes.
Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, the helpers, `drop_repeated_relu` and, in Python, the trace).
## Write a transform
`Transform::from_function(name, fn)` in C++ and `Transform(name, fn)` in Python make a transform from a function. `fn` gets the graph that `optimize()` is running and edits it with the graph edits. In C++ it returns a `Result`, and a failed one ends the run. In Python it returns None. The `@transforms.transform` decorator makes a transform from the function it decorates, named after the function, and `@transforms.transform(name, runs_after=...)` sets the name and the order. The name is how the transform reads in the report and in errors, so it may be neither empty nor the name of one of the runtime's transforms.
`optimize()` runs each entry of the list at its place, once per iteration, and stops at the first iteration that changes nothing. Each entry has one row in the report, and its `applications` counts the iterations in which the transform changed the graph. A transform changes the graph when one of its edits does, so `count_nodes` below, which only reads the graph, counts none, though it runs in both iterations and records the node count each time. A handle you made equals itself and the copies the report holds.
```cpp title="a_transform.cpp"
int main() {
ModelGraph graph = trace_model();
// A transform you write: a name, and the function optimize() calls at its place in the list.
const Transform drop = Transform::from_function("drop_repeated_relu", drop_repeated_relu);
// One that reads the graph and changes nothing: it records the node count each time it runs.
std::vector counts;
const Transform count_nodes = Transform::from_function("count_nodes", [&counts](ModelGraph& g) -> Result {
counts.push_back(g.nodes().size());
return {};
});
const OptimizeReport report = optimize_with(graph, {count_nodes, drop});
std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str());
// Add(Relu(x), Neg(x)) | count_nodes 0, drop_repeated_relu 1
std::printf("%lld | %zu %zu\n", static_cast(report.iterations), counts.at(0), counts.at(1));
// 2 | 5 4: count_nodes ran in both iterations
std::printf("%s\n", report.transforms.at(1).transform == drop ? "true" : "false"); // true
return 0;
}
```
```python title="a_transform.py"
drop = Transform("drop_repeated_relu", drop_repeated_relu) # a name, and the function optimize() calls
counts: list[int] = []
@transforms.transform # a transform named after the function it decorates
def count_nodes(g: crt.graph.ModelGraph) -> None:
counts.append(len(g.nodes())) # it reads the graph and changes nothing
report = graph.optimize([count_nodes, drop])
print(formula_of(graph), rows(report))
# Add(Relu(x), Neg(x)) [('count_nodes', 0), ('drop_repeated_relu', 1)]
print(report.iterations, counts) # 2 [5, 4]: count_nodes ran in both iterations
print(report.transforms[1].transform == drop) # True
```
Write a transform as a function when the rewrite reads more of the graph than a pattern states, or edits parts of it that a rule does not, such as the graph's inputs and outputs.
## Order the transforms
`runs_after` lists the transforms, the runtime's or yours, that a transform must follow in any list that holds both. It is the third argument of `Transform::from_function` and the `runs_after=` keyword of `Transform` and of the decorator. `optimize()` checks the list before anything runs. A list that places a transform ahead of one it must follow is refused with the code name `INVALID_ARGUMENT`, and the message names both entries and the order to use. A transform the list does not hold asks nothing. Here `drop_repeated_relu` follows `remove_double_neg`, which can expose a Relu over a Relu.
```cpp title="order.cpp"
int main() {
ModelGraph graph = trace_model();
// remove_double_neg can expose a Relu over a Relu (Relu(Neg(Neg(Relu(x))))), so drop_repeated_relu runs after it.
const Transform drop =
Transform::from_function("drop_repeated_relu", drop_repeated_relu, {transforms::remove_double_neg()});
try {
optimize_with(graph, {drop, transforms::remove_double_neg()}); // a list is checked before anything runs
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// INVALID_ARGUMENT | optimize: transforms[0] (drop_repeated_relu) must run after remove_double_neg, which the list places later, at transforms[1]; list remove_double_neg ahead of it
}
const OptimizeReport report = optimize_with(graph, {transforms::remove_double_neg(), drop});
std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str());
// Add(Relu(x), Neg(x)) | remove_double_neg 0, drop_repeated_relu 1
return 0;
}
```
```python title="order.py"
# remove_double_neg can expose a Relu over a Relu (Relu(Neg(Neg(Relu(x))))), so drop_repeated_relu runs after it.
drop = Transform("drop_repeated_relu", drop_repeated_relu, runs_after=[transforms.remove_double_neg()])
try:
graph.optimize([drop, transforms.remove_double_neg()]) # a list is checked before anything runs
except crt.InvalidArgumentError as error:
print(error.code_name, "|", error)
# INVALID_ARGUMENT | optimize: transforms[0] (drop_repeated_relu) must run after remove_double_neg, which the list places later, at transforms[1]; list remove_double_neg ahead of it
report = graph.optimize([transforms.remove_double_neg(), drop])
print(formula_of(graph), rows(report))
# Add(Relu(x), Neg(x)) [('remove_double_neg', 0), ('drop_repeated_relu', 1)]
```
Declare the order when a transform depends on another one's result, so that no list can run them the other way.
## What a transform may not do
Inside a transform, its graph serves the graph edits and every query. It refuses `optimize()`, `finalize()`, `to()` and `attach_kv_cache()` with the code name `FAILED_PRECONDITION`, since `optimize()` is still working on it. The graph's inputs and outputs change only through their own edits (`add_input`, `add_output`, `remove_output`, `rename_input` and `rename_output`), and every other edit keeps them as they are. The graph also refuses an edit from another thread, and moving from the graph or assigning over it leaves both graphs as they were.
A transform that fails ends the run, and the graph is as `optimize()` found it, with the changes of every transform in the run undone. In C++ a transform fails by returning a failed `Result` or by throwing. `optimize()` raises that failure with its status and its code name, and the message starts with the entry that failed, as in `optimize: transforms[1] (no_neg) failed:`. In Python, `optimize()` raises the exception the transform raised, as that same object. The finalize refusal comes back as `finalize()` raised it, and an exception class of your own comes back as itself.
```cpp title="refusals.cpp"
int main() {
ModelGraph graph = trace_model();
// Inside a transform, its own graph refuses finalize(): optimize() is still working on the graph.
const Transform finalizes = Transform::from_function("finalizes", [](ModelGraph& g) -> Result {
g.finalize();
return {};
});
try {
optimize_with(graph, {finalizes});
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// FAILED_PRECONDITION | optimize: transforms[0] (finalizes) failed: finalize: the transform 'finalizes' is running on this graph inside optimize(); call finalize() after optimize() returns
}
// A failure ends the run: optimize() raises it with its status and code, and every change of the run is undone.
const Transform no_neg = Transform::from_function("no_neg", [](ModelGraph& g) -> Result {
if (g.find_nodes(OpCode::Neg).empty()) return {};
return Result(ClikaRT::Status::Unsupported, "the graph still computes a Neg", "NEG_NOT_SERVED");
});
try {
optimize_with(graph, {Transform::from_function("drop_repeated_relu", drop_repeated_relu), no_neg});
} catch (const Error& error) {
std::printf("%s | %s\n", error.code_name().c_str(), error.what());
// NEG_NOT_SERVED | optimize: transforms[1] (no_neg) failed: the graph still computes a Neg
}
std::printf("%s\n", formula_of(graph).c_str()); // Add(Relu(Relu(x)), Neg(x)): the bypassed Relu is back
return 0;
}
```
```python title="refusals.py"
# Inside a transform, its own graph refuses finalize(): optimize() is still working on the graph.
@transforms.transform
def finalizes(g: crt.graph.ModelGraph) -> None:
g.finalize()
try:
graph.optimize([finalizes])
except crt.InvalidArgumentError as error: # the refusal finalize() raised, as it raised it
print(error.code_name, "|", error)
# FAILED_PRECONDITION | finalize: the transform 'finalizes' is running on this graph inside optimize(); call finalize() after optimize() returns
class NegNotServed(Exception):
"""The target the graph is built for runs no Neg."""
# An exception a transform raises ends the run: optimize() raises that same exception, and every change of the
# run is undone.
@transforms.transform
def no_neg(g: crt.graph.ModelGraph) -> None:
if g.find_nodes(OpCode.Neg):
raise NegNotServed("the graph still computes a Neg")
try:
graph.optimize([Transform("drop_repeated_relu", drop_repeated_relu), no_neg])
except NegNotServed as error:
print(type(error).__name__, "|", error) # NegNotServed | the graph still computes a Neg
print(formula_of(graph)) # Add(Relu(Relu(x)), Neg(x)): the Relu that drop_repeated_relu bypassed is back
```
Fail a transform when it meets a graph it cannot handle: the run ends, and the graph stays as `optimize()` found it.
## Build one operator with add_node
`add_node(code, inputs, attributes)` adds an operator of `code`, built the way its public op builds it, so the op's own checks apply. `inputs` holds its values, one per input port in the op's order, with `std::nullopt` (None in Python) for an optional operand left out. `attributes` holds its settings under the names a node's attributes report (`Node::attributes()` in C++, `node.attributes` in Python). A setting left out takes the op's default, an integer also serves a number, and an enumerated setting takes its enumerator's name. Nothing reads the new operator's outputs until an edit wires them in, and `finalize()` drops a node that nothing reads by then. A code that no single public op call builds, such as an operator with several outputs or a quantized one, is refused, and `insert` builds it from the public ops that compute it.
Here `fold_neg_into_sub` turns `a + Neg(b)` into `a - b`. The Sub takes the Add's own settings, `alpha` and `activation`, which the two ops name alike. `replace_all_uses_with` moves the Add's reads to the Sub, the graph's output among them, and the Add and the Neg go.
```cpp title="add_node.cpp"
int main() {
ModelGraph graph = trace_model();
// a + Neg(b) is a - b: a Sub with the Add's own settings takes the Add's place, and the Neg goes once nothing
// reads it.
const Transform fold = Transform::from_function("fold_neg_into_sub", [](ModelGraph& g) -> Result {
for (const Node& add : g.find_nodes(OpCode::Add)) {
const std::optional neg = add.input(1)->producer();
if (!neg.has_value() || neg->op_code() != OpCode::Neg) continue;
// The Sub reads a and b, with the Add's settings (alpha, activation) under their own names.
const Node sub = g.add_node(OpCode::Sub, {add.input(0), neg->input(0)}, add.attributes());
g.replace_all_uses_with(*add.output(0), *sub.output(0));
g.remove_node(add);
if (neg->out_degree() == 0) g.remove_node(*neg);
}
return {};
});
const OptimizeReport report = optimize_with(graph, {fold});
std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str());
// Sub(Relu(Relu(x)), x) | fold_neg_into_sub 1
return 0;
}
```
```python title="add_node.py"
# a + Neg(b) is a - b: a Sub with the Add's own settings takes the Add's place, and the Neg goes once nothing
# reads it.
@transforms.transform
def fold_neg_into_sub(g: crt.graph.ModelGraph) -> None:
for add in g.find_nodes(OpCode.Add):
neg = add.input(1).producer()
if neg is None or neg.op_code != OpCode.Neg:
continue
sub = g.add_node(OpCode.Sub, [add.input(0), neg.input(0)], add.attributes) # alpha and activation
g.replace_all_uses_with(add.output(0), sub.output(0))
g.remove_node(add)
if neg.out_degree() == 0:
g.remove_node(neg)
report = graph.optimize([fold_neg_into_sub])
print(formula_of(graph), rows(report)) # Sub(Relu(Relu(x)), x) [('fold_neg_into_sub', 1)]
```
Build with `add_node` when the new operator is one call of a public op, with settings you read off the graph or choose yourself.
## Splice in a function with insert and replace
`insert(fn, inputs)` traces `fn`, a function over the public ops, over one tensor per input, each with that value's dtype, device and shape, and adds the operators it records, reading the values given. In C++ `fn` returns a `std::vector`, and in Python a Tensor or a sequence of Tensors. `insert` returns the values `fn` returns, which nothing reads until an edit wires them in, as `add_node`'s outputs are wired. `replace(old_outputs, fn, inputs)` is `insert` followed by the wiring and the removal: every read of `old_outputs[i]` reads new value `i`, the graph's outputs included, and the operators that computed the old outputs from the inputs go, with every producer only they read. An operator an input is computed from stays. `fn` computes values, so it writes no input in place and does not edit the graph.
Here two transforms put one Relu in place of a Relu over a Relu. `with_insert` wires the new Relu in with `replace_all_uses_with` and removes the pair itself, and `with_replace` does all of that in one call. Each rewrites one pair per call and returns, since its edits remove nodes that its loop still holds, and the next iteration finds the next pair.
```cpp title="insert_replace.cpp"
int main() {
// Relu(Relu(v)) is Relu(v): a function over the public ops.
const ClikaRT::graph::TraceFunction relu_of = [](const std::vector& in) {
return std::vector{ops::relu(in[0])};
};
// insert() adds the operators relu_of records over the values given, and returns their values, which nothing
// reads until an edit wires them in.
const Transform with_insert =
Transform::from_function("with_insert", [&relu_of](ModelGraph& g) -> Result {
for (const Node& outer : g.find_nodes(OpCode::Relu)) {
const std::optional inner = outer.input(0)->producer();
if (!inner.has_value() || inner->op_code() != OpCode::Relu) continue;
const std::vector made = g.insert(relu_of, {*inner->input(0)});
g.replace_all_uses_with(*outer.output(0), made.at(0));
g.remove_node(outer);
g.remove_node(*inner);
return {}; // one pair per call: the next iteration finds the next one
}
return {};
});
// replace() does all of that in one call: the operators that compute the old output from the inputs go, the
// outer Relu and the inner one only it reads.
const Transform with_replace =
Transform::from_function("with_replace", [&relu_of](ModelGraph& g) -> Result {
for (const Node& outer : g.find_nodes(OpCode::Relu)) {
const std::optional inner = outer.input(0)->producer();
if (!inner.has_value() || inner->op_code() != OpCode::Relu) continue;
g.replace({*outer.output(0)}, relu_of, {*inner->input(0)});
return {};
}
return {};
});
ModelGraph graph = trace_model();
ModelGraph other = trace_model();
const std::string inserted = rows(optimize_with(graph, {with_insert}));
const std::string replaced = rows(optimize_with(other, {with_replace}));
std::printf("%s | %s\n", formula_of(graph).c_str(), inserted.c_str()); // Add(Relu(x), Neg(x)) | with_insert 1
std::printf("%s | %s\n", formula_of(other).c_str(), replaced.c_str()); // Add(Relu(x), Neg(x)) | with_replace 1
return 0;
}
```
```python title="insert_replace.py"
def relu_of(tensors: list[crt.Tensor]) -> crt.Tensor:
return crt.relu(tensors[0]) # Relu(Relu(v)) is Relu(v): a function over the public ops
# insert() adds the operators relu_of records over the values given, and returns their values, which nothing
# reads until an edit wires them in.
@transforms.transform
def with_insert(g: crt.graph.ModelGraph) -> None:
for outer in g.find_nodes(OpCode.Relu):
inner = outer.input(0).producer()
if inner is not None and inner.op_code == OpCode.Relu:
(made,) = g.insert(relu_of, [inner.input(0)])
g.replace_all_uses_with(outer.output(0), made)
g.remove_node(outer)
g.remove_node(inner)
return # one pair per call: the next iteration finds the next one
# replace() does all of that in one call: the operators that compute the old output from the inputs go, the
# outer Relu and the inner one only it reads.
@transforms.transform
def with_replace(g: crt.graph.ModelGraph) -> None:
for outer in g.find_nodes(OpCode.Relu):
inner = outer.input(0).producer()
if inner is not None and inner.op_code == OpCode.Relu:
g.replace([outer.output(0)], relu_of, [inner.input(0)])
return
other = trace_model()
inserted = rows(graph.optimize([with_insert]))
replaced = rows(other.optimize([with_replace]))
print(formula_of(graph), inserted) # Add(Relu(x), Neg(x)) [('with_insert', 1)]
print(formula_of(other), replaced) # Add(Relu(x), Neg(x)) [('with_replace', 1)]
```
Use `replace` when the new values are a computation over values the graph already has, and `insert` when you wire the result in yourself.
## Write a rule
A rule rewrites the occurrences of a `Pattern`, which names the operators to find, one pattern node each, and the edges between them. `Pattern::chain` (in Python, `Pattern.chain` or `Pattern([...])`) builds a straight line, and `add_node` and `add_edge` build any other shape, as the next section does. `transforms::rewrite(name, pattern, replacement, condition)` in C++ and `transforms.rewrite(name, pattern, replacement, condition)` in Python make the rule, a transform like any other. Each iteration searches the graph for the pattern once and takes the occurrences in the order the search returns them. The rule replaces each occurrence its condition accepts, every one when there is no condition, with the values `replacement(match, tensors)` computes from one traced tensor per value the occurrence reads, traced as `replace` traces its function. Every occurrence is rewritten in the same iteration, so the rule's row counts that iteration once, and the next iteration searches again, so an occurrence a replacement creates is rewritten then. In Python, the `@transforms.rule(pattern, condition=..., name=..., runs_after=...)` decorator makes a rule from the replacement it decorates, named after the function unless given a name.
The match is a `RewriteMatch`. `nodes` holds the matched node for each pattern node, in the order the nodes were added, with `std::nullopt` (None in Python) for an optional node the occurrence lacks. `inputs` holds the values the occurrence reads from outside it, each once, and the replacement gets one tensor per entry, in that order. `outputs` holds the values it makes that a node outside it reads or the graph returns, and the replacement returns one new value per entry. Here the replacement also records what it reads of each occurrence.
```cpp title="rules.cpp"
int main() {
// y = Relu(Relu(x)) + Relu(Relu(Neg(x))): two Relus over a Relu.
ModelGraph pairs = trace_model([](const std::vector& in) {
const Tensor first = ops::relu(ops::relu(in[0]));
const Tensor second = ops::relu(ops::relu(ops::neg(in[0])));
return std::vector{ops::add(first, second)};
});
std::printf("%s\n", formula_of(pairs).c_str()); // Add(Relu(Relu(x)), Relu(Relu(Neg(x))))
std::vector seen; // what the replacement reads of each occurrence
const OpCode pair[] = {OpCode::Relu, OpCode::Relu};
const Transform collapse_relu = transforms::rewrite(
"collapse_relu", Pattern::chain(pair), [&seen](const RewriteMatch& match, const std::vector& in) {
seen.push_back(std::to_string(match.nodes.size()) + " nodes, reads " + formula(match.inputs.at(0)) +
", " + std::to_string(match.outputs.size()) + " output");
return std::vector{ops::relu(in[0])}; // one Relu over the value the pair reads
});
const std::string report = rows(optimize_with(pairs, {collapse_relu}));
std::sort(seen.begin(), seen.end());
for (const std::string& occurrence : seen) std::printf("%s\n", occurrence.c_str());
// 2 nodes, reads Neg(x), 1 output
// 2 nodes, reads x, 1 output
std::printf("%s | %s\n", formula_of(pairs).c_str(), report.c_str());
// Add(Relu(x), Relu(Neg(x))) | collapse_relu 1
return 0;
}
```
```python title="rules.py"
# y = Relu(Relu(x)) + Relu(Relu(Neg(x))): two Relus over a Relu.
def two_pairs(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
x = inputs[0]
return [crt.relu(crt.relu(x)) + crt.relu(crt.relu(crt.neg(x)))]
seen: list[tuple[list[str], list[str], int]] = [] # what the replacement reads of each occurrence
def one_relu(match: RewriteMatch, tensors: list[crt.Tensor]) -> crt.Tensor:
nodes = [node.op_code.name for node in match.nodes]
seen.append((nodes, [formula(value) for value in match.inputs], len(match.outputs)))
return crt.relu(tensors[0]) # one Relu over the value the pair reads
collapse_relu = transforms.rewrite("collapse_relu", Pattern.chain([OpCode.Relu, OpCode.Relu]), one_relu)
pairs = trace_model(two_pairs)
print(formula_of(pairs)) # Add(Relu(Relu(x)), Relu(Relu(Neg(x))))
report = pairs.optimize([collapse_relu])
print(sorted(seen)) # [(['Relu', 'Relu'], ['Neg(x)'], 1), (['Relu', 'Relu'], ['x'], 1)]
print(formula_of(pairs), rows(report)) # Add(Relu(x), Relu(Neg(x))) [('collapse_relu', 1)]
```
Write a rule when the rewrite is local: a fixed arrangement of operators, a condition on the match, and a replacement computed from the values it reads.
## The occurrences a rule skips
A rule skips an occurrence, and its condition never sees it, in three cases: an earlier replacement in the same iteration removed one of its nodes, the occurrence is not convex, or nothing outside it reads a value it makes. An occurrence is not convex when a path leaves it and comes back, a path no replacement can keep, and `ModelGraph::is_convex` answers the same question. Every other occurrence goes to `condition(match)`, which decides whether to replace it, by the truth of its result in Python. With no condition, the rule replaces every occurrence.
Which occurrence an earlier replacement removes depends on the order the search returns them, so no program here shows that case. The program below builds its pattern node by node, the Add as the root and the Neg feeding its second operand through `add_edge`, and counts the condition's calls. The condition accepts an Add with its default settings whose value is the one the occurrence hands on. On the page's graph the condition is asked once and the fold runs. On `n = Neg(x); y = Relu(n) + n`, the path from the Neg through the Relu leaves the occurrence and comes back, so the rule skips it and the condition is never asked.
```cpp title="skipped_occurrences.cpp"
int main() {
// Add(a, Neg(b)), built node by node: the Add is the root, and the Neg feeds its second operand.
Pattern pattern;
const Pattern::NodeId add = pattern.add_node(OpCode::Add);
const Pattern::NodeId neg = pattern.add_node(OpCode::Neg);
pattern.add_edge(neg, add, 0, 1); // the Neg's output 0 feeds the Add's input 1
int asked = 0; // how many occurrences the condition is asked about
// a + Neg(b) is a - b for an Add with its default settings, when its value is the one the occurrence hands on.
const auto plain_sum = [&asked](const RewriteMatch& match) {
++asked;
const Node& total = *match.nodes[0];
const std::optional alpha = total.attribute("alpha");
const double* scale = alpha.has_value() ? std::get_if(&alpha->value) : nullptr;
return match.outputs.size() == 1 && scale != nullptr && *scale == 1.0 &&
total.fused_activation() == ops::Activation::Identity;
};
const Transform fold_neg_into_sub = transforms::rewrite(
"fold_neg_into_sub", std::move(pattern),
[](const RewriteMatch&, const std::vector& in) {
return std::vector{ops::sub(in[0], in[1])}; // match.inputs holds a, then b
},
plain_sum);
ModelGraph graph = trace_model();
std::string report = rows(optimize_with(graph, {fold_neg_into_sub}));
std::printf("%s | %s | asked %d\n", formula_of(graph).c_str(), report.c_str(), asked);
// Sub(Relu(Relu(x)), x) | fold_neg_into_sub 1 | asked 1
// n = Neg(x), y = Relu(n) + n: the path from the Neg through the Relu leaves the occurrence and comes back in.
asked = 0;
ModelGraph looped = trace_model([](const std::vector& in) {
const Tensor n = ops::neg(in[0]);
return std::vector{ops::add(ops::relu(n), n)};
});
report = rows(optimize_with(looped, {fold_neg_into_sub}));
std::printf("%s | %s | asked %d\n", formula_of(looped).c_str(), report.c_str(), asked);
// Add(Relu(Neg(x)), Neg(x)) | fold_neg_into_sub 0 | asked 0
return 0;
}
```
```python title="skipped_occurrences.py"
# Add(a, Neg(b)), built node by node: the Add is the root, and the Neg feeds its second operand.
pattern = Pattern()
add = pattern.add_node(OpCode.Add)
neg = pattern.add_node(OpCode.Neg)
pattern.add_edge(neg, add, to_port=1)
asked: list[str] = [] # the Add of each occurrence the condition is asked about
# a + Neg(b) is a - b for an Add with its default settings, when its value is the one the occurrence hands on.
def plain_sum(match: RewriteMatch) -> bool:
total = match.nodes[0]
asked.append(total.name)
return (len(match.outputs) == 1 and total.attribute("alpha") == 1
and total.attribute("activation") == "Identity")
@transforms.rule(pattern, condition=plain_sum)
def fold_neg_into_sub(match: RewriteMatch, tensors: list[crt.Tensor]) -> crt.Tensor:
return tensors[0] - tensors[1] # match.inputs holds a, then b
report = graph.optimize([fold_neg_into_sub])
print(formula_of(graph), rows(report), len(asked)) # Sub(Relu(Relu(x)), x) [('fold_neg_into_sub', 1)] 1
# n = Neg(x), y = Relu(n) + n: the path from the Neg through the Relu leaves the occurrence and comes back in.
def looped(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
n = crt.neg(inputs[0])
return [crt.relu(n) + n]
asked.clear()
loop = trace_model(looped)
report = loop.optimize([fold_neg_into_sub])
print(formula_of(loop), rows(report), len(asked)) # Add(Relu(Neg(x)), Neg(x)) [('fold_neg_into_sub', 0)] 0
```
Write a condition for each fact the replacement relies on, such as a matched node's settings, since the rule itself checks only the three cases above.
## The same API as the runtime's
A rule you write with this API can give the same graph as one of the runtime's own transforms. `clamp_to_relu` turns a Clamp with a zero floor and no ceiling into a Relu, and the program writes it as a rule over one Clamp node. The rule's condition reads the Clamp's settings: a floor held as the literal zero, no ceiling, and one input, since a bound given as a tensor is an input of its own. Both runs turn the first Clamp into a Relu and keep the one with a ceiling, so the two graphs return the same formula. Their node names differ, since `replace` names the new Relu afresh where the runtime's transform keeps the Clamp's name.
```cpp title="same_api.cpp"
int main() {
// y = Clamp(x, 0) + Clamp(x, 0, 6): a zero floor alone, then a floor and a ceiling.
const ClikaRT::graph::TraceFunction clamps = [](const std::vector& in) {
const Tensor floored = ops::clamp(in[0], 0.0);
const Tensor bounded = ops::clamp(in[0], 0.0, 6.0);
return std::vector{ops::add(floored, bounded)};
};
// A Clamp with a zero floor held as its literal, and no ceiling, is a Relu.
const auto zero_floor = [](const RewriteMatch& match) {
const std::optional lower = match.nodes[0]->attribute("min");
const std::optional upper = match.nodes[0]->attribute("max");
const ClikaRT::Scalar* bound = lower.has_value() ? std::get_if(&lower->value) : nullptr;
const bool zero = bound != nullptr && ((bound->is_float() && bound->float_value() == 0.0) ||
(bound->is_int() && bound->int_value() == 0));
const bool no_ceiling = !upper.has_value() || std::holds_alternative(upper->value);
return match.inputs.size() == 1 && zero && no_ceiling;
};
Pattern one_clamp;
one_clamp.add_node(OpCode::Clamp);
const Transform clamp_to_relu_rule = transforms::rewrite(
"clamp_to_relu_rule", std::move(one_clamp),
[](const RewriteMatch&, const std::vector& in) { return std::vector{ops::relu(in[0])}; },
zero_floor);
ModelGraph by_runtime = trace_model(clamps);
ModelGraph by_rule = trace_model(clamps);
const std::string runtime_rows = rows(optimize_with(by_runtime, {transforms::clamp_to_relu()}));
const std::string rule_rows = rows(optimize_with(by_rule, {clamp_to_relu_rule}));
std::printf("%s | %s\n", formula_of(by_runtime).c_str(), runtime_rows.c_str());
// Add(Relu(x), Clamp(x)) | clamp_to_relu 1
std::printf("%s | %s\n", formula_of(by_rule).c_str(), rule_rows.c_str());
// Add(Relu(x), Clamp(x)) | clamp_to_relu_rule 1
std::printf("%s\n", formula_of(by_rule) == formula_of(by_runtime) ? "true" : "false"); // true
return 0;
}
```
```python title="same_api.py"
# y = Clamp(x, 0) + Clamp(x, 0, 6): a zero floor alone, then a floor and a ceiling.
def clamps(inputs: list[crt.Tensor]) -> list[crt.Tensor]:
x = inputs[0]
return [crt.clamp(x, 0.0) + crt.clamp(x, 0.0, 6.0)]
# A Clamp with a zero floor held as its literal, and no ceiling, is a Relu.
def zero_floor(match: RewriteMatch) -> bool:
settings = match.nodes[0].attributes
return len(match.inputs) == 1 and settings.get("min") == 0 and settings.get("max") is None
clamp_to_relu_rule = transforms.rewrite("clamp_to_relu_rule", Pattern([OpCode.Clamp]),
lambda match, tensors: crt.relu(tensors[0]), zero_floor)
by_runtime, by_rule = trace_model(clamps), trace_model(clamps)
runtime_rows = rows(by_runtime.optimize([transforms.clamp_to_relu()]))
rule_rows = rows(by_rule.optimize([clamp_to_relu_rule]))
print(formula_of(by_runtime), runtime_rows) # Add(Relu(x), Clamp(x)) [('clamp_to_relu', 1)]
print(formula_of(by_rule), rule_rows) # Add(Relu(x), Clamp(x)) [('clamp_to_relu_rule', 1)]
print(formula_of(by_rule) == formula_of(by_runtime)) # True
```
Write your own version of one of the runtime's transforms when you need a variant of it, such as a condition of your own, and compare the graphs the two give.
## Finalize and run the transformed graph
After the transforms, `finalize()` makes the graph runnable, and `run()` computes with it, as for any graph. A list of your own can start from the default list: `default_transforms()` returns it, and a transform added at its end runs after the runtime's. Each program checks the result against a reference it computes itself, and the Python program also imports numpy as `np` for that.
```cpp title="after_the_transforms.cpp"
int main() {
ModelGraph graph = trace_model();
// The runtime's default list, with a transform of your own at its end.
std::vector list = graph.default_transforms();
list.push_back(Transform::from_function("drop_repeated_relu", drop_repeated_relu));
optimize_with(graph, list);
graph.finalize(); // the graph runs from here on, and refuses edits
const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F};
const std::vector results = graph.run({Tensor::from_data(x.data(), {2, 3}, ClikaRT::DataType::Float32)});
const std::vector y = results.front().reshape({-1}).item_as_vec();
bool matches = y.size() == x.size();
for (std::size_t i = 0; matches && i < x.size(); ++i) {
matches = y[i] == (x[i] > 0.0F ? x[i] : 0.0F) - x[i]; // the reference, Relu(x) - x, by hand
}
for (std::size_t i = 0; i < y.size(); ++i) std::printf("%s%g", i == 0 ? "" : " ", static_cast(y[i]));
std::printf("\n%s\n", matches ? "true" : "false");
// 0 2 0 4 0 6
// true
return 0;
}
```
```python title="after_the_transforms.py"
# The runtime's default list, with a transform of your own at its end.
graph.optimize(graph.default_transforms() + [Transform("drop_repeated_relu", drop_repeated_relu)])
graph.finalize() # the graph runs from here on, and refuses edits
x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32) # numpy as the data entry
(y,) = graph.run([crt.tensor(x)])
print(y.numpy().tolist()) # [[0.0, 2.0, 0.0], [4.0, 0.0, 6.0]] (numpy as the data exit)
print(np.array_equal(y.numpy(), np.maximum(x, 0) - x)) # True (numpy states the reference)
```
Check a graph your transforms rewrote against a reference like this one before you serve it. [Optimize and finalize](optimize-and-finalize.mdx) covers the runtime's transforms and their defaults, [Edit a graph](edit-a-graph.mdx) the edits a transform makes, and [Query a graph](query-a-graph.mdx) the reads a transform or a condition can make.
---
# System requirements
Platforms, accelerators, host software, memory and storage for running ClikaRT.
Source: https://docs.clika.io/clikart/system-requirements.md
ClikaRT ships one distribution per platform. Every distribution includes the CPU backend; accelerator backends are included where the platform has them and load on demand at run time.
## Platforms
| Platform | Architecture | Backends |
| --- | --- | --- |
| Linux | x86_64 | CPU, CUDA, Vulkan |
| Linux | arm64 | CPU, CUDA, Vulkan |
| Android | arm64-v8a | CPU, Vulkan |
| Windows | x86_64 | CPU, Vulkan (CUDA on request) |
| Windows | arm64 | CPU, Vulkan |
| macOS | Apple silicon | CPU, Metal |
## Host software
| Requirement | Detail |
| --- | --- |
| C++ toolchain | A C++17 compiler for your own code; the library itself has no compiler requirement |
| CMake | 3.19 or newer |
| Linux glibc | 2.28 or newer (manylinux_2_28 compatible) |
| macOS | 14 (Sonoma) or newer, on Apple silicon |
| NVIDIA driver | The driver alone: the runtime's CUDA images carry the CUDA runtime and cuBLASLt inside them, so no CUDA toolkit install and no `LD_LIBRARY_PATH`. A driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image; the Linux archives carry both, and the Python wheel comes in one flavor per image |
| Vulkan | A driver supporting Vulkan 1.2 or newer |
| Python | CPython 3.10 to 3.14 for the `clika-runtime` wheel, which carries the runtime and `clika_runtime.modelverse`, on Linux, macOS and Windows x64; CPython 3.11 to 3.14 on Windows arm64. The wheel is an LZMA-compressed zip, which `pip` reads and `uv pip` does not |
| License credential | The `CLIKA1-...` credential your project's license carries, in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes. Every process that runs an operator needs one ([License credential](getting-started/get-clikart.mdx#license-credential)) |
## Accelerators
| Backend | Hardware |
| --- | --- |
| CUDA | NVIDIA GPUs from compute capability 7.0 (Volta) upward, Jetson included |
| Vulkan | Desktop and mobile GPUs with a conformant Vulkan driver (NVIDIA, AMD, Intel, Qualcomm Adreno) |
| Metal | Apple silicon |
## Memory
A running model needs memory in proportion to the size of its weights file. The minimums below were measured on devices running the platform's performance benchmark (up to 8 concurrent requests, prompts up to about 2K tokens), with models whose weights are up to about 2.5 GB. The KV cache drives the peak, so fewer concurrent requests or a shorter context need less, and more requests, a longer context or a larger model can need more than these rules give.
| Backend | Platforms | Minimum memory |
| --- | --- | --- |
| CPU | Linux x86_64, Linux arm64, Windows x86_64, macOS | 2 GB of RAM plus about 3.5 times the model's file size |
| CUDA on a discrete NVIDIA GPU | Linux x86_64 | 1 GB of system RAM plus about 0.5 times the model's file size, in addition to GPU memory for the model |
| CUDA on Jetson | Linux arm64 | 2 GB plus about 4 times the model's file size, in the memory the CPU and GPU share |
| Metal on Apple silicon | macOS | no rule: the one measurement, on a 32 GB machine, read about 11 to 14 GB for models from about 0.9 to 2.4 GB, nearly the same whatever the model's size, which does not give a minimum; a rule needs the benchmark on a 16 GB and an 8 GB Mac |
| Vulkan on an integrated GPU | Linux x86_64, Windows x86_64 | no rule: the host's resident-memory counter leaves out the GPU driver's own allocations, so a figure read that way understates the need; a rule needs a GPU-memory counter beside it |
The file size is the size of the model's weights on disk as the platform lists the model. The measurements used models stored as bf16 safetensors; a quantized file of the same model is smaller on disk, and these rules were not measured for it.
## Storage
### Engine package
The engine package is the ClikaRT build the platform delivers to a device for a benchmark or a model deployment. The platform delivers it to Linux, macOS and Windows x86_64 devices; Windows arm64 devices receive none. Where it can, the platform sends only the part of the engine the device uses. The device keeps the download in its cache next to the extracted tree, so plan for both together. They live under the device agent's staging directory: `/var/lib/clika-runtime-agent/staging` for a system-wide agent on Linux and macOS, `%ProgramData%\Clika\DeviceAgent\staging` on Windows, otherwise `~/.clika-rt/staging`.
| Device | Download | Installed | Both together |
| --- | --- | --- | --- |
| Linux x86_64, no GPU | the dialog's figure | about 320 MB | the download plus about 320 MB |
| Linux x86_64, NVIDIA GPU, driver 580 or newer | the dialog's figure | about 1.6 GB | the download plus about 1.6 GB |
| Linux x86_64, NVIDIA GPU, older driver | the dialog's figure | about 3.0 GB | the download plus about 3.0 GB |
| Linux x86_64, another GPU (AMD, Intel): the whole package | about 1.96 GB | about 4.1 GB | about 6.1 GB |
| Linux arm64, no GPU | the dialog's figure | about 210 MB | the download plus about 210 MB |
| Linux arm64, with the Vulkan backend | the dialog's figure | about 350 MB | the download plus about 350 MB |
| Linux arm64, NVIDIA GPU, driver 580 or newer | the dialog's figure | about 1.8 GB | the download plus about 1.8 GB |
| Linux arm64, NVIDIA GPU, older driver | the dialog's figure | about 3.6 GB | the download plus about 3.6 GB |
| Linux arm64, the whole package | about 2.41 GB | about 5.1 GB | about 7.5 GB |
| Windows x86_64, no GPU | the dialog's figure | about 220 MB | the download plus about 220 MB |
| Windows x86_64, with a GPU: the whole package | about 159 MB | about 370 MB | about 530 MB |
| Windows arm64, the whole package | about 154 MB | about 330 MB | about 480 MB |
| macOS, Apple silicon: the whole package | about 31 MB | about 170 MB | about 200 MB |
| Android arm64, the whole package | about 64 MB | about 310 MB | about 380 MB |
The installed figures are the release archive's tree, summed, less the backend images a row does not carry: on Linux x86_64 the CUDA 13 image is about 1.2 GB, the CUDA 12 image about 2.5 GB and the Vulkan backend about 150 MB of the 4.1 GB; on Linux arm64 about 1.4 GB, 3.3 GB and 150 MB of the 5.1 GB; on Windows the Vulkan backend is about 150 MB of the tree; on macOS the Metal backend about 40 MB. A whole package's download is the release archive itself, whose size the platform's dialog and `clika-cli runtime-sdk list` show beside it; a composed download (one backend's worth of the tree) is smaller, and the dialog shows its size in the same place, so a row that reads "the dialog's figure" takes it from there.
The first time the platform prepares a part of the engine for one kind of device, that one job or deployment receives the whole package for the platform instead (on Linux x86_64, the "another GPU" row). Each model deployment keeps its own extracted copy, and after an engine update the device also keeps the previous version's download for a while.
### Models
Required storage grows with the models you pull, by each model's file size. Benchmarks keep their models in a cache of up to 20 GB by default and remove the oldest first. A model deployment keeps its model until the deployment is deleted with its model removed. A model used by both is stored twice.
## Platform software
The device agent takes about 25 MB of disk. `clika-cli` takes about 15 MB of disk and about 70 MB of RAM while it runs.
---
# ClikaRT CLI
ClikaRT CLI is the command line interface that infers, serves and benchmarks popular models, built on top of Modelverse, the CLIKA model library, and running them on ClikaRT.
Source: https://docs.clika.io/modelverse.md
ClikaRT CLI is a command line interface that lets you infer, serve or benchmark popular models, one command each. It is built on top of Modelverse, the CLIKA model library: models packaged so that [ClikaRT](/clikart) loads and runs them as they are, a catalog of registered model families covering language, vision, audio and multimodal models, and a C++ library when you want the same machinery inside your own application. The executable is `clikart-cli`: one command takes a model name to generated text, one more serves it over HTTP, one more benchmarks it on the device it runs on.
## Why ClikaRT CLI
For the AI developer. The checkpoints you already use are the input. ClikaRT CLI resolves a Hugging Face Hub repo id, a pasted Hugging Face URL or a local directory to model files, matches them to a registered family, and runs them; ONNX exports, GGUF quantizations and safetensors checkpoints all load through ClikaRT unchanged. `clikart-cli prompt "..."` is a working generation before you have written any code, and every knob you expect (sampling, system prompt, context length, KV cache mode) is a flag.
For the backend engineer. A model becomes an OpenAI-compatible endpoint in one command. `clikart-cli serve` hosts `/v1/chat/completions` with streaming, plus embeddings, transcription and the other engine routes a model family provides, and existing OpenAI clients point at it by changing one base URL. `clikart-cli` is built for scripts: stdout carries only the payload, diagnostics go to stderr, and the exit codes follow a fixed four-value contract. No Python runs anywhere.
For the embedded developer. A model family ships quantized variants, and you pick the one that fits the device. A GGUF repo with ten quantizations is a selector away (`/:Q6_K`), the option table with file sizes prints before anything downloads, and the same model runs wherever ClikaRT runs, from a workstation GPU to a phone.
For the defense, healthcare and finance developer. Nothing here requires a network at run time. Fetch a model on a connected machine, move the directory, and point the CLI at it; a local directory is a first-class model source, and an offline flag makes any network touch an error instead of a surprise. The install is one archive extracted into one directory, runtime and Modelverse together, and the models arrive the same way.
For the business. Every model in the catalog has a known license: the catalog records who published each family and under what terms. You do not have to vet checkpoints from unknown sources. Models come straight from Hugging Face by repository id, so the models your team already uses work as-is. Quantized variants run the same model on cheaper hardware, which lowers serving cost. And you do not have to build inference for the popular models yourself. ClikaRT CLI already runs them.
## What you get
- **The catalog.** Dozens of registered model families, from Llama, Qwen and Gemma through Whisper, CLIP, DETR and Depth Anything. Each family declares its input and output modalities, the checkpoint formats it matches, and the commands it can run. `clikart-cli list` prints it.
- **The executable.** `clikart-cli` inspects (`info`), downloads (`fetch`) and runs models. Which commands a model supports (prompt, serve, transcribe, embed, bench and more) depends on its family, discovered per model.
- **The server.** An OpenAI-compatible HTTP server with streaming chat completions, a built-in web chat page and a health probe, plus per-modality routes for transcription, embeddings, depth, detection and segmentation.
- **The library.** The surface behind all of it, in C++ with Python and Kotlin bindings: fetch a snapshot, load a runnable model, build a serving pipeline, or mount your own engine on the server. For C++, one `find_package(Modelverse CONFIG)` integrates it; the bindings arrive through their package managers.
- **The packaging.** One archive per platform, the ClikaRT runtime and Modelverse inside it, extracted and run in place. A manifest pins the exact runtime each build linked against, so a mismatched pair refuses with a readable error. No installer and no downloads at run time.
## Where to go next
- [Getting started](getting-started/index.md): install the archive and run your first model.
- [How-to guides](how-to/index.md): problem-oriented recipes, from quantization selection to offline deployment.
- [Model requirements](model-requirements.md): memory and device figures per model variant.
- [Additional examples](examples.md): the example programs, one per subsystem.
- [ClikaRT](/clikart): the runtime underneath, with its own tutorial and API reference.
---
# Additional examples
The Modelverse example programs in the release's examples archive, standalone CMake projects against the installed package, from the catalog probe to vision, OCR, reranking and vision-language chat.
Source: https://docs.clika.io/modelverse/examples.md
The Modelverse examples are standalone `find_package(Modelverse CONFIG)` projects on the public API only, one shared `README.md` walk-through beside them, and together they cover the library surface the [tutorial](getting-started/first-model/01-pick-a-model.md) meets through the `clikart-cli` executable. They live in the release's examples archive, `ClikaRT--examples.tar.xz` ([Get ClikaRT](/clikart/getting-started/get-clikart) names the download), under `cpp/modelverse/`; the release archive's own `examples/src` carries the runtime's examples ([the ClikaRT catalog](/clikart/examples)). Building them needs the release archive for your platform, extracted, with `CLIKART_BUNDLE_DIR` naming its directory: the one `cmake/` directory there carries both products' packages.
Read top to bottom; each row assumes a little of the ones above it.
| Example | What it shows |
| --- | --- |
| `init_model` | The registered catalog and identity resolution: list the families, resolve a model's identity from a local directory or a Hugging Face Hub repo id. The snapshot fetches configs and companions only, so identity costs no weight download. |
| `00_generate` | The minimal end-to-end path: explicit snapshot, registry match, generative pipeline, one prompt, text. The program [part 4 of the tutorial](getting-started/first-model/04-use-it-from-code.mdx) builds. |
| `01_serve` | The ServeAPI in process: implement the `ChatEngine` interface, mount it on `ServeApi`, and round-trip one `/v1/chat/completions` request through the built-in HTTP client. Runs with no arguments and no checkpoint, so it is also the fastest server smoke test. |
| `02_custom_node` | Your own `ClikaRT::runtime::Model` node composed with the library's serving nodes in one pipeline: a synthesized two-layer llama generates and a short custom node post-processes the reply, with the request surface unchanged. Offline, runs with no arguments. The worked version against a real model is [Add your own node to a model pipeline](how-to/add-a-pipeline-node.md). |
| `03_conversational_pipeline` | A conversational AI as one pipeline: wav bytes in, the reply waveform out, through speech to text, the chat template, a chat model and text to speech across eight nodes; real models, so it fetches on first run and takes a question wav plus the voice-reference wav the speech model requires. The guide is [A conversational AI as one pipeline](how-to/conversational-ai-pipeline.md). |
| `04_classify_zero_shot` | Zero-shot image classification: the labels you name scored against a picture by cosine in a dual-tower checkpoint's shared space (SigLIP); the picture drawn by the program; no arguments. |
| `05_detect_objects` | Object detection over the checkpoint's own label set (DETR): boxes in the picture's pixels above a confidence; no arguments. |
| `06_detect_by_phrase` | Open-vocabulary detection (OWL-ViT): the prompt's phrases encoded once, a box per thing found, labeled with its phrase; no arguments. |
| `07_estimate_depth` | Monocular depth (Depth Anything V2): one value per pixel at the picture's size, metric or relative; no arguments. |
| `08_read_text` | Text recognition (GLM-OCR): a page of block letters drawn with the operators, read back as text through the registry's text-recognition door; no arguments. |
| `09_rerank` | Reranking (Qwen3-Reranker): one relevance score per document against a query, printed best first; no arguments. |
| `10_describe_a_picture` | Vision-language chat (Qwen3-VL): one question about a picture, the picture riding the user turn; greedy, no arguments. |
A program that runs with no arguments and prints the same text on every run sits beside its recorded output (`.out`), the text the release printed; a program that takes a checkpoint or a clip, or prints a port, is built and not recorded.
## Build and run
From the extracted `examples/` directory:
```bash
cmake -S cpp/modelverse -B build-examples \
-DModelverse_DIR="$CLIKART_BUNDLE_DIR/cmake"
cmake --build build-examples
```
```bash
build-examples/init_model # catalog only, offline
build-examples/init_model -r openai/whisper-large-v3-turbo # + identity from Hugging Face
build-examples/00_generate Qwen/Qwen2.5-0.5B-Instruct "Once upon a time, in a port town by a cold sea,"
build-examples/01_serve
build-examples/02_custom_node # offline, no arguments
build-examples/05_detect_objects # the vision, OCR, reranking and chat programs: no arguments, the checkpoint from the local cache
```
`01_serve` prints its bound port, answers one request against itself, and exits `PASS`, which makes it the one to run first when checking a new machine:
```text
serving on 127.0.0.1:41627
HTTP 200
choice: echo: Say hello.
01_serve: PASS
```
Every program runs compute, so it needs the license credential in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes; without one a call is refused with the code name `LICENSE_FAILED`. Where models come from, cache placement and the offline path are the same for the examples as for `clikart-cli`; the examples' shared `README.md` restates them next to the code.
---
# First steps
New to Modelverse? Start here. What the model library is, how to install it, and your first model from catalog to served endpoint.
Source: https://docs.clika.io/modelverse/getting-started.md
New to Modelverse? This section is where to start. It gives enough orientation to hold the product in your head, an install you can verify in minutes, and a first model that goes from the catalog to a served endpoint. Read it in order:
1. **[Modelverse at a glance](overview.mdx)**: what the model library is and is not, and the five-minute mental model.
2. **[Get Modelverse](get-modelverse.md)**: pick your platform, get the download and verify commands.
3. **[Quick install](installation.md)**: extract the archive and prove it works in two commands.
4. **Tutorial series**: four parts, each a complete step. [Pick a model](first-model/01-pick-a-model.md) (the catalog, identity, what a download would cost), [fetch and prompt](first-model/02-fetch-and-prompt.md) (weights on disk, first generated text), [serve and chat](first-model/03-chat-and-serve.md) (an OpenAI-compatible endpoint with a built-in chat page), and [use it from your code](first-model/04-use-it-from-code.mdx) (the same model inside your own program).
5. **Adding your own model**: the author-side series, three parts on the smallest real model (an embedding model: text in, vector out). [A model from scratch](own-model/01-a-model-from-scratch.md) (a checkpoint is a directory you can write by hand), [join the catalog](own-model/02-join-the-catalog.md) (register a family of your own), and [make it embed](own-model/03-make-it-embed.md) (the smallest real backend, the standard commands in two lines). Read it when you bring your own architecture; the first series does not depend on it.
6. **[What to read next](next-steps.md)**: where to go once it runs.
## How the Modelverse docs are layered
- **This section** orients: condensed, in reading order, concepts explained where they first appear.
- **[How-to guides](../how-to/index.md)**: problem-oriented recipes, one per "how do I X", with [additional examples](../examples.md) as the end-to-end reading inside it.
- **[ClikaRT](/clikart)** documents the runtime underneath, including the API reference the C++ surface builds on and the [system requirements](/clikart/system-requirements) Modelverse inherits.
---
# Serve and chat
Host the model behind an OpenAI-compatible HTTP endpoint, hold a conversation on the built-in chat page, and call it with curl.
Source: https://docs.clika.io/modelverse/getting-started/first-model/chat-and-serve.md
The model generates on demand; this part keeps it loaded. `serve` hosts the model behind an OpenAI-compatible server with a built-in chat page. One command, no configuration files:
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct serve
```
```text
[2026-10-04 07:56:43.933] [modelverse] [info] serving on http://127.0.0.1:8000 (model=Qwen2.5-0.5B-Instruct); open it in a browser for the chat page
[2026-10-04 07:56:43.933] [modelverse] [info] endpoints: POST /v1/chat/completions /v1/messages /v1/messages/count_tokens /v1/audio/transcriptions; GET /v1/models /health /props / (web UI) /dashboard
```
The defaults bind `127.0.0.1:8000`; `--host 0.0.0.0` opens it to the network and `--port` moves it. Three things are now running:
- **The API.** `POST /v1/chat/completions`, streaming (server-sent events) and non-streaming, in the OpenAI request and response shape.
- **A web chat page.** `http://127.0.0.1:8000/` serves a built-in chat UI from inside the binary (`--no-web-ui` disables it).
- **A health probe.** `GET /health` answers `{"status":"ok"}`, for load balancers and scripts.
## Chat in the browser
The chat page is the conversation surface: a scrolling transcript with an input box, the assistant reply streaming token by token as the pipeline decodes it. The conversation carries its history, so follow-up questions see earlier turns. The same conversation is available to any OpenAI client through the API, where the system message is the first `messages` entry and the sampling knobs (`temperature`, `top_p`) ride each request.
One templated turn from the terminal stays `prompt`, part 2's command. `prompt`, `serve` and `bench` are the text-generation surface, and there is no terminal `chat` verb: `clikart-cli chat` is refused by name, `error: family 'qwen' (text->text) provides no chat (it provides: prompt, serve, bench, mm_bench)`.
## Call it over HTTP
From a second terminal:
```bash
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Say hello in French."}],
"max_tokens": 32
}'
```
```text
{"id":"chatcmpl-1791068204-0","object":"chat.completion","created":1791068204,"model":"Qwen2.5-0.5B-Instruct","choices":[{"index":0,"message":{"role":"assistant","content":"Bonjour! C'est un plaisir de vous rencontrer. Comment puis-je vous aider aujourd'hui ?"},"finish_reason":"stop"}],"usage":{"prompt_tokens":34,"completion_tokens":20,"total_tokens":54,"prompt_tokens_details":{"cached_tokens":0}}}
```
The `id`, `created` and `usage` values are your run's own. Add `"stream": true` and the response arrives as server-sent events, one delta per chunk, the way OpenAI clients expect. An existing client needs one change:
```python
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
reply = client.chat.completions.create(
model="qwen",
messages=[{"role": "user", "content": "Say hello in French."}],
)
print(reply.choices[0].message.content)
```
The endpoint surface is bigger than chat: a Whisper model's `serve` answers `POST /v1/audio/transcriptions`, an embedding model's answers `POST /v1/embeddings`, and so on per family. [Serve a model over the OpenAI and Anthropic APIs](../../how-to/serve-openai-compatible.md) has the full route table and the operational flags.
Next: [part 4](04-use-it-from-code.mdx), the same model inside your own program.
---
# Fetch and prompt
Download the model snapshot, generate your first text, and control sampling from the command line.
Source: https://docs.clika.io/modelverse/getting-started/first-model/fetch-and-prompt.md
Part 1 established what the model is; this part downloads it and makes it generate. At the end you have the weights on disk in a directory you control and a repeatable one-shot generation command.
Generation is compute, so the credential from [Quick install](../installation.md#2-place-the-license-credential) has to be in place before the `prompt` command below. `fetch` downloads files and needs none.
## Fetch the snapshot
`fetch` downloads a source's files (companions first, then the weights, one progress bar per file) and prints exactly one thing on stdout: the local path the download landed at.
```bash
clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct
```
```text
/home/you/.cache/huggingface/hub/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775
```
Progress, timing and the file count print to stderr, so the path is safe to capture:
```bash
MODEL_DIR=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct)
```
By default snapshots land in the standard Hugging Face Hub cache, shared with other Hugging Face tooling on the machine; `--cache-dir models` keeps them in a directory of your choosing instead. Fetching a source that is already cached verifies and returns immediately, so scripts can fetch unconditionally. A gated repo needs a token. The CLI reads it the way the other Hugging Face tooling does: `HF_TOKEN` in the environment, else the token a Hub login stored (`hf auth login` writes it to `$HF_HOME/token`; `HF_TOKEN_PATH` names another file). There is no token flag, deliberately, so a token can never land in shell history. Qwen2.5 0.5B Instruct is not gated, so its fetch needs no token; for a gated repo, accept its license once on the repo's Hub page, provide a token, and the fetch proceeds (without them it refuses by name: `'/' is gated or private on https://huggingface.co; set $HF_TOKEN or log in to the hub`).
Two related commands you already have: `fetch --dry` prints the total of what the fetch downloads without downloading it (`info --dry` adds the file listing behind that total), and a local directory used as a source skips fetching entirely.
One refusal worth meeting on purpose: `fetch` takes checkpoints in the formats from part 1 (safetensors, torch containers, GGUF) and refuses anything else before a byte of weights moves. An ONNX-only export, for example:
```text
$ clikart-cli fetch onnx-community/Qwen2-0.5B-Instruct-ONNX
error: 'onnx-community/Qwen2-0.5B-Instruct-ONNX' ships its weights only in formats the loaders do not read: ONNX (onnx/model.onnx, onnx/model_bnb4.onnx, onnx/model_fp16.onnx and 5 more). The loaders read safetensors, GGUF and the torch container
```
## The first generation
`prompt` runs one templated generation of the positional message: the model's own chat template wraps your text, the pipeline decodes, and the generated text is the stdout payload. The template is why a question gets an answer here: a raw, untemplated completion would CONTINUE the question's shape instead (part 4 shows that mode from the library). Prompt answers one shot; a conversation runs against the served model in part 3.
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct prompt "The capital of France is"
```
```text
Paris.
```
The first run loads the weights (a progress bar on stderr); repeat runs on a warm cache start in seconds. The same invocation works with `"$MODEL_DIR"` in place of the repo id, which is the fully offline form.
## The knobs are flags
Sampling and context are controlled per invocation. The ones you will use first:
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct prompt "Name three rivers." \
--max-new-tokens 32 \
--temperature 0.2 \
--seed 7
```
```text
Three rivers that I can name for you are the Yangtze River, the Yellow River, and the Pearl River. Each of these rivers is significant in Chinese
```
- `--max-new-tokens` caps the generation length; `--max-seq` caps the whole context.
- `--temperature` shapes sampling; `--greedy` disables it for the single most likely continuation, and `--seed` makes a sampled run repeatable.
- A conversation's knobs move into the request once the model is served: the system message is the first `messages` entry, and `temperature`/`top_p` ride each OpenAI request ([part 3](03-chat-and-serve.md)).
- `--device cuda` places the model explicitly; the default picks the best available device, and `clikart-cli devices` shows the candidates. `--device auto` names that default: the same pick, falling back down the accelerator order (CUDA, TPU, Metal, Vulkan) when a load fails, and the CPU last. `GET /props` on a served model reports which device it landed on.
A vision-capable model takes media the same way. With the flagship multimodal family the message and the image travel together:
```bash
clikart-cli google/gemma-4-E2B-it prompt "What is on the sign?" --image photo.jpg
```
The root commands' options also have a file form: `generate-template` writes a JSON file of every root option at its default, `--template ` loads one, and explicit flags win over it (a family's own verbs, like `prompt` and `serve`, carry their options as flags only). [Script clikart-cli](../../how-to/script-the-cli.md) covers that workflow.
Next: [part 3](03-chat-and-serve.md), the model behind an HTTP endpoint, with a conversation on its built-in chat page.
---
# Pick a model
Read the catalog, resolve a model's identity, and see what a download would cost, all before any weights move.
Source: https://docs.clika.io/modelverse/getting-started/first-model/pick-a-model.md
This tutorial takes one model from the catalog to a served endpoint in four parts, each a complete session, and ends with the same model running inside a C++ program. Core concepts are explained where they first appear. This part picks the model and learns everything about it without downloading a single weight.
It assumes the install directory exists and `clikart-cli` is on your `PATH` (see [Quick install](../installation.md)). The model is Qwen2.5 0.5B Instruct, small enough to run on any machine in the [system requirements](/clikart/system-requirements); every command works the same with any other model in the catalog.
## The catalog knows the families
`list` prints every registered model family: its modalities, the commands it provides, and whether it is runnable on this build. No network is involved; the catalog is compiled into the `clikart-cli` executable.
```bash
clikart-cli list
```
```text
registered model families (47)
bert runnable
* Input Modalities: text
* Output Modalities: embedding
* Valid Combos: text -> embedding
* Model Types: bert
* Commands: embed, similar, bench, serve
* Web UI: generic page
* Vendor: Google
* License page: https://github.com/google-research/bert/blob/master/LICENSE
------------------------------------------
...
```
A family is the model architecture Modelverse knows how to run; a model you fetch is a checkpoint of that family. The `Commands` row is the contract for parts 2 and 3: whatever it lists is what that model can do.
## A source names a model
Everything model-specific starts from a source: a Hugging Face Hub repo id (`/`), a pasted Hugging Face URL, or a local directory. Modelverse runs checkpoints in the formats model publishers ship on the Hub: safetensors, torch containers (`pytorch_model.bin`), and GGUF. (ONNX exports are a different lane: the ClikaRT runtime runs ONNX models directly; its how-to guide covers that path.) `info` resolves a source's identity:
```bash
clikart-cli info Qwen/Qwen2.5-0.5B-Instruct
```
```text
Qwen/Qwen2.5-0.5B-Instruct
family=qwen model_type=qwen2 architecture=Qwen2ForCausalLM
variant: dense
components: tokenizer=yes image=no audio=no video=no
license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT
model: Qwen/Qwen2.5-0.5B-Instruct provider: Qwen source: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct
license: apache-2.0 https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/blob/main/LICENSE
```
Identity resolution reads configuration files only. `clikart-cli` downloads a few KB of JSON, matches it against the registered families, and reports what it found; weights do not move. This is deliberate: you can interrogate a 70B model from a laptop. The last three lines are the licensing report: the family's license page, the checkpoint's provider and source, and the checkpoint's own license with whether the repository is gated. Read them before you ship what the model produces.
## `--dry` shows what a fetch would cost
Add `--dry` and `info` also lists the files a fetch downloads, with sizes, split into companions (configs, tokenizer) and weights, and says how much of the repository it leaves behind (other weight formats, files the fetch never takes):
```bash
clikart-cli info Qwen/Qwen2.5-0.5B-Instruct --dry
```
```text
Qwen/Qwen2.5-0.5B-Instruct
family=qwen model_type=qwen2 architecture=Qwen2ForCausalLM
variant: dense
components: tokenizer=yes image=no audio=no video=no
license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT
model: Qwen/Qwen2.5-0.5B-Instruct provider: Qwen source: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct
license: apache-2.0 https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/blob/main/LICENSE
companions:
config.json 659 B
generation_config.json 242 B
merges.txt 1.6 MiB
tokenizer.json 6.7 MiB
tokenizer_config.json 7.1 KiB
vocab.json 2.6 MiB
weights:
model.safetensors 942 MiB
total: 953 MiB (weights 942 MiB, 7 files)
repository: 953 MiB in 10 files; the rest is not fetched (other weight formats, files the fetch never takes)
(dry; nothing downloaded)
```
Nothing here is Qwen-specific. `info google/gemma-4-E2B-it` reports `variant: dense + multimodal`; `info openai/whisper-large-v3-turbo` reports an audio component. Some repos ship several weight options to choose between; [Run a specific GGUF quantization](../../how-to/run-a-gguf-quantization.md) covers picking one.
## The model's own commands
`clikart-cli` has four commands that work without naming a model (`list`, `devices`, `info` and `fetch`; `generate-template`, their file-form helper, rides beside them). Every other command runs on a model you name: pass the source and `--help`, and `clikart-cli` resolves its family and prints that family's own command surface:
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct --help
```
```text
Qwen/Qwen2.5-0.5B-Instruct
family=qwen text->text
usage: clikart-cli Qwen/Qwen2.5-0.5B-Instruct [options]
commands the 'qwen' family provides for this model
options:
-h, --help show this help and exit
--verbose, -v debug diagnostics (the effective-options report)
--quiet, -q warnings and errors only
--no-color plain output (also honored: $NO_COLOR, TERM=dumb, non-tty)
commands:
prompt one templated generation of the positional user message
serve run the http server (OpenAI and Anthropic APIs)
bench the family-owned benchmark flow
mm_bench the multi-modality benchmark flow (media axes over the serving
path)
```
Asking a model for a command its family does not provide refuses precisely and names what it does provide, with exit code 3 ([Script clikart-cli](../../how-to/script-the-cli.md) has the full exit contract).
Next: [part 2](02-fetch-and-prompt.md), the weights arrive and the model speaks.
---
# Use it from your code
Link the Modelverse library and run the same model inside your own program: snapshot, load, pipeline, text.
Source: https://docs.clika.io/modelverse/getting-started/first-model/use-it-from-code.md
{/* CERTIFICATION: the C++ sample compiles and runs against the pinned release bundle (linux_x86_64, CPU); the output block is that run's capture. load.max_seq mirrors the CLI's 4096 default (the library otherwise keeps the checkpoint's full window and sizes the KV cache from it). The Kotlin arm is a staged binding surface and carries its own in-tab marker. */}
Everything the `clikart-cli` executable did in parts 1 through 3 is a library call. This part writes the smallest program that does what `prompt` does: resolve a snapshot, load the runnable model, and generate.
## The project
For C++, the install directory you already have is also the SDK (the headers, the library, and the ClikaRT runtime beside them): the platform's one archive for your platform carries the runtime and Modelverse together ([Get Modelverse](../get-modelverse.md)), so there is no second download. A two-file CMake project is the whole setup. The C++17 toolchain from the [ClikaRT quick install](/clikart/getting-started/installation) prerequisites is the one requirement.
```cmake title="CMakeLists.txt"
cmake_minimum_required(VERSION 3.19)
project(hello_modelverse LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
find_package(Modelverse CONFIG REQUIRED)
add_executable(hello main.cpp)
target_link_libraries(hello PRIVATE Modelverse::modelverse)
# Copy the Modelverse library and the ClikaRT runtime next to the binary,
# so the program runs from the build directory as-is.
modelverse_stage_runtime(TARGET hello)
```
`find_package(Modelverse CONFIG)` wires everything: it defines the one link target (`Modelverse::modelverse`), and finds the ClikaRT runtime installed in the same root (a `ClikaRT::ClikaRT` you already provide is respected instead).
Python arrives in the `clika-runtime` wheel, which carries the model library as `clika_runtime.modelverse`; Kotlin arrives in the one Maven artifact, `io.clika:clika-runtime`, which carries the model library as `io.clika.modelverse` (the artifact's zip, extracted, is the Maven repository the build names). Both are downloads of the platform ([Get ClikaRT](/clikart/getting-started/get-clikart)):
```bash
pip install ./clika_runtime-0.6.4-cp313-cp313-manylinux_2_28_x86_64.whl # Python: the wheel, the model library inside
# Kotlin/Gradle: implementation("io.clika:clika-runtime:0.6.4")
```
## The program
```cpp title="main.cpp"
#include
#include "clika_modelverse/hub/hub.h"
#include "clika_modelverse/registry/model_registry.h"
#include "clika_modelverse/runtime/serving_pipeline.h"
#include "clika_modelverse/runtime/text_nodes.h"
namespace mv = clika_modelverse;
namespace mvr = clika_modelverse::runtime;
int main() {
// 1. Model files, exactly as `fetch` gets them. An already-fetched or
// local directory passes through untouched.
mv::hub::SnapshotOptions snap;
snap.cache_dir = "models"; // beside the program; empty = the shared Hugging Face cache
const mv::hub::SnapshotResult snapped = mv::hub::snapshot(
"Qwen/Qwen2.5-0.5B-Instruct", snap);
// 2. The registry matches the snapshot to its family and returns the
// runnable model, on the device you name (CPU by default).
mv::LoadOptions load;
load.max_seq = 4096; // the CLI's default context cap; 0 keeps the
// checkpoint's full window and sizes the KV cache from it
mv::GenerativeModel model =
mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load);
// 3. A serving pipeline around it: tokenizer -> decoder -> detokenizer,
// the same assembly `prompt` and `serve` run on.
mv::generation::GenerationConfig gen = model.defaults();
gen.max_new_tokens = 64;
gen.stop = {"."}; // raw completion: stop at the first sentence end
mvr::PipelineOptions opts;
mvr::GenerativePipeline pipe = mvr::build_generative_pipeline(model, gen, opts);
// 4. One request through it.
ClikaRT::runtime::Request req;
req.inputs.set("prompt", mvr::string_to_byte_tensor(
"Once upon a time, in a port town by a cold sea,"));
const auto sid = pipe.pipeline->enqueue(std::move(req));
const ClikaRT::runtime::Response resp = pipe.pipeline->await(sid);
std::printf("%s\n", mvr::byte_tensor_to_string(resp.outputs.get("text")).c_str());
return 0;
}
```
```python title="main.py"
import clika_runtime.modelverse as mv
# 1. and 2. Model files and the runnable model in one call: the snapshot
# downloads into the hub cache (a local directory passes through untouched),
# the registry matches it to its family and loads it on the device named.
model = mv.AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", device="cpu")
# 3. and 4. The generation: text in, text out; the keyword arguments are the
# decode knobs over the checkpoint's own defaults.
print(model.generate("Once upon a time, in a port town by a cold sea,", max_new_tokens=64, stop=["."]))
```
The model served in part 3 answers any OpenAI client as well (`clikart-cli serve`, then a client pointed at `http://127.0.0.1:8000/v1`), which is the path for a program that must not load the model itself.
{/* CERTIFICATION: the Kotlin arm shows the Android binding's shape as the pinned release's Kotlin package
declares it (Modelverse.load, AutoModelForCausalLM.fromPretrained, LoadOptions, generate); it is not compiled
by the docs build. */}
```kotlin title="Chat.kt"
import io.clika.modelverse.AutoModelForCausalLM
import io.clika.modelverse.LoadOptions
import io.clika.modelverse.Modelverse
// On a worker thread of the app, never the main thread.
fun reply(context: android.content.Context, license: String): String {
// 0. The runtime, the model library and the bridge, once per process,
// with the license credential placed before any model runs.
Modelverse.load(context, license)
// 1. and 2. Model files and the runnable model in one call: a hub
// repository id downloads into the app's cache (a snapshot directory
// or a .gguf file passes through untouched), and the registry loads
// it on the CPU. The context length sizes the key-value cache for a phone.
val model = AutoModelForCausalLM.fromPretrained(
"Qwen/Qwen2.5-0.5B-Instruct",
LoadOptions(contextLength = 4096),
)
// 3. and 4. The generation: text in, text out, blocking for the reply;
// chat(messages, config, listener) is the streaming form.
return model.use { it.generate(prompt = "Once upon a time, in a port town by a cold sea,") }
}
```
Build and run the C++ project:
```bash
cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake"
cmake --build build
./build/hello
```
```text
there lived a young woman named Akira.
```
The raw pipeline CONTINUES text; there is no chat template in the loop, which is the visible difference from `prompt` in part 2 (templated, answers a question). That is why this program feeds it a story opener and stops at the first sentence end: give a raw completion a question and it rambles on in the question's shape. A program that answers a question renders the conversation through the model's chat template first, as the executable does; the next section is that step.
## Answer a question
A chat model's prompt format ships with the model as a template, and the tokenizer renders a `messages` array in the OpenAI API's shape into the exact prompt string; [Tokenize and chat templates](/clikart/how-to/tokenize-and-chat-templates#render-a-conversation-with-the-models-chat-template) is the contract. Never hand-build the role markers.
`Tokenizer::apply_chat_template` renders the messages. In the program above, the rendered string takes the story opener's place as the `prompt` input, and the `.` stop goes, so the pipeline answers as the assistant.
```cpp title="chat_template.cpp"
int main() {
Tokenizer tok = Tokenizer::from_huggingface(model_dir());
Json messages = Json::array();
Json system = Json::object();
system["role"] = "system";
system["content"] = "You are a concise assistant.";
messages.push_back(std::move(system));
Json user = Json::object();
user["role"] = "user";
user["content"] = "What does a tokenizer do?";
messages.push_back(std::move(user));
const std::string prompt = tok.apply_chat_template(messages);
std::printf("=== rendered prompt ===\n%s\n=======================\n", prompt.c_str());
// The template writes the prompt's own bos/eos framing, so encode_chat adds
// no special tokens on top of it.
const std::vector ids = tok.encode_chat(messages, /*add_generation_prompt=*/true);
std::printf("encode_chat produced %zu tokens\n", ids.size());
return 0;
}
```
`chat` renders and generates in one call, and `render` is the rendered prompt alone; the conversation is yours to carry, so append the reply and the next call sees the whole history.
```python title="chat_messages.py"
def main() -> None:
model = mv.AutoModelForCausalLM.from_pretrained(resolve_model(), device="cpu", offline=True)
messages = [
{"role": "system", "content": "You answer in one short sentence."},
{"role": "user", "content": "Why is the sky blue?"},
]
# A checkpoint without a chat template reads a plain prompt instead, so a
# program that may meet either one says which it got.
if not model.has_chat_template:
print("BLOCKED: this checkpoint ships no chat template")
sys.exit(3)
# render is the prompt text the template produced, the same string
# apply_chat_template returns on the transformers side.
print(f"=== rendered prompt ===\n{model.render(messages)}\n=======================")
reply = model.chat(messages, max_new_tokens=32, temperature=0.0)
print(f"=== reply ===\n{reply}\n=============")
# The conversation is the caller's: append the reply and the next call
# carries the whole history.
messages.append({"role": "assistant", "content": reply})
followup = model.chat(messages + [{"role": "user", "content": "In fewer words?"}],
max_new_tokens=32, temperature=0.0)
print(f"=== follow-up ===\n{followup}\n=================")
assert reply.strip(), "the first reply is empty"
assert followup.strip(), "the follow-up reply is empty"
assert messages[-1]["role"] == "assistant", "the history did not gain the assistant turn"
if __name__ == "__main__":
main()
```
## What the four steps are
The program is the executable's anatomy laid bare, and each step is independently useful:
- **The snapshot call** resolves any source (a Hugging Face Hub repo id, URL, or local directory) to a directory of model files. Point it at a directory you shipped with your application and no network code is ever built in.
- **The registry load** is the catalog from part 1 as a function: identity match, factory, weights onto the device. Where the weights load is `LoadOptions::where`, a `Stream` or a `Device`: a Device (or nothing; the `device` field then decides) loads through a fresh stream of the model's own, a Stream loads through yours, behind whatever it already carries. The load synchronizes the weights, so their stream stops mattering once it returns; each inference call then runs on the stream its input tensors arrive on (or the calling thread's stream when they carry none), which is what lets two models fed from two streams overlap. The CLI's `--device` fills the same option.
- **The pipeline build** assembles the serving pipeline. It is a [ClikaRT serving-runtime](/clikart) pipeline underneath, so sessions, continuous batching and the async model documented there apply unchanged.
- **The request** is the generation loop. `serve` from part 3 is this loop behind HTTP; the library's `ServeApi` lets you mount the same engines, or your own, in-process.
Failures follow the runtime's error model per language: C++ returns values directly and raises `ClikaRT::Error`, Python raises `clika_runtime.ClikaRTError` with the same code name, and Kotlin throws `ModelverseException` with its `codeName`; a served model reports a failure as the HTTP status the client raises.
You have taken a model from the catalog to your own binary. [What to read next](../next-steps.md).
---
# Get Modelverse
Pick your platform's ClikaRT archive; Modelverse ships inside it, with the download and verify commands.
Source: https://docs.clika.io/modelverse/getting-started/get-modelverse.md
Modelverse ships inside the ClikaRT release archives: one archive per platform, carrying the ClikaRT runtime and Modelverse together (the `clikart-cli` executable, the library, the public headers, and the CMake packages for both products). Installing ClikaRT installs Modelverse; there is no separate Modelverse download and no second install step. The archive comes from your CLIKA Platform deployment: [Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) covers the web dialog, the platform CLI and the MCP tools that hand it out, with the license key the runtime needs beside it. The commands start from that archive and the checksum shown with it.
The archive comes from the CLIKA platform's **Download ClikaRT** dialog, the platform CLI (`clika-cli runtime-sdk download`) or the MCP tools ([Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) walks each surface); the dialog shows the archive's SHA-256 beside its download button. The release the dialog offers is the deployment's own pin, which may differ from the release these pages are written for (`0.6.4`); a download's own `README.md` is the authority for the release it came from, its commands and file names included. Take the one file for your platform, check it against that digest, and extract:
```bash
echo " ClikaRT_linux_x86_64-0.6.4.tar.xz" | sha256sum -c
tar -xf ClikaRT_linux_x86_64-0.6.4.tar.xz
export MODELVERSE_INSTALL_DIR="$PWD/ClikaRT_linux_x86_64-0.6.4"
```
`MODELVERSE_INSTALL_DIR` is the extracted directory, the same one the ClikaRT documentation calls `CLIKART_BUNDLE_DIR`; one directory carries both products. (macOS checks with `shasum -a 256 -c` over the same line. Windows, in PowerShell: `(Get-FileHash ).Hash -eq ""`, then `tar -xf `.)
## License credential
Running a model needs the license credential CLIKA issued for your deployment, because Modelverse runs on the ClikaRT runtime and the runtime runs compute under a license. The credential is the `CLIKA1-...` text your project's license shows as its license key ([Runtime licenses](/platform/concepts/runtime-licenses) is where it comes from). Set it for the process, as the credential text or as the path of a file holding it:
```bash
export CLIKA_RT_LICENSE=CLIKA1-...
"$MODELVERSE_INSTALL_DIR"/bin/clikart-cli prompt "Hello" --device cpu
```
`bin/clikart-license-init ` stores it once under your user account instead, and the variable wins when both are present. Without a valid credential the model commands are refused with the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a license past its end date.
From Python, `clika_runtime.modelverse` takes the credential from the process's environment the same way. [License the runtime](/clikart/how-to/license-the-runtime) is the whole contract for both products.
Per platform, the archive to pick and the check that it runs on your machine:
| OS | Architecture | Archive | Verify with | Notes |
| --- | --- | --- | --- | --- |
| Linux | x86_64 | `ClikaRT_linux_x86_64-.tar.xz` | `bin/clikart-cli devices` | CPU always; CUDA with the NVIDIA driver alone, Vulkan with a Vulkan 1.2 driver |
| Linux | arm64 | `ClikaRT_linux_arm64-.tar.xz` | `bin/clikart-cli devices` | CPU · CUDA · Vulkan, as above |
| Windows | x86_64 | `ClikaRT_windows_x86_64-.zip` | `bin\clikart-cli.exe devices` | CPU · Vulkan |
| Windows | arm64 | `ClikaRT_windows_arm64-.zip` | `bin\clikart-cli.exe devices` | CPU · Vulkan |
| macOS | Apple silicon | `ClikaRT_macos_arm64-.tar.xz` | `bin/clikart-cli devices` | CPU · Metal (ships with macOS) |
| Android | arm64-v8a | `ClikaRT_android_arm64-.tar.xz` | push the directory to the device, run over `adb shell` | CPU · Vulkan |
The Linux archives carry both CUDA images, and each image is self-contained: the CUDA runtime and cuBLASLt are inside it, so a CUDA machine needs its NVIDIA driver and nothing else, no CUDA toolkit install and no `LD_LIBRARY_PATH`. The runtime loads the image the driver serves (a driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image).
From Python, Modelverse ships inside the runtime's `clika-runtime` wheel as `clika_runtime.modelverse` ([Get ClikaRT](/clikart/getting-started/get-clikart#python-wheels) picks the wheel for your machine), so one install gives both:
```bash
pip install ./clika_runtime-0.6.4-cp313-cp313-manylinux_2_28_x86_64.whl
python -c "import clika_runtime.modelverse as mv; print(mv.__version__)"
```
The subpackage loads the runtime before its own extension, and a runtime and a model library from different releases refuse to import, naming both versions. The `clikart-cli` executable comes from the same wheel, which installs it as a console script, and from the archive's `bin/`. The C++ library ships in the archive only.
Every archive carries `dist.json` at its top level, which describes the platform build itself (the target it serves, the backends inside, the library layout), and `VERSION`, the release it belongs to. The runtime and the model library in one archive are built together; a model library paired with a runtime of another release is refused with a readable error naming both versions, never undefined behavior.
Accelerators, drivers and hardware are the runtime's: the [ClikaRT system requirements](/clikart/system-requirements) are the authority.
Then continue with [Quick install](installation.md), which picks up at the extracted directory.
---
# Quick install
Extract the release archive and prove Modelverse works in two commands, no toolchain required.
Source: https://docs.clika.io/modelverse/getting-started/installation.md
The fast path from nothing to a running `clikart-cli` executable. It needs no compiler and no runtime dependencies; everything below runs the extracted directory in place.
## 0. Prerequisites
A terminal and `tar` with xz support (Linux, macOS and Windows 10+ have it out of the box); the archive carries the ClikaRT runtime and Modelverse together, so there is no toolchain to install. Running a model needs a credential your platform issued for the project; [License the runtime](/clikart/how-to/license-the-runtime) says where `clikart-cli` reads it. (A C++17 toolchain becomes relevant only in [part 4 of the tutorial](first-model/04-use-it-from-code.mdx), when you embed Modelverse in your own program.)
Running a model needs memory and disk that grow with the model's file size; the [ClikaRT system requirements](/clikart/system-requirements#memory) give the minimum memory per backend, measured at the platform's benchmark settings.
## 1. Extract the archive
One archive per platform ([Get Modelverse](get-modelverse.md) has the download and checksum commands), one directory out of it, and that directory is the whole install:
```bash
tar -xf ClikaRT_linux_x86_64-.tar.xz
export MODELVERSE_INSTALL_DIR="$PWD/ClikaRT_linux_x86_64-"
ls "$MODELVERSE_INSTALL_DIR"
```
The `ls` is the success check: `bin/`, `cmake/`, `include/`, `lib/`, `dist.json` and `VERSION` are there, and `README-Modelverse.md` sits beside the runtime's own `README.md`. The runtime's libraries and the Modelverse library share `lib/`, which is exactly where `bin/clikart-cli` looks for them. A new terminal loses the `export`; re-run it there.
## 2. Place the license credential
Modelverse runs on the ClikaRT runtime, and the runtime runs compute under a license. Put the `CLIKA1-...` credential your project's license shows in the environment, or store it once under your user account:
```bash
export CLIKA_RT_LICENSE=CLIKA1-... # this shell, and the programs it starts
"$MODELVERSE_INSTALL_DIR"/bin/clikart-license-init CLIKA1-... # or once, for this user account
```
[License the runtime](/clikart/how-to/license-the-runtime) has the whole contract.
## 3. Ask it about this machine
The fastest proof the install works: it loads the runtime, probes the compute backends, and prints what it found. Nothing gets installed and no network is touched.
```bash
"$MODELVERSE_INSTALL_DIR"/bin/clikart-cli devices
```
```text
== compute devices on this machine ==
CPU available
CUDA available
Vulkan available
device name driver runtime memory_gb sm_count compute_capability warp max_threads_per_block shared_mem_kb tensor_cores fp16 bf16 fp8_e4m3 status
CPU:0 AMD - - 61.95013 6 - 1 0 0 no yes yes no ok
CUDA:0 NVIDIA GeForce RTX 4080 13.2 13.3 15.569031 76 8.9 32 1024 48 yes yes yes yes ok
Vulkan:0 NVIDIA GeForce RTX 4080 595.71.05 1.3.0 15.9921875 0 - 32 1024 0 yes yes yes no ok
```
Your table shows your hardware; `CPU available` alone is a pass. A backend listed as `compiled, not loaded` was built into this dist but found no usable device or driver on this machine. The `status` column reads `ok` for a device this build can run work on; a device that is present but unusable under this build names the reason there instead of failing at first use.
## 4. Print the catalog
The second command proves the model registry is intact, still with no network:
```bash
"$MODELVERSE_INSTALL_DIR"/bin/clikart-cli list
```
```text
registered model families (52 runnable; 10 hidden, --all shows all)
chatterbox runnable
* Input Modalities: text
* Output Modalities: audio
* Valid Combos: text -> audio
* Commands: speak, serve, bench
* Web UI: generic page
* Vendor: Resemble AI
* License page: https://github.com/resemble-ai/chatterbox/blob/master/LICENSE
-----------------------------------------
...
```
The list is long; `--full` adds each family's description. For convenience, put `bin/` on your `PATH`; the commands in the rest of the documentation assume `clikart-cli` resolves.
```bash
export PATH="$MODELVERSE_INSTALL_DIR/bin:$PATH"
```
## If something failed
- `tar: xz: Cannot exec` or `xz: command not found`: install `xz-utils` (Debian/Ubuntu) or `xz`.
- `No such file or directory` on the binary: the archive does not match this machine; `cat "$MODELVERSE_INSTALL_DIR/dist.json"` names the platform it was built for.
- A shared-library error on start: the binary finds its libraries in `lib/` relative to itself, so run it inside the extracted directory; copying `bin/clikart-cli` out alone breaks that lookup.
- A runtime-mismatch error on start: the model library is built against the runtime of its own release and refuses another, naming both versions; a directory hand-mixed from two releases is refused. Extract one archive, unmodified, and it clears.
The install works. [Pick a model](first-model/01-pick-a-model.md).
---
# What to read next
Where the documentation goes after the tutorial: how-to guides, model requirements, the examples, and the ClikaRT docs underneath.
Source: https://docs.clika.io/modelverse/getting-started/next-steps.md
You installed the archive, fetched a model, served it, and ran it from C++. The rest of the documentation, in a useful reading order:
- **[How-to guides](../how-to/index.md)**: problem-oriented recipes past the tutorial, from picking a GGUF quantization to running with no network at all. [Additional examples](../examples.md) sit inside it: the standalone example programs, from the catalog probe to an in-process server.
- **[Your own models and nodes](../how-to/register-your-own-family.md)**: the how-to group for bringing your own architecture, from the full registration contract to [composing your own pipeline node](../how-to/add-a-pipeline-node.md).
- **[Model requirements](../model-requirements.md)**: the memory and device figures per model variant, for sizing a deployment before you fetch anything.
- **[ClikaRT](/clikart)**: the runtime underneath. Its tutorial explains tensors, devices and the async model the pipelines run on; its API reference covers every public name your C++ program touches; its [system requirements](/clikart/system-requirements) are the platform authority.
---
# Modelverse at a glance
What the Modelverse model library is and is not, and the five-minute mental model.
Source: https://docs.clika.io/modelverse/getting-started/overview.md
Modelverse is the CLIKA model library, built on the [ClikaRT](/clikart) runtime. It is a catalog of model families, the `clikart-cli` executable, which inspects, fetches and runs any model in it, and the library behind both (C++ first, with Python and Kotlin bindings over the same surface). It is not a training framework, not a conversion pipeline, and not a hosted service. Models arrive as the checkpoint files their authors published; Modelverse knows which family they belong to and what that family can do.
## The mental model
Five ideas carry the whole product, in the order you meet them.
1. **The catalog is a registry of model families.** A family (Llama, Qwen, Whisper, CLIP, ...) declares its input and output modalities, the checkpoints it matches, and the commands it provides. The catalog is built into `clikart-cli`; listing it needs no network.
```bash
clikart-cli list
```
2. **A model is a source you name.** A Hugging Face Hub repo id, a pasted Hugging Face URL, or a local directory all name a model. Identity resolution reads config files only, so asking what something is never downloads weights.
```bash
clikart-cli info Qwen/Qwen2.5-0.5B-Instruct
clikart-cli info ./my-model-dir
```
A GGUF repo that ships several quantizations takes a selector, `/:Q6_K`; with several options and no selection, `clikart-cli` refuses and prints the option table instead of guessing.
3. **The `clikart-cli` executable has four commands that work without naming a model: `list`, `devices`, `info`, and `fetch`.** Every other command runs on a model you name: `clikart-cli `. Which commands a model supports depends on what kind of model it is: a text-generation model has `prompt`, `serve`, and `bench`; a speech model has `transcribe`. `clikart-cli --help` prints the commands a model supports:
```bash
clikart-cli --help # that model's commands
clikart-cli prompt "The capital of France is"
```
4. **Serving is one of those commands.** `serve` hosts the model behind an OpenAI-compatible HTTP server, with streaming chat completions, a built-in web chat page, and per-modality routes (transcription, embeddings, depth and more). Existing OpenAI clients connect by changing their base URL.
```bash
clikart-cli serve --port 8000
```
5. **Everything `clikart-cli` does, the library does.** The executable is a thin layer over the `clika_modelverse` library: a snapshot call fetches, the registry loads a runnable model, a serving pipeline generates. Your application makes the same two calls in its own language:
{/* CERTIFICATION: the Python arm's two lines were run against the pinned
release's clika_modelverse wheel and loaded a GenerativeModel. The C++ arm
and the CLI are the other shipped surfaces. The Kotlin arm shows the
Modelverse Kotlin binding's shape; its artifact ships with the releases
that carry it. */}
```cpp
mv::hub::SnapshotResult snapped = mv::hub::snapshot("Qwen/Qwen2.5-0.5B-Instruct", opts);
mv::GenerativeModel model = mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load);
```
```python
snapped = mv.hub.snapshot("Qwen/Qwen2.5-0.5B-Instruct", cache_dir="models")
model = mv.ModelRegistry.builtin().load_generative(snapped.local_dir)
```
```kotlin
val snapped = Hub.snapshot("Qwen/Qwen2.5-0.5B-Instruct", cacheDir = "models")
val model = ModelRegistry.builtin().loadGenerative(snapped.localDir)
```
The [tutorial series](first-model/01-pick-a-model.md) turns these into working sessions, one idea per part.
## Platforms
Modelverse ships inside the ClikaRT release archive, one archive per platform: Linux (x86_64, arm64), Windows (x86_64, arm64), macOS and Android, and from Python it ships inside the `clika-runtime` wheel as `clika_runtime.modelverse`. Extract the archive and run in place; the manifest inside pins the exact runtime this build linked against, so a mismatched pair refuses with a readable error. The [ClikaRT system requirements](/clikart/system-requirements) apply unchanged; Modelverse adds no requirements of its own.
---
# A model from scratch
Write an embedding model's checkpoint by hand (config, tokenizer, random weights), load it through the registry, and compare two texts, all offline.
Source: https://docs.clika.io/modelverse/getting-started/own-model/a-model-from-scratch.md
This series is for adding your own model, as opposed to running the catalog's. It walks the same ground the [first tutorial](../first-model/01-pick-a-model.md) covered from the consumer side, now from the author's side, in three parts with a running result each. The worked model is an embedding model on purpose: text goes in, one vector comes out, and there is no machinery beyond that idea. This part demystifies the checkpoint itself: a model is a directory of files, and you can write one by hand.
The program below synthesizes a tiny embedding checkpoint (an 8-word vocabulary, hidden size 32, a gemma3-shaped text encoder), loads it through the registry, and compares three texts. Nothing downloads and the weights are random; random weights still embed, they are only bad at it, and that is enough to see every file a model needs and where each one enters. Same two-file CMake project as [part 4 of the first tutorial](../first-model/04-use-it-from-code.mdx); only `main.cpp` changes.
## The files an embedding model needs
- `config.json` names the architecture (`model_type`, `architectures`) and its geometry (layers, heads, hidden size). Identity resolution reads exactly this.
- `tokenizer.json` turns text into token ids; the toy one below is a word-level vocabulary of eight entries.
- The weights (`model.safetensors` here) are tensors under the names the architecture expects.
- An embedding checkpoint additionally ships its module chain, the sentence-transformers layout: `modules.json` lists the stages (transformer, pooling, dense heads, normalize), and each configurable stage carries its own small directory.
## The program
```cpp title="main.cpp"
#include
#include
#include
#include
#include
#include
#include
#include "clika_modelverse/modules/pooling.h"
#include "clika_modelverse/registry/model_registry.h"
namespace fs = std::filesystem;
namespace mv = clika_modelverse;
using ClikaRT::DataType;
using ClikaRT::Tensor;
// The whole identity: model_type and architectures are what the registry
// matches; the rest is geometry the loader shapes the encoder from.
constexpr const char* kConfigJson = R"({
"model_type": "gemma3_text",
"architectures": ["Gemma3TextModel"],
"hidden_size": 32,
"num_hidden_layers": 4,
"num_attention_heads": 4,
"num_key_value_heads": 2,
"head_dim": 8,
"vocab_size": 8,
"intermediate_size": 64,
"max_position_embeddings": 128,
"rms_norm_eps": 1e-6,
"rope_theta": 1000000.0,
"rope_local_base_freq": 10000.0,
"sliding_window": 16,
"sliding_window_pattern": 2,
"query_pre_attn_scalar": 8,
"use_bidirectional_attention": true
})";
// A word-level tokenizer: eight words, one id each.
constexpr const char* kTokenizerJson = R"({
"version": "1.0",
"pre_tokenizer": {"type": "WhitespaceSplit"},
"model": {"type": "WordLevel",
"vocab": {"a": 0, "b": 1, "c": 2, "d": 3, "e": 4, "f": 5, "g": 6, "[eos]": 7},
"unk_token": "a"},
"decoder": null,
"added_tokens": [{"id": 7, "content": "[eos]", "special": true, "normalized": false}]
})";
// The module chain: transformer -> mean pooling -> two dense heads -> normalize.
constexpr const char* kModulesJson = R"([
{"idx": 0, "name": "0", "path": "", "type": "sentence_transformers.models.Transformer"},
{"idx": 1, "name": "1", "path": "1_Pooling", "type": "sentence_transformers.models.Pooling"},
{"idx": 2, "name": "2", "path": "2_Dense", "type": "sentence_transformers.models.Dense"},
{"idx": 3, "name": "3", "path": "3_Dense", "type": "sentence_transformers.models.Dense"},
{"idx": 4, "name": "4", "path": "4_Normalize", "type": "sentence_transformers.models.Normalize"}
])";
constexpr const char* kPoolingJson = R"({
"word_embedding_dimension": 32,
"pooling_mode_cls_token": false,
"pooling_mode_mean_tokens": true,
"pooling_mode_max_tokens": false,
"pooling_mode_mean_sqrt_len_tokens": false,
"pooling_mode_weightedmean_tokens": false,
"pooling_mode_lasttoken": false
})";
Tensor randn(std::mt19937& rng, std::vector shape) {
std::int64_t n = 1;
for (const std::int64_t d : shape) n *= d;
std::normal_distribution dist(0.0f, 0.2f);
std::vector host(static_cast(n));
for (float& v : host) v = dist(rng);
return Tensor::from_data(host.data(), shape, DataType::Float32);
}
// One dense head: its own directory with a config and one weight.
void write_dense(std::mt19937& rng, const fs::path& dir, std::int64_t in,
std::int64_t out) {
fs::create_directories(dir);
std::ofstream(dir / "config.json")
<< R"({"in_features": )" << in << R"(, "out_features": )" << out
<< R"(, "bias": false, "activation_function": "torch.nn.modules.linear.Identity"})";
ClikaRT::NamedTensors sd;
sd.set("linear.weight", randn(rng, {out, in}));
ClikaRT::io::save_safetensors(sd, (dir / "model.safetensors").string());
}
// Random weights under the names the gemma3-shaped encoder expects (bare
// keys, no prefix); the fixed seed keeps every run identical.
fs::path write_snapshot() {
const fs::path dir = fs::temp_directory_path() / "my_first_model";
fs::create_directories(dir);
std::ofstream(dir / "config.json") << kConfigJson;
std::ofstream(dir / "tokenizer.json") << kTokenizerJson;
std::ofstream(dir / "modules.json") << kModulesJson;
fs::create_directories(dir / "1_Pooling");
std::ofstream(dir / "1_Pooling" / "config.json") << kPoolingJson;
constexpr std::int64_t kVocab = 8, kHidden = 32, kLayers = 4, kHeads = 4;
constexpr std::int64_t kKvHeads = 2, kHeadDim = 8, kFfn = 64;
constexpr std::int64_t kDenseMid = 48, kDim = 16;
std::mt19937 rng(20260831);
ClikaRT::NamedTensors sd;
const auto put = [&](const std::string& name, std::vector shape) {
sd.set(name, randn(rng, std::move(shape)));
};
put("embed_tokens.weight", {kVocab, kHidden});
put("norm.weight", {kHidden});
for (std::int64_t l = 0; l < kLayers; ++l) {
const std::string p = "layers." + std::to_string(l) + ".";
put(p + "self_attn.q_proj.weight", {kHeads * kHeadDim, kHidden});
put(p + "self_attn.k_proj.weight", {kKvHeads * kHeadDim, kHidden});
put(p + "self_attn.v_proj.weight", {kKvHeads * kHeadDim, kHidden});
put(p + "self_attn.o_proj.weight", {kHidden, kHeads * kHeadDim});
put(p + "self_attn.q_norm.weight", {kHeadDim});
put(p + "self_attn.k_norm.weight", {kHeadDim});
put(p + "mlp.gate_proj.weight", {kFfn, kHidden});
put(p + "mlp.up_proj.weight", {kFfn, kHidden});
put(p + "mlp.down_proj.weight", {kHidden, kFfn});
put(p + "input_layernorm.weight", {kHidden});
put(p + "post_attention_layernorm.weight", {kHidden});
put(p + "pre_feedforward_layernorm.weight", {kHidden});
put(p + "post_feedforward_layernorm.weight", {kHidden});
}
ClikaRT::io::save_safetensors(sd, (dir / "model.safetensors").string());
write_dense(rng, dir / "2_Dense", kHidden, kDenseMid);
write_dense(rng, dir / "3_Dense", kDenseMid, kDim);
return dir;
}
int main() {
// 1. A checkpoint is a directory; this one is yours, written just now.
const fs::path dir = write_snapshot();
std::printf("wrote %s\n", dir.string().c_str());
// 2. The registry reads config.json, matches architecture
// "Gemma3TextModel" to the gemma-embedding family, and returns the
// runnable model, exactly as for a fetched checkpoint.
mv::EmbeddingModel model =
mv::ModelRegistry::builtin().load_embedding(dir.string(), {});
// 3. Embed three texts: one [3, 16] Float32 tensor, one row per text
// (the module chain pools, projects to 16 dims, and normalizes).
const std::string texts[] = {"a b c", "a b d", "f g"};
const Tensor rows = model.embed(texts);
std::printf("embeddings: %s\n", rows.to_string().c_str());
// 4. Compare them: the cosine matrix is unit rows times their own
// transpose. The explicit normalize keeps the demo self-contained
// (the chain's Normalize stage already produced unit rows).
const Tensor unit = mv::modules::l2_normalize_rows(rows);
const Tensor sim = ClikaRT::ops::matmul(unit, unit, {}, std::nullopt,
false, /*transpose_b=*/true);
const std::vector s = ClikaRT::ops::reshape(sim, {9}).item_as_vec();
std::printf("close pair (a b c ~ a b d): %.3f\n", s[1]);
std::printf("far pair (a b c ~ f g): %.3f\n", s[2]);
return 0;
}
```
```bash
cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake"
cmake --build build
./build/hello
```
```text
[2026-09-22 00:30:20.206] [modelverse] [info] hub: '/tmp/my_first_model' already on disk; loading from /tmp/my_first_model
[2026-09-22 00:30:20.208] [gemma-embedding] [info] embedding encoder on CPU:-1 (dim 16)
wrote /tmp/my_first_model
embeddings: Tensor(shape=[3, 16], dtype=Float32, device=CPU, numel=48, data=[0.03671, -0.04182, -0.2858, 0.1808, 0.1262, 0.399, ...])
close pair (a b c ~ a b d): 0.400
far pair (a b c ~ f g): -0.324
```
Random weights, so the numbers mean little; what matters is the shape of what happened. A directory you wrote from scratch went through the same registry, loader and embedding surface as a fetched checkpoint would, and three texts became three vectors you can compare, because a checkpoint is nothing more than these files.
## What the registry did with it
`config.json`'s `model_type` and `architectures` are the identity keys; the architecture `Gemma3TextModel` matched the built-in `gemma-embedding` family, and that family's loader shaped the encoder from the geometry fields and bound your `model.safetensors` names to it. The tokenizer file turned each text into ids before the encoder saw them, and `modules.json` told the loader what follows the encoder: mean pooling, the two dense projections, and the final normalize. Delete the directory and nothing else remembers it.
Your own architecture will not say `"architectures": ["Gemma3TextModel"]`, and then the match fails; that refusal, and fixing it by registering a family of your own, is [part 2](02-join-the-catalog.md).
---
# Your architecture joins the catalog
Rename the toy model's architecture so nothing matches it, then register a family of your own and watch the same directory resolve.
Source: https://docs.clika.io/modelverse/getting-started/own-model/join-the-catalog.md
Part 1 rode the built-in `gemma-embedding` family. Your real architecture has its own name, and the catalog does not know it; this part makes the failure visible and then fixes it the way every built-in family fixes it, by self-registration.
## Break the match first
In part 1's `kConfigJson`, rename the identity keys the way a Hugging Face checkpoint of your own architecture would name them:
```json
"model_type": "my_model",
"architectures": ["MyModelModel"],
```
Rebuild and run, and the registry refuses precisely; the raised `ClikaRT::Error` carries:
```text
wrote /tmp/my_first_model
no model template registered for model_type='my_model' / architecture='MyModelModel'
```
Nothing else changed. The files are fine; the catalog has no entry whose identity keys match them.
## Register the family
A family is one translation unit that pushes a `ModelRegistration` into the shared registry at static-initialization time, and a TU compiled into your own program registers exactly like a built-in one. Add this file to the project and add it to `add_executable`:
```cpp title="my_model_family.cpp"
#include "clika_modelverse/models/registration.h"
namespace {
using namespace clika_modelverse;
// The identity template: what a resolved checkpoint of this family IS.
// The registry fills architecture, model_type and variant from config.json.
class MyModel final : public Model {
public:
std::string_view family() const noexcept override { return "my-model"; }
};
constexpr std::string_view kModelTypes[] = {"my_model"};
constexpr std::string_view kArchitectures[] = {"MyModelModel"};
// Constructing the registrar is the whole hookup; there is no list to edit.
const ModelFamilyRegistrar kRegistrar{
[] {
ModelRegistration registration{};
registration.metadata.family = "my-model";
registration.metadata.vendor = "my-org";
registration.metadata.description = "the tutorial's own model family";
registration.metadata.input_combos = kTextOnlyCombos;
registration.metadata.output_combos = kTextOnlyCombos;
registration.hf_model_types = kModelTypes;
registration.hf_architectures = kArchitectures;
registration.factory = make_model();
return registration;
}(),
};
} // namespace
```
To watch it resolve, have `main` ask for identity instead of embeddings for a moment:
```cpp
const std::unique_ptr matched = mv::ModelRegistry::builtin().open(dir.string());
std::printf("%s resolved: family=%s model_type=%s\n", dir.string().c_str(),
std::string(matched->family()).c_str(),
std::string(matched->model_type()).c_str());
```
```text
wrote /tmp/my_first_model
/tmp/my_first_model resolved: family=my-model model_type=my_model
```
## What you registered, and what you did not
The registration so far is identity-only: `metadata` (the family's name, its publisher, its declared modalities), the match keys, and the `factory` that builds the identity object. That is enough for the directory to resolve, for the family to appear in a catalog listing, and for every runnable command to refuse with a precise message naming what the family provides (nothing yet). An incomplete declaration (no family name, no modalities, no identity factory) is refused at process start, not papered over downstream.
Running is a separate concern on the same registration, and that split is deliberate: identity must stay cheap (no weights) and total (every checkpoint of yours resolves), while running is opt-in per capability. Wiring it is [part 3](03-make-it-embed.md); the field-by-field reference for everything a registration can carry is [Register your own model family](../../how-to/register-your-own-family.md).
---
# Make it embed
Give the my-model family the smallest real embedding backend, opt into the similar and serve commands, and compare texts through your own family.
Source: https://docs.clika.io/modelverse/getting-started/own-model/make-it-embed.md
Part 2's family resolves but cannot run. For an embedding family the whole running contract is one small interface, `EmbeddingBackend` (`clika_modelverse/models/base/embedding_model.h`): token ids in, embedding rows out, plus the width, the pooling declaration and the device. This part implements the smallest real backend, wires the factory, and opts into the standard commands.
## The smallest real backend
The toy backend embeds each text as the mean of its tokens' embedding-table rows, L2-normalized. That is a real embedding model (the bag-of-words baseline retrieval systems start from), and it needs exactly one weight:
```cpp title="my_model_backend.cpp"
#include "clika_modelverse/models/base/embedding_model.h"
namespace {
using namespace clika_modelverse;
class MyModelBackend final : public EmbeddingBackend {
public:
MyModelBackend(ClikaRT::Tensor table, ClikaRT::Device device)
: table_(std::move(table)), device_(device) {}
ClikaRT::Tensor embed_ids(
const std::vector>& sequences) const override {
std::vector rows;
rows.reserve(sequences.size());
for (const std::vector& ids : sequences) {
const ClikaRT::Tensor picked = ClikaRT::ops::index_select(table_, 0, ids);
rows.push_back(ClikaRT::ops::mean(picked, /*dim=*/0));
}
return ClikaRT::ops::l2_normalize(ClikaRT::ops::stack(rows), /*dim=*/-1);
}
std::int64_t embedding_dim() const override { return table_.shape()[1]; }
Pooling pooling() const override { return Pooling::Mean; }
bool normalizes() const override { return true; }
ClikaRT::Device device() const override { return device_; }
private:
ClikaRT::Tensor table_;
ClikaRT::Device device_;
};
} // namespace
```
A real encoder runs its layers between the lookup and the pooling; where those layers go, and how a checkpoint's names bind onto them, is the [full contract's](../../how-to/register-your-own-family.md) territory. The interface does not change with the depth: however sophisticated the encoder, it enters the family as this same `EmbeddingBackend`. One optional declaration to know: `max_concurrent_sessions()` says how many requests one handle's forward may run at once, and its default of 1 is the safe answer for a backend that keeps per-request state; a forward that is a pure function of its inputs and weights can declare itself unbounded.
## The registration delta
The factory reads the snapshot, builds the backend from the one weight it uses, and wraps it with the snapshot's tokenizer; the registration gains the factory and the command surface:
```cpp
ClikaRT::Result build_my_model(const std::string& dir,
const Model& identity,
const LoadOptions& options) {
auto weights = ClikaRT::io::load_safetensors(dir + "/model.safetensors");
auto table = weights.get("embeddings.word_embeddings.weight").to(options.device);
auto backend = std::make_unique(std::move(table), options.device);
auto tokenizer = ClikaRT::tokenizer::from_huggingface(dir);
return EmbeddingModel(std::move(backend), std::move(tokenizer),
std::string(identity.family()),
std::string(identity.model_type()));
}
// The command surface, composed from the same public builders every
// embedding family uses: `similar` and `serve`, two lines.
FamilyApp build_my_model_cli(const FamilyCliContext& context) {
FamilyApp fam{batteries::cli::make_family_app(
context, "commands the 'my-model' family provides"), {}};
batteries::cli::similar_command(fam, context);
batteries::cli::serve_command(fam, context,
batteries::serving::generic_serve);
return fam;
}
```
```cpp
registration.cli = CliSurface{build_my_model_cli, my_model_cli_verbs()};
registration.factory = make_model();
registration.build_embedding = build_my_model;
```
## Run it
Part 1's `main` works again unchanged (it never named a family, only a directory), now through your own:
```text
wrote /tmp/my_first_model
embeddings: Tensor(shape=[3, 32], dtype=Float32, device=CPU:0, numel=96, ...)
close pair (a b c ~ a b d): 0.667
far pair (a b c ~ f g): -0.041
```
The close pair scores closer than in part 1, and honestly so: mean-pooled bags of shared words ARE similar, which is exactly what this backend measures. And because the commands came from the shared battery, your family now serves them like any catalog family: `similar` compares texts from the command line, and `serve` answers `POST /v1/embeddings` and `POST /v1/similarity` ([the route table](../../how-to/serve-openai-compatible.md)).
## Where the full contract lives
Everything real that the toy skipped is registration fields on the same struct, documented in [Register your own model family](../../how-to/register-your-own-family.md) and field-by-field in `clika_modelverse/models/registration.h`: encoders with layers and their weight binding, the generative, speech and reranking factories with their GGUF twins, `probe_snapshot` for repos without a `config.json`, and custom verb surfaces. For custom logic AROUND a catalog model rather than a model of your own, [Add your own node to a model pipeline](../../how-to/add-a-pipeline-node.md) is the shorter road.
---
# How-to guides
Problem-oriented recipes. Each guide answers one "how do I X" with a worked session and its real output.
Source: https://docs.clika.io/modelverse/how-to.md
Practical guides covering common tasks. Each guide answers one concrete "how do I X" with a worked session: real commands or real code, with the output they produce. Read the [tutorial](../getting-started/first-model/01-pick-a-model.md) first; the guides assume its ground (sources, fetching, the family-owned commands) and go deeper on one problem at a time, in any order.
## Models and weights
- [Run a specific GGUF quantization](run-a-gguf-quantization.md): pick one weight option of a repo that ships many, by tag, glob or exact file.
- [Run fully offline](run-fully-offline.md): fetch on a connected machine, move the directory, and make any network touch an error.
## Running and serving
- [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md): the full route table, streaming, operational flags, and pointing existing clients at it.
- [Transcribe audio](transcribe-audio.md): speech-to-text from the command line and over HTTP.
- [Translate text](translate-text.md): neural machine translation from the command line and over HTTP, with the model's own language codes.
- [Speak text](speak-text.md): text-to-speech with voice references and the synthesis knobs.
- [Estimate depth, segment and detect](depth-segment-detect.md): the three image verbs, one artifact per input, open-vocabulary detection included.
- [Benchmark a model on this machine](bench-a-model.md): the `bench` flow, its sweep axes and how to read the structured report.
- [Embed, compare and rerank](embed-and-rerank.md): dense vectors, similarity matrices and cross-encoder reranking, for texts and images.
## Scripting and integration
- [Script clikart-cli](script-the-cli.md): the exit-code contract, the stdout/stderr split, option templates and shell completion.
- [Add Modelverse to an existing CMake project](existing-cmake-project.md): `find_package` against the installed directory, staging the runtime beside your binary, and picking the dist on cross builds.
## Your own models and nodes
The custom-model track at its full depth; the gentle guided walk is the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md).
- [Register your own model family](register-your-own-family.md): what a snapshot directory must contain, identity matching, and the registration that makes your architecture runnable.
- [Add your own node to a model pipeline](add-a-pipeline-node.md): compose custom pre- or post-processing with the zoo's serving nodes, on the same request surface.
- [A conversational AI as one pipeline](conversational-ai-pipeline.md): speech-to-text, a chat model and text-to-speech composed into a single pipeline, wav bytes in and the spoken answer out.
## Complete programs
- [Additional examples](../examples.md): the standalone example programs, one per subsystem.
Sizing questions (which variant fits which device) live in [Model requirements](../model-requirements.md). For every public C++ name, the ClikaRT API reference on [/clikart](/clikart).
---
# Add your own node to a model pipeline
Compose custom pre- or post-processing with the zoo's serving nodes in one pipeline, on the same request surface.
Source: https://docs.clika.io/modelverse/how-to/add-a-pipeline-node.md
Your application needs logic the model does not have: redaction, templating, routing, scoring, any transform of what goes in or comes out. In Modelverse that logic is not a wrapper around the pipeline; it is a node inside it. The zoo's serving nodes and your code meet on one contract, `ClikaRT::runtime::Model`, and `ClikaRT::runtime::Pipeline::create` composes any mix of them into one pipeline with one request/response surface.
A node implements three members: `schema()` declares its input and output tensors by name, `phases()` declares its execution phases, and `Phase_RunOnce` does the work. The zoo's tokenizer, generative decoder and detokenizer implement exactly the same three, which is why yours can stand beside them without an adapter layer.
## The program
Same two-file project as [part 4 of the tutorial](../getting-started/first-model/04-use-it-from-code.mdx); only `main.cpp` changes. The custom node here uppercases the reply, standing in for any post-processing you own:
```cpp title="main.cpp"
#include
#include
#include
#include
#include
#include
#include
#include "clika_modelverse/generation/generation_config.h"
#include "clika_modelverse/hub/hub.h"
#include "clika_modelverse/registry/model_registry.h"
#include "clika_modelverse/runtime/decoder_node.h"
#include "clika_modelverse/runtime/pipeline_io.h"
#include "clika_modelverse/runtime/text_nodes.h"
namespace mv = clika_modelverse;
namespace mvr = clika_modelverse::runtime;
namespace crt = ClikaRT::runtime;
using ClikaRT::DataType;
// The detokenizer's reply routes here under this name instead of going
// straight out.
constexpr const char* kDraftText = "draft_text";
// Everything a user writes: declare I/O, declare a phase, do the work.
class ShoutNode final : public crt::Model {
public:
ShoutNode() {
schema_.inputs.push_back(ClikaRT::spec::TensorSpec{
kDraftText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim},
/*optional=*/false});
schema_.outputs.push_back(ClikaRT::spec::TensorSpec{
mvr::kText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim},
/*optional=*/false});
phases_.push_back(crt::PhaseSpec{"shout", crt::PhaseKind::RunOnce});
}
const crt::ModelSchema& schema() const override { return schema_; }
ClikaRT::Span phases() const override { return phases_; }
void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override {
std::string text = mvr::byte_tensor_to_string(ctx.inputs->get(kDraftText));
std::transform(text.begin(), text.end(), text.begin(), [](unsigned char c) {
return static_cast(std::toupper(c));
});
ctx.outputs->set(mvr::kText, mvr::string_to_byte_tensor(text));
}
private:
crt::ModelSchema schema_;
std::vector phases_;
};
int main() {
// The zoo half, exactly as in the tutorial: snapshot, registry, model.
mv::hub::SnapshotOptions snap;
snap.cache_dir = "models";
const mv::hub::SnapshotResult snapped = mv::hub::snapshot(
"Qwen/Qwen2.5-0.5B-Instruct", snap);
mv::LoadOptions load;
load.max_seq = 4096; // the CLI's default context cap, as in part 4
mv::GenerativeModel model =
mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load);
mv::generation::GenerationConfig gen = model.defaults();
gen.max_new_tokens = 24;
gen.stop = {"."}; // raw completion: one sentence is the demo
// The zoo's three serving nodes, constructed directly.
mvr::TokenizerNode tok(model.tokenizer(), /*add_special_tokens=*/false);
mvr::GenerativeDecoderModel dec(model.provider(), gen, /*max_active=*/1,
mv::KVCacheMode::Continuous,
&model.tokenizer(),
/*prefill_chunk_tokens=*/0);
mvr::DetokenizerNode detok(model.tokenizer());
ShoutNode shout; // yours
// One pipeline over all four. Edges wire by name; the one rename (the
// detokenizer's `text` becomes `draft_text`) routes the reply through
// the custom node instead of straight out.
std::vector steps(4);
steps[0].name = "tokenize";
steps[0].model = &tok;
steps[1].name = "decode";
steps[1].model = &dec;
steps[2].name = "detokenize";
steps[2].model = &detok;
steps[2].output_map.emplace_back(mvr::kText, kDraftText);
steps[3].name = "shout";
steps[3].model = &shout;
// The pipeline's public request surface: what callers set and read.
crt::ModelSchema ext;
ext.inputs.push_back(ClikaRT::spec::TensorSpec{
mvr::kPrompt, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim},
/*optional=*/false});
ext.outputs.push_back(ClikaRT::spec::TensorSpec{
mvr::kText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim},
/*optional=*/false});
crt::Pipeline pipeline = crt::Pipeline::create(std::move(steps), std::move(ext));
// The same Request/Response surface every pipeline speaks.
crt::Request req;
req.inputs.set(mvr::kPrompt, mvr::string_to_byte_tensor(
"Once upon a time, in a port town by a cold sea,"));
const crt::SessionId sid = pipeline.enqueue(std::move(req));
const crt::Response resp = pipeline.await(sid);
std::printf("%s\n", mvr::byte_tensor_to_string(resp.outputs.get(mvr::kText)).c_str());
pipeline.shutdown();
return 0;
}
```
```text
THERE LIVED A YOUNG GIRL NAMED KAITO.
```
One build flag matters here: the custom node derives from the serving runtime's `Model`, so its translation unit compiles with `-fno-rtti`, matching how the library builds; without it the link fails on `typeinfo for ClikaRT::runtime::Model`. The [ClikaRT integration guide](/clikart/how-to/existing-cmake-project) covers the flag and which consumers need it.
## What the composition rests on
- **Edges wire by tensor name.** Each step's outputs feed the next step's same-named inputs; `output_map` renames one edge where names must differ. The single rename above is the whole routing change: without it the detokenizer's `text` would leave the pipeline directly.
- **The external schema is the caller's contract.** Only what it declares is settable and readable from outside; everything between the nodes stays internal. Requests and responses are exactly the ones part 4 used, so a caller cannot tell a customized pipeline from a stock one.
- **Placement in the chain is yours.** A node before the tokenizer transforms the prompt (templating, redaction); a node after the detokenizer transforms the reply, as here; a fully custom model runs beside a zoo model in the same pipeline.
The `02_custom_node` program in [Additional examples](../examples.md) is this guide's offline twin: it synthesizes a toy checkpoint at run time, so the composition runs in moments with no download and no arguments. When HTTP is the goal, mount an engine on the built-in server instead ([Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md)); a pipeline node changes what a model computes, an engine changes what the server serves. And when the model itself is yours rather than the catalog's, that is the other half of the custom track: the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md), with [Register your own model family](register-your-own-family.md) as its full contract.
---
# Benchmark a model on this machine
The bench flow: serving-shaped sweep cells, the structured report, and how to read its columns before committing a deployment.
Source: https://docs.clika.io/modelverse/how-to/bench-a-model.md
Whether a model meets your latency and throughput targets is a property of this machine, this quantization and this serving shape, and `bench` measures exactly that. Every runnable text family provides it, the flow drives the same serving pipeline `prompt` and `serve` run on, and the result is a structured table, never ad-hoc prints.
## The default sweep
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct bench --device cuda
```
The default sweep runs three serving-shaped cells (balanced `128:128`, prefill-heavy `2048:128`, decode-heavy `128:1024`) over the default concurrency ladder (1, 4, 8), each cell strictly one at a time so cells never contend with each other. One CSV row per cell and concurrency, the last column naming what the row measured:
```text
isl,osl,concurrency,first_call_ms,ttft_ms,prefill_tok_s,itl_p50_ms,itl_p99_ms,prefill_chunk_ms_p50,prefill_chunk_ms_max,decode_tok_s,speedup_vs_c1,batching_efficiency,admitted_tok,retires,peak_active_bytes,num_allocs,num_inflight_park_waits,status,reason,notes
128,128,1,685.7974,14.375232,8904.204,4.709812,11.978596,0,0,194.13823,1,1,112,3,3422640128,59180,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,128,4,937.06464,25.81384,4969.912,6.5161915,15.847513,0,0,552.3547,2.845162,0.7112905,448,12,3746307584,60982,6,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,128,8,745.5794,36.65184,3492.3213,4.9885244,10.634405,0,0,1434.3369,7.388225,0.92352813,896,24,4083747328,93121,12,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,1,820.59033,67.9156,30155.074,5.575619,8.472811,0,0,159.87682,1,1,2032,3,4083747328,65325,20,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,4,1258.1548,245.00293,8360.729,7.7572827,15.318566,0,0,414.64108,2.5935037,0.6483759,8128,12,4192593408,67393,41,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,8,1308.3411,464.48145,4409.434,6.383371,11.002449,0,0,787.79083,4.927487,0.61593586,16256,24,4744088576,99717,63,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,1,4980.8936,10.374589,12337.838,4.823417,8.758105,0,0,190.86696,1,1,112,3,4744088576,498596,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,4,7377.311,26.746096,4836.729,6.6391444,13.775986,0,0,592.14624,3.102403,0.77560073,448,12,4744088576,508606,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,8,5360.8525,39.64144,3232.0154,5.407993,9.28407,0,0,1437.6843,7.5323896,0.9415487,896,24,4744088576,767141,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
```
The numbers are one machine's run (one workstation GPU, this model at its checkpoint dtype, bf16, cold prompts, the default); yours are the point. Stdout carries the CSV, ready for a spreadsheet or a script; diagnostics ride stderr. A failed cell reports `error` in `status` without invalidating the rows around it, and the `reason` column carries the model's own refusal or the cell's first failure in words (empty on an `ok` row). With `--output`, the report streams to disk as rows complete, so a run killed at its deadline keeps every finished cell.
## Shape the sweep to your workload
- `--cells 512:256,4096:64`: your own `isl:osl` list, replacing the default three. `--isl`/`--osl` set a single cell directly.
- `--concurrency 1,4,16`: the parallel-session ladder; `speedup_vs_c1` and `batching_efficiency` in the report tell you what the added concurrency actually bought.
- `--shared-prefix 256`: how many prompt tokens every session shares (a system prompt, a RAG preamble); with it the `admitted_tok` and `retires` columns show the paged prefix cache engaging. Every LLM cell runs cold prompts by default (a fresh random prompt per session every iteration, so a paged prefix cache serves no repeat; `--unique-prompts` names that default and stays accepted). `--repeat-prompts` is the warm protocol: the warm-up and every iteration reuse one prompt set, so `prefill_tok_s` reads cache admits, and the row's `notes` say so.
- `--warmup` and `--iters`: untimed full passes of each cell before it measures, then timed passes averaged. The default is two warm-up passes and one timed iteration; `--warmup 0` times a cold cell. A cell still cold after the default passes is a defect worth reporting, not a reason to raise the count.
- The load knobs apply unchanged: `--kvcache paged|continuous` (each verb has its own default and `bench` runs paged), `--prefill-chunk`, `--step-token-budget`, a GGUF weight selector on the source, `--max-seq`.
- `--memory-stages FILE` appends the runtime's memory-pool counters to a JSON file at each stage of the model's life (after the weights bind, after the one-forward warm-up, after the first generation), one row per device, with the checkpoint's size. It is the footprint diagnostic behind the figures on [Model requirements](../model-requirements.md).
- `--output PREFIX` writes the report to files beside the terminal view; `--profile` adds the per-op profile on stderr (the summary opens with `Profiling Report Summary:`, its `top_ops:` block ranks op groups by self time, and the per-op table carries an `AvgSelf(us)` column, the per-op average that excludes nested spans), `--profile-pipeline` the per-node one.
## Reading the columns
`ttft_ms` is the enqueue-to-first-token wall, the number an interactive user feels; `first_call_ms` is the cell's very first call, which pays the kernel loads a fresh process owes and is kept out of `ttft_ms` for that reason. `prefill_tok_s` is prompt ingest; `decode_tok_s` is the aggregate generation rate across sessions. The inter-token percentiles `itl_p50_ms`/`itl_p99_ms` sample decode-only gaps; on a chunked-prefill run the chunk boundary walls report separately (`prefill_chunk_ms_p50`/`_max`), so a decode tail is never polluted by an ingest wall. When the p99 sits far above the p50, the deployment story is a batching or cache-pressure story, and the concurrency ladder narrows down which. Three columns carry the memory story: `peak_active_bytes` is the runtime pool's high-water mark for the cell, `num_allocs` the allocation count behind it, and `num_inflight_park_waits` the number of times a request waited for a slot rather than proceeding, a direct read on admission pressure. The `-v` diagnostics carry the load and peak-memory figures the [Model requirements](../model-requirements.md) page's estimates are checked against.
## Audio and media axes
A speech-to-text family's `bench` sweeps audio length instead of token shapes (`--audio-seconds 5,30,120`, with the same `--concurrency` ladder), reporting ingest as encoder frames over the end-to-end transcribe wall. Multimodal families add `mm_bench`, the media sweep over the serving path: `--images` per request, `--image-resolution`, `--video-frames`, with `effective_isl` reporting what the prefill actually ingested once media expansions counted, so a media-heavy row's ingest rate stays comparable to a text row's.
Benchmark the quantization you intend to ship ([Run a specific GGUF quantization](run-a-gguf-quantization.md)); a `Q4_K_M` and a `Q8_0` of the same model can sit on different sides of a latency target. The serving flags that shaped a good bench row carry directly onto `serve` ([Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md)).
---
# A conversational AI as one pipeline
Compose speech-to-text, a chat model and text-to-speech into a single runtime::Pipeline: wav bytes in, the spoken answer out.
Source: https://docs.clika.io/modelverse/how-to/conversational-ai-pipeline.md
[Add your own node to a model pipeline](add-a-pipeline-node.md) put one custom node beside the zoo's text trio. This guide composes a complete application the same way: a conversational AI as one `ClikaRT::runtime::Pipeline`, a spoken question in as wav bytes, the spoken answer out. Eight steps, three of them the zoo's serving nodes, five of them yours:
| step | node | in -> out |
| --- | --- | --- |
| listen | yours | wav bytes -> whisper features |
| transcribe | yours | features -> question text |
| template | yours | question -> the chat-templated prompt |
| tokenize | zoo `TokenizerNode` | prompt -> input_ids |
| decode | zoo `GenerativeDecoderModel` | input_ids -> tokens, continuous-batched |
| detokenize | zoo `DetokenizerNode` | tokens -> answer text |
| speak | yours | answer text -> waveform and its sample rate |
| post | yours | waveform -> the peak-normalized waveform |
The complete program is `examples/cpp/modelverse/03_conversational_pipeline.cpp` in the examples tree.
## The models
Three registry loads. The TTS is a voice-cloning family and requires a voice reference; any short spoken wav works. `max_seq` sizes the decoder's KV pool, and a conversation turn needs nowhere near a 32768-token context window, so cap it:
```cpp
constexpr const char* kSttRepo = "openai/whisper-large-v3-turbo";
constexpr const char* kLlmRepo = "Qwen/Qwen2.5-1.5B-Instruct";
constexpr const char* kTtsRepo = "ResembleAI/chatterbox-flash";
constexpr std::int64_t kLlmMaxSeq = 8192;
```
```cpp
const mv::ModelRegistry& registry = mv::ModelRegistry::builtin();
mv::LoadOptions load;
load.device = device.value();
mv::LoadOptions llm_load = load;
llm_load.max_seq = kLlmMaxSeq;
const mv::SttModel stt = unwrap(registry.load_stt(kSttRepo, load));
const mv::GenerativeModel llm = unwrap(registry.load_generative(kLlmRepo, llm_load));
const mv::TtsModel tts = unwrap(registry.load_tts(kTtsRepo, load));
```
**Placement.** `--device` puts all three models on one device (the CPU by default, `cuda` on request). A single request runs the steps in sequence, so the models share that device without contention.
## The custom nodes
A node is the same three-member `runtime::Model` contract the [custom-node guide](add-a-pipeline-node.md) walks: `schema()` names the input and output tensors, `phases()` declares one `RunOnce` phase, `Phase_RunOnce` does the work. The speech nodes wrap the model handles they are given; the tensor specs they share are three small helpers over `spec::TensorSpec` (a byte string, a Float32 tensor of dynamic dims, an Int32 scalar):
```cpp
spec::TensorSpec bytes_spec(const char* name) {
return spec::TensorSpec{name, DataType::UInt8, {spec::TensorSpec::kDynamicDim},
/*optional=*/false};
}
spec::TensorSpec f32_spec(const char* name, std::size_t rank) {
return spec::TensorSpec{name, DataType::Float32,
std::vector(rank, spec::TensorSpec::kDynamicDim),
/*optional=*/false};
}
spec::TensorSpec i32_scalar_spec(const char* name) {
return spec::TensorSpec{name, DataType::Int32, {}, /*optional=*/false};
}
```
The listen node turns the request's encoded wav bytes into the whisper feature tensor through the model's own preprocessor configuration:
```cpp
class ListenNode final : public crt::Model {
public:
explicit ListenNode(const mv::SttModel& stt) : stt_(&stt) {
this->schema_.inputs.push_back(bytes_spec(kAudioIn));
this->schema_.outputs.push_back(f32_spec(kFeatures, 3));
this->phases_.push_back(crt::PhaseSpec{"listen", crt::PhaseKind::RunOnce});
}
const crt::ModelSchema& schema() const override { return this->schema_; }
ClikaRT::Span phases() const override { return this->phases_; }
void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override {
const std::vector wav =
unwrap(ctx.inputs->get(kAudioIn)).item_as_vec();
ctx.outputs->set(kFeatures, this->stt_->features_from_bytes(wav));
}
private:
const mv::SttModel* stt_;
crt::ModelSchema schema_;
std::vector phases_;
};
```
The transcribe node is the same shape over `SttModel::transcribe_features`, and the speak node wraps `TtsModel::synthesize` with the required voice reference, emitting the waveform plus its sample rate as two outputs:
```cpp
void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override {
mv::SynthesizeOptions opts;
opts.voice = this->voice_;
const mv::SynthesizedAudio audio = this->tts_->synthesize(
mvr::byte_tensor_to_string(unwrap(ctx.inputs->get(kAnswerText))), opts);
const std::int32_t rate = audio.sample_rate;
ctx.outputs->set(kSpeech, unwrap(Tensor::from_data(
audio.samples.data(), {static_cast(audio.samples.size())},
DataType::Float32)));
ctx.outputs->set(kSpeechRate, unwrap(Tensor::from_data(&rate, {}, DataType::Int32)));
}
```
The post node peak-normalizes the waveform in tensor math, so it runs wherever the waveform lives; the sample rate rides from speak straight to the pipeline output. Request inputs are read-only on the node side, so the two ops that read `speech` allocate their results and only those are written in place:
```cpp
void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override {
const Tensor speech = unwrap(ctx.inputs->get(kSpeech));
Tensor peak = unwrap(ops::amax(unwrap(ops::abs(speech))));
CLIKA_CHECK(ops::clamp_(peak, kPeakFloor));
Tensor normalized = unwrap(ops::div(speech, peak));
CLIKA_CHECK(ops::mul_(normalized, kPeakTarget));
ctx.outputs->set(kAudioOut, std::move(normalized));
}
```
One build flag: a translation unit deriving from `runtime::Model` compiles with `-fno-rtti`, matching how the library builds.
## The chat template is the caller's job
A pipeline consumes prompt bytes as-is; there is no hidden templating between nodes. The template node renders the transcript into the checkpoint's own chat framing, generation turn appended, so the decoder answers the question instead of continuing it:
```cpp
void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override {
ClikaRT::json::Json messages = ClikaRT::json::Json::array();
ClikaRT::json::Json turn = ClikaRT::json::Json::object();
turn["role"] = ClikaRT::json::Json("user");
turn["content"] = ClikaRT::json::Json(
mvr::byte_tensor_to_string(unwrap(ctx.inputs->get(kTranscript))));
messages.push_back(std::move(turn));
ctx.outputs->set(mvr::kPrompt, mvr::string_to_byte_tensor(unwrap(
this->tokenizer_->apply_chat_template(messages, /*add_generation_prompt=*/true))));
}
```
Skip this step and an instruction-tuned model greedy-decodes an untemplated instruction straight to its end-of-turn token; the reply is empty and nothing errors. The template node is where that goes right.
## Wiring
The zoo trio constructs directly. A `PipelineStep` names the step, points at its node (non-owning; you keep the node alive), carries the step's execution knobs, and two rename maps: `input_map` from a pipeline name to a node input, `output_map` from a node output to a pipeline name. Edges otherwise wire by name, so one rename routes the detokenizer's generic `text` onto the conversational `answer_text` edge, where the speak node picks it up, and the decode step's scheduler admits `kMaxActive` sessions at a time:
```cpp
crt::PipelineStep step(const char* name, crt::Model& model) {
crt::PipelineStep st;
st.name = name;
st.model = &model;
return st;
}
```
```cpp
mvr::TokenizerNode tokenize(llm.tokenizer(), /*add_special_tokens=*/false);
mvr::GenerativeDecoderModel decode(llm, gen, kMaxActive, mv::KVCacheMode::Continuous,
/*prefill_chunk_tokens=*/0);
mvr::DetokenizerNode detokenize(llm.tokenizer());
SpeakNode speak(tts, voice_path);
PostNode post;
std::vector steps;
steps.push_back(step("listen", listen));
steps.push_back(step("transcribe", transcribe));
steps.push_back(step("template", prompt));
steps.push_back(step("tokenize", tokenize));
steps.push_back(step("decode", decode));
steps.back().exec_config.scheduler.max_active_sessions = kMaxActive;
steps.push_back(step("detokenize", detokenize));
steps.back().output_map.emplace_back(mvr::kText, kAnswerText);
steps.push_back(step("speak", speak));
steps.push_back(step("post", post));
crt::ModelSchema ext;
ext.inputs.push_back(bytes_spec(kAudioIn));
ext.outputs.push_back(f32_spec(kAudioOut, 1));
ext.outputs.push_back(i32_scalar_spec(kSpeechRate));
ext.outputs.push_back(bytes_spec(kAnswerText));
crt::Pipeline pipeline = crt::Pipeline::create(std::move(steps), std::move(ext));
```
The external schema is the caller's whole contract: only `audio` is settable from outside; `audio_out`, `speech_rate` and `answer_text` are readable; every edge between the nodes stays internal.
## One spoken question
The request surface is the one every pipeline speaks. The response carries the answer's text, its normalized waveform and the waveform's sample rate; `io::save_audio` writes the wav file (`io::encode_audio` yields the same bytes in memory, for a serving payload):
```cpp
crt::Request request;
request.inputs.set(kAudioIn, bytes_to_tensor(question));
const crt::Response response =
unwrap(pipeline.await(unwrap(pipeline.enqueue(std::move(request)))));
const std::string answer =
mvr::byte_tensor_to_string(unwrap(response.outputs.get(kAnswerText)));
const Tensor reply = unwrap(response.outputs.get(kAudioOut));
const std::int32_t rate =
unwrap(response.outputs.get(kSpeechRate)).item();
CLIKA_CHECK(ClikaRT::io::save_audio(reply, rate, reply_path));
```
With a spoken "What is the capital of France?" as `question.wav` and any short spoken clip as `voice.wav` (`espeak-ng -v en-us -w question.wav "What is the capital of France?"` synthesizes one when no recording is at hand), `03_conversational_pipeline question.wav voice.wav reply.wav --device cuda` prints:
```text
answer: The capital of France is Paris.
spoken: 61440 samples at 24000 Hz -> reply.wav
```
The sample count varies from run to run; the TTS decodes with sampling.
## Streaming
`enqueue` takes `RequestCallbacks` for event-driven consumption. The contract from `runtime/serving.h`: `on_chunk` runs per streamed chunk and the last clean-finish chunk carries `final = true`; `on_complete` runs once, at finalize, with the response; both are contained at the boundary, so a throw from one never unwinds into the engine; with callbacks set, no blocking `await` is needed. The request opts in with `streaming = true`:
```cpp
crt::RequestCallbacks cbs;
cbs.on_chunk = [&](const crt::ResponseChunk& c) {
++chunks;
if (c.final) saw_final_marker = true;
events.push_back(c.final ? 1 : 0);
};
cbs.on_complete = [&](ClikaRT::Result) {
++completes;
events.push_back(2);
};
crt::Request r = audio_request(question_wav);
r.streaming = true;
crt::SessionId sid = 0;
sid = pipe.enqueue(std::move(r), std::move(cbs));
```
Callbacks run on the engine's pump thread; do not block in them.
## Concurrent sessions
The decode step's `max_active_sessions` sizes its continuous batch; sessions beyond it queue rather than fail. With `max_active` 4, eight requests enqueued before any is awaited:
```cpp
std::vector ids;
for (std::size_t i = 0; i < args.burst_n; ++i) {
ids.push_back(unwrap(pipe.enqueue(audio_request(question_wav))));
}
std::size_t good = 0;
for (const crt::SessionId sid : ids) {
try {
if (carries_answer(text_of(unwrap(pipe.await(sid)), convo::kAnswerText))) ++good;
} catch (const ClikaRT::Error& e) {
std::printf("[burst ] session %llu FAILED: %s\n",
static_cast(sid), error_detail(e).c_str());
}
}
```
Where a session silently returning a wrong answer under a burst would be costly, verify batched answers or serve one session at a time (`max_active` 1, the example's setting); Clika/ClikaRT#2058 is the tracker for that failure on the CUDA paged cache.
## Rebuilding, and the KV pool
A live `GenerativeDecoderModel` holds its full KV pool on the device; two of them hold two pools, an out-of-memory at the second build on a 16 GB card. Scope the first pipeline and its nodes so they destruct before a rebuild (the example calls `pipeline.shutdown()` before it returns), and let `max_seq` size the pool for the conversation you serve.
## What the composition buys
The composed pipeline builds its executors, its KV pool and its warmed kernels once and reuses them every turn: about 7 seconds per session, against about 22 seconds for the same three models driven in sequence with a fresh pipeline per turn. What a manual driver that kept its pipeline still would not get is the continuous batching.
---
# Estimate depth, segment and detect
The three image verbs: per-pixel depth maps, semantic segmentation masks, and open-vocabulary detection, from the command line and over HTTP.
Source: https://docs.clika.io/modelverse/how-to/depth-segment-detect.md
The vision families answer three per-image questions: how far away everything is (`depth`), what class each pixel belongs to (`segment`), and where the named things are (`detect`). The three verbs share their media handling (any image format the decoder accepts, one artifact per input) and their serving story, so this guide works all three.
## Estimate depth
A depth family (Depth Anything and its relatives) maps each pixel to distance:
```bash
clikart-cli depth-anything/Depth-Anything-V2-Small-hf depth photo.jpg
```
```text
photo.jpg: 64x64 min=1.877290 max=4.294917
```
The payload is one line per image: the map's extent and its value range (relative maps: larger means closer). Artifacts opt in with `--output `, one per image, and each payload line then ends with ` -> `. The default `--output-format raw` writes the Float32 map as `_depth.npy`, the form a downstream consumer wants; `--output-format colormap` renders `_depth.png` instead, and `--colormap` picks its lookup table (`spectral` or `gray`).
## Segment an image
A segmentation family labels every pixel with a class:
```bash
clikart-cli facebook/detr-resnet-50-panoptic segment photo.jpg
```
The payload is the class-mask document: which classes appear, their pixel share, and the rendered overlay's path. `--mask-output mask.npy` also writes the raw class-id mask (UInt8, `[H, W]`) as a sidecar for programmatic consumers, and `--confidence` adds per-label confidence to the listing.
## Detect, with an open vocabulary or a trained class set
The open-vocabulary detectors (OWL-ViT, Grounding DINO) find objects you NAME at run time through `--labels`; the closed-set DETR line (DETR, Deformable DETR, D-FINE, RF-DETR, RT-DETR) detects its checkpoint's trained classes and takes no label prompt. The open-vocabulary form:
```bash
clikart-cli google/owlvit-base-patch32 detect photo.jpg \
--labels "a red bicycle, a street sign" --confidence 0.3
```
```text
label score box (x1, y1, x2, y2)
a red bicycle 0.812 (118, 204, 371, 468)
a street sign 0.644 (402, 61, 455, 152)
```
One row per hit, in score order. The labels encode once and every image scores against them, so a many-image sweep pays the text encoding a single time; `--max-detections` caps the rows per image.
## Over HTTP
Each verb has its serving twin in [the route table](serve-openai-compatible.md): `POST /v1/depth` (multipart image in, depth artifacts out), `POST /v1/segmentations` (image in, class-mask document out), and `POST /v1/detections` (image plus an open-vocabulary `prompt` in, box/score/label rows out):
```bash
clikart-cli google/owlvit-base-patch32 serve --port 8000
```
```bash
curl -s http://127.0.0.1:8000/v1/detections \
-F image=@photo.jpg -F prompt="a red bicycle, a street sign"
```
The operational story (binding, health, admission control) is the serve guide's, unchanged. Model sizing lives in [Model requirements](../model-requirements.md); the embedding-side image comparisons (which image is LIKE which, rather than what is IN one) are [Embed, compare and rerank](embed-and-rerank.md).
---
# Embed, compare and rerank
Dense vectors, similarity matrices and cross-encoder reranking from the command line and over HTTP: the retrieval building blocks.
Source: https://docs.clika.io/modelverse/how-to/embed-and-rerank.md
Retrieval systems stand on three operations: turn inputs into dense vectors, compare vectors, and rerank candidates against a query with a model that reads both together. Modelverse's embedding and reranking families provide all three as commands, and the same models serve them over HTTP.
## Compare texts
Any text-embedding family provides `similar`: the positional texts embed in one batch, and the payload names the inputs by index and prints their cosine matrix:
```bash
clikart-cli Qwen/Qwen3-Embedding-0.6B-GGUF:Q8_0 similar \
"How do I reset my password?" \
"Password reset instructions" \
"Quarterly revenue rose 4 percent"
```
```text
[0] How do I reset my password?
[1] Password reset instructions
[2] Quarterly revenue rose 4 percent
scores (cosine; 1 = identical direction):
[0] [1] [2]
[0] 1.0000 0.7265 0.2591
[1] 0.7265 1.0000 0.3705
[2] 0.2591 0.3705 1.0000
```
`--top-k N` switches the rendering to a per-input ranking (the nearest N others per row), the form you want when the list is long; `--json` emits the same result as a document for scripts.
## Compare images, or texts against images
A dual-tower family (CLIP, SigLIP) embeds texts and images into one space, so `similar` grows two more forms: with `--image` attachments the matrix is texts against images, and with only images it is images against images:
```bash
clikart-cli google/siglip-base-patch16-224 similar \
"a photo of a cat" "a photo of a dog" --image pet1.jpg --image pet2.jpg
```
For the raw vectors, `embed` takes images and prints a summary row per input (vector width, L2 norm, a quantized-vector signature, and each row's cosine against the first), with `--output PREFIX` writing `PREFIX.csv` (one row per image, the full vector as columns) and `PREFIX.json` (run configuration, summary and vectors) for whatever indexes them next:
```bash
clikart-cli google/siglip-base-patch16-224 embed catalog/*.jpg --output catalog_vectors
```
## Rerank documents against a query
Embedding similarity is a coarse first pass; a reranker is a cross-encoder that reads the query and each document together and scores actual relevance. The reranking families provide `rerank`, with every document scored through one packed forward:
```bash
clikart-cli Qwen/Qwen3-Reranker-0.6B rerank \
"how to reset a password" \
"Password reset instructions" \
"Changing your username" \
"Quarterly revenue rose 4 percent"
```
```text
0.9231 [0] Password reset instructions
0.4106 [1] Changing your username
0.0312 [2] Quarterly revenue rose 4 percent
```
`--instruction` prepends the task instruction the checkpoint was trained with, when your use differs from the default retrieval phrasing. The everyday pipeline is both stages in order: `similar --top-k` over the corpus to shortlist, `rerank` over the shortlist to decide.
## Over HTTP
The same models serve the retrieval routes ([the full route table](serve-openai-compatible.md)):
```bash
clikart-cli Qwen/Qwen3-Embedding-0.6B-GGUF:Q8_0 serve --port 8000
```
```bash
curl -s http://127.0.0.1:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input": ["How do I reset my password?", "Password reset instructions"]}'
```
`POST /v1/embeddings` answers in the OpenAI shape, so `client.embeddings.create(...)` works against the base URL unchanged; `POST /v1/similarity` returns the cosine matrix directly, saving the round trip through raw vectors when the comparison is all you need.
Embedding models are small ([Model requirements](../model-requirements.md) carries the figures), and the GGUF quantization selector on the source works here exactly as in [Run a specific GGUF quantization](run-a-gguf-quantization.md).
---
# Add Modelverse to an existing CMake project
find_package against the installed directory, staging the runtime beside your binary, and picking the dist on cross builds.
Source: https://docs.clika.io/modelverse/how-to/existing-cmake-project.md
Your project already builds; this guide adds the model zoo to it without restructuring anything. The installed directory ([Quick install](../getting-started/installation.md)) is the SDK: one `find_package`, one link target, C++17 or newer.
## The integration
```cmake
find_package(Modelverse CONFIG REQUIRED PATHS "$ENV{MODELVERSE_INSTALL_DIR}/cmake")
target_link_libraries(my_app PRIVATE Modelverse::modelverse)
```
That wires the public headers, your platform's `libclika_modelverse` with every model family registered, and the ClikaRT runtime installed in the same root. `PATHS` keeps the location out of the project files; setting `Modelverse_DIR` or extending `CMAKE_PREFIX_PATH` on the configure line works identically, which is the usual shape in CI:
```bash
cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake"
```
## Run from your build directory
The libraries live in the install root, and your freshly built binary does not know where that is. `modelverse_stage_runtime` copies the Modelverse library and the ClikaRT runtime next to the executable, so it runs from the build directory as-is, and the staged layout is exactly what you ship:
```cmake
modelverse_stage_runtime(TARGET my_app) # beside the binary
modelverse_stage_runtime(TARGET my_app DESTINATION lib) # or a subdirectory
```
## When you already link ClikaRT
A project already using the runtime keeps its arrangement: if `ClikaRT::ClikaRT` exists before `find_package(Modelverse ...)` runs, Modelverse links against it instead of resolving the runtime from the install root. The one constraint is a real pairing: each Modelverse build is built against the runtime of its own release, and a mismatched pair refuses at load with a readable error naming both versions rather than misbehaving later.
## Cross builds pick a dist
The package selects the dist (platform/arch flavor) matching your target with the same detection ClikaRT ships, so a cross build that already configures ClikaRT correctly needs nothing extra. When the detection cannot see your intent, name the dist:
```bash
cmake -S . -B build-android \
-DCMAKE_TOOLCHAIN_FILE="$NDK/build/cmake/android.toolchain.cmake" \
-DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake" \
-DMODELVERSE_DIST=android-arm64
```
An install built for a different platform fails at configure time, naming what it carries (`dist.json` records the platform); the fix is extracting the archive matching your target ([Get Modelverse](../getting-started/get-modelverse.md)) and pointing the configure at it.
The program to write once the target links is [part 4 of the tutorial](../getting-started/first-model/04-use-it-from-code.mdx); for the runtime-only integration story (no model zoo), ClikaRT's own guide on [/clikart](/clikart) covers the same ground one layer down.
---
# Register your own model family
Teach the catalog an architecture it does not know, from the snapshot layout through identity matching to runnable commands.
Source: https://docs.clika.io/modelverse/how-to/register-your-own-family.md
The catalog covers the common architectures; your in-house model is not in it. Modelverse's answer is self-registration: a family is a translation unit that pushes a `ModelRegistration` into the shared registry at static-initialization time, and every built-in family registers exactly this way. There is no central list to edit. You compile one registration TU into your program against the public headers, and `list`, `info` and the loaders treat your family like any other. This page is the contract at full depth; for the guided walk from a hand-written checkpoint to a running family of your own, start with the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md); this page's worked example is a text-generation family, the detail-heavy case the tutorial deliberately avoids.
## What a snapshot must contain
A model is a directory of files, and identity resolution reads only its configuration:
- `config.json` with `model_type` and `architectures`; these two values are the identity keys a registration matches on.
- The processor files the model needs (`tokenizer.json` and friends for text; preprocessor configs for image and audio).
- The weights (`*.safetensors`, ONNX, or GGUF); untouched at identity time.
A repo that ships no `config.json` can still be matched: a registration may carry a `probe_snapshot` function that recognizes the family from the file set alone and answers the `model_type` to route as.
## The registration, identity first
One complete TU. With only this much, `list` shows the family (identity-only), `info` resolves your checkpoints to it, and every runnable command refuses precisely, naming what the family provides:
```cpp title="my_model_family.cpp"
#include "clika_modelverse/models/registration.h"
namespace {
using namespace clika_modelverse;
// The identity template: what a resolved checkpoint of this family IS.
// The registry fills architecture, model_type and variant from config.json.
class MyModel final : public Model {
public:
std::string_view family() const noexcept override { return "my-model"; }
};
// The identity keys, matched against config.json.
constexpr std::string_view kModelTypes[] = {"my_model"};
constexpr std::string_view kArchitectures[] = {"MyModelForCausalLM"};
// The Hugging Face pipeline tags the family's checkpoints carry, by the
// constants of models/base/pipeline_tags.h (the registrar refuses a literal
// outside that vocabulary).
constexpr std::string_view kPipelineTags[] = {pipeline_tags::kTextGeneration};
// Static-init self-registration: constructing the registrar is the whole hookup.
const ModelFamilyRegistrar kRegistrar{
[] {
ModelRegistration registration{};
registration.metadata.family = "my-model";
registration.metadata.vendor = "my-org";
registration.metadata.pipeline_tags = kPipelineTags;
registration.metadata.description = "the in-house my_model language models";
registration.metadata.input_combos = kTextOnlyCombos;
registration.metadata.output_combos = kTextOnlyCombos;
registration.hf_model_types = kModelTypes;
registration.hf_architectures = kArchitectures;
registration.factory = make_model();
return registration;
}(),
};
} // namespace
```
Compile that TU into any program that links `Modelverse::modelverse` (the command-dispatch functions in the library serve programs that embed them; the shipped `clikart-cli` binary knows only the built-in families). An incomplete declaration (no family name, no vendor, no pipeline tag or one outside the Hugging Face vocabulary, no modalities, no identity factory) is refused at process start, not papered over downstream.
```text
my-model identity-only
* Input Modalities: text
* Output Modalities: text
* Commands: none (identity-only)
* Vendor: my-org
```
## Making it runnable
Two more fields turn identity into execution:
- **`build_generative`** materializes the runnable model from a resolved snapshot: read the config, adapt the checkpoint's tensor names where they differ from what your forward expects (`ClikaRT::io::TensorsAdapter`, rename-at-lookup, no payload copies), assemble the forward from the `clika_modelverse/modules/` building blocks (attention, decoder stack, projections, rope cache) or your own `nn::Module`, and return a `GenerativeModel` wrapping it with the snapshot's tokenizer and generation defaults. The GGUF twin, `build_generative_gguf`, receives the already-loaded file so identity read and weight bind share one mapping.
- **`cli`** declares the family's commands. The battery builders compose the standard text surface in one line, and with it your family's `prompt`, `serve` and `bench` behave exactly as the tutorial documents them:
```cpp
registration.cli = CliSurface{batteries::cli::llm_cli, batteries::cli::llm_cli_verbs()};
registration.factory = make_model();
registration.build_generative = build_my_model;
```
A family wanting a different surface composes its own `FamilyApp` from the same public builders (`batteries/cli/prompt.h`, `serve.h`, `bench.h`), one verb per builder; the advertised verb list and the App are pinned to agree by the registration contract.
Custom behavior around an EXISTING family needs none of this: [Add your own node to a model pipeline](add-a-pipeline-node.md) composes your logic with any catalog model. The complete field-by-field contract, including the speech, embedding and reranking factories, is `clika_modelverse/models/registration.h` in the installed headers.
## The nearest family to start from
A checkpoint no family claims is refused by name: `clikart-cli info ` prints the `model_type` and `architectures[0]` it read and that no template serves them, and `clikart-cli list` prints every family with the keys it claims, so the family whose definition the checkpoint's `config.json` shares is the one to read first. A composite whose `config.json` carries a nested `text_config` naming a registered family composes that family's text encoder through the text-backbone seam (`build_text_backbone` on the registration), with no forward of its own for the tower. A decoder driven by the embeddings of an encoder you wrote has no such seam in this release: its forward is written from the modules, as the section above does for a text-generation family.
---
# Run a specific GGUF quantization
Pick one weight option of a repo that ships many, by tag, glob or exact file, and see the size table before anything downloads.
Source: https://docs.clika.io/modelverse/how-to/run-a-gguf-quantization.md
A GGUF repo usually publishes one model in several quantizations, from a 2-bit file that fits a phone to an 8-bit one that wants a workstation. Modelverse treats those as weight options of one source: it never guesses between them, the option table with sizes prints before any download, and a selector names the one you want.
## See the options
Ask with `--dry`. A source with several weight options leads with the option table:
```bash
clikart-cli info Qwen/Qwen3-4B-GGUF --dry
```
```text
Qwen/Qwen3-4B-GGUF
family=qwen model_type= architecture=qwen3
variant: dense
components: tokenizer=no image=no audio=no video=no
license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT
model: Qwen/Qwen3-4B-GGUF provider: Qwen source: https://huggingface.co/Qwen/Qwen3-4B-GGUF
license: apache-2.0 https://huggingface.co/Qwen/Qwen3-4B-GGUF/blob/main/LICENSE
'Qwen/Qwen3-4B-GGUF' ships 5 weight options; pick one:
option files size selector
Qwen3-4B-Q4_K_M 1 2.3 GiB :Q4_K_M
Qwen3-4B-Q5_0 1 2.6 GiB :Q5_0
Qwen3-4B-Q5_K_M 1 2.7 GiB :Q5_K_M
Qwen3-4B-Q6_K 1 3.1 GiB :Q6_K
Qwen3-4B-Q8_0 1 4.0 GiB :Q8_0
pick by tag (":Q6_K" / --weights Q6_K; must match one option), by glob (":*Q8*"), by exact file (org/repo/.gguf, a /blob|/resolve URL, or --weights .gguf), or ":latest" for the hub default
total: nothing is fetched until one option is selected
repository: 14.7 GiB in 9 files; the rest is not fetched (other weight formats, files the fetch never takes)
(dry; nothing downloaded)
```
Running the bare source refuses with the same table and exit code 2; the refusal is the answer to "what is there", not an obstacle.
## Select one
A selector rides the source after a colon, and the same grammar works on every command the model provides:
```bash
clikart-cli Qwen/Qwen3-4B-GGUF:Q6_K prompt "The capital of France is"
```
The refusal's own closing line is the whole grammar; the four forms, most specific last:
- **A tag**, `Qwen/Qwen3-4B-GGUF:Q6_K`. An exact option name resolves outright; otherwise a case-insensitive contains-match runs over the option names and each option's first file name. The everyday form.
- **A glob**, `Qwen/Qwen3-4B-GGUF:*Q4*`. Must match exactly one option; an ambiguous selector refuses with every `name (file)` candidate listed.
- **`:latest`**, the repo's default option, an explicit opt-in rather than a silent fallback.
- **An exact file**, `Qwen/Qwen3-4B-GGUF/Qwen3-4B-Q6_K.gguf`, a pasted Hugging Face Hub `blob`/`resolve` URL, or a unique trailing part of the file name. Unambiguous, the right form for scripts that must never drift when the repo adds options.
`fetch` takes the same selection, either in the source or as `--weights Q6_K`, and prints the selected weight file's local path as its payload:
```bash
WEIGHTS=$(clikart-cli fetch Qwen/Qwen3-4B-GGUF --weights Q6_K)
```
## Pick the compute dtype
On a quantized checkpoint, `--dtype` (`float16`, `bfloat16`, `float32`) selects the compute dtype: the precision every quantized weight decodes to and the activations and KV cache run at, behind the generative, embedding and reranker loads. The option pins the whole tree, and the token table of a quantized checkpoint stays quantized in memory, decoding each gathered row at the selected dtype, so a multi-gigabyte f32 table is never materialized. Absent, the checkpoint's own dtype rules. On a dense checkpoint the option is the weight cast applied at load: a llama checkpoint takes `float16`, `bfloat16` or `float32`, an object-detection checkpoint `float16` or `bfloat16`; a dense checkpoint of any other family refuses the option by name.
```bash
clikart-cli Qwen/Qwen3-4B-GGUF:Q6_K prompt "The capital of France is" --dtype float16
```
## What counts as one option
An option is a group, not always a file: a multi-part split (`-00001-of-00009.gguf`) is one option named by its stem (the table's `files` column counts its parts), and a subdirectory holding one quantization's parts is one option named by the subdirectory. Splits are listed with a `(split; not served yet)` note and are selectable by name, but loading one refuses precisely; the loader serves single-file GGUF. Sidecar files (mmproj, imatrix) are excluded from the options and the table says how many it left out.
Which quantization fits which device is a sizing question; [Model requirements](../model-requirements.md) carries the figures. The loading path underneath (block-quantized weights decoded inside the kernels, served at their on-disk footprint) is ClikaRT's, documented in its GGUF guide on [/clikart](/clikart).
---
# Run fully offline
Fetch on a connected machine, move the model directory, and make any network touch an error instead of a surprise.
Source: https://docs.clika.io/modelverse/how-to/run-fully-offline.md
Modelverse needs a network exactly once per model, and not necessarily on the machine that runs it. A local directory is a first-class source everywhere a repo id is, and an offline switch turns any accidental network touch into an error. This is the deployment shape for air-gapped and restricted networks: the release archive installs by extraction alone, no package manager, and the models arrive the same way the archive did.
## Stage on a connected machine
Fetch into a directory you control (not the shared cache), so the result is a self-contained folder:
```bash
clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct --cache-dir staged
```
```text
staged/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775
```
The payload path on stdout is the directory to ship. `fetch --dry` first shows what will be downloaded and how big it is; for a GGUF source, select the quantization at staging time (`--weights Q6_K`) so only the variant you deploy moves. A gated repo needs a token on this machine only (`HF_TOKEN`, or the token a Hub login stored); the token never travels with the files.
## Move it, run it
Copy the directory however files reach the target (rsync over an approved channel, physical media; the `cp` below stands in for that transfer). The cache uses the Hugging Face hub layout, so capture the payload path instead of assuming it; fetching an already-cached source verifies and returns the same path immediately, which makes the capture free. On the offline machine, the copied directory is the source:
```bash
SNAPSHOT=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct --cache-dir staged)
cp -r "$SNAPSHOT" ./Qwen2.5-0.5B-Instruct
clikart-cli ./Qwen2.5-0.5B-Instruct prompt "The capital of France is"
```
A local directory passes through resolution untouched, so `info`, `prompt`, `serve` and the library's snapshot call all accept it identically. No part of model resolution knows or cares that the machine has no route out. The license credential is the one thing that does, and the section below is that one thing.
On an Android phone the same move works in two shapes. An app's cache root is its own cache directory (`ClikaRtAndroid.load` sets `XDG_CACHE_HOME` to it), and the model library's hub cache sits under it in the same layout as on the workstation, so a snapshot directory copied there serves the app with no network. The `clikart-cli` executable pushed to the phone over `adb` takes `--cache-dir ` for a cache you pushed beside it, or the copied snapshot directory itself as the source.
## Make offline a guarantee
Trust but verify: add `--offline` and any operation that would touch the network fails with a readable error instead of hanging on a dead route:
```bash
clikart-cli ./Qwen2.5-0.5B-Instruct serve --offline
```
`--offline` also goes before the command, where it covers the commands that name no model: `clikart-cli --offline info `, `--offline fetch`, `--offline list`. The environment forms hold the same guarantee process-wide, useful under systemd or in a container where flags are out of reach: `CLIKA_MODELVERSE_OFFLINE=1` (Modelverse's own switch) or `HF_HUB_OFFLINE=1` (honored for compatibility with other Hugging Face Hub tooling). With the switch set, a repo-id source still works when the snapshot is already in the cache; resolution is cache-only. A cached GGUF repository resolves its identity offline from its own weight file, so a GGUF source needs no hub metadata on the target either.
## The credential has to be the offline kind
The runtime runs compute under a license, so the target machine needs a credential like any other ([Get Modelverse](../getting-started/get-modelverse.md#license-credential)). The kind matters here and nowhere else. An offline license is verified locally against Clika's signature and needs no route out, which is what makes it the one to ship to an air-gapped target. An online license needs a route to the platform, so a machine with no route cannot use one. Ask your platform for an offline license, and carry it to the target the way you carry the model directory.
```bash
export CLIKA_RT_LICENSE=/srv/clika/clikart.license # the credential text, or a file holding it
clikart-cli ./Qwen2.5-0.5B-Instruct prompt "The capital of France is" --offline
```
`bin/clikart-license-init ` stores it once under the user account instead, which suits a service account that starts without a shell profile. An offline license carries an end date, and [Runtime licenses](/platform/concepts/runtime-licenses) covers what happens on it and how to re-issue.
## What to verify on the target
Three commands, no network, in order: `clikart-cli devices` proves the runtime loads and sees the hardware, `clikart-cli info ./` proves the model resolves, and a one-line `prompt` proves end to end, the credential included. If the documentation should travel too, the docs site you are reading has a self-contained offline bundle (the "Offline Docs" button in the navigation bar).
Sizing the model to the offline hardware is the usual question in these deployments; [Model requirements](../model-requirements.md) and [Run a specific GGUF quantization](run-a-gguf-quantization.md) together answer it.
---
# Script clikart-cli
The contracts that make clikart-cli scriptable: exit codes, the stdout/stderr split, option templates and shell completion.
Source: https://docs.clika.io/modelverse/how-to/script-the-cli.md
The `clikart-cli` executable is built to be driven by scripts, and the contracts below are stable interfaces, not implementation details. A script that branches on them keeps working across releases.
## stdout is the payload, stderr is everything else
Every command writes exactly its payload to stdout: generated text, a transcript, the catalog, a fetched path, CSV. Banners, progress bars, timings, warnings and errors all go to stderr. Capturing a result is therefore safe by construction:
```bash
MODEL_DIR=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct)
ANSWER=$(clikart-cli "$MODEL_DIR" prompt "Two plus two is" --greedy --max-new-tokens 4)
```
`-q` mutes stderr down to warnings and errors, `-v` raises it to debug (including the effective-options report showing which value came from which source), and `--no-color` strips styling (also honored automatically for non-TTY output, `NO_COLOR`, and `TERM=dumb`). The runtime's own `[ClikaRT]` log lines follow the same rule: the level tag is colored only when stderr is a terminal, `NO_COLOR` is unset or empty, and `TERM` is not `dumb`, so a captured stderr holds no escape bytes. None of these change the payload.
Machine-readable forms exist where the payload is tabular: `devices --json` and `devices --csv` print the hardware report as a document instead of the human view, `check --csv` prints the fit verdict table (one row per device, every term of the law a column) as CSV, `list --json` prints the registry as one document (per family: the vendor, the Hugging Face pipeline tags it serves, the modalities, the served and refused model types and GGUF architectures, the advertised verbs, whether it is runnable), and `info --json` / `info --csv` print the supported-models document, or its table as CSV, for one source (a one-entry document), several sources or `@FILE` (one source per line): per model the family, the resolved commit, the parameter count and where it was read from, the license, the gate, and one row per weight variant with the files a fetch lands and their sizes, which a catalog generated from the registry reads. The grammar is `clikart-cli info [--weights SEL] [--json | --csv] [--schema] [--dry]`: one source with no form flag prints the human description, `--schema` prints the document's JSON Schema and needs no source, and a source the hub refuses is a warning on stderr while the document carries the rest and the exit code is 1.
## The exit-code contract
Four values, fixed:
| Exit | Meaning | Example |
| --- | --- | --- |
| 0 | success | the payload is on stdout |
| 1 | runtime failure | download interrupted, weights failed to load |
| 2 | usage error | unknown flag, missing argument, unselected weight options |
| 3 | the model cannot do this here | a verb the family does not provide, an identity-only family, an unsupported variant |
The distinction between 2 and 3 is the useful one: 2 means fix the invocation, 3 means fix the model choice. A retry loop should retry 1 and never 2 or 3.
```bash
SRC=openai/whisper-large-v3-turbo # the model this script expects
FILE=meeting.wav # the input it transcribes
if ! clikart-cli "$SRC" transcribe "$FILE" > out.txt; then
case $? in
2) echo "bad invocation, check the flags" >&2; exit 2 ;;
3) echo "$SRC has no transcribe; pick a speech-to-text model" >&2; exit 3 ;;
*) echo "runtime failure, retrying" >&2 ;;
esac
fi
```
## Options as a file
`generate-template` writes a JSON file with every root option at its default (null where there is no built-in default); `--template` loads one, and explicit flags always win over it. This is how a deployment pins its configuration in version control instead of a growing alias:
```bash
clikart-cli generate-template defaults.json
clikart-cli info "$SRC" --template defaults.json
clikart-cli fetch "$SRC" --template defaults.json --cache-dir /data/models
```
A family's own verbs are their own contract; the template governs the root verbs. Secrets stay out of both: the Hugging Face Hub token comes from `HF_TOKEN` in the environment or from the token a Hub login stored, never a flag and never a template key, so neither shell history nor a committed template can leak it.
## Completion
`clikart-cli` generates its own shell completion:
```bash
source <(clikart-cli completion bash) # zsh and fish work the same way
```
Completion covers the root verbs and flags; model commands are discovered per source at run time, which completion cannot see, and that is the one place tab-completion ends and ` --help` takes over.
---
# Serve a model over the OpenAI and Anthropic APIs
Host any runnable model behind the built-in HTTP server, stream completions, and point existing OpenAI and Anthropic clients at it.
Source: https://docs.clika.io/modelverse/how-to/serve-openai-compatible.md
Any model whose family provides `serve` becomes an HTTP endpoint in one command, speaking the OpenAI request and response shapes and, for chat, the Anthropic Messages API beside them. Existing OpenAI and Anthropic SDK code connects by changing its base URL; nothing else about the client changes.
## Start the server
```bash
clikart-cli Qwen/Qwen2.5-0.5B-Instruct serve --host 0.0.0.0 --port 8000
```
```text
[2026-10-04 07:59:45.310] [modelverse] [info] serving on http://0.0.0.0:8000 (model=Qwen2.5-0.5B-Instruct); open it in a browser for the chat page
[2026-10-04 07:59:45.310] [modelverse] [info] endpoints: POST /v1/chat/completions /v1/messages /v1/messages/count_tokens /v1/audio/transcriptions; GET /v1/models /health /props / (web UI) /dashboard
[2026-10-04 07:59:45.310] [modelverse] [info] authentication: none; every route is open
```
The load flags from `prompt` apply unchanged (`--device`, `--max-seq`, a GGUF weight selector on the source); the operational flags are the server's own:
- `--host` and `--port`: the bind address, `127.0.0.1:8000` by default. Loopback by default is deliberate; expose it with `--host 0.0.0.0` when you mean to.
- `--no-web-ui`: disable the built-in chat page on `/`.
- `--cors`: answer cross-origin requests, for a browser front end on another origin.
- `--api-key KEY`: require this key on every route but `/health`, `/v1/health`, `GET /` and `GET /dashboard`, sent as `Authorization: Bearer ` (the OpenAI clients' form) or `x-api-key: ` (the Anthropic clients' form). A request without it, or with another key, answers 401 in the client's own error envelope with a `WWW-Authenticate: Bearer` challenge. Without the flag every route is open, so set it whenever the server listens beyond loopback; a key on the command line is readable in the process list. The two pages open without a key (a browser navigation carries no header); the chat page sends the key typed into its API key field on every call, and the dashboard reads the same stored key. **The key is a best-effort gate, not a security boundary:** it keeps a stray client on your network from using the model, and that is all it does. Access control for a server people you do not know can reach (identities, per-user keys and their rotation, rate limits, TLS, audit logs) belongs to an infrastructure layer in front of the serving process, a reverse proxy or an API gateway; the serving process itself stays a single-tenant engine behind it.
- `--max-active` and `--max-queue`: admission control. Requests past the queue limit are shed and surface as HTTP 503 after retries with backoff, so an overloaded server degrades loudly instead of stalling silently.
- `--max-upload-mib`: the largest request body accepted, 256 MiB by default. A larger upload answers 413 naming the size and the limit. Audio sizing: a 60-minute 16 kHz mono WAV is 111 MiB.
- `--read-timeout`: seconds a connection may stay silent while its request is read, 30 by default. A stalled read closes the connection rather than holding a slot.
- `--tool-parser`: how tool calls are read out of the model's text. By default the server picks, from the call formats the model's family declares, the one the checkpoint's chat template writes. The flag names another parser (`serve --help` lists them), or `none` turns extraction off. With no parser in use, a tool call stays in the reply as plain text. `/props` reports the parser in use as `tool_parser`, and `tool_parser_source` says where it came from: `family` or `option`.
## The routes
What the server answers depends on the model's family; every serving model carries the common rows.
| Route | Serves | Model families |
| --- | --- | --- |
| `POST /v1/chat/completions` | chat, streaming (SSE) and non-streaming | text and multimodal generators |
| `POST /v1/messages` | chat in the Anthropic Messages API shape, streaming (SSE) and non-streaming | text and multimodal generators |
| `POST /v1/messages/count_tokens` | the input token count the same `/v1/messages` request would ingest (text only) | text and multimodal generators |
| `POST /v1/audio/transcriptions` | multipart audio file to transcript; `response_format` `json` (the default), `text`, `verbose_json` (the task, the language, the duration and the segments with their seconds and tokens), `srt` or `vtt` (one cue per segment); `language` the hint, `temperature` the decode's sampling (0, the default, is greedy; above 0 the decode samples under a fixed seed); `timestamp_granularities[]` takes `segment` (word timestamps are not produced); `prompt` is refused (no decode here conditions on a text prompt) | speech-to-text (Whisper) |
| `POST /v1/audio/speech`; `POST`, `GET`, `DELETE /v1/voices` | text in, waveform out, in a voice registered first through the voice registry ([Speak text](speak-text.md)); `response_format` `wav` (the default) or `pcm` always, and `mp3`, `opus`, `aac` or `flac` when the host provides the `ffmpeg` tool (`CLIKA_RT_FFMPEG` names its path, else it is taken from the PATH; the tool is spawned, never linked); `GET /props` lists the served set as `speech_formats`, and the start-up log says which | text-to-speech (Chatterbox, Qwen3-TTS) |
| `POST /v1/embeddings` | texts to dense vectors; `encoding_format` `float` (the default) or `base64` (each row as the base64 of its float32 little-endian bytes); `dimensions` shortens every row to its leading entries renormalized to unit length, refused past the model's own count | embedding models |
| `POST /v1/similarity` | texts to a cosine matrix | embedding models |
| `POST /v1/translations` | texts plus a language pair | translation models |
| `POST /v1/images/generations`, `POST /v1/videos/generations` | prompt to a generated image or clip; `response_format` `b64_json` (the default) or `url`: the file is kept in memory under an unguessable id for an hour (at most 64 files, 256 MiB in all; the oldest leave first) and served at `GET /v1/files/` with no key, the id being the credential as a signed URL's is | diffusion models |
| `POST /v1/depth` | multipart image to depth artifacts | depth estimators |
| `POST /v1/segmentations` | multipart image to a class-mask document | segmentation models |
| `POST /v1/detections` | multipart image to box/score/label rows (an open-vocabulary detector also takes a `prompt`) | detectors |
| `GET /v1/models`, `GET /props` | what is loaded and how it is configured; the model listing answers the OpenAI shape (`object`, `created`, `owned_by`, `license`), or the Anthropic one (`type`, `display_name`, `created_at`, `has_more`) to a request carrying `anthropic-version` or `x-api-key` | all |
| `GET /health`, `GET /v1/health` | `{"status":"ok"}`, the machine probe | all |
| `GET /` | the built-in web chat page | all, unless `--no-web-ui` |
| `GET /dashboard` | the live request dashboard | all |
| any other path | 404 in the client's own envelope: the OpenAI one (`code` `unknown_url`, the path named) or, under `/v1/messages`, the Messages one (`not_found_error`) | all |
## Call it with the OpenAI API
Non-streaming, from anything that can POST JSON:
```bash
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "One-line haiku about rain."}]}'
```
Streaming adds `"stream": true` and reads server-sent events, one `data:` line per delta, `data: [DONE]` last, the shape OpenAI clients already parse. Which means the SDKs work as-is:
```python
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
for chunk in client.chat.completions.create(
model="qwen",
messages=[{"role": "user", "content": "One-line haiku about rain."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="", flush=True)
```
A multimodal model takes OpenAI-shaped image content parts in the same route; a Whisper model is called with a multipart file upload instead ([Transcribe audio](transcribe-audio.md) shows both of its forms).
## Call it with the Anthropic API
`POST /v1/messages` takes the Messages API request body on the same server. `max_tokens` is required, as it is in the Anthropic API. The `anthropic-version` header is optional; `2023-06-01` is the version served, and any other value is refused. A request may name any model, and the reply names the one the server loaded.
```bash
curl -s http://127.0.0.1:8000/v1/messages \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-d '{"model": "qwen", "max_tokens": 256,
"messages": [{"role": "user", "content": "One-line haiku about rain."}]}'
```
The Anthropic SDK takes the server's address as its base URL, without `/v1`. Any API key works, because the server checks none:
```python
from anthropic import Anthropic
client = Anthropic(base_url="http://127.0.0.1:8000", api_key="unused")
with client.messages.stream(
model="qwen",
max_tokens=256,
messages=[{"role": "user", "content": "One-line haiku about rain."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
```
A streamed reply follows the Messages API event order. It opens with `message_start` and a `ping`. Each content block then streams as `content_block_start`, its `content_block_delta` events and `content_block_stop`. The reply closes with `message_delta`, which carries the stop reason and the output usage, and then `message_stop`. There is no `[DONE]` line. A `ping` also goes out whenever the stream has been silent for 5 seconds.
`max_tokens` is a cap. When the context window runs out first, the reply stops there with `stop_reason` `model_context_window_exceeded`. A prompt that fills the window on its own is refused with a message that starts `prompt is too long:`. The other stop reasons are `end_turn`, `max_tokens`, `stop_sequence` (with the matched `stop_sequence`) and `tool_use`. In `usage`, `input_tokens` counts the prompt tokens the prefix cache did not serve, `cache_read_input_tokens` counts the ones it did, and `output_tokens` counts the reply.
Errors come back in the Anthropic error envelope, with `request_id` in the body, and every answer carries a `request-id` header. A refusal is a 400 `invalid_request_error` whose message starts with the path of the field it names. The other statuses:
- 404 `not_found_error`: a path under `/v1/messages` that no route serves, such as the batches API.
- 413 `request_too_large`: a body past `--max-upload-mib`.
- 529 `overloaded_error`: the server is at capacity, for example when admission control sheds the request. The answer carries `retry-after` and `x-should-retry: true`.
- 503 `api_error`: the device faulted. Restart the server; the answer carries `x-should-retry: false`.
- 500 `api_error`: the generation failed.
The `anthropic-version` header is read when present and taken as the served version when absent (a client that omits it is answered, not refused).
### What the Messages API route serves, accepts and refuses
| Request | What the server does |
| --- | --- |
| text; base64 images (JPEG, PNG, GIF, WebP) to a multimodal model | served |
| `system` as a string or as text blocks | served |
| `stop_sequences` (up to 4, each up to 128 bytes), `temperature`, `top_p`, `top_k` | served |
| custom tools, `tool_choice` `auto` or `none`, `tool_use` and `tool_result` blocks | served; calls come back as `tool_use` blocks when a parser reads the model's call format (the family's own by default, or the one `--tool-parser` names) |
| `thinking` `enabled` (a `budget_tokens` of at least 1024, below `max_tokens`), `adaptive` or `disabled`; `output_config.effort` at a level the model declares | served by a model that reasons |
| a final assistant turn (a response prefill: text only, with no trailing whitespace) | served; the reply carries only the new text |
| `cache_control` | accepted, no effect; the server's prefix cache needs no marker |
| `metadata`, `service_tier`, `output_config.task_budget` | accepted, no effect |
| `context_management` | accepted, not applied; the history renders in full |
| a tool's `strict` and `defer_loading`; the tool-search server tools | accepted; arguments are not schema-constrained, and every tool is offered up front |
| an effort level the model does not declare | accepted; the model's default applies |
| fields the server does not know, in a request that sends `anthropic-beta` | ignored |
| `tool_choice` `any` or `tool` (a forced tool call) | refused |
| other server tools, `container`, `output_config.format` (structured output) | refused |
| `document`, `search_result` and other block types the server does not serve | refused |
| an image `file` source or an image URL (the server fetches no remote payload) | refused |
| `max_tokens` 0 | refused |
| fields the server does not know, in a request without `anthropic-beta` | refused: `: Extra inputs are not permitted` |
| a response prefill with `thinking` `enabled` or `adaptive`, or to a model that always reasons; `thinking` `disabled` on a model that always reasons and declares no effort levels | refused |
| an image in a `count_tokens` request | refused; it counts text only |
The server logs `context_management`, `strict`, `defer_loading`, the tool-search tools and an effort level the model does not declare once each, at Info, naming the field.
## Use it from Claude Code
Claude Code speaks the Messages API, so it runs against this server once its environment variables point it there. Start the server with a context window that holds Claude Code's prompt, and put your own device in `--device`:
```bash
clikart-cli Qwen/Qwen3-30B-A3B-Instruct-2507 serve --device vulkan:0 --max-seq 65536
```
Then start Claude Code with these variables:
```bash
unset ANTHROPIC_API_KEY
export ANTHROPIC_BASE_URL=http://127.0.0.1:8000
export ANTHROPIC_AUTH_TOKEN=unused
export ANTHROPIC_MODEL=Qwen3-30B-A3B-Instruct-2507
export ANTHROPIC_DEFAULT_HAIKU_MODEL=Qwen3-30B-A3B-Instruct-2507
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
export API_TIMEOUT_MS=600000
claude
```
- `ANTHROPIC_BASE_URL` is the server's address, without `/v1`.
- `ANTHROPIC_AUTH_TOKEN` is the server's `--api-key` value, or any value when the server runs without one. Claude Code sends it as an `Authorization: Bearer` header, and it ranks above `ANTHROPIC_API_KEY` in Claude Code's [credential order](https://code.claude.com/docs/en/authentication). Unsetting `ANTHROPIC_API_KEY` keeps an Anthropic key out of a session that talks to another server.
- `ANTHROPIC_MODEL` names the main model, and `ANTHROPIC_DEFAULT_HAIKU_MODEL` names the one Claude Code uses for background tasks. Behind a custom base URL, Claude Code passes any model name through unchecked, and this server answers every request with the model it loaded, so both name the served model, the id `/props` reports as `model_id`.
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is the window Claude Code assumes for the model. For a model name it does not recognize, Claude Code can assume a window that differs from the one this server serves, so set it to the window `/props` reports as `n_ctx`. Claude Code then compacts the conversation before it reaches the server's limit ([model configuration](https://code.claude.com/docs/en/model-config)).
- `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` turns off Claude Code's other traffic, such as auto-updates, telemetry and error reporting, for a network that reaches only this server.
- `API_TIMEOUT_MS` is how long Claude Code waits for a reply, in milliseconds. 600000 (10 minutes) is its default; raise it when the first prefill of Claude Code's prompt takes longer than that on your device.
To check the connection, run `/status` in Claude Code. Its `Anthropic base URL` line shows this server's address, and its `Auth token` line names `ANTHROPIC_AUTH_TOKEN` ([verify the connection](https://code.claude.com/docs/en/llm-gateway-connect#verify-the-connection)).
Claude Code sends its whole system prompt and every tool definition with each request, about 16,000 tokens before the conversation starts, so the served window must hold them. The server caps the window at 4096 tokens by default, and with that cap Claude Code refuses the request itself with "Prompt is too long". `--max-seq` sets the window, and `/props` reports it as `n_ctx`. The tools reach the model through its chat template, and its calls come back to Claude Code as `tool_use` blocks when a parser reads the model's call format: the family's own by default, or the one `--tool-parser` names (above). `/props` shows which parser is in use.
With `ANTHROPIC_BASE_URL` pointing at a host other than `api.anthropic.com`, Claude Code turns off features such as Remote Control and server-managed settings ([feature availability](https://code.claude.com/docs/en/feature-availability)). Anthropic's gateway documentation states that Anthropic "doesn't support routing Claude Code to non-Claude models through any gateway" ([LLM gateways](https://code.claude.com/docs/en/llm-gateway)).
## Operate it
`GET /health` is the readiness probe for a load balancer. `GET /props` reports the loaded model and its effective options, the remote twin of `-v`'s effective-options report; its `device` field names the device the model actually landed on, which is what you check after `--device auto` or a fallback. Its `routes` list gives every API route the server answers, each with its method and its path, plus `api` (`openai` or `anthropic`) on a route that speaks one of those two APIs. The dashboard on `/dashboard` shows the same routes and the live requests; `/api/requests` is its JSON feed. Everything the server logs goes to stderr like the rest of the `clikart-cli` executable, so systemd or a container runtime captures it without configuration.
For the same server inside your own process (your engines, your routes, the built-in UI swapped for your page), the `01_serve` program in [Additional examples](../examples.md) is the smallest complete consumer of the library's `ServeApi`.
---
# Speak text
Text-to-speech from the command line and over HTTP, with voice references and the synthesis knobs.
Source: https://docs.clika.io/modelverse/how-to/speak-text.md
A text-to-speech family turns text into a waveform. The catalog's runnable TTS families (Chatterbox, Qwen3-TTS) provide `speak`: the model synthesizes, the wav lands on disk, and the payload on stdout is the file's path, ready to pipe into the next command. The runnable Chatterbox checkpoint is `ResembleAI/chatterbox-flash` (the original `ResembleAI/chatterbox` snapshot ships a flow estimator the runtime does not serve, and the tool's refusal names the served one). CSM is cataloged for identity resolution today (`info` recognizes it); its runnable path is not shipped yet.
## From the command line
```bash
clikart-cli ResembleAI/chatterbox-flash speak \
"The install works. Pick a model from the catalog." \
--voice reference.wav --output hello.wav
```
```text
hello.wav
```
Play it with anything; the file is 16-bit PCM at 24 kHz. The flags:
- `--output PATH`: the wav path (`speech.wav` by default). The path is the stdout payload, so `aplay "$(clikart-cli speak "..." )"` works as a one-liner.
- `--voice PATH`: REQUIRED for Chatterbox, a reference voice whose timbre the model matches (the family ships no default voice; `speak` without it refuses and says so). Where a family ships named voices, `--voices-dir` points at the set.
- The synthesis knobs are the family's own (` speak --help` lists them); Chatterbox exposes `--exaggeration` (expressiveness), `--steps` (decode steps, quality against speed) and `--temperature`.
- The load knobs apply unchanged: `--device`, `--cache-dir`, `--offline`.
## Over HTTP
The same model serves `POST /v1/audio/speech` (text in, waveform out) behind a voice registry, rows in [the route table](serve-openai-compatible.md): `POST /v1/voices` registers a reference clip under a name, `GET /v1/voices` lists the registered voices, `DELETE /v1/voices/` frees a slot, and a speech request names one of them. Chatterbox ships no default voice, so a speech request before any registration is refused (`voice 'default' is not registered; POST /v1/voices first`).
```bash
clikart-cli ResembleAI/chatterbox-flash serve --port 8000
```
Register the reference clip once, as a multipart upload with `name` (the voice id) and `file` (the audio); the reply names the clip's decoded rate and duration:
```bash
curl -s http://127.0.0.1:8000/v1/voices -F name=reference -F file=@reference.wav
```
```text
{"name":"reference","sample_rate":16000,"duration_seconds":11.0}
```
Then ask for speech in that voice; the reply body is the wav itself (`audio/wav`, 16-bit PCM at 24 kHz):
```bash
curl -s http://127.0.0.1:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "The install works.", "voice": "reference"}' \
-o hello.wav
```
OpenAI SDK clients use their `audio.speech.create(...)` call against the base URL with `voice` set to a registered name; the server's built-in web page for a TTS model is a speech console rather than a chat window.
The round trip with [Transcribe audio](transcribe-audio.md) is the natural smoke test: speak a sentence, transcribe the wav, and compare the text. Sizing lives in [Model requirements](../model-requirements.md).
---
# Transcribe audio
Speech-to-text with a Whisper model, from the command line and as an OpenAI-compatible HTTP endpoint.
Source: https://docs.clika.io/modelverse/how-to/transcribe-audio.md
A speech-to-text family provides `transcribe`, and the workflow is the same one the tutorial used for text: resolve the model, run its command, get the payload on stdout. This guide transcribes a file locally, then serves transcription over HTTP.
## From the command line
```bash
clikart-cli openai/whisper-large-v3-turbo transcribe meeting.wav
```
```text
Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week.
```
The transcript is the entire stdout payload, ready to redirect into a file. Decoding progress and timing ride stderr. The audio loader decodes WAV, FLAC, MP3 and OGG (Vorbis) and resamples to the model's expected rate, so a file in any of those containers transcribes as is. A container outside that set (OGG Opus, M4A/AAC, WMA) is refused by name, `audio decode: : OGG (Opus) is not supported; this build decodes WAV, FLAC, MP3, OGG (Vorbis)`, with exit code 1; a damaged file, or a file that is not audio at all, is refused the same way, with the file named and the first bytes it actually found quoted back. The web UI is not bound by this set: it decodes a recording or an upload in the browser and sends 16 kHz WAV, so anything the browser plays transcribes.
The load knobs from the tutorial apply unchanged: `--device cuda` places the encoder and decoder, `--cache-dir` controls where the snapshot lives, and a local directory works as the source for offline machines. Whisper checkpoints come in sizes from `whisper-tiny` (fits anywhere, fastest, roughest) to `whisper-large-v3` (the accuracy reference); [Model requirements](../model-requirements.md) has the figures. A clip longer than one 30 s window is windowed and transcribed whole.
## Over HTTP
The same model serves the OpenAI transcription route:
```bash
clikart-cli openai/whisper-large-v3-turbo serve --port 8000
```
Call it as a multipart upload, the OpenAI shape:
```bash
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F file=@meeting.wav \
-F model=whisper
```
```text
{"text":" Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week."}
```
OpenAI SDK clients use their `audio.transcriptions.create(...)` call against the server's base URL, unchanged. The server-side flags and operations story (binding, health, admission control) is the common one in [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md).
For concurrent callers, `--concurrent-sessions N` admits N requests at once: Whisper keeps one encoder and mints decoder-only replicas over the same weights, encodes every window on the shared encoder with no lock held, and locks one replica for the decode alone, so N uploads run without a whole-request lock. The library form is `LoadOptions::concurrent_sessions`.
For transcription inside your own process, the library's `SttModel` and the transcription engine mount on the same `ServeApi` the `01_serve` example demonstrates ([Additional examples](../examples.md)).
---
# Translate text
Neural machine translation from the command line and over HTTP, with the model's own language codes.
Source: https://docs.clika.io/modelverse/how-to/translate-text.md
A translation family provides `translate`, and the workflow is the transcription guide's sibling: resolve the model, run its command, get the payload on stdout. This guide translates sentences locally, then serves translation over HTTP. The model is `facebook/m2m100_418M`, a sequence-to-sequence translation model covering a hundred languages under two-letter codes; its model card is where its license and its language inventory are read, and the runtime prints the license at fetch, at load and in `info`. A chat model that translates (Hy-MT2 among them) answers through its `prompt` command instead, since its family provides no `translate`.
## From the command line
```bash
clikart-cli facebook/m2m100_418M translate --to ko \
"The installation is complete." "Pick a model from the catalog."
```
```text
설치가 완료되었습니다.
카탈로그에서 모델을 선택합니다.
```
Each source text prints as one line, in input order; the payload is only the translations, so redirecting into a file gives one translation per line. Decoding progress and timing ride stderr. The flags:
- `--to LANG` (required): the target language code from the model's own inventory; M2M-100 uses two-letter codes (`en`, `ko`, `fr`, `de`, `ja`, `zh`, ...), and the server's `/props` lists them under `languages`.
- `--from LANG`: the source language code; unset means the family's default source, which read the English above.
- `--max-new-tokens N`: the decode budget in new tokens (`0`, the default, keeps the family default, clamped to the model's window).
- The load knobs apply unchanged: `--device`, `--cache-dir`, `--offline`.
## When the budget binds first
A translation that stops at its decode budget rather than at the model's end token still exits 0 and still prints its payload; what marks it is one warning on stderr naming the limit that bound it:
```bash
clikart-cli facebook/m2m100_418M translate --to ko --max-new-tokens 8 \
"The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon."
```
```text
회의는 10시에 시작되며
```
```text
[2026-10-09 15:05:25.113] [m2m-100] [warning] translation stopped at the 8-token budget (--max-new-tokens); output may be incomplete
```
The output above is cut on purpose: an 8-token budget cannot hold the sentence, and the warning is the demonstration. A script that must catch truncation without parsing stderr uses the HTTP route below, where each row carries `finish_reason`.
## Over HTTP
The same model serves `POST /v1/translations` ([the route table](serve-openai-compatible.md)):
```bash
clikart-cli facebook/m2m100_418M serve --port 8000
```
```bash
curl -s http://127.0.0.1:8000/v1/translations \
-H "Content-Type: application/json" \
-d '{"texts": ["Hello.", "The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon."],
"source_lang": "en", "target_lang": "ko", "max_new_tokens": 8}'
```
```text
{"model":"m2m-100","source_lang":"en","target_lang":"ko","translations":[{"text":"안녕하세요","finish_reason":"stop"},{"text":"회의는 10시에 시작되며","finish_reason":"length"}],"usage":{"prompt_tokens":34,"completion_tokens":10,"total_tokens":44}}
```
Batch through the `texts` array; the response carries one translation per input, in order. Each row's `finish_reason` is the machine-readable form of the budget story above: `stop` means the model produced its end token, `length` means the decode budget cut the translation short (the small `max_new_tokens` above forces one of each). Omit `max_new_tokens` to keep the family default. `source_lang` is optional, as `--from` is.
The server-side flags and operations story (binding, health, admission control) is the common one in [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md). Sizing and quantization work as everywhere else ([Model requirements](../model-requirements.md)); for speech instead of text as the input, [Transcribe audio](transcribe-audio.md) chains into this guide.
---
# Model requirements
Memory and device figures per model variant, for sizing a deployment before anything downloads.
Source: https://docs.clika.io/modelverse/model-requirements.md
What a model needs to run is mostly decided before you fetch it: the variant and quantization fix the weight size, and the context length fixes the working memory on top. This page carries the figures for the models the documentation uses and the classes of hardware they fit; `clikart-cli info --dry` gives the on-disk number for any source, and the `bench` command measures the rest on your machine. `--memory-stages ` on any verb writes the runtime's memory-pool counters per device at each stage of the model's life, which is the measured form of the Run column.
How to read the tables: **Disk** is the snapshot size (`fetch --dry` total). **Run** is the resident memory serving one session at the default context; longer contexts and concurrent sessions add KV cache on top, roughly linear in tokens. **Fits** names the smallest sensible device class: phone (4 GB), laptop or small board (8 GB), workstation GPU (8 GB VRAM and up). A model's license is not restated here: the model card on the Hugging Face Hub is its source of truth, and the runtime shows it to you where it matters, as `clikart-cli info ` prints the family's license page and the checkpoint's own license link, and the same lines print at fetch and at load.
## Text generation
| Model | Variant | Disk | Run | Fits |
| --- | --- | --- | --- | --- |
| Qwen2.5 0.5B Instruct | fp16 (safetensors) | 0.9 GB | 1.2 GB | phone |
| Qwen3 0.6B | fp16 (safetensors) | 1.4 GB | 1.9 GB | phone |
| Qwen3 4B | GGUF Q6_K | 3.3 GB | 4.3 GB | laptop |
| Qwen3 4B | GGUF Q8_0 | 4.3 GB | 5.4 GB | workstation GPU, 8 GB |
The quantization ladder trades memory for fidelity in predictable steps: Q8_0 is near-fp16, Q6_K is the everyday default, Q4_K_M is the last stop before quality degrades noticeably in chat use. `--kv-quant` shrinks the per-token cache the same way when long contexts dominate the budget.
## Multimodal
| Model | Variant | Disk | Run | Fits |
| --- | --- | --- | --- | --- |
| Gemma 4 E2B IT (vision) | fp16 (safetensors) | 9.5 GB | 12.4 GB | workstation GPU, 16 GB |
| Gemma 4 E2B IT (vision) | GGUF Q4_0 (QAT) | 4.0 GB | 5.3 GB | laptop |
A multimodal run holds the vision tower alongside the language model, and these figures include it: the GGUF row counts the projector file (`mmproj`) that carries the tower beside the weights.
## Speech-to-text
| Model | Variant | Disk | Run | Fits |
| --- | --- | --- | --- | --- |
| Whisper base | fp16 (safetensors) | 0.3 GB | 0.6 GB | phone |
| Whisper large-v3-turbo | fp16 (safetensors) | 1.6 GB | 2.2 GB | laptop |
| Whisper large-v3 | fp16 (safetensors) | 3.1 GB | 4.0 GB | laptop |
Transcription memory is stable with audio length (the model works in 30-second windows); throughput, not memory, is what a faster device buys.
## Embeddings
| Model | Variant | Disk | Run | Fits |
| --- | --- | --- | --- | --- |
| Qwen3 Embedding 0.6B | GGUF Q8_0 | 0.7 GB | 1.1 GB | phone |
Embedding models run one forward per input with no generation loop, so batch size is the only memory knob that matters in practice.
## After a model closes
Closing a model returns its memory to the runtime's pool, not to the operating system: the pool keeps the bytes cached so the next allocation is cheap, and releases them under memory pressure or when asked. A program that closes one model and loads another therefore holds both until it asks in between: `ClikaRT::device::release_cached_memory(device)` from C++, `clika_runtime.empty_cache(device)` from Python, `Device.releaseCachedMemory()` from Kotlin. The Kotlin Modelverse binding makes the call at a model's close; a C++ or Python program makes it itself, after the close and before the next load, and `memory_stats(device)` shows the pool's cached bytes before and after.
## Reading a figure you do not see here
For any other model: `clikart-cli check ` answers the question before anything is downloaded. It reads the model's documents and its file listing and prints one row per device: the weight bytes, the cache bytes per token of context, the cache dtype, the context length, the activation estimate, the total need against the device's memory, whether it fits, and the largest sequence length that fits (`--max-seq N` judges another length; `--assume-device phone=8G` judges a device that is not this machine; `--csv` for a script). `info --dry` prints the exact disk cost, and `devices` prints what this machine offers. When the estimate is close to the device's limit, measure with `bench` before committing the deployment; it reports peak resident memory along with throughput.
On a board whose GPU and CPU share one memory (a Jetson, a phone), the device total `check` prices against is the whole machine's memory, which the operating system and every other process share, and the law takes no headroom off; so apply your own margin there, and size against what is free when the model starts rather than the total. `--memory-stages FILE` on any verb writes what a load holds on each device at each stage of the model's life, the measured form of the figure. A second model loaded beside the first takes its memory from the same pool.
```bash
clikart-cli check Qwen/Qwen2.5-1.5B-Instruct --max-seq 4096 --assume-device phone=8G
```
For a catalog of models, `clikart-cli info --json` prints
one document: each model (the base repository) with its variants, the base
weights and every quantization among the sources, read from the hub's
documents, listings and GGUF headers with nothing downloaded. Per variant it
carries the quant tag as the file names it, the precision, the bits per
weight, the files with their sizes and digests, what a runtime must decode
to serve it (`requires`), the fit terms `check` prices from, and whether
this build serves it. `info --schema` prints the JSON Schema the document
validates against; a served model's `/props` carries the same block as
`variant`, so a client matches the served weights to a catalog row.
```bash
clikart-cli info unsloth/Qwen3-1.7B-GGUF --json
clikart-cli info @sources.txt --json # one source per line
clikart-cli info --schema
```
---
# Supported models
The models the release serves, read off its supported-models document: one row per model with its weight variants and the backends the release's run record covers.
Source: https://docs.clika.io/modelverse/supported-models.md
The release ships its supported-models document beside the archives: `supported_models.json`, and `supported_models.csv`, the same document flat. It lists 109 checkpoints over 56 model families with 312 weight variants. This page lists the 80 of them whose weights carry a permissive license, the document's own class; the model card is where a checkpoint's license is read, and the runtime prints it at fetch, at load and in `info`. The whole document is a download of the platform with the release ([Get ClikaRT](/clikart/getting-started/get-clikart)).
A model is named by its Hugging Face repository id, which `clikart-cli fetch ` and `AutoModel.from_pretrained(id)` take as they are. **Parameters** and **Context** are read off the checkpoint. **Weights** lists the weight variants the release serves for the model, base first, with their format. **Proven on** names the backends the release's own run record covers for the model; an empty cell is a model the release lists without a recorded run.
## Text generation and chat
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [arianraje/qwen3-4b-gdn-hybrid-stage1-align-dtfix](https://huggingface.co/arianraje/qwen3-4b-gdn-hybrid-stage1-align-dtfix) | `qwen3-next` | about 4.5 B | 40,960 | `BF16` (safetensors) | CPU, CUDA, Metal |
| [DavidAU/Qwen3-MOE-4x0.6B-2.4B-Writing-Thunder](https://huggingface.co/DavidAU/Qwen3-MOE-4x0.6B-2.4B-Writing-Thunder) | `qwen` | about 1.5 B | 40,960 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [microsoft/phi-2](https://huggingface.co/microsoft/phi-2) | `phi` | about 2.8 B | 2,048 | `F16` (safetensors) | CPU, CUDA, Metal |
| [microsoft/Phi-3-mini-4k-instruct](https://huggingface.co/microsoft/Phi-3-mini-4k-instruct) | `phi` | about 3.8 B | 4,096 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct) | `phi` | about 3.8 B | 131,072 | `IQ1_M` (gguf), `IQ1_S` (gguf), `IQ2_XS` (gguf), `IQ3_XS` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q3_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan |
| [moonshotai/Kimi-Linear-48B-A3B-Instruct](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct) | `kimi-linear` | about 49.1 B | | `BF16` (safetensors) | CPU |
| [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) | `gpt-oss` | about 20.9 B | 131,072 | `U8` (safetensors), `MXFP4` (gguf) | CPU, CUDA, Vulkan |
| [OuteAI/Lite-Oute-1-300M](https://huggingface.co/OuteAI/Lite-Oute-1-300M) | `mistral` | about 300 M | 4,096 | `F32` (safetensors) | CPU, CUDA, Metal |
| [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) | `qwen` | about 494 M | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) | `qwen` | about 752 M | 40,960 | `BF16` (gguf), `F8_E4M3` (safetensors), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q2_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ1_M` (gguf), `UD-IQ1_S` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_XXS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan |
| [Qwen/Qwen3-0.6B-Base](https://huggingface.co/Qwen/Qwen3-0.6B-Base) | `qwen` | about 596 M | 32,768 | `Qwen3-Embedding-0.6B-f16` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan |
| [SupraLabs/Supra2-100M-Base](https://huggingface.co/SupraLabs/Supra2-100M-Base) | `qwen` | about 101 M | 2,048 | `F32` (safetensors) | CPU, CUDA, Vulkan |
| [zai-org/GLM-4-9B-0414](https://huggingface.co/zai-org/GLM-4-9B-0414) | `glm` | about 9.4 B | 32,768 | `BF16` (safetensors), `IQ2_M` (gguf), `IQ3_M` (gguf), `IQ3_XS` (gguf), `IQ3_XXS` (gguf), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q2_K_L` (gguf), `Q3_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q3_K_XL` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_L` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_L` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q6_K_L` (gguf), `Q8_0` (gguf), `THUDM_GLM-4-9B-0414-bf16` (gguf) | CPU, CUDA, Vulkan, Metal |
| [zai-org/GLM-4.5-Air](https://huggingface.co/zai-org/GLM-4.5-Air) | `glm` | about 110.5 B | 131,072 | `BF16` (safetensors) | CPU |
## Vision-language chat
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) | `deepseek_ocr` | about 3.3 B | 32,768 | `BF16` (safetensors) | CPU, CUDA, Metal |
| [deepseek-ai/DeepSeek-OCR](https://huggingface.co/deepseek-ai/DeepSeek-OCR) | `deepseek_ocr` | about 3.3 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal |
| [deepseek-ai/DeepSeek-OCR-2](https://huggingface.co/deepseek-ai/DeepSeek-OCR-2) | `deepseek_ocr` | about 3.4 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal |
| [deepseek-community/DeepSeek-OCR-2](https://huggingface.co/deepseek-community/DeepSeek-OCR-2) | `deepseek_ocr` | about 3.4 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal |
| [florence-community/Florence-2-base](https://huggingface.co/florence-community/Florence-2-base) | `florence2` | about 232 M | 1,024 | `F16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [google/gemma-4-26B-A4B](https://huggingface.co/google/gemma-4-26B-A4B) | `gemma` | about 26.5 B | 262,144 | `BF16` (safetensors) | CPU, CUDA |
| [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) | `gemma` | about 25.8 B | 262,144 | `BF16` (safetensors) | |
| [google/gemma-4-26B-A4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized) | `gemma` | about 26.5 B | 262,144 | `BF16` (safetensors) | CUDA |
| [google/gemma-4-31B](https://huggingface.co/google/gemma-4-31B) | `gemma` | about 32.7 B | 262,144 | `BF16` (safetensors) | CPU, CUDA |
| [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) | `gemma` | about 31.3 B | 262,144 | `BF16` (safetensors) | CPU, CUDA, Vulkan |
| [google/gemma-4-31B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-unquantized) | `gemma` | about 32.7 B | 262,144 | `BF16` (safetensors), `I32` (safetensors), `gemma-4-31B_q4_0-it` (gguf) | CPU, CUDA, Vulkan |
| [meta-models/Muse-Glimmer-30B](https://huggingface.co/meta-models/Muse-Glimmer-30B) | `muse-glimmer` | about 29.8 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan |
| [microsoft/Florence-2-base](https://huggingface.co/microsoft/Florence-2-base) | `florence2` | about 232 M | 1,024 | `F16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [Qwen/Qwen2-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen2-VL-2B-Instruct) | `qwen-vl` | about 2.2 B | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [Qwen/Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) | `qwen-vl` | about 2.1 B | 262,144 | `BF16` (safetensors), `F16` (gguf), `Q4_K_M` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal |
| [Qwen/Qwen3-VL-30B-A3B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct) | `qwen-vl` | about 31.1 B | 262,144 | `F8_E4M3` (safetensors) | CPU, CUDA, Vulkan |
| [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) | `qwen3-5` | about 873 M | 262,144 | `BF16` (safetensors), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_XXS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan, Metal |
| [Qwen/Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) | `qwen3-5-moe` | about 36 B | 262,144 | `BF16` (gguf), `I32` (safetensors), `MXFP4_MOE` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_S` (gguf), `UD-IQ3_XXS` (gguf), `UD-IQ4_NL` (gguf), `UD-IQ4_XS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_L` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_S` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan |
| [thinkingmachines/Inkling-Small](https://huggingface.co/thinkingmachines/Inkling-Small) | `inkling` | about 266 B | | `U8` (safetensors) | CPU |
## Multimodal (any to any)
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it) | `gemma` | about 12 B | 262,144 | `BF16` (safetensors) | CPU, CUDA, Vulkan |
| [google/gemma-4-12B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized) | `gemma` | about 12 B | 262,144 | `gemma-4-12b-it-qat-q4_0` (gguf) | CPU, CUDA, Vulkan |
| [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) | `gemma` | about 5.1 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [google/gemma-4-E2B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-unquantized) | `gemma` | about 5.1 B | 131,072 | `gemma-4-E2B_q4_0-it` (gguf) | CPU, CUDA, Vulkan |
| [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) | `gemma` | about 8 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [google/gemma-4-E4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized) | `gemma` | about 7.9 B | 131,072 | `gemma-4-E4B_q4_0-it` (gguf) | CPU, CUDA, Vulkan |
## Speech to text
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [nvidia/canary-1b-v2](https://huggingface.co/nvidia/canary-1b-v2) | `canary` | about 979 M | | `F32` (safetensors) | |
| [nvidia/parakeet-ctc-0.6b](https://huggingface.co/nvidia/parakeet-ctc-0.6b) | `parakeet` | about 609 M | | `F32` (safetensors) | |
| [nvidia/parakeet-rnnt-0.6b](https://huggingface.co/nvidia/parakeet-rnnt-0.6b) | `parakeet` | about 617 M | | `F32` (safetensors) | |
| [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | `parakeet` | about 627 M | | `F32` (safetensors) | |
| [openai/whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo) | `whisper` | about 809 M | | `whisper-large-v3-turbo-q4_0` (gguf), `whisper-large-v3-turbo-q4_1` (gguf), `whisper-large-v3-turbo-q8_0` (gguf) | CPU, CUDA |
| [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny) | `whisper` | about 38 M | 448 | `F32` (safetensors), `F16` (gguf), `Q4_K_M` (gguf), `Q5_K_M` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal |
## Text to speech
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [sesame/csm-1b](https://huggingface.co/sesame/csm-1b) | `csm` | about 1.6 B | 2,048 | `F32` (safetensors) | CPU, CUDA, Metal |
## Translation
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [tencent/Hy-MT2-1.8B](https://huggingface.co/tencent/Hy-MT2-1.8B) | `hunyuan` | about 2 B | 262,144 | `BF16` (safetensors), `Q4_K_M` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal |
| [tencent/Hy-MT2-30B-A3B](https://huggingface.co/tencent/Hy-MT2-30B-A3B) | `hy-v3` | about 30.1 B | 262,144 | `F8_E4M3` (safetensors) | CPU, CUDA, Vulkan |
## Text embedding
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [TaylorAI/bge-micro-v2](https://huggingface.co/TaylorAI/bge-micro-v2) | `bert` | about 17 M | 512 | `F16` (safetensors) | Metal |
## Feature extraction
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) | `qwen-embedding` | about 596 M | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Image embedding
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [facebook/dinov2-small](https://huggingface.co/facebook/dinov2-small) | `dinov2` | about 22 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Reranking
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [Qwen/Qwen3-Reranker-0.6B](https://huggingface.co/Qwen/Qwen3-Reranker-0.6B) | `qwen-reranker` | about 596 M | 40,960 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Object detection
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [facebook/detr-resnet-50](https://huggingface.co/facebook/detr-resnet-50) | `detr` | about 42 M | 1,024 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [microsoft/table-transformer-detection](https://huggingface.co/microsoft/table-transformer-detection) | `detr` | about 29 M | 1,024 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [PekingU/rtdetr_r18vd](https://huggingface.co/PekingU/rtdetr_r18vd) | `rt-detr` | about 20 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [PekingU/rtdetr_v2_r18vd](https://huggingface.co/PekingU/rtdetr_v2_r18vd) | `rt-detr` | about 20 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [Roboflow/rf-detr-nano](https://huggingface.co/Roboflow/rf-detr-nano) | `rf-detr` | about 30 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [ustc-community/dfine-nano-coco](https://huggingface.co/ustc-community/dfine-nano-coco) | `d-fine` | about 4 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Open-vocabulary object detection
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [google/owlv2-base-patch16](https://huggingface.co/google/owlv2-base-patch16) | `owlvit` | about 155 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [google/owlvit-base-patch32](https://huggingface.co/google/owlvit-base-patch32) | `owlvit` | about 153 M | 16 | `F32` (safetensors) | CPU, CUDA, Metal |
| [IDEA-Research/grounding-dino-tiny](https://huggingface.co/IDEA-Research/grounding-dino-tiny) | `grounding-dino` | about 172 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [openmmlab-community/mm_grounding_dino_tiny_o365v1_goldg](https://huggingface.co/openmmlab-community/mm_grounding_dino_tiny_o365v1_goldg) | `grounding-dino` | about 173 M | 512 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Zero-shot image classification
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [google/siglip-base-patch16-224](https://huggingface.co/google/siglip-base-patch16-224) | `siglip` | about 203 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [google/siglip2-base-patch16-naflex](https://huggingface.co/google/siglip2-base-patch16-naflex) | `siglip` | about 375 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [wkcn/TinyCLIP-ViT-8M-16-Text-3M-YFCC15M](https://huggingface.co/wkcn/TinyCLIP-ViT-8M-16-Text-3M-YFCC15M) | `clip` | about 23 M | 77 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Image segmentation
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [facebook/detr-resnet-50-panoptic](https://huggingface.co/facebook/detr-resnet-50-panoptic) | `detr-panoptic` | | 1,024 | `torch` (torch) | CPU, CUDA, Vulkan, Metal |
| [Roboflow/rf-detr-seg-nano](https://huggingface.co/Roboflow/rf-detr-seg-nano) | `rf-detr-seg` | about 34 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Depth estimation
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [depth-anything/DA3-Small](https://huggingface.co/depth-anything/DA3-Small) | `da3` | about 34 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [depth-anything/Depth-Anything-V2-Small-hf](https://huggingface.co/depth-anything/Depth-Anything-V2-Small-hf) | `depth-anything` | about 25 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
| [Intel/dpt-large](https://huggingface.co/Intel/dpt-large) | `dpt` | about 342 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Image to text
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [zai-org/GLM-OCR](https://huggingface.co/zai-org/GLM-OCR) | `glm` | about 1.3 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Text to image
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [Qwen/Qwen-Image](https://huggingface.co/Qwen/Qwen-Image) | `qwenimage` | about 20.4 B | | `BF16` (safetensors) | CPU, CUDA, Vulkan |
| [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | `z-image` | about 6.2 B | | `F32` (safetensors) | CPU, CUDA, Vulkan |
## Audio-language chat
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [Qwen/Qwen2-Audio-7B-Instruct](https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct) | `qwen2-audio` | about 8.4 B | | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Image to 3D
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [depth-anything/DA3-BASE](https://huggingface.co/depth-anything/DA3-BASE) | `da3` | about 135 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal |
## Text to video
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [Wan-AI/Wan2.1-T2V-1.3B-Diffusers](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) | `wan` | about 1.4 B | | `F32` (safetensors) | CPU, CUDA, Vulkan |
| [Wan-AI/Wan2.2-TI2V-5B-Diffusers](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers) | `wan` | about 5 B | | `F32` (safetensors) | CPU, CUDA, Vulkan |
## Other
| Model | Family | Parameters | Context | Weights | Proven on |
| --- | --- | --- | --- | --- | --- |
| [facebook/m2m100_418M](https://huggingface.co/facebook/m2m100_418M) | `m2m-100` | | 1,024 | `torch` (torch) | CPU, CUDA, Vulkan, Metal |
| [fromziro/ZeroS-v0.1-150M](https://huggingface.co/fromziro/ZeroS-v0.1-150M) | `qwen3-5` | about 152 M | 2,048 | `F32` (safetensors) | CPU, CUDA, Metal |
| [rpatel622/mamba2-130m-hf-Q8_0-GGUF](https://huggingface.co/rpatel622/mamba2-130m-hf-Q8_0-GGUF) | `mamba2` | about 168 M | | `mamba2-130m-hf-q8_0` (gguf), `mamba2-130m-q8_0` (gguf) | CPU, CUDA, Vulkan |
---
# Platform
The CLIKA Platform runs benchmarks on real devices, keeps a fleet of them under management, serves models from them, and issues the licenses ClikaRT needs.
Source: https://docs.clika.io/platform.md
The CLIKA Platform is where a team measures models on the hardware they will actually ship on. You register your devices with the platform, point it at a model, and it runs the benchmark on each device and brings the numbers back: accuracy per test, tokens per second, time to first token, peak memory. The same platform keeps those devices under management afterwards, runs models on them as long-lived endpoints, and issues the licenses a ClikaRT runtime needs.
One deployment carries three surfaces over the same data: a web application, a command-line tool (`clika-cli`), and an MCP server that lets Claude drive the platform for you in plain language.
## Why the Platform
For the ML engineer. A benchmark answers a question about a model on a device, and both halves matter. Pick a model from the Hugging Face Hub or from your project's model list, pick the quality tests to run (MMLU, GSM8K, HumanEval+, IFEval, and the rest of the catalog), pick the devices, and the platform runs every model on every device and lines the results up side by side. You do not choose a runner, a container image, or a script: the platform maps each model and device pair to the job definition that is confirmed to work on that platform and architecture.
For the team that owns the devices. A device runs one agent, which dials out to the platform over a single gRPC connection and needs no inbound port. From then on the device reports health every 15 seconds, accepts benchmark jobs and long-running services, exposes a terminal and a file browser, and can be reached through a remote desktop where the platform's providers can start one. Jetsons, x86 workstations, Windows machines, Raspberry Pi class boards and Android phones join the same fleet.
For the operator running the deployment. The platform runs on your infrastructure, in your cloud or on your own hardware. An organization holds users, roles and projects; a project holds the devices, models, benchmarks and licenses of one piece of work. API keys carry a subset of their owner's permissions, so an integration gets exactly the access it needs. Every authorization decision is recorded.
For whoever ships ClikaRT. A runtime credential issued by the platform is what a ClikaRT runtime presents to prove it may run. The platform issues those credentials per project, in an online form for runtimes that can reach the network and an offline bundle for the ones that cannot, and can revoke or rotate either one.
## What you get
- **Devices.** A fleet with live health, hardware inventory, tags and per-device job history. Registration is one install command; the agent does the rest.
- **Models.** A model is a reference to a Hugging Face repository plus the task it performs. Registering one is metadata only, so no weights move until a run needs them. The Models page also carries the Modelverse catalog, the models CLIKA has confirmed to run well on ClikaRT, each one click from registered.
- **Benchmarks.** A benchmark run fans out over every model and device you picked, and returns one comparable result set: quality scores, throughput and latency, memory, and the per-sample inputs and outputs behind the score.
- **Jobs and services.** A job is one benchmark execution on one device, with artifact push, setup, script, teardown and result collection. A service is a process the platform keeps running on a device and restarts when it dies.
- **Model serving.** Run a model on your own devices as an OpenAI-compatible endpoint the platform proxies for you.
- **Artifacts.** Versioned, checksummed files (datasets, weights, engine bundles, scripts) that the platform pushes to devices and deduplicates on arrival.
- **Organizations, projects and access.** Members and roles at the organization, membership and resources per project, scoped API keys for integrations.
- **Licensing.** The runtime credentials your ClikaRT projects issue, update and revoke, plus activation of the deployment's own platform license.
- **CLI and Claude.** Every documented operation is a CLI command and an MCP tool, so the same workflow runs from a terminal, from CI, or from a conversation with Claude.
## Where to go next
- [Getting started](getting-started/index.md): the five-minute mental model, then a tutorial that takes you from signing in to reading your first benchmark result.
- [Concepts](concepts/index.md): one page each for organization, project, device, model, artifact, benchmark, job, service and deployment.
- [How-to guides](how-to/index.md): job and service definition YAML, members and API keys, license activation, reading results, downloading the ClikaRT SDK.
- [Using the CLI](cli/index.md): the platform from a terminal, with every `clika-cli` command, its arguments and its flags.
- [Using with AI (MCP)](mcp/index.md): connecting Claude Desktop, Claude Code, Codex and other MCP clients to the platform.
---
# Using the CLI
clika-cli, the CLIKA Platform command-line client: install it, log in, and drive every platform operation from a terminal or a CI job.
Source: https://docs.clika.io/platform/cli.md
`clika-cli` is the CLIKA Platform from a terminal. It talks to your deployment's API and does the same things the web application does, so a benchmark you would start by clicking can be started from a shell script or a CI job, a device you would inspect on its page can be queried from your prompt, and an AI assistant can be given the platform as a set of tools. One binary covers the whole platform: devices, benchmarks, jobs, artifacts, models, projects and licenses are all subcommands, each generated from the deployment's own API description, so the CLI and the platform never disagree about what exists.
New to it? [Get started with the CLI](get-started.md) takes you from nothing to your first commands in about ten minutes. The rest of this page is the reference: how to install it, how it authenticates, how the commands are shaped, and where each group of commands is documented.
`clika-dm` is the sibling binary for a CLIKA Device Management deployment. It is built from the same source and behaves identically; only the set of resources differs. Everything on these pages describes `clika-cli`.
## What you can do with it
- **Drive the platform from scripts.** List devices, launch a benchmark, wait for it, and fail the build if a leg did not complete.
- **Keep definitions in git.** [`export`](apply-export.md) writes a job or service definition as YAML, [`apply`](apply-export.md) puts it back. The file you commit is the file the platform runs.
- **Reach into a device.** Run a command on it, push a file to it, or open an SSH session, without leaving your terminal.
- **Give an AI assistant access to the platform.** The [`mcp`](mcp.md) subcommand turns the CLI into a Model Context Protocol server that Claude Code, Cursor and similar tools can call.
## Install
Every deployment serves the binaries built from its own commit, so the CLI you install always matches the orchestrator it talks to. The channel lives at `/files/installation/cli/`, where `` is the address of your deployment (for example `https://platform.clika.io`).
**The download channel requires authentication.** That is deliberate, because the binaries are part of the deployment rather than a public download. You authenticate with one of three things:
| Credential | Where it comes from | Best for |
| --- | --- | --- |
| Download token | The web app, under **Settings** then **Developer access**. Valid for 15 minutes. | A one-off install on your own machine. |
| API key | The web app, same tab. Starts with `clika_` and does not expire until you revoke it. | CI runners and unattended installs. |
| Session token | An existing browser session. | Rarely needed by hand. |
### macOS and Linux
The **Developer access** settings tab prints a ready-to-paste one-liner with the token already filled in. It looks like this:
```
curl -fsSL "https://platform.clika.io/files/installation/cli/install.sh?token=" | CLIKA_BASE_URL="https://platform.clika.io" CLIKA_DOWNLOAD_TOKEN="" sh
```
The installer detects your operating system and CPU architecture, downloads the matching binary, checks it against the channel's `checksums.txt`, and installs it into `/usr/local/bin` when that is writable and `~/.local/bin` otherwise. Set `CLIKA_INSTALL_DIR` to choose a different directory. With an API key instead of a download token, set `CLIKA_API_KEY` in place of `CLIKA_DOWNLOAD_TOKEN`.
### Windows
On Windows the same tab prints a PowerShell one-liner, which runs the installer's PowerShell twin, `install.ps1`:
```
$env:CLIKA_BASE_URL='https://platform.clika.io'; $env:CLIKA_DOWNLOAD_TOKEN=''; irm 'https://platform.clika.io/files/installation/cli/install.ps1?token=' | iex
```
It runs in Windows PowerShell 5.1 and PowerShell 7 and needs no administrator rights. It checks the download against `checksums.txt`, installs `clika-cli.exe` into `%LOCALAPPDATA%\Programs\clika\bin` (`CLIKA_INSTALL_DIR` chooses another directory), adds that directory to your user `PATH`, and puts it on the current session's `PATH` so the CLI runs straight away. On Windows on Arm it installs the native arm64 build. Running it again upgrades the CLI in place.
Both installers also take `CLIKA_MCP`, a comma-separated list of AI clients (`claude-code`, `claude-desktop`, `codex`) to register the CLI's MCP server with once it is installed. [Get started with the CLI](get-started.md#register-your-ai-clients-in-the-same-step) shows the commands.
### Downloading one binary by hand
Any platform can skip the installer and fetch a single asset. The asset name is `clika-cli--`, with `.exe` appended on Windows:
```
curl -fsSL -H "Authorization: Bearer clika_..." -o clika-cli "https://platform.clika.io/files/installation/cli/clika-cli-linux-amd64" && chmod +x clika-cli
```
The channel publishes these assets:
| Operating system | Architecture | Asset |
| --- | --- | --- |
| Linux | x86_64 | `clika-cli-linux-amd64` |
| Linux | ARM64 | `clika-cli-linux-arm64` |
| macOS | Apple silicon | `clika-cli-darwin-arm64` |
| Windows | x86_64 | `clika-cli-windows-amd64.exe` |
| Windows | ARM64 | `clika-cli-windows-arm64.exe` |
Alongside them the channel serves `latest.json` (the published version), `checksums.txt` and `signatures.txt`, which is what [`self-update`](self-update.md) reads.
## First login
The CLI signs in with an API key: create one in the web app under **Settings** then **Developer access**, then run `login` once. It prompts for the key without echoing it, and saves the deployment address and the key to a profile file that every later command reads:
```
clika-cli --base-url https://platform.clika.io login
```
Email and password sign-in is for the web application only; the CLI never asks for a password. For CI, skip `login` and set `CLIKA_BASE_URL` and `CLIKA_API_KEY` in the environment instead.
Then check that it worked:
```
clika-cli devices list
```
Full detail on credential types, profiles and multiple deployments is on the [authentication and profiles](authentication.md) page.
### Where credentials are stored
`login` writes `$XDG_CONFIG_HOME/clika-cli/.json`, which on a default Linux or macOS setup is `~/.config/clika-cli/default.json`. The file is created with mode `0600` (readable only by you) and holds one deployment's base URL and credential. Use `--profile ` to keep several deployments side by side.
## Global flags
These flags work on every command. Put them anywhere on the command line.
| Flag | Type | Default | What it means |
| --- | --- | --- | --- |
| `--base-url` | string | from the profile | Address of the orchestrator to talk to, for example `https://platform.clika.io`. Reads `CLIKA_BASE_URL` when the flag is absent. |
| `--api-key` | string | from the profile | A `clika_` API key, the CLI's only credential. Pass `--api-key=`, or a bare `--api-key` to be prompted without echo; when standard input is piped the value is read from it. Reads `CLIKA_API_KEY`. |
| `--profile` | string | `default` | Which saved profile to load credentials from, or save them to. |
| `-o`, `--output` | string | `table` | Output format: `table`, `json` or `yaml`. |
| `--color` | string | `auto` | When to emit ANSI color: `auto` (only when writing to a terminal), `always`, or `never`. |
| `--no-color` | bool | `false` | Turn color off. A non-empty `NO_COLOR` environment variable does the same. |
| `--insecure-tls` | bool | `false` | Skip TLS certificate verification. Only for on-premise development stacks that use a throwaway certificate authority. |
| `-y`, `--yes` | bool | `false` | Answer confirmation prompts with yes. Required for destructive deletes when standard input is not a terminal. |
| `-h`, `--help` | bool | `false` | Print help for the command and exit. |
| `-v`, `--version` | bool | `false` | Print the CLI version and exit. |
Credentials resolve in this order, highest priority first:
1. The `--api-key` and `--base-url` flags.
2. The `CLIKA_API_KEY` and `CLIKA_BASE_URL` environment variables.
3. The saved profile.
## Output formats
Every command that prints a response honors `-o`/`--output`.
| Value | What you get |
| --- | --- |
| `table` (default) | A column table for any list response. Common resources (devices, jobs, artifacts, models, benchmark groups, job definitions, service definitions) have hand-picked columns; every other list derives its columns from the fields the rows actually carry, headed by the JSON field names. A single resource, or a shape the CLI does not recognise, prints as pretty JSON. Timestamps render as ages such as `3m ago`, and byte counts render with units. |
| `json` | Pretty-printed JSON of the whole response, pagination envelope included. |
| `yaml` | The same document as YAML. |
Paginated endpoints wrap their array in `{items, total, page, page_size}`. Table mode unwraps that so you see only the rows; `json` and `yaml` keep it, so a script can still read the page state.
An empty list in table mode writes a note such as `No devices found` to standard error and prints nothing to standard output, which keeps pipes clean. In `json` and `yaml` the empty document is printed as it came back.
Color marks resource status only, and only when standard output is a terminal, so piped output never carries escape codes:
```
clika-cli devices list
clika-cli devices list -o json | jq '.items[].name'
clika-cli jobs list --color=always | less -R
```
Generated commands also accept `--raw`, which prints the response body byte for byte with no formatting at all.
## Names instead of UUIDs
Most resources have both a UUID and a human name. Anywhere a command takes an id, you can type the name instead:
```
clika-cli devices get jetson-01
clika-cli job-definitions delete llm-latency
```
The CLI lists the collection once and matches on the `name` field. An argument that already looks like a UUID is used directly, with no extra request. Name lookup is enabled for devices, jobs, models, artifacts, job definitions, service definitions and benchmark groups, and names inside an [`apply`](apply-export.md) document resolve the same way.
If two resources share a name, the CLI stops and lists the candidate ids rather than picking one, so a script never silently targets the wrong resource.
## Deletes ask first
Deleting a device also deletes its job history, so `devices delete` and `devices batch-delete` prompt for confirmation:
```
$ clika-cli devices delete jetson-01
delete device "jetson-01" (1f0a...)? (this also deletes its job history) [y/N]:
```
Pass `-y`/`--yes` to skip the prompt. When standard input is not a terminal (a CI job, a pipeline), the CLI refuses to proceed without `--yes`, so a script has to state destructive intent explicitly. Other deletes accept `--yes` as a no-op, so you can pass it uniformly.
Every delete reports what it removed, echoing both the name you typed and the id it resolved to:
```
$ clika-cli job-definitions delete smoke-test
deleted job definition "smoke-test" (7a31...)
```
Under `-o json`, `-o yaml` or `--raw` that line moves to standard error and standard output carries the response body, so scripts keep parsing.
## Exit codes
| Code | Meaning |
| --- | --- |
| `0` | The command succeeded. |
| `1` | The command failed. The reason is printed to standard error, prefixed with `error:`. A benchmark watch whose legs did not all complete also exits `1`. |
| the remote code | `devices exec` and `devices command` propagate the exit code of the command that ran on the device, the way `ssh` does. |
| `130` | You interrupted a credential prompt with Ctrl-C. |
## The command tree
Running `clika-cli --help` groups the tree into five sections.
| Section | Commands | Documented in |
| --- | --- | --- |
| Getting started | `login`, `logout`, `self-update`, `version` | [Authentication and profiles](authentication.md), [self-update](self-update.md) |
| Resources | 60 groups generated from the deployment's API description, one per resource family | The pages listed below |
| Declarative | `apply`, `export` | [apply and export](apply-export.md) |
| Advanced | `call`, `mcp`, `tools` | [MCP server and generic dispatch](mcp.md) |
| Additional | `completion`, `help` | This page |
The reference pages are grouped and ordered to follow the [concepts](../concepts/index.md), so a command sits where the thing it acts on is explained:
| Group | Pages |
| --- | --- |
| Organization and project | [authentication and profiles](authentication.md), [organizations, projects and access](projects-and-orgs.md) |
| Device | [devices](devices.md), [transfers](transfers.md), [events, metrics and alerts](monitoring.md) |
| Benchmark | [benchmarks](benchmarks.md) |
| Model deployment | [model deployment](model-deployment.md) |
| Runtime licenses | [licensing](licensing.md) |
| Artifact | [artifacts and models](artifacts.md) |
| Job | [jobs](jobs.md), [job definitions](job-definitions.md) |
| Service | [services and service definitions](services.md) |
| Tools and automation | [apply and export](apply-export.md), [MCP and generic dispatch](mcp.md), [cloud instances](cloud.md), [keeping the CLI current](self-update.md), [other groups](other-resources.md), [cheat sheet](cheatsheet.md) |
The CLI has no platform administration surface: an API key never carries the administrator step-up, so platform administration is done in the admin dashboard of the web app.
Because the resource commands are generated, the mapping from the API to the command line is mechanical and worth knowing:
| In the API | On the command line |
| --- | --- |
| A path parameter such as `{id}` | A positional argument, `` or `` |
| A query parameter | A flag, `--` |
| A request body | `--body ''` or `--body-file