# CLIKA documentation, full text > ClikaRT is the inference runtime, ClikaRT CLI the command line interface that infers, serves and benchmarks popular models on it (built on Modelverse, the CLIKA model library), and the Platform the service that runs benchmarks on devices and issues the runtime licenses. Every page below is a plain address on this site; the site also offers a self-contained offline copy (the "Offline Docs" button). Every page of `https://docs.clika.io/llms.txt`, in its order; each begins with its title, its one-line description and its address. --- # ClikaRT ClikaRT is CLIKA's inference runtime, a C++ library with Python and Kotlin bindings, that runs models on CPUs, GPUs and other hardware accelerators through one public API. Source: https://docs.clika.io/clikart.md ClikaRT is the CLIKA inference runtime: a C++ library, with Python and Kotlin bindings, that loads models and runs them on CPUs, GPUs and other hardware accelerators through one public API. Link one library and include one header, or import one package, and the same code runs on every backend the runtime ships. ## Why ClikaRT For the AI developer. The API is PyTorch-shaped C++, and the same shape reaches Python in full and Kotlin at the inference level through the [bindings](bindings.md). Tensors move with `.to(device)`, and the built-in operators chain the way you expect (`.reshape()`, `.relu()`, `.matmul()`). What PyTorch does not give you is where this runs. The same code covers CPU, CUDA, Vulkan and Metal (mobile GPU via Vulkan, Apple silicon via Metal), and the backend is picked at run time, not at build time. Library built-ins provide the convenience of PyTorch, but you can implement whatever you want, as close to the hardware as you like; the library is built with freedom as a first-class citizen. Serving, tokenizers and pipelines live in the same library, so a model checkpoint becomes a served endpoint. For the backend engineer. No AI background is required to serve a model well. Pull a packaged model from [Modelverse](/modelverse), load it, and hand requests to the serving runtime; batching, device placement and precision are the runtime's job, not model configuration details you need to understand. The library behaves like normal C++, which inference stacks usually do not. Integration is one `find_package`, one link target, C++17, and no Python in the ship path. For the embedded developer. Small targets are a first-class platform. A complete Android deployment is the Kotlin artifact's Android AAR, roughly 140 MB to download (the core library, its Vulkan backend and the model library inside), and the same code runs on a phone's GPU, on Apple silicon and on Jetson, from the same bundle with no device-specific build. The device carries no Python and no package manager, only the libraries your binary links. For the defense, healthcare and finance developer. Everything runs on your hardware. User data never leaves the device, and the self-contained bundle installs and builds where there is no network at all. For the business. Nothing external to install, resolve, or keep in sync. The bundle is self-contained, and every third-party component it embeds is listed in its `licenses/` directory. There is no dependency tree to audit, no third-party library update that breaks your team's product, and no copyleft surprise for legal. Hardware stays a choice, not a commitment. TensorRT runs only on NVIDIA GPUs; ClikaRT runs the same models on NVIDIA, AMD, Intel and Qualcomm GPUs and on plain CPUs, years-old hardware included. Run your state-of-the-art models on older or legacy machines to reduce inference and cloud costs. ## What you get - **Tensors and operators.** `ClikaRT::Tensor` plus the `ClikaRT::ops` library: element-wise math, reductions, convolutions, attention, indexing. This is the operator set a model needs. - **Backends.** CPU always; CUDA, Vulkan and Metal where the platform has them. Accelerator backends load on demand; unavailable ones are absent. - **Quantized weights.** Weight-only quantization schemes (`nn::QLinearWoQ`, GGUF block formats, FP8, NF4 and more) decode inside the kernels. - **Model loading.** safetensors, GGUF, ONNX and NumPy through `ClikaRT::io`; tokenizers, chat templates and processors alongside. - **Serving.** A runtime with sessions, continuous batching and pipelines; an HTTP server and client; a CLI framework. - **Bindings.** Two levels over the one C++ public surface: Python (a wheel) binds all of it; Kotlin (a desktop library and an Android AAR) binds what runs a ready model, under the model library's own names. [Language bindings](bindings.md) says what each level promises. ## Where to go next - [Getting started](getting-started/index.md): install the bundle and write your first program. - [System requirements](system-requirements.md): supported platforms and hardware. - [Examples](examples.md): the public example projects, one topic each. - [API reference](api/index.md): every public namespace, class and function. --- # API reference The ClikaRT public C++ API, by namespace and class. Source: https://docs.clika.io/clikart/api.md The public API of ClikaRT is the `ClikaRT::` namespace, declared by the headers under `ClikaRT/`. Include `ClikaRT/clika_rt.h` for everything, or the individual headers named on each page. ## Namespaces | Name | Description | | --- | --- | | [`ClikaRT::cli`](./ClikaRT/cli/index.md) | | | [`ClikaRT::device`](ClikaRT/device-namespace/index.md) | | | [`ClikaRT::dtype`](./ClikaRT/dtype/index.md) | The dtype-system helpers: classification predicates, size arithmetic, and the readable name. [`DataType`](./index.md#DataType) itself stays at [`ClikaRT`](./index.md)`::` (the one name every signature spells); everything ABOUT a dtype lives here. | | [`ClikaRT::encoding`](./ClikaRT/encoding/index.md) | | | [`ClikaRT::env`](./ClikaRT/env/index.md) | | | [`ClikaRT::graph`](./ClikaRT/graph/index.md) | | | [`ClikaRT::http`](./ClikaRT/http/index.md) | | | [`ClikaRT::io`](./ClikaRT/io/index.md) | | | [`ClikaRT::json`](./ClikaRT/json/index.md) | | | [`ClikaRT::logging`](./ClikaRT/logging/index.md) | | | [`ClikaRT::nn`](./ClikaRT/nn/index.md) | | | [`ClikaRT::ops`](./ClikaRT/ops/index.md) | | | [`ClikaRT::placement`](./ClikaRT/placement/index.md) | | | [`ClikaRT::platform`](./ClikaRT/platform/index.md) | | | [`ClikaRT::processor`](./ClikaRT/processor/index.md) | | | [`ClikaRT::progress`](./ClikaRT/progress/index.md) | | | [`ClikaRT::quant`](./ClikaRT/quant/index.md) | The quantized-checkpoint import taxonomy: the parsed `quantization_config` facts and the routing enums the split factories consume. | | [`ClikaRT::random`](./ClikaRT/random/index.md) | | | [`ClikaRT::regex`](./ClikaRT/regex/index.md) | | | [`ClikaRT::runtime`](./ClikaRT/runtime/index.md) | | | [`ClikaRT::spec`](./ClikaRT/spec/index.md) | | | [`ClikaRT::tables`](./ClikaRT/tables/index.md) | | | [`ClikaRT::threading`](./ClikaRT/threading/index.md) | | | [`ClikaRT::tokenizer`](./ClikaRT/tokenizer/index.md) | | ## Classes | Name | Description | | --- | --- | | [`Device`](./ClikaRT/Device.md) | A compute device: a backend API plus a device index (e.g. the `1` in CUDA device 1). Default-constructs to the whole-machine host CPU. | | [`EagerScope`](./ClikaRT/EagerScope.md) | While alive, forces EAGER execution on the calling thread: every op runs its kernel and produces a real value before dispatch returns, even when nested inside a [`TracingScope`](./ClikaRT/TracingScope.md). Use it to guard a region that must compute concrete values (a metric, a running average, a control-flow decision read back to the host) so a caller who wrapped the surrounding code in a [`TracingScope`](./ClikaRT/TracingScope.md) cannot turn those ops into un-materialized placeholders. Outside any scope execution is already eager, so this is a no-op there. Non-copyable, non-movable. | | [`Error`](./ClikaRT/Error.md) | Thrown by the throwing API overloads and by [Result::value\_or\_throw()](./ClikaRT/Result.md#value_or_throw). | | [`FakeTensor`](./ClikaRT/FakeTensor.md) | A value type (no out-of-line surface): it holds a `vector<`[`spec::SymInt`](./ClikaRT/spec/SymInt.md)`>`, a dtype, a copyable [`Stream`](./ClikaRT/Stream.md), and a flag, all ABI-stable, so it crosses by layout. | | [`NamedTensors`](./ClikaRT/NamedTensors.md) | An insertion-agnostic `name → `[`Tensor`](./ClikaRT/Tensor.md) map. Copyable and movable (a copy shares each entry's storage; the tensors are refcount handles). A moved-from [NamedTensors](./ClikaRT/NamedTensors.md) is empty: every read answers empty / false / 0, and [`set`](./ClikaRT/NamedTensors.md#set) fills it again (a dict moved into `load_state_dict` stays a usable object). | | [`ProfileSession`](./ClikaRT/ProfileSession.md) | A profiling session: start one or more captures, then export the results. Move-only. The session must outlive any in-flight async work it captured (workers that finish after a capture ends still record into it). | | [`ProfileSummary`](./ClikaRT/ProfileSummary.md) | The counters a session accumulated, as plain values: the typed form of [`ProfileSession::summary_text()`](./ClikaRT/ProfileSession.md#summary_text) and of the `summary` object in `report.json`, field for field, so a consumer reads numbers instead of parsing text. Durations are nanoseconds; counts are events recorded inside the session's captures. | | [`QTensor`](./ClikaRT/QTensor.md) | A quantized weight, viewed with its scheme. A value type over a refcounted payload handle; copying a [`QTensor`](./ClikaRT/QTensor.md) never copies weight bytes. | | [`Result`](./ClikaRT/Result.md) | Holds either a success value of type T or a ([Status](./index.md#Status), message) failure. Move-only. Check [ok()](./ClikaRT/Result.md#ok) before reading [value()](./ClikaRT/Result.md#value), or use [value\_or\_throw()](./ClikaRT/Result.md#value_or_throw). \[\[nodiscard\]\]: a call that returns a [`Result`](./ClikaRT/Result.md) and discards it swallows the failure it may carry; the compiler warns at such a site; consume it ([`unwrap`](./index.md#unwrap), [`CLIKART_CHECK`](./index.md#CLIKART_CHECK), a named read) or discard deliberately with a `(void)` cast. | | [`Result\`](./ClikaRT/Result.void.md) | Outcome of a fallible operation that yields no value on success (e.g. registering a route, writing a file). Check [`ok()`](./ClikaRT/Result.void.md#ok), or call [`value_or_throw()`](./ClikaRT/Result.void.md#value_or_throw) to raise [`ClikaRT::Error`](./ClikaRT/Error.md) on failure in your own TU. \[\[nodiscard\]\]: a statement-position effect call that ignores its result swallows the failure; the compiler warns; consume it ([`CLIKART_CHECK`](./index.md#CLIKART_CHECK), [`unwrap`](./index.md#unwrap), a named read) or discard deliberately with a `(void)` cast. | | [`Scalar`](./ClikaRT/Scalar.md) | A scalar operand that keeps its KIND across the boundary: a bool, an integer (any width, stays integral), or a double. What it buys over a bare double parameter: an integer scalar against an integer tensor stays in the integer domain (`ops::sub(2, int64_tensor)` is Int64, exact at any magnitude), where a double would force weak-float promotion. Every ctor is implicit on purpose. Pass the bare literal. Distinct from [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md): a [`Scalar`](./ClikaRT/Scalar.md) never carries a tensor, which is what keeps a two-tensor call (`ops::add(a, b)`) unambiguous against the tensor-first overloads. | | [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md) | A ClikaRT-level value type (NOT op-layer machinery): one optional slot that carries a scalar (a bool, an integer or a double) OR a tensor. Every ctor is implicit on purpose. Pass a bare `bool`, `double`, an integer, a [`Tensor`](./ClikaRT/Tensor.md), or `std::nullopt` straight to a [`ScalarOrTensor`](./ClikaRT/ScalarOrTensor.md) parameter; spelling the wrap at a call site (`ops::ScalarOrTensor(x)`) is redundant noise. [`ops::ScalarOrTensor`](./ClikaRT/ops/ScalarOrTensor.md) remains a valid spelling via the alias below. | | [`Span`](./ClikaRT/Span.md) | A non-owning view over [`size()`](./ClikaRT/Span.md#size) contiguous elements of `T`, the C++17 stand-in for `std::span` (see the file note above). It never allocates, copies, or owns: the viewed memory must outlive every copy of the view. `T`'s constness is the access law: [`Span`](./ClikaRT/Span.md)`` reads, [`Span`](./ClikaRT/Span.md)`` writes through to the caller's memory. | | [`Stream`](./ClikaRT/Stream.md) | A handle to a device execution stream. Cheap to copy (it is an identifier, not the stream's resources). | | [`StreamOrDevice`](./ClikaRT/StreamOrDevice.md) | A value type: it holds a copyable [`Stream`](./ClikaRT/Stream.md) handle and/or a [`Device`](./ClikaRT/Device.md), so it crosses the ABI by layout; [`resolve_impl`](./ClikaRT/StreamOrDevice.md#resolve) is its one library entry (the placement law lives in the runtime). | | [`SynchronousScope`](./ClikaRT/SynchronousScope.md) | While alive, every op dispatched on the calling thread COMPLETES before the dispatch call returns, whichever stream runs it: the moment a call returns, its result is ready. Thread-local (it changes nothing for other threads and throttles no stream), so it opens anywhere, on a busy lane included; it trades away the throughput that asynchronous pipelining buys, for as long as it lives. Scopes nest. Non-copyable, non-movable. | | [`SynchronousStreamScope`](./ClikaRT/SynchronousStreamScope.md) | While alive, every op dispatched on the calling thread WAITS for completion before the dispatch call returns, and `stream` runs at most one task at a time, so the moment a dispatch returns, the result is ready ([`Stream::query_idle()`](./ClikaRT/Stream.md#query_idle) is true). Deterministic, ideal for debugging / per-op stepping; it trades away the throughput that async pipelining buys. The stream's prior limit is restored on destruction. Open it only on a quiescent stream (synchronize first). Non-copyable, non-movable. | | [`Tensor`](./ClikaRT/Tensor.md) | | | [`TracingScope`](./ClikaRT/TracingScope.md) | While alive, the calling thread builds a LAZY graph: ops return `Unscheduled` placeholder tensors carrying lineage, and NO kernel runs. Call [`Tensor::synchronize()`](./ClikaRT/Tensor.md#synchronize) to materialize the graph in one pass. Outside the scope execution is eager (the default). Non-copyable, non-movable. | ## Enumerations ### enum Status {#Status}
`enum class` **`Status`** `:` `std::int32_t`
Outcome of an operation. | Enumerator | Value | Description | | --- | --- | --- | | `Ok` | `0` | | | `InvalidArgument` | | | | `NotFound` | | | | `Unsupported` | | | | `Internal` | | | | `Unavailable` | | The work did not run: it was shed under load, ran past its deadline, or was canceled. Retry later (HTTP 503). |
Declared in `ClikaRT/common/result.h`, line 93
### enum DataType {#DataType}
`enum class` **`DataType`** `:` `std::int32_t`
The element format of a tensor's values. `int32`-backed for a stable ABI. The set covers booleans, signed/unsigned integers (including sub-byte widths), the IEEE-style floats, and the narrow floating formats used by quantized models (FP8 / FP6 / FP4). `Undefined` is the unset value. ABI note: the numeric values are part of the public ABI: new types append at the end; existing ones never reorder or drop. | Enumerator | Value | Description | | --- | --- | --- | | `Undefined` | `0` | | | `Bool` | | | | `Int2` | | | | `Int4` | | | | `Int8` | | | | `Int16` | | | | `Int32` | | | | `Int64` | | | | `UInt1` | | | | `UInt2` | | | | `UInt4` | | | | `UInt8` | | | | `UInt16` | | | | `UInt32` | | | | `UInt64` | | | | `Float16` | | | | `BFloat16` | | | | `Float32` | | | | `Float64` | | | | `Float8_E4M3` | | | | `Float8_E5M2` | | | | `Float8_E4M3FNUZ` | | | | `Float8_E5M2FNUZ` | | | | `Float8_E8M0` | | | | `Float6_E2M3` | | | | `Float6_E3M2` | | | | `Float4_E2M1` | | |
Declared in `ClikaRT/compute/data_type.h`, line 21
## Type aliases ### using OptionalTensor {#OptionalTensor}
`using` **`OptionalTensor`** `=` [`Tensor`](./ClikaRT/Tensor.md)
A tensor argument that may be omitted. It is just [`Tensor`](./ClikaRT/Tensor.md); pass a default [`Tensor`](./ClikaRT/Tensor.md)`{}` (an undefined tensor) to mean "not provided". The alias documents, at a signature, that the parameter is optional (e.g. a norm's `weight`/`bias`, an attention `mask`). Mirrors the role of `c10::optional<`[`Tensor`](./ClikaRT/Tensor.md)`>` in ATen, using the undefined-tensor sentinel [ClikaRT](./index.md) already carries.
Declared in `ClikaRT/compute/tensor.h`, line 950
## Functions ### is\_capability\_decline() {#is_capability_decline}
`bool` **`is_capability_decline`**(`std::string_view` `code_name`) `noexcept`
True when `code_name` (an [`Error::code_name()`](./ClikaRT/Error.md#code_name) or [`Result::code_name()`](./ClikaRT/Result.md#code_name)) names a capability decline: a case the runtime does not serve, whether the operation itself, the device, or the dtype, memory layout, quantization scheme, source and destination pair, geometry or parameter value it was asked for. It is the one failure a caller may serve another way (another device, another route); every other failure belongs to the operation, and so does a backend that runs nothing at all (`BACKEND_UNAVAILABLE`, `BACKEND_NOT_LOADED`, `BACKEND_DRIVER_ABSENT`). False for an empty or unknown name; names compare exactly, case included.
Declared in `ClikaRT/common/result.h`, line 160
### status\_of() {#status_of}
[`Status`](./index.md#Status) **`status_of`**(`std::string_view` `code_name`) `noexcept`
The coarse [`Status`](./index.md#Status) a failure named `code_name` carries: the `status()` the runtime reports beside that `code_name()`, so a binding or a log reader can class a failure from its name alone. [`Status::Internal`](./index.md#Status) for an empty or unknown name; names compare exactly, case included.
Declared in `ClikaRT/common/result.h`, line 166
### unwrap(T) {#unwrap}
`template <``typename T`, `typename std::enable_if>::value, int>::type` `=` `0``>`
`constexpr` `T&` **`unwrap`**(`T&` `value`) `noexcept`
The lvalue identity: the argument itself, by reference (no copy).
Declared in `ClikaRT/common/result.h`, line 495
### unwrap(T) {#unwrap-2}
`template <``typename T`, `typename std::enable_if>::value &&!std::is_lvalue_reference::value, int>::type` `=` `0``>`
`constexpr` `detail::unwrap_plain_t` **`unwrap`**(`T&&` `value`)
The rvalue identity: the argument's VALUE (one move), never a reference into it. A reference-returning arm would hand a range-for or an `auto&&` binding the storage of a temporary that dies at the end of the range-init full-expression; returning the value makes that binding own what it reads.
Declared in `ClikaRT/common/result.h`, line 507
### unwrap(Result\) {#unwrap-3}
`template <``typename T``>`
`T` **`unwrap`**([`Result`](./ClikaRT/Result.md)`&&` `r`)
The [`Result`](./ClikaRT/Result.md)`` read: move the value out, or raise on failure.
Declared in `ClikaRT/common/result.h`, line 513
### unwrap(Result\) {#unwrap-4}
`void` **`unwrap`**([`Result`](./ClikaRT/Result.md)`&&` `r`)
The [`Result`](./ClikaRT/Result.md)`` read: no value; run the failure check only.
Declared in `ClikaRT/common/result.h`, line 515
### GetVersionInfo() {#GetVersionInfo}
`std::string` **`GetVersionInfo`**()
Human-readable [ClikaRT](./index.md) version, e.g. "0.1.0".
Declared in `ClikaRT/common/version.h`, line 13
### quantized\_view() {#quantized_view}
[`QTensor`](./ClikaRT/QTensor.md) **`quantized_view`**([`Tensor`](./ClikaRT/Tensor.md) `tensor`)
Declared in `ClikaRT/compute/q_tensor.h`, line 92
### make\_quantized() {#make_quantized}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized`**(
    [`Tensor`](./ClikaRT/Tensor.md) `payload`,
    `std::string_view` `scheme`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`,
    [`OptionalTensor`](./index.md#OptionalTensor) `global_scale` `=` `{}`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 114
### make\_quantized\_mxfp4() {#make_quantized_mxfp4}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_mxfp4`**(
    [`Tensor`](./ClikaRT/Tensor.md) `blocks`,
    [`Tensor`](./ClikaRT/Tensor.md) `scales`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 122
### make\_quantized\_mxfp8() {#make_quantized_mxfp8}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_mxfp8`**(
    [`Tensor`](./ClikaRT/Tensor.md) `codes`,
    [`Tensor`](./ClikaRT/Tensor.md) `e8m0_scale`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 148
### make\_quantized\_nvfp4() {#make_quantized_nvfp4}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_nvfp4`**(
    [`Tensor`](./ClikaRT/Tensor.md) `codes`,
    [`Tensor`](./ClikaRT/Tensor.md) `sub_scales`,
    [`OptionalTensor`](./index.md#OptionalTensor) `global_scale`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 166
### make\_quantized\_fp8() {#make_quantized_fp8}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_fp8`**(
    [`Tensor`](./ClikaRT/Tensor.md) `codes`,
    [`Tensor`](./ClikaRT/Tensor.md) `scale`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 195
### make\_quantized\_fp8\_blocked() {#make_quantized_fp8_blocked}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_fp8_blocked`**(
    [`Tensor`](./ClikaRT/Tensor.md) `codes`,
    [`Tensor`](./ClikaRT/Tensor.md) `scale`,
    `std::int64_t` `block_size`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 243
### make\_quantized\_affine() {#make_quantized_affine}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_affine`**(
    [`Tensor`](./ClikaRT/Tensor.md) `codes`,
    [`Tensor`](./ClikaRT/Tensor.md) `scales`,
    [`Tensor`](./ClikaRT/Tensor.md) `biases`,
    `std::int64_t` `group_size`,
    `std::int64_t` `bits`,
    [`Span`](./ClikaRT/Span.md)`` `logical_shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 284
### make\_quantized\_gptq() {#make_quantized_gptq}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_gptq`**(
    [`Tensor`](./ClikaRT/Tensor.md) `qweight`,
    [`Tensor`](./ClikaRT/Tensor.md) `qzeros`,
    [`Tensor`](./ClikaRT/Tensor.md) `scales`,
    [`OptionalTensor`](./index.md#OptionalTensor) `g_idx`,
    `std::int64_t` `bits`,
    `std::int64_t` `group_size`,
    [`quant::QuantizationConfig::ZerosConvention`](./ClikaRT/quant/QuantizationConfig.md#ZerosConvention) `zeros_convention`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 405
### make\_quantized\_bnb4() {#make_quantized_bnb4}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_bnb4`**(
    [`Tensor`](./ClikaRT/Tensor.md) `packed`,
    [`Tensor`](./ClikaRT/Tensor.md) `absmax`,
    [`OptionalTensor`](./index.md#OptionalTensor) `nested_absmax`,
    [`OptionalTensor`](./index.md#OptionalTensor) `nested_quant_map`,
    [`OptionalTensor`](./index.md#OptionalTensor) `offset`,
    [`OptionalTensor`](./index.md#OptionalTensor) `quant_map`,
    `std::string_view` `quant_type`,
    `std::int64_t` `blocksize`,
    `std::int64_t` `nested_blocksize`,
    [`Span`](./ClikaRT/Span.md)`` `shape`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 421
### make\_quantized\_awq() {#make_quantized_awq}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_awq`**(
    [`Tensor`](./ClikaRT/Tensor.md) `qweight`,
    [`Tensor`](./ClikaRT/Tensor.md) `qzeros`,
    [`Tensor`](./ClikaRT/Tensor.md) `scales`,
    `std::int64_t` `bits`,
    `std::int64_t` `group_size`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 431
### make\_quantized\_ct\_pack() {#make_quantized_ct_pack}
[`QTensor`](./ClikaRT/QTensor.md) **`make_quantized_ct_pack`**(
    [`Tensor`](./ClikaRT/Tensor.md) `packed`,
    [`Tensor`](./ClikaRT/Tensor.md) `scales`,
    `std::int64_t` `bits`,
    `std::int64_t` `group_size`
)
Declared in `ClikaRT/compute/q_tensor.h`, line 448
### operator+() {#operator-plus}
[`Tensor`](./ClikaRT/Tensor.md) **`operator+`**(`double` `scalar`, `const` [`Tensor`](./ClikaRT/Tensor.md)`&` `t`)
`scalar + t`, elementwise, the double-scalar mirror of `t + scalar`.
Declared in `ClikaRT/compute/tensor.h`, line 960
### operator\*() {#operator-mul}
[`Tensor`](./ClikaRT/Tensor.md) **`operator*`**(`double` `scalar`, `const` [`Tensor`](./ClikaRT/Tensor.md)`&` `t`)
`scalar * t`, elementwise.
Declared in `ClikaRT/compute/tensor.h`, line 962
### operator-() {#operator-minus}
[`Tensor`](./ClikaRT/Tensor.md) **`operator-`**(`double` `scalar`, `const` [`Tensor`](./ClikaRT/Tensor.md)`&` `t`)
`scalar - t`, elementwise: each element subtracted FROM the scalar.
Declared in `ClikaRT/compute/tensor.h`, line 964
### operator/() {#operator-div}
[`Tensor`](./ClikaRT/Tensor.md) **`operator/`**(`double` `scalar`, `const` [`Tensor`](./ClikaRT/Tensor.md)`&` `t`)
`scalar / t`, elementwise: the scalar divided BY each element.
Declared in `ClikaRT/compute/tensor.h`, line 966
### operator\<\<() {#operator-lshift}
`std::ostream&` **`operator<<`**(`std::ostream&` `os`, `const` [`Tensor`](./ClikaRT/Tensor.md)`&` `t`)
[Stream](./ClikaRT/Stream.md) a tensor's `to_string()` summary (shape, dtype, device, first values) to an `ostream`, so `std::cout << t << '\n'` works. Header-only; forwards to the (infallible) `to_string()`; reads values for any dtype, integers included.
Declared in `ClikaRT/compute/tensor.h`, line 971
## Platform notes {#platform_notes} ### macOS: JIT acceleration and the Hardened Runtime {#macos_jit} On macOS, [ClikaRT](./index.md)'s CPU runtime can accelerate some workloads by compiling specialized kernels at run time. The host application must be allowed to map executable memory: an application built with the Hardened Runtime needs the `com.apple.security.cs.allow-jit` entitlement. Without it, [ClikaRT](./index.md) detects the restriction at startup and serves its standard kernels; results are identical, and only the acceleration is unavailable. ## `ClikaRT/cli/cli.h` {#cli-h} ```cpp #include ``` Umbrella for the public command-line parser ([`ClikaRT::cli`](./ClikaRT/cli/index.md)): typed options / flags / positionals, subcommand routing, display + mutually-exclusive groups, auto help/usage, env-var fallbacks, and shell-completion generation. Every fallible call returns `Result`; the parser never throws, never exits, and never mutates argv. ## `ClikaRT/clika_rt.h` {#clika_rt-h} ```cpp #include ``` [ClikaRT](./index.md) public API umbrella header: one include for the whole public C++ API, the [`ClikaRT`](./index.md)`::` namespace with one sub-namespace per module, and the one link target `ClikaRT::ClikaRT` (`find_package(ClikaRT CONFIG)`). `AGENTS.md` beside these headers is the lookup index: what each module gives, which header to open for a task, and the gotchas. The modules, each under its own directory and namespace: `compute/` (`Tensor`, `Device`, `Stream`, the dtypes, the execution scopes, and `ops::`, the operator library), `nn/` (the module tree a model is built from, the packed leaves, the KV cache), `graph/` (`ModelGraph`, trace, compile, the queries and transforms), `io/` (checkpoints, media, the ONNX file), `runtime/` (the serving runtime: nodes, executors, pipelines), `tokenizer/`, `processor/`, `json/`, `template/`, `regex/`, `http/`, `cli/`, `progress/`, `profiler/`, `tables/`, `threading/`, `logging/`, `encoding/`, `platform/` and `common/` (the result type, `Span`, the environment names, the version). Three laws every program meets first. Every fallible call has two spellings: the distribution's default build returns values and raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure, a build with the `Result` surface returns `Result`; [`ClikaRT::unwrap`](./index.md#unwrap), [`CLIKART_TRY`](./index.md#CLIKART_TRY) and [`CLIKART_CHECK`](./index.md#CLIKART_CHECK) read the same under both, and a failure's `code_name()` is the name to branch on ([`common/result.h`](./index.md#result-h)). Operators are asynchronous: an `ops::` call queues its work and returns, and a host read (`to_string`, `item`, `item_as_vec`) waits for the value ([`compute/stream.h`](./ClikaRT/Stream.md), [`compute/scope.h`](./index.md#scope-h)). Every process that runs an operator or a model needs the license credential, `CLIKA_RT_LICENSE` or the per-user file, before its first call ([`common/env_vars.h`](./ClikaRT/env/index.md#env_vars-h)). ## `ClikaRT/common/macros.h` {#macros-h} ```cpp #include ``` ### Macros #### \#define CLIKART\_IS\_WINDOWS {#CLIKART_IS_WINDOWS}
`#define` **`CLIKART_IS_WINDOWS`** `0`
Declared in `ClikaRT/common/macros.h`, line 12
#### \#define CLIKART\_IS\_ANDROID {#CLIKART_IS_ANDROID}
`#define` **`CLIKART_IS_ANDROID`** `0`
Declared in `ClikaRT/common/macros.h`, line 18
#### \#define CLIKART\_IS\_LINUX {#CLIKART_IS_LINUX}
`#define` **`CLIKART_IS_LINUX`** `0`
Declared in `ClikaRT/common/macros.h`, line 24
#### \#define CLIKART\_IS\_MACOS {#CLIKART_IS_MACOS}
`#define` **`CLIKART_IS_MACOS`** `0`
Declared in `ClikaRT/common/macros.h`, line 40
#### \#define CLIKART\_IS\_IOS {#CLIKART_IS_IOS}
`#define` **`CLIKART_IS_IOS`** `0`
Declared in `ClikaRT/common/macros.h`, line 41
#### \#define CLIKART\_IS\_APPLE {#CLIKART_IS_APPLE}
`#define` **`CLIKART_IS_APPLE`** `(CLIKART_IS_MACOS || CLIKART_IS_IOS)`
Declared in `ClikaRT/common/macros.h`, line 44
#### \#define CLIKART\_IS\_X86\_64 {#CLIKART_IS_X86_64}
`#define` **`CLIKART_IS_X86_64`** `0`
Declared in `ClikaRT/common/macros.h`, line 49
#### \#define CLIKART\_IS\_ARM64 {#CLIKART_IS_ARM64}
`#define` **`CLIKART_IS_ARM64`** `0`
Declared in `ClikaRT/common/macros.h`, line 55
#### \#define CLIKART\_PUBLIC\_EXPORT {#CLIKART_PUBLIC_EXPORT}
`#define` **`CLIKART_PUBLIC_EXPORT`**
Declared in `ClikaRT/common/macros.h`, line 79
#### \#define CLIKART\_LOCAL {#CLIKART_LOCAL}
`#define` **`CLIKART_LOCAL`**
Declared in `ClikaRT/common/macros.h`, line 88
#### \#define CLIKART\_STATIC\_C\_API {#CLIKART_STATIC_C_API}
`#define` **`CLIKART_STATIC_C_API`**
Declared in `ClikaRT/common/macros.h`, line 97
#### \#define CLIKART\_HAS\_EXCEPTIONS {#CLIKART_HAS_EXCEPTIONS}
`#define` **`CLIKART_HAS_EXCEPTIONS`** `0`
Declared in `ClikaRT/common/macros.h`, line 112
## `ClikaRT/common/result.h` {#result-h} ```cpp #include ``` Value-returned result type for the [ClikaRT](./index.md) API. Holds a value on success or a Status + message on failure. Every fallible public method returns `Result` and never throws on its own. You turn a failure into an exception on the calling side with `value_or_throw()` (which raises [`ClikaRT::Error`](./ClikaRT/Error.md)), or inspect `ok()` / `status()` / `message()` and never pay for exceptions at all. Works with exceptions disabled: a consumer compiling with `-fno-exceptions` (or defining `CLIKART_NO_EXCEPTIONS`) gets a `value_or_throw()` that reports the failure to stderr and `std::abort()`s instead of throwing; the inspecting API (`ok()` / `status()` / `message()`) is unaffected. ── The error-handling vocabulary: each name's role ───────────────────── Two LAYERS live in this header and they are not duplicates: The LIBRARY's own boundary shims (not for consumer code): - [`CLIKART_RESULT(T)`](./index.md#CLIKART_RESULT) / [`CLIKART_UNWRAP(expr)`](./index.md#CLIKART_UNWRAP): how the public headers' inline wrappers adapt the library's `Result` to the built error-handling shape ([`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE)). Consumer code never writes these. - [`CLIKART_INPLACE_RESULT(T)`](./index.md#CLIKART_INPLACE_RESULT) / [`CLIKART_INPLACE_UNWRAP(expr, self)`](./index.md#CLIKART_INPLACE_UNWRAP): the same adaptation for the IN-PLACE wrappers (the write-through ops and mutating tensor methods): under the value-returning surface they keep the historical `T&` return (the mutated `self`/`out`); under the `Result`-returning surface they return the impl's `Result` straight through. [`CLIKART_CHECK`](./index.md#CLIKART_CHECK) is the caller's mode-stable spelling for these calls. The caller vocabulary (each spelling compiles and behaves identically whichever shape the library was built with): - `ClikaRT::unwrap(x)`: read a value; extracts a `Result` (raising on failure) and forwards anything else unchanged, so the same call-site text serves both shapes. - Implicit extraction: `T x = fn(...);` / `g(fn(...))` compiles under BOTH shapes: a TEMPORARY `Result` converts to `T`, raising on failure exactly as `unwrap`. Call sites written against the value-returning surface keep compiling when a build opts into the `Result` surface, so a codebase adopts explicit handling gradually rather than all at once. Rvalue-only (a NAMED `Result` is read through `ok()`/`value()`/`unwrap`), and excluded for `bool` payloads (`operator bool` tests OKNESS everywhere, never the payload; read a bool payload through `unwrap`/`value()`). Note `auto x = fn(...);` still binds the `Result` itself; spell the type (or `unwrap`) to extract. - [`CLIKART_TRY(expr)`](./index.md#CLIKART_TRY): the CAPTURE idiom, "hand me data, never throw": yields a `Result` whatever happens, catching [`ClikaRT::Error`](./ClikaRT/Error.md) AND any other exception (nothing propagates). - [`CLIKART_TRY_OR_RETURN(var, expr)`](./index.md#CLIKART_TRY_OR_RETURN) / [`CLIKART_CHECK(expr)`](./index.md#CLIKART_CHECK): the PROPAGATE idiom for `Result`-returning consumer functions: on failure they RETURN the error (status + message + code\_name verbatim) to the enclosing function's caller; only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted; foreign exceptions pass through untouched. ### Macros #### \#define CLIKART\_DETAIL\_HAS\_CXXABI {#CLIKART_DETAIL_HAS_CXXABI}
`#define` **`CLIKART_DETAIL_HAS_CXXABI`** `0`
Declared in `ClikaRT/common/result.h`, line 74
#### \#define CLIKART\_USE\_RESULT\_TYPE {#CLIKART_USE_RESULT_TYPE}
`#define` **`CLIKART_USE_RESULT_TYPE`** `0`
Declared in `ClikaRT/common/result.h`, line 87
#### \#define CLIKART\_TRY {#CLIKART_TRY}
`#define` **`CLIKART_TRY(...)`** `(::ClikaRT::detail::try_capture([&]() { return (__VA_ARGS__); }))`
CLIKART\_TRY: opt back INTO Result-style error handling. The public API returns values directly and raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure. When you would rather inspect a `Result` than catch, wrap the call: ```text Result r = CLIKART_TRY(ops::matmul(a, b)); if (!r.ok()) { log(r.message()); return; } use(r.value()); ``` Yields a `Result` where `T` is the (decayed) type the expression produces: `Result` for a `void` call, `Result` for a `Tensor`- or `Tensor&`-returning one. An expression that already produces a `Result` passes through as that same `Result` (never `Result>`), so the idiom reads identically whichever error-handling shape the library was built with. The expression is evaluated exactly once. A [`ClikaRT::Error`](./ClikaRT/Error.md) is captured with its status; any other exception is captured as `Status::Internal`; nothing propagates. This is the inverse of the internal unwrap-or-propagate idiom: it CATCHES exceptions at the call site and hands you data.
Declared in `ClikaRT/common/result.h`, line 581
#### \#define CLIKART\_RESULT {#CLIKART_RESULT}
`#define` **`CLIKART_RESULT(...)`** `__VA_ARGS__`
CLIKART\_RESULT / CLIKART\_UNWRAP: the two-mode public boundary shape. A fallible public method delegates to a `Result`-returning `*_impl` in the library and adapts that `Result` for the caller. These macros pick the adaptation at build time from [`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE) (the [`CLIKART_USE_RESULT_TYPE`](./index.md#CLIKART_USE_RESULT_TYPE) CMake option, baked into `build_info.h`); the SAME header source compiles both ways, no per-method edit to flip. [`CLIKART_UNWRAP`](./index.md#CLIKART_UNWRAP) **supplies its own `return`**, so the wrapper body is just the unwrap; one macro serves both a value and a `void` boundary (a `void`-returning function may `return` a `void` expression): [CLIKART\_RESULT(Tensor)](./index.md#CLIKART_RESULT) to(Device d) const \{ [CLIKART\_UNWRAP(to\_impl(d))](./index.md#CLIKART_UNWRAP); \} [CLIKART\_RESULT(void)](./index.md#CLIKART_RESULT) write(...) const \{ [CLIKART\_UNWRAP(write\_impl(...))](./index.md#CLIKART_UNWRAP); \} - **CLIKART\_USE\_RESULT\_TYPE == 0** (default): the return type is the bare value `T` (`void`), and the unwrap is `return (expr).value_or_throw();`; the wrapper raises [`ClikaRT::Error`](./ClikaRT/Error.md) on failure, caller-side. Byte-behaviour-identical to the historical value-returning surface. - **CLIKART\_USE\_RESULT\_TYPE == 1**: the return type is `Result` (`Result`) and the unwrap is `return (expr);`; the impl's `Result` passes straight through; the wrapper never throws and the caller inspects `.ok()` / `.status()` / `.value()`. [`CLIKART_RESULT`](./index.md#CLIKART_RESULT) is variadic so a comma-bearing type ([`CLIKART_RESULT`](./index.md#CLIKART_RESULT)`(std::array)`, [`CLIKART_RESULT`](./index.md#CLIKART_RESULT)`(std::vector)`) is not split into two macro arguments. Boundaries that transform the impl's result before returning it (an in-place op returning `Tensor&` / `*this`, a templated readback, a contained callback) keep their explicit `value_or_throw()`; they are not `Result`-shaped and are the deliberate carve-outs (same family as the user-callback types).
Declared in `ClikaRT/common/result.h`, line 617
#### \#define CLIKART\_UNWRAP {#CLIKART_UNWRAP}
`#define` **`CLIKART_UNWRAP(expr)`** `return (expr).value_or_throw()`
Declared in `ClikaRT/common/result.h`, line 618
#### \#define CLIKART\_INPLACE\_RESULT {#CLIKART_INPLACE_RESULT}
`#define` **`CLIKART_INPLACE_RESULT(...)`** `__VA_ARGS__&`
CLIKART\_INPLACE\_RESULT / CLIKART\_INPLACE\_UNWRAP: the boundary shims for the IN-PLACE wrapper family (a write-through op returning its `out`, a mutating tensor method returning `*this`). The impl twins all return `Result`; the wrapper manufactures the reference: inline [CLIKART\_INPLACE\_RESULT(Tensor)](./index.md#CLIKART_INPLACE_RESULT) relu\_(Tensor\& self) \{ [CLIKART\_INPLACE\_UNWRAP(impl::relu\_(self), self)](./index.md#CLIKART_INPLACE_UNWRAP); \} - **CLIKART\_USE\_RESULT\_TYPE == 0** (default): the return type is `T&` and the unwrap is `(expr).value_or_throw(); return self;`, the historical reference-returning surface, raising [`ClikaRT::Error`](./ClikaRT/Error.md) caller-side on failure. - **CLIKART\_USE\_RESULT\_TYPE == 1**: the return type is `Result` and the unwrap is `return (expr);`; the impl's `Result` passes straight through (the caller already holds the buffer it handed in). The caller's mode-stable spelling for these calls is [`CLIKART_CHECK(...)`](./index.md#CLIKART_CHECK); reference-chaining is a value-surface-only idiom.
Declared in `ClikaRT/common/result.h`, line 643
#### \#define CLIKART\_INPLACE\_UNWRAP {#CLIKART_INPLACE_UNWRAP}
`#define` **`CLIKART_INPLACE_UNWRAP(expr, self)`** `(expr).value_or_throw(); return self`
Declared in `ClikaRT/common/result.h`, line 644
#### \#define CLIKART\_DETAIL\_CONCAT2 {#CLIKART_DETAIL_CONCAT2}
`#define` **`CLIKART_DETAIL_CONCAT2(a, b)`** `a##b`
Declared in `ClikaRT/common/result.h`, line 648
#### \#define CLIKART\_DETAIL\_CONCAT {#CLIKART_DETAIL_CONCAT}
`#define` **`CLIKART_DETAIL_CONCAT(a, b)`** `CLIKART_DETAIL_CONCAT2(a, b)`
Declared in `ClikaRT/common/result.h`, line 649
#### \#define CLIKART\_TRY\_OR\_RETURN {#CLIKART_TRY_OR_RETURN}
`#define` **`CLIKART_TRY_OR_RETURN(var, expr)`** `auto CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__) = \ ::ClikaRT::detail::propagate_capture([&]() { return (expr); }); \ if (!CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__).ok()) \ return {CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__).status(), \ CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__).message(), \ CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__).code_name()}; \ var = std::move(CLIKART_DETAIL_CONCAT(clikart_try_state_, __LINE__).value())`
CLIKART\_TRY\_OR\_RETURN: the PROPAGATE idiom's value form, for `Result`-returning consumer functions. ```text Result mean_of(Tensor t) { CLIKART_TRY_OR_RETURN(auto m, ClikaRT::ops::mean(t)); return m.item(); } ``` Evaluates `expr` exactly once; on success declare-assigns (or assigns) `var` from the value; on a LIBRARY failure returns the error (status, message, and the fine `code_name` verbatim, through the three-argument `Result` constructor) to the enclosing function's caller. Only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted; a foreign exception passes through untouched. The SAME text serves both error-handling shapes: under the value-returning surface `expr` yields the bare value (a failure throws and is converted here); under the `Result`-returning surface `expr` yields a `Result` that passes through unflattened. Error timing is the settle-time law (see `unwrap`): failures surface at the read site, and a settled failure re-reports identically on re-read. Expands to multiple statements; use it in statement context (never as an unbraced `if` body), one per source line. With exceptions disabled a failure already aborted inside the library, so a reached call only ever sees success.
Declared in `ClikaRT/common/result.h`, line 673
#### \#define CLIKART\_CHECK {#CLIKART_CHECK}
`#define` **`CLIKART_CHECK(expr)`** `do { \ auto clikart_check_state_ = \ ::ClikaRT::detail::propagate_capture([&]() { return (expr); }); \ if (!clikart_check_state_.ok()) \ return {clikart_check_state_.status(), \ clikart_check_state_.message(), \ clikart_check_state_.code_name()}; \ } while (false)`
CLIKART\_CHECK: the PROPAGATE idiom's effect form: `expr` is run for its effect (an in-place op, a write, a registration) and any LIBRARY failure returns the error (status/message/`code_name` verbatim) to the enclosing function's caller; only [`ClikaRT::Error`](./ClikaRT/Error.md) is converted, foreign exceptions pass through. The mode-stable spelling for the in-place `Tensor&`-returning ops ([`CLIKART_CHECK(ops::relu_(x))`](./index.md#CLIKART_CHECK)`;`) and for any `Result` call under the `Result`-returning surface. Single statement (safe as an `if` body); settle-time + sticky re-read semantics as on `unwrap`.
Declared in `ClikaRT/common/result.h`, line 691
## `ClikaRT/common/version.h` {#version-h} ```cpp #include ``` [ClikaRT](./index.md) version information. ## `ClikaRT/compute/scope.h` {#scope-h} ```cpp #include ``` RAII execution-mode scopes. [ClikaRT](./index.md) executes asynchronously by default; ops dispatch and return, and the work runs later. These scopes change that for the calling thread, for their lifetime: force every op to complete before dispatch returns (deterministic / debug), or switch to lazy graph-building (tracing). ## `ClikaRT/http/http.h` {#http-h} ```cpp #include ``` HTTP umbrella: the client ([`client.h`](./ClikaRT/http/DownloadOptions.md)) and the server ([`server.h`](./ClikaRT/http/index.md#server-h)) together. Include this to get both sides; include the role header directly ([`ClikaRT/http/client.h`](./ClikaRT/http/DownloadOptions.md) or [`ClikaRT/http/server.h`](./ClikaRT/http/index.md#server-h)) to name the side you use. ## `ClikaRT/nn/nn.h` {#nn-h} ```cpp #include ``` [ClikaRT](./index.md) public NN-module umbrella; include this one header for the whole `nn` surface. New public modules are added HERE (and only here); consumers and [`clika_rt.h`](./index.md#clika_rt-h) never enumerate the individual headers. A translation unit that defines a class deriving from `nn::Module` is compiled with run-time type information disabled, as the library is. ## `ClikaRT/profiler/profiler.h` {#profiler-h} ```cpp #include ``` Capture a profile of [ClikaRT](./index.md) compute work and export it: a Chrome trace you can open in `chrome://tracing` / Perfetto, a per-op summary, or the raw report JSON. While a capture is active, every op dispatched on the capturing thread (and on the streams it drives) is timed automatically; you do not instrument individual calls. Work that other threads dispatch is outside the capture: a serving pipeline runs its nodes on its own executor threads, so a capture opened around a pipeline request records only the capture's own markers. Profile a pipeline through the `clikart-cli` CLI's `bench --profile`, which captures where the nodes run; a model driven on the calling thread (a detector, a transcriber, an embedder) profiles through this session directly. The shape is RAII: ```text ClikaRT::ProfileSession session; { auto capture = session.start_capture("decode"); // ... run ops / drive streams ... } // capture ends here session.save("profile_out"); // chrome_trace.json + summary.txt + ... printf("%s\n", ClikaRT::unwrap(session.summary_text()).c_str()); const ClikaRT::ProfileSummary s = ClikaRT::unwrap(session.summary()); // the same numbers, typed printf("%zu allocations\n", s.total_allocations); ``` --- # Kotlin API The ClikaRT Kotlin binding, documented from its sources. Source: https://docs.clika.io/clikart/api-kt.md The [`clika-runtime`](./clika-runtime/index.md) module, package by package. | Package | | --- | | [`io.clika.modelverse`](./clika-runtime/io.clika.modelverse/index.md) | | [`io.clika.runtime`](./clika-runtime/io.clika.runtime/index.md) | --- # clika_runtime The clika_runtime module, introspected from the installed wheel. Source: https://docs.clika.io/clikart/api-py.md clika-runtime: the Python package of the ClikaRT on-device inference runtime. Importing this package loads the runtime library shipped in-package (under ``clika_runtime/lib/``). The Modelverse model library ships in the same package as :mod:`clika_runtime.modelverse` (``import clika_runtime.modelverse as mv``); importing it loads the runtime first and resolves it from here. The top level publishes :class:`Tensor`, :class:`Size`, :class:`Device`, :class:`Stream`, the dtype objects (``float32``, ``bfloat16``, ...), the tensor factories (``tensor``, ``zeros``, ``randn``, ...), every operator of :mod:`clika_runtime.ops` as a free function (``clika_runtime.matmul(a, b)``), the placement and execution scopes (``device``, ``stream``, ``synchronous``, ``tracing``, ``eager``, ``meta_init``), ``eval`` / ``async_eval`` / ``synchronize``, ``save`` / ``load`` / ``load_with_metadata``, ``compile`` / ``trace``, the device functions, and the exception classes. Modes are strings (``approximate="tanh"``); the typed enumerations stay under ``clika_runtime._core.ops``. Models are built with :mod:`clika_runtime.nn`; :mod:`clika_runtime.torch` (imported on first use, never as a side effect) exchanges tensors and modules with PyTorch. Execution is asynchronous by default: an operator returns at once and its work rides the calling thread's current lane on the input's device. A value settles when it is read (``numpy()``, ``item()``, ``tolist()``, ``print``), when ``eval(*trees)`` or ``synchronize()`` is called, or inside a ``synchronous()`` region where every operator completes before it returns. ``tracing()`` records operations instead of running them until ``eval`` materializes them; ``eager()`` restores the default inside a tracing region. Reading the value of a traced or storage-free tensor raises :class:`ClikaRTError`. Every process that runs an operator or a model needs the license credential CLIKA issued for the project: ``CLIKA_RT_LICENSE`` in the environment (the ``CLIKA1-...`` text, or the path of a file holding it), set before this package is imported, else the per-user file the ``clikart-license-init`` console script writes once. Without one the first call raises :class:`ClikaRTError` with the code name ``LICENSE_FAILED``; every error carries a stable ``code_name`` beside its message, and the code name is the thing to branch on (:mod:`clika_runtime.errors`). A model from a model hub loads through :mod:`clika_runtime.modelverse`: ``mv.AutoModelForCausalLM.from_pretrained(repo_id)`` downloads the snapshot into the hub cache and loads it, ``mv.snapshot_download(repo_id)`` downloads without loading, and a local directory is a source everywhere a repository id is. A model written here from the operators starts from :mod:`clika_runtime.nn` (``nn.Module``, ``nn.Linear``, ``nn.KVCache``) and the fused attention operator; the documentation's examples walk one end to end. | Name | Kind | | --- | --- | | [`ClikaRTError`](./ClikaRTError.md) | class | | [`CompiledFunction`](./CompiledFunction.md) | class | | [`CompiledModule`](./CompiledModule.md) | class | | [`Device`](./Device.md) | class | | [`DeviceProperties`](./DeviceProperties.md) | class | | [`InternalError`](./InternalError.md) | class | | [`InvalidArgumentError`](./InvalidArgumentError.md) | class | | [`NotFoundError`](./NotFoundError.md) | class | | [`OutOfMemoryError`](./OutOfMemoryError.md) | class | | [`QTensor`](./QTensor.md) | class | | [`Size`](./Size.md) | class | | [`Stream`](./Stream.md) | class | | [`Tensor`](./Tensor.md) | class | | [`TensorSpec`](./TensorSpec.md) | class | | [`TracedGraph`](./TracedGraph.md) | class | | [`UnavailableError`](./UnavailableError.md) | class | | [`UnsupportedError`](./UnsupportedError.md) | class | | [`dtype`](./dtype.md) | class | | [`finfo`](./finfo.md) | class | | [`iinfo`](./iinfo.md) | class | | [functions](./functions.md) | module functions | --- # Language bindings ClikaRT has two API levels: FULL, the whole public surface, for C++ and Python; INFERENCE, everything that runs a ready model, for Kotlin. Every binding compiles against the C++ public headers. Source: https://docs.clika.io/clikart/bindings.md ClikaRT is a C++ library, and every other language reaches it through a binding compiled against the same public C++ headers a C++ program includes. There is no separate C layer in between: what a binding can do is exactly what the C++ API does, at one of two levels. ## The two levels | Level | Languages | What it covers | |---|---|---| | **FULL** | C++, Python | every public header: the runtime (tensors, operators, devices and streams, model loading, graphs and transforms, tokenizers and processors, serving) and the Modelverse model library | | **INFERENCE** | Kotlin | what runs, serves and measures a ready model: loading it from a file or a hub snapshot, the registry and its model cards, generation and chat with streaming and cancel, the tokenizers and processors the model needs, image, audio and video input, the task pipelines, serving, the benchmark and the fit check of a model, the device, hardware and memory facts, the host utilities (JSON, tables, templates, regular expressions), errors, logging, the profiler and progress, and the Android entry | A FULL language binds every public header. An INFERENCE language binds what runs a ready model and declines the rest by design: building or editing a model, the `nn` modules, graph queries and transforms, tracing and compiling your own code, and ONNX export are FULL-level work, done from C++ or Python. A model you author reaches an app as a served model or a compiled graph, never as Kotlin source. ## What each binding carries The rows are the surfaces a program reaches for; a cell names the member where the answer is partial. | Surface | C++ | Python | Kotlin | | --- | --- | --- | --- | | tensors, operators, dtypes, in-place forms | yes | yes | yes (`Ops`, one function per operator) | | devices, streams, execution scopes (synchronous, tracing, placement) | yes | yes | yes | | `nn` modules (`Linear`, `Conv`, `KVCache`, fused projections, `load_state_dict`) | yes | yes | no | | ONNX: open a model, compile it, optimize, run by position or name | yes | yes | yes (`OnnxModel`, `ModelGraph`) | | graph query and edit, transforms, `trace` and `compile` of your own code, ONNX export | yes | yes | no | | readers: safetensors, `.npy`, GGUF, images, audio, video | yes | yes | images, audio, video and `.npy` (`Io`, `VideoReader`); no safetensors or GGUF reader | | tokenizer, chat template, streaming decode | yes | yes | yes | | image, audio and video processors | yes | yes | yes | | the serving runtime: `FunctionModel`, `Executor`, `Pipeline`, batching | yes | yes | yes | | HTTP client and server, JSON, regular expressions, templates, tables | yes | yes | yes | | downloading a model from the hub | `hub::snapshot` (the model library) | `mv.snapshot_download` | `Modelverse.snapshot`, and inside `fromPretrained` | | the model library: generate, chat, serve, pipelines | yes | yes | yes | | the benchmark and the fit check of a model (`bench`, `check`) | yes | yes | yes (`BenchReport`, `FitReport`, `maxContextLength`) | | speech to text, text to speech, vision, translation | yes | yes | yes, as the handles `SttModel`, `TtsModel`, `VisionModel` and `TranslateModel` | | text embedding and reranking | yes | yes | no | | image and video generation | yes | yes | no | | PyTorch interoperation (`torch.compile` backend, DLPack exchange) | no | yes | no | | the Android entry (`ClikaRtAndroid.load`, memory-pressure trim, logcat) | no | no | yes | ## What every binding promises - **The same names.** The INFERENCE binding carries the model library's own object names (`AutoConfig`, `AutoModel`, `AutoProcessor`, `AutoTokenizer`, `fromPretrained`, `generate`, `pipeline(task)`), so a model loads, generates and serves under the names the C++ and Python surfaces use. - **The same errors.** A failure arrives as a typed error carrying the status, the stable code name and the message, in each language's own error type. - **The same version law.** A binding is built against one release of the runtime and refuses to load another, naming both versions. - **One runtime library per process.** The Python wheel carries the runtime library once for both products; the Kotlin artifact's Android variant carries it too, and a desktop JVM program takes it from the release archive's `lib/`. Nothing ships a second copy. ## The artifacts Every language ships one artifact that carries both products, the runtime and the Modelverse model library. Each is a download of the platform ([Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk)), taken from the same release. | Language | Level | How you get it | |---|---|---| | C++ | FULL | the release archive for your platform: the headers, the libraries, `find_package(ClikaRT CONFIG)` and `find_package(Modelverse CONFIG)` | | Python | FULL | the `clika-runtime` wheel for your CPython version and platform, installed from the file (`pip install `), with Modelverse inside as `clika_runtime.modelverse` | | Kotlin | INFERENCE | the `io.clika:clika-runtime` Maven artifact in `clika-runtime-maven-.zip`: one coordinate with an Android AAR variant and a desktop JVM jar variant, carrying `io.clika.runtime` and `io.clika.modelverse`. The AAR carries the runtime libraries, so an Android app needs nothing else; the desktop jar carries the classes and the JNI bridges and loads the runtime libraries from the release archive of the same platform, its `lib/` named on `java.library.path` | ## Where to go next - C++: the [API reference](api/index.md), every public namespace, class and function. - Python: [Use ClikaRT from Python](how-to/use-clikart-from-python.mdx), the wheel's shape from the NumPy boundary to models as `nn.Module`. - Kotlin: [Deploy to mobile](getting-started/first-program/06-deploy-to-mobile.mdx) in the tutorial, [Package ClikaRT in an Android app](how-to/package-clikart-in-an-android-app.mdx), and the Modelverse part [Use it from code](/modelverse/getting-started/first-model/use-it-from-code). --- # Additional examples The example programs that ship with a release: the runtime's own CMake projects inside every desktop archive, and the examples archive with complete C++, Python, Kotlin and Android programs beside their recorded output. Source: https://docs.clika.io/clikart/examples.md Two sets of example programs ship with a release, and both are downloads of the platform ([Get ClikaRT](getting-started/get-clikart.mdx)). The release archive for a desktop platform carries the runtime's own examples under `examples/src`: each is a standalone `find_package(ClikaRT CONFIG)` project on the public API only, each builds its topic up one chapter at a time, starting at `00_hello_world`, and every example directory has its own `README.md` walk-through; this page catalogs that set first. The examples archive, `ClikaRT--examples.tar.xz`, is the other: complete programs in C++, Python and Kotlin, three Android apps and a voice translator, each beside its recorded output, under one `examples/` directory whose `README.md` and `AGENTS.md` index them. ## The runtime's own examples, inside the archive Read roughly top to bottom; each row assumes a little of the ones above it. | Example | What it shows | Chapters | | --- | --- | --- | | `version` | The smallest consumer: link the bundle, print `GetVersionInfo()` | `00_hello_world` | | `compute` | The tensor/op engine: tensors and dtypes, device properties, data movement, the `ops::` library, the async model, zero-copy `.to(device)`, hardware probes, a distributed matmul, the profiler, cast chains | `00_hello_world` · `01_tensors` · `02_devices` · `03_data_movement` · `04_operators` · `05_async` · `06_zero_copy` · `07_hardware` · `08_distributed_matmul` · `09_profiler` · `10_cast_chain` | | `async` | The async execution model in depth: dispatch vs ready, safe host reads, `on_complete`, the synchronous scope, tracing | `00_dispatch_vs_ready` · `01_safe_reads` · `02_on_complete` · `03_sync_scope` · `04_tracing_scope` | | `runtime` | The serving runtime: nodes and phases, per-session state, continuous batching, pipelines, vision models, the interface layer, a two-stage model, an encoder with a stateful decoder | `00_functional_api` · `01_model_api` · `02_stateful_functional_api` · `03_stateful_model_api` · `04_pipeline` · `05_simple_vision_model` · `06_interface` · `07_two_stage_vision_model` · `08_encoder_and_stateful_decoder` | | `nn` | Neural-network modules: make with plain counts, bind weights, forward; Linear, Conv, and a KV cache bound straight into the attention op | `00_linear` · `01_conv` · `02_attention_kvcache` | | `processor` | Image and audio pre-processing: resize, rescale and normalize a generated image; log-mel features from a synthesized waveform | `00_image` · `01_audio` | | `json` | A small JSON value for payloads: parse, typed getters, build with `operator[]`, `dump` round trip | `00_hello_world` | | `errors` | The failure vocabulary: `unwrap`, `CLIKART_TRY` capture, `CLIKART_TRY_OR_RETURN` propagation, and `code_name()`, the stable machine-readable failure name | `00_result` | | `cli` | Typed argument parsing: options and flags, subcommands, validators, shell completion | `00_hello_world` · `01_args` · `02_subcommands` · `03_completion` | | `logging` | Structured logging: levels, the level filter, named subsystem loggers | `00_hello_world` · `01_levels` · `02_named` | | `progress` | Terminal progress: bars (bytes, ETA, rate), spinners, multi-bar groups | `00_hello_world` · `01_bytes_eta_rate` · `02_spinner` · `03_group` | | `templating` | Jinja2-compatible rendering: variables, logic, filters, chat prompts | `00_hello_world` · `01_logic` · `02_filters` · `03_chat_prompts` | | `tokenizer` | Text to token ids and back: HF `tokenizer.json`, byte offsets, batching and vocab, chat templates | `00_hello_world` · `01_offsets` · `02_batch_and_vocab` · `03_huggingface` | | `tables` | A columnar dataframe: CSV, dtypes, select/filter/sort, group-by/join, transforms, GPU acceleration, a row cursor | `00_hello_world` · `01_columns_and_dtypes` · `02_csv` · `03_select_filter_sort` · `04_groupby_join` · `05_transform` · `06_acceleration` · `07_rows` | | `tables_benchmark` | The pandas-vs-ClikaRT performance sibling of `tables` | one standalone project | | `io` | Loading data and weights: NumPy `.npy`, safetensors, GGUF, images, and the ONNX model stack ([guide](how-to/run-an-onnx-model.mdx)) | `00_hello_world` · `01_safetensors` · `02_gguf` · `03_image` · `04_onnx` | | `http_server` | An HTTP service: routing, JSON APIs, middleware, a templated site, SSE streaming, an image-upload endpoint that runs compute | `00_hello_world` · `01_routing` · `02_json_api` · `03_middleware` · `04_serve_a_website` · `05_sse` · `06_image_compute` | | `download` | A download tool combining CLI, progress bar and HTTP client | `00_hello_world` · `01_multi_file` | | `flash_attention` | A hand-written CUDA kernel on ClikaRT streams: upstream Flash Attention ported, not rewritten, dispatched through the custom-op interface | one project (`main.cpp` + the ported kernel) | ### Build and run ```bash cmake -S "$CLIKART_BUNDLE_DIR/examples/src/compute" -B build-compute \ -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build-compute ./build-compute/compute_00_hello_world ``` Or build them all at once from the top-level project: `cmake -S "$CLIKART_BUNDLE_DIR/examples/src" -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake"`. The top-level project adapts to the distribution it is built against; it skips an example the distribution cannot build rather than failing. On platforms where the archive carries pre-built example binaries (`examples/bin`), run them in place; they find the libraries through a relative rpath, with no library paths to set. ## The examples archive `ClikaRT--examples.tar.xz` extracts to one `examples/` directory. Every program in it runs on its own against a release, and every tutorial and how-to program sits beside its recorded output (`.out`), the text the release printed when the recording was made. `examples/README.md` is the front door, one table per language; `examples/AGENTS.md` is the index written for a reader or an AI assistant: the facts a first run needs, a task table, the commands, the download call in each language, and what each binding carries. Nothing in the archive reaches outside it. Each language's programs name the one distribution they need, taken from the same release: | Directory | Programs | Needs | Build and run | | --- | --- | --- | --- | | `cpp/` | `clika_rt/`, the engine one topic per directory in chapters; `tutorial/`, the getting-started programs; `howto/`, one directory per how-to page; `modelverse/`, programs over the model library ([the Modelverse catalog](/modelverse/examples)) | the release archive for your platform, extracted, with `CLIKART_BUNDLE_DIR` naming its directory | `cmake -S cpp -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" && cmake --build build`, then the binary under `build/`; `cpp/README.md` | | `python/` | `clika_rt/`, self-asserting chapters over the wheel, from the NumPy boundary through `nn.Module`, ONNX compile-and-run, tokenizers, processors and tracing; `tutorial/` and `howto/` as the C++ tree has them, with the Python-only pages (a model authored in Python, PyTorch interoperation, pytrees) | the `clika-runtime` wheel for your interpreter and platform, installed with `pip install ` | `python3 `; `python/README.md` | | `kotlin/` | `clika_rt/`, instrumented tests that run the operator surface and the device facts on a connected Android device; `tutorial/`, the getting-started programs on a desktop JVM; `howto/`, the vision-task programs on a desktop JVM | the Kotlin artifact, `clika-runtime-maven-.zip`, extracted anywhere and named to gradle as `-PclikaRtMavenRepo=`; a desktop run also takes the release archive's `lib/` on `java.library.path` | `gradle connectedAndroidTest -PclikaRtMavenRepo=` in a chapter's directory; `kotlin/README.md` | | `android/` | three sample apps, one screen each: `hello` loads the runtime and prints the version, the backends and one operator's result; `chat` loads a chat model and streams its replies; `serve` hosts a loaded model for the network from a foreground service | the Kotlin artifact, as above, and an Android SDK | one gradle build beside `shared/`; `android/README.md` | | `voice_translate/` | a live voice translator as one Compose Multiplatform app for Android and the desktop JVM: speech to text, translation and text to speech over the model library's Kotlin binding | the Kotlin artifact and, per platform, the release archives the build lays out | `voice_translate/README.md` | Every program here runs compute, so every one of them needs a license credential in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes ([Get ClikaRT](getting-started/get-clikart.mdx#license-credential)); without one a call is refused with the code name `LICENSE_FAILED`. A model program names a checkpoint from the local hub cache or takes one as an argument; `clikart-cli fetch ` downloads it once, with no credential. The tutorial programs this documentation shows ([Your first program](getting-started/first-program/01-your-first-program.mdx) onward) are the archive's `cpp/tutorial/`, `python/tutorial/` and `kotlin/tutorial/` programs, with the output recorded beside them. [Use ClikaRT from Python](how-to/use-clikart-from-python.mdx) walks the first steps of the Python lane; [Language bindings](bindings.md) says what each binding carries. --- # First steps New to ClikaRT? Start here. What the runtime is, how to install it, and the tutorial series. Source: https://docs.clika.io/clikart/getting-started.md New to ClikaRT? This section is where to start. It gives enough orientation to hold the whole library in your head, an install you can verify in minutes, and a first program that grows into a real pipeline. Read it in order: 1. **[ClikaRT at a glance](overview.mdx)**: what the runtime is and is not, who it is for, and the five-minute mental model. 2. **[Get ClikaRT](get-clikart.mdx)**: pick your platform, get the download and verify commands. 3. **[Quick install](installation.md)**: prerequisites, the bundle, and two ways to prove it works. 4. **Tutorial series**: six parts, each a complete step, with core concepts explained where they first appear. [Your first program](first-program/01-your-first-program.mdx), [tensors and operators](first-program/02-tensors-and-operators.mdx), [devices and the async model](first-program/03-devices-and-async.mdx), [your first pipeline](first-program/04-your-first-pipeline.mdx) (weights from disk, compute on the best device, results back), [serve it](first-program/05-serve-it.mdx) (the same model behind the serving runtime, answering requests), and [deploy to mobile](first-program/06-deploy-to-mobile.mdx) (the part 4 program on a real phone, unchanged). 5. **[What to read next](next-steps.md)**: where to go once it runs. ## How the ClikaRT docs are layered - **This section** orients: condensed, in reading order, concepts woven in. - **[How-to guides](/how-to/index.md)**: problem-oriented recipes, one per "how do I X", with [additional examples](/examples.md) as the end-to-end reading inside it. - **[API reference](/api/index.md)**: the full reference per language: C++, Python and Kotlin. [System requirements](/system-requirements.md) sits alongside, for the platform and hardware tables. --- # Complete installation Per-platform toolchains, IDE setup, and the full failure catalog. Under construction. Source: https://docs.clika.io/clikart/getting-started/complete-installation.md This page will carry the full installation reference. Until it lands, [Quick install](installation.md) covers the fast path and [system requirements](/system-requirements.md) lists the platforms, toolchains and drivers. Planned contents: - Per-OS toolchains in depth: Linux distributions, macOS and Xcode, Windows and Visual Studio. - Platform notes: Windows, macOS, Android. - IDE setup. - Adding ClikaRT to an existing CMake project (also planned as a how-to guide). - The full failure catalog, beyond quick install's four lines. --- # Deploy to mobile Cross-compile the part 4 program for Android, push it over adb, and run it on the phone's GPU. The code does not change. Source: https://docs.clika.io/clikart/getting-started/first-program/deploy-to-mobile.md **Your code does not change.** The program from [part 4 of the series](04-your-first-pipeline.mdx), byte for byte, compiled for Android and pushed to a phone, prints the same numbers there that it printed on your workstation, on the phone's GPU when it has one and on its CPU otherwise. This page is the complete path: cross-compile, push, run, and the few lines of Kotlin an app adds around it. You need the bundle (its `android-arm64` dist ships the CPU and Vulkan backends), the Android NDK, and a phone with USB debugging enabled. ## 1. Cross-compile Same project, same `main.cpp`. The configure line adds the NDK's toolchain file and the ABI; the bundle's CMake package selects the Android dist on its own: ```bash cmake -S . -B build-android \ -DCMAKE_TOOLCHAIN_FILE="$ANDROID_NDK/build/cmake/android.toolchain.cmake" \ -DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 \ -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build-android ``` ```text -- ClikaRT 0.6.4: dist android-arm64 (backends: cpu;vulkan) ``` ## 2. Push and run The binary plus the two libraries from the Android dist go to the device; the phone needs nothing else installed: ```bash adb shell mkdir -p /data/local/tmp/clikart adb push build-android/hello \ "$CLIKART_BUNDLE_DIR/lib/libClikaRT.so" \ "$CLIKART_BUNDLE_DIR/lib/libClikaRT_vulkan.so" \ /data/local/tmp/clikart/ adb shell "cd /data/local/tmp/clikart && chmod +x hello && \ CLIKA_RT_LICENSE=CLIKA1-... LD_LIBRARY_PATH=. ./hello" ``` The program runs compute, so it needs a license credential like any other ClikaRT program, and a binary under `adb shell` reads it from the environment of the shell that starts it ([Get ClikaRT](../get-clikart.mdx#license-credential)). Push the credential to a file on the device and give `CLIKA_RT_LICENSE` that path when you would rather keep it off the command line. On a Galaxy S24 Ultra: ```text y = Tensor(shape=[2, 4], dtype=Float32, device=Vulkan:0, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...]) row means = [4.25, 4.25] (expected 8*1*0.5 + 0.25 = 4.25) ``` The workstation printed `device=CUDA:0`; the phone prints `device=Vulkan:0`. Same program, same numbers. `pick_device()` from [part 3](03-devices-and-async.mdx) found the phone's GPU the same way it found the workstation's; on this phone the part 3 discovery loop lists: ```text CPU 0: ARM Vulkan 0: Adreno (TM) 750 ``` ## 3. From an app: a few lines of Kotlin ClikaRT is a C++ library, and an Android app reaches it through the app's own native code. Kotlin loads your library and calls your function; ClikaRT stays on the native side: ```kotlin object Pipeline { init { System.loadLibrary("pipeline") } external fun run(): String } ``` ```cpp title="pipeline_jni.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; extern "C" JNIEXPORT jstring JNICALL Java_com_example_app_Pipeline_run(JNIEnv* env, jobject) { const Tensor x = Tensor::ones({2, 8}, DataType::Float32); const Tensor w = Tensor::full({4, 8}, 0.5, DataType::Float32); const Tensor y = ops::relu(ops::linear(x, w)); return env->NewStringUTF(y.to_string().c_str()); } ``` The CMake side adds one target next to `hello`: ```cmake add_library(pipeline SHARED pipeline_jni.cpp) target_link_libraries(pipeline PRIVATE ClikaRT::ClikaRT) ``` Beneath the JNI boundary the native side is the C++ library you built for arm64. Kotlin code can also skip the custom library entirely: the `io.clika:clika-runtime` artifact ships its own JNI bridge and every runtime library, and calls the runtime directly. Python stays on the server and desktop side. An app has no shell to export a variable from and no per-user license file, so the credential is an argument instead. `ClikaRtAndroid.load(context, license = "CLIKA1-...")` places it for the process before the binding loads. A credential given there replaces one already in the process environment, and the default, `null`, leaves the environment as it is. Ship the credential the way you ship any other secret your app needs at start-up. What the app ships beside the library, the build levels it declares (`compileSdk` 36, `minSdk` 28, the Android Gradle plugin 9.3.2 on gradle 9.5 or newer), the model library's own `libClikaRT_modelverse.so` when the app runs a model from the library, and the two facts it places before the first call (the CPU worker count `CLIKA_RT_NUM_THREADS`, read once before the first compute, and the cache root) are on [Package ClikaRT in an Android app](../../how-to/package-clikart-in-an-android-app.mdx). Every platform works this way: one bundle, one toolchain file, the same code. A Jetson needs no cross-compile at all: it is an arm64 Linux machine and runs the linux-arm64 build directly. The [system requirements](/system-requirements.md) table lists the platforms and their backends. The tutorial ends here; [what to read next](../next-steps.md). --- # Devices and the async model Discover the machine's devices, place tensors on them, and understand dispatch vs ready, and why host reads are always safe. Source: https://docs.clika.io/clikart/getting-started/first-program/devices-and-async.md The same program from parts 1-2 runs unchanged on a GPU. Data placement is the only new ingredient. This part adds device discovery and `.to(device)`, then explains the execution model behind every `ops::` call you have made so far. ## The program ```cpp title="main.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::device::ComputeAPI; namespace device = ClikaRT::device; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; // The best device this machine has, probed at run time. An unavailable // backend is a fact, not an error: the predicate answers false. namespace { Device pick_device() { if (device::is_cuda_available()) return Device::cuda(); if (device::is_vulkan_available()) return Device::vulkan(); if (device::is_metal_available()) return Device::metal(); return Device::cpu(); } } // namespace int main() { // What is this machine carrying? enumerate_devices lists the concrete // handles a backend exposes; get_device_properties describes one. for (ComputeAPI api : {ComputeAPI::CPU, ComputeAPI::CUDA, ComputeAPI::Vulkan, ComputeAPI::Metal}) { for (Device dev : device::enumerate_devices(api)) { const device::DeviceProperties p = device::get_device_properties(dev); std::printf("%-7s %d: %s\n", device::compute_api_name(api), dev.index, p.name.c_str()); } } const Device dev = pick_device(); std::printf("running on %s\n", device::compute_api_name(dev.api)); // .to(device) moves data; ops run where their inputs live. const Tensor a = Tensor::ones({512, 512}, DataType::Float32).to(dev); // This call DISPATCHES the matmul and returns. The kernel runs in its // own time; nothing here waits for it. const Tensor c = ops::matmul(a, a); // A host read is where the wait lands: it synchronizes first, so you // never observe unfinished bytes. Every element is 512 (= K). std::printf("every element = %.0f\n", ops::amax(c).item()); return 0; } ``` ```python title="main.py" import clika_runtime as crt def main() -> None: # The best device this machine has: Device.gpu() probes the available # accelerators and FALLS BACK to the CPU when none is present; it never # fails, so the same script runs everywhere. dev = crt.Device.gpu() print(f"running on {dev!r}") # Factories take the device directly; ops run where their inputs live. a = crt.ones(512, 512, device=dev) # This call DISPATCHES the matmul and returns. The kernel runs in its # own time; nothing here waits for it. c = a @ a # A host read is where the wait lands: it synchronizes first, so you # never observe unfinished bytes. Every element is 512 (= K). print(f"every element = {c.amax().item():.0f}") if __name__ == "__main__": main() ``` ```kotlin title="Main.kt" import io.clika.runtime.Backends import io.clika.runtime.ClikaRt import io.clika.runtime.ComputeApi import io.clika.runtime.Device import io.clika.runtime.Ops import io.clika.runtime.Tensors import java.util.Locale // The best device this machine has, probed at run time. A backend that is // not in this build, or whose devices do not serve, is a fact the probe // answers, never an error. private fun pickDevice(): Device { for (api in listOf(ComputeApi.CUDA, ComputeApi.VULKAN, ComputeApi.METAL)) { if (Backends.isAvailable(api) && Backends.deviceStatus(api) == null && Backends.deviceCount(api) > 0) { return Device(api, 0) } } return Device(ComputeApi.CPU) } fun main() { ClikaRt.load() // What is this machine carrying? Backends.devices lists the devices a // backend exposes; Device.properties() describes one. for (api in listOf(ComputeApi.CPU, ComputeApi.CUDA, ComputeApi.VULKAN, ComputeApi.METAL)) { if (!Backends.isAvailable(api)) continue for (dev in Backends.devices(api)) { println("%-7s %d: %s".format(Locale.ROOT, api.label, dev.index, dev.properties().name)) } } val dev = pickDevice() println("running on ${dev.api.label}") // Placement is a factory argument; operators run where their inputs live. val a = Tensors.ones(longArrayOf(512, 512), where = dev) // This call DISPATCHES the matmul and returns. The kernel runs in its // own time; nothing here waits for it. val c = Ops.matmul(a, a) // A host read is where the wait lands: item() synchronizes first, so // you never observe unfinished bytes. Every element is 512 (= K). val m = Ops.amax(c) println("every element = ${"%.0f".format(Locale.ROOT, m.item())}") listOf(a, c, m).forEach { it.release() } } ``` On a machine with an NVIDIA GPU: ```text CPU 0: AMD CUDA 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition CUDA 1: NVIDIA RTX PRO 6000 Blackwell Workstation Edition Vulkan 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition Vulkan 1: NVIDIA RTX PRO 6000 Blackwell Workstation Edition running on CUDA every element = 512 ``` The same binary on a CPU-only machine lists only the CPU and runs there; no rebuild, no configuration. ## Devices and backends A `Device` is a backend API plus a zero-based index. `Device::cuda(1)` is the second CUDA GPU, and `Device::cpu()` is the default everything starts on. Every distribution has the CPU backend compiled in; CUDA, Vulkan and Metal are shared libraries the runtime loads on demand, the first time something asks. `is_backend_available(api)` (and the shorthands `is_cuda_available()` and friends) answers whether that load works on this machine; `enumerate_devices(api)` returns the concrete handles, and an empty list is a valid answer, not an error. This is why the same binary runs everywhere. Absent hardware costs you a branch, not a build configuration. `.to(device)` returns the tensor moved (a no-op copy if it is already there), and operators run on the device their inputs live on; there is no global "current device" to set. ## Dispatch is not execution ClikaRT is asynchronous by nature. An `ops::` call **dispatches** work and returns; the kernel runs and the result becomes **ready** in its own time. On the default CPU stream ops happen to run inline, which is why parts 1-2 never confronted this. On an accelerator, or a worker stream made with `Stream::create`, the dispatch returns first and the compute overlaps with your code. `t.synchronize()` blocks until `t`'s pending work is done; `t.on_complete(callback)` is the push-style equivalent, firing when the result is ready. ## Reading across the async boundary Two rules cover every host read: 1. **Reads wait for you.** `item()`, `item_as_vec()` and `const_data_ptr()` synchronize before handing back bytes. A read issued right after dispatching heavy work blocks until the result is real. The wait moves into the read; it never disappears. You can never observe garbage through the public read surface. 2. **Pointers are device pointers.** `const_data_ptr()` addresses the buffer on the tensor's own device. On a CUDA tensor that is CUDA memory; call `t.to(Device::cpu())` first to read it on the host. (`item` / `item_as_vec` do the host transfer for you.) So the failure mode is never corruption, it is a surprise stall, a "cheap" read that waited for a matmul. When latency matters, choose where the wait lands: an explicit `synchronize()`, an `on_complete` callback, or a read whose cost you have accepted. The `async` [example project](/examples.md) measures all of this with timers. Next: [part 4](04-your-first-pipeline.mdx). It loads weights from disk, computes, and reads results back. --- # Serve it Wrap part 4's model in the serving runtime: a schema, a lambda serving the forward pass, and an executor answering requests. Source: https://docs.clika.io/clikart/getting-started/first-program/serve-it.md Part 4 ended with a program that loads a checkpoint and computes once. This part puts the same model behind the serving runtime: declare its input/output contract, serve the forward pass with a lambda, and hand requests to an executor. The tutorial ends where the overview's pitch ends, a model checkpoint answering requests. ## The program ```cpp title="main.cpp" #include #include #include #include #include namespace rt = ClikaRT::runtime; using ClikaRT::DataType; using ClikaRT::NamedTensors; using ClikaRT::Tensor; using ClikaRT::spec::TensorSpec; namespace io = ClikaRT::io; namespace ops = ClikaRT::ops; namespace { // The layer's geometry, the checkpoint's names, and the edge names the schema // declares and every request and response addresses. A misspelled edge name // routes nothing, silently, so each is spelled once. constexpr std::int64_t kFeatures = 8; constexpr std::int64_t kOutputs = 4; constexpr std::int64_t kBatch = 2; constexpr const char* kWeight = "mlp.weight"; constexpr const char* kBias = "mlp.bias"; constexpr const char* kX = "x"; constexpr const char* kY = "y"; } // namespace int main() { // The checkpoint from part 4, re-created so this program stands alone. NamedTensors weights; weights.set(kWeight, Tensor::full({kOutputs, kFeatures}, 0.5, DataType::Float32)); weights.set(kBias, Tensor::full({kOutputs}, 0.25, DataType::Float32)); const std::string ckpt = (std::filesystem::temp_directory_path() / "first_program.safetensors").string(); io::save_safetensors(weights, ckpt); // Load it and declare the model's I/O contract: batches of 8 features // in, batches of 4 activations out. kDynamicDim leaves the batch open. const NamedTensors loaded = io::load_safetensors(ckpt); const Tensor w = loaded.get(kWeight); const Tensor b = loaded.get(kBias); rt::ModelSchema schema; schema.inputs = {TensorSpec{kX, DataType::Float32, {TensorSpec::kDynamicDim, kFeatures}, false}}; schema.outputs = {TensorSpec{kY, DataType::Float32, {TensorSpec::kDynamicDim, kOutputs}, false}}; // Part 4's forward pass, served by a lambda. No subclass needed. rt::FunctionModel model{schema}; model.on_run_once("run", [&](rt::PhaseContext& ctx) { const Tensor x = ctx.inputs->get(kX); ctx.outputs->set(kY, ops::relu(ops::linear(x, w, b))); }); rt::Executor exec = rt::Executor::create(model); // borrowed; keep `model` alive // One request in, one response out: the serving loop in miniature. rt::Request req; req.inputs.set(kX, Tensor::ones({kBatch, kFeatures}, DataType::Float32)); const rt::Response resp = exec.await(exec.enqueue(std::move(req))); std::printf("y = %s\n", resp.outputs.get(kY).to_string().c_str()); exec.shutdown(); return 0; } ``` {/* CERTIFICATION: the arm as shown is verified against the pinned release's cp313 wheel (tools/tutorial_check.py over the program the page embeds; trace -> optimize -> finalize -> run, matching its recording). The pipeline executor over graphs (clika_runtime._core.runtime: Pipeline, Request, Response, enqueue/wait) is bound but has no published example chapter in the pinned tree (chapters end at 15_trace_and_compile), so the arm stops at the finalized ModelGraph. Extend it to the executor when the upstream chapter lands. */} The Python lane serves a traced graph: `crt.trace` captures the forward pass as a `ModelGraph`, and after `optimize()` and `finalize()` the finalized graph is the servable unit. The pipeline executor over graphs is bound but not yet published, so this arm stops at the graph. ```python title="main.py" import tempfile import clika_runtime as crt import clika_runtime.nn.functional as F def main() -> None: # The checkpoint from part 4, re-created so this program stands alone. ckpt = f"{tempfile.gettempdir()}/first_program.safetensors" crt.io.save_safetensors({ "mlp.weight": crt.full((4, 8), 0.5), "mlp.bias": crt.full((4,), 0.25), }, ckpt) loaded = crt.io.load_safetensors(ckpt) w, b = loaded["mlp.weight"], loaded["mlp.bias"] # The traced callable takes and returns LISTS of tensors. crt.trace runs # it once over data-free stand-ins and captures the operator graph, which # comes back as recorded: optimize() runs the graph optimizer and # finalize() readies the graph to serve. def forward(ins: list) -> list: return [F.relu(F.linear(ins[0], w, b))] g = crt.trace(forward, example_inputs=[crt.ones(2, 8)]) g.optimize() g.finalize() # A request in, a response out. (y,) = g.run([crt.ones(2, 8)]) print(f"y = {y}") if __name__ == "__main__": main() ``` The Kotlin binding does not carry the serving runtime; the C++ arm is the serving story today. ```text y = Tensor(shape=[2, 4], dtype=Float32, device=CPU, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...]) ``` The same 4.25s part 4 computed, produced this time by an executor answering a request. ## Schemas, lambdas, executors `ModelSchema` is the model's I/O contract: named `TensorSpec`s with dtype and dims, and `kDynamicDim` leaves a dimension open, so one served model accepts any batch size. `FunctionModel` serves a callable against that schema with no subclass; the lambda reads its inputs and sets its outputs through the `PhaseContext`. `Executor::create` borrows the model (keep it alive) and turns it into a queue: `enqueue` accepts a `Request`, `await` blocks for its `Response`, and `shutdown` drains the queue. Sessions, continuous batching and pipelines build on this same executor; the `runtime` [example project](/examples.md) walks each one. ## One flag for the serving runtime The serving runtime is built without RTTI, so a target that uses `runtime::` adds one line to part 1's CMake (the bundle's own runtime examples set the same flag): ```cmake target_compile_options(hello PRIVATE -fno-rtti) ``` Next: [part 6](06-deploy-to-mobile.mdx), the same tutorial program on a real phone. --- # Tensors and operators Factories and dtypes, the ops:: library, views, operator sugar, and reading values back to the host. Source: https://docs.clika.io/clikart/getting-started/first-program/tensors-and-operators.md Part 1 made one tensor; this part covers the compute vocabulary you will use everywhere: building tensors, transforming them with `ops::`, and reading values back. Same project as [part 1](01-your-first-program.mdx); only `main.cpp` changes. ## The program ```cpp title="main.cpp" #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; int main() { // Factories build tensors from a shape and a dtype. const Tensor threes = Tensor::full({2, 3}, 3.0, DataType::Float32); // from_data copies host bytes into a tensor of the given shape + dtype. const float host[6] = {0, 1, 2, 3, 4, 5}; const Tensor x = Tensor::from_data(host, {2, 3}, DataType::Float32); // ops:: free functions return their result directly. Elementwise math // broadcasts, and a scalar binds wherever a tensor does. const Tensor y = ops::add(ops::mul(x, 2.0), threes); // y = 2x + 3 // Shape ops are views: the same bytes behind a new layout, no copy. const Tensor yt = ops::permute(y, {1, 0}); // 2x3 -> 3x2 const Tensor flat = ops::reshape(y, {6}); // Tensor carries operator and method sugar over the same ops, so a // chain reads like the math it computes. const Tensor m = (y - 3.0).abs().max(); // Reading back: to_string() for a summary, item() for the single // element of a one-element tensor, item_as_vec() for a 0-D/1-D tensor. std::printf("y = %s\n", y.to_string().c_str()); std::printf("y^T = %s\n", yt.to_string().c_str()); std::printf("max|y - 3| = %.0f\n", m.item()); const std::vector v = flat.item_as_vec(); std::printf("flat = ["); for (std::size_t i = 0; i < v.size(); ++i) std::printf("%s%.0f", i ? ", " : "", v[i]); std::printf("]\n"); return 0; } ``` ```python title="main.py" import clika_runtime as crt def main() -> None: # Factories build tensors directly; dtypes are attributes on the package. threes = crt.full((2, 3), 3.0) x = crt.arange(6, dtype=crt.float32).reshape(2, 3) # Elementwise math reads as operators; scalars broadcast. y = 2 * x + threes # y = 2x + 3 # Shape methods are views: the same bytes behind a new layout, no copy. yt = y.permute(1, 0) # 2x3 -> 3x2 flat = y.reshape(-1) # Method chains: |y - 3| reduced to its global maximum. amax is the # global reduce; max(dim) is the dim-wise form returning values and # indices. m = (y - 3).abs().amax() # Reading back: repr(t) is the summary, item() reads a scalar, and # numpy() on a CPU tensor is a zero-copy view when an array is wanted. print(f"y = {y}") print(f"y^T = {yt}") print(f"max|y - 3| = {m.item():.0f}") v = flat.numpy() print("flat = [" + ", ".join(f"{e:.0f}" for e in v) + "]") if __name__ == "__main__": main() ``` ```kotlin title="Main.kt" import io.clika.runtime.ClikaRt import io.clika.runtime.Ops import io.clika.runtime.Tensors import java.util.Locale fun main() { ClikaRt.load() // Factories build tensors from a shape or from host values; float32 is // the default dtype, and Tensors.of copies the values in. val threes = Tensors.full(longArrayOf(2, 3), 3.0) val x = Tensors.of(floatArrayOf(0f, 1f, 2f, 3f, 4f, 5f), longArrayOf(2, 3)) // Ops carries one function per operator. Elementwise math broadcasts, // and a number binds wherever a tensor does. val scaled = Ops.mul(x, 2.0) val y = Ops.add(scaled, threes) // y = 2x + 3 // Shape operators are views: the same bytes behind a new layout, no copy. val yt = Ops.permute(y, longArrayOf(1, 0)) // 2x3 -> 3x2 val flat = Ops.reshape(y, longArrayOf(6)) // A chain of operators reads like the math it computes: max|y - 3| = 10. val d = Ops.sub(y, 3.0) val ad = Ops.abs(d) val m = Ops.amax(ad) // Reading back: summary() renders the shape, dtype, device and values; // item() reads the one element of a one-element tensor; toFloatArray() // copies a tensor's values out. println("y = ${y.summary()}") println("y^T = ${yt.summary()}") println("max|y - 3| = ${"%.0f".format(Locale.ROOT, m.item())}") println("flat = [${flat.toFloatArray().joinToString(", ") { "%.0f".format(Locale.ROOT, it) }}]") listOf(x, threes, scaled, y, yt, flat, d, ad, m).forEach { it.release() } } ``` ```text y = Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[3, 5, 7, 9, 11, 13]) y^T = Tensor(shape=[3, 2], dtype=Float32, device=CPU, numel=6, data=[3, 9, 5, 11, 7, 13]) max|y - 3| = 10 flat = [3, 5, 7, 9, 11, 13] ``` ## Dtypes `DataType` names the element format. The everyday set is `Float32`, `Float16`, `BFloat16`, `Float64`, the signed and unsigned integer widths (`Int8` ... `Int64`, `UInt8` ... `UInt64`) and `Bool`; beyond it are the sub-byte integers (`Int4`, `Int2`) and the narrow float families (FP8, FP6, FP4) that quantized models use. `ClikaRT::data_type_name(t.dtype())` prints one; `t.to(DataType::Float16)` casts. Factories take the dtype explicitly. Nothing defaults behind your back. ## Copies are handles A `Tensor` copy is a cheap reference to the same underlying data, not a deep copy. Writes through one copy are visible through the others, and the data stays alive as long as any copy does. For independent data, build a fresh tensor (a factory or `from_data`, which copies the source bytes and does not retain the pointer). ## The `ops::` library Every operator is a free function in `ClikaRT::ops`, taking tensors and returning a tensor: elementwise math, reductions, matrix products, convolutions, attention, indexing. This is the operator set a model needs. Shape ops (`reshape`, `permute`, `narrow`) return **views**, metadata over the source's bytes, no copy. For the common ones, `Tensor` adds sugar. Arithmetic operators (`y - 3.0`) and chainable methods (`.abs()`, `.relu()`, `.max()`, `.matmul(...)`) forward to the same `ops::` functions with the same error contract as part 1. The [API reference](/api/index.md) documents every operator. ## Reading values back Three host-side reads, in increasing weight. `to_string()` is an infallible summary (shape, dtype, device, first values) for logging. `item()` reads the single element of a one-element tensor, typically a reduction result. `item_as_vec()` reads all elements of a 0-D or 1-D tensor, contiguous and dtype-matched. A wrong `T` or a wrong element count raises `ClikaRT::Error`. Every one of these reads is also a synchronization point; [part 3](03-devices-and-async.mdx) explains what that means. Next: [part 3](03-devices-and-async.mdx), devices, and the asynchrony you have been using without noticing. --- # Your first pipeline Load weights and input from disk with ClikaRT::io, compute on the best device, read the results back. Source: https://docs.clika.io/clikart/getting-started/first-program/your-first-pipeline.md Part 4 puts the pieces together into the shape every real ClikaRT program has: **load weights and input from disk, place them on a device, compute, read results back**. It is self-contained. Stage 1 writes the files it needs, standing in for a real export. ## The program ```cpp title="main.cpp" #include #include #include #include #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::NamedTensors; using ClikaRT::Tensor; namespace io = ClikaRT::io; namespace ops = ClikaRT::ops; namespace { // The layer's geometry and the checkpoint's names, each spelled once. constexpr std::int64_t kFeatures = 8; // columns of x, columns of W constexpr std::int64_t kOutputs = 4; // rows of W, the activations out constexpr std::int64_t kBatch = 2; constexpr const char* kWeight = "mlp.weight"; constexpr const char* kBias = "mlp.bias"; Device pick_device() { // from part 3 namespace device = ClikaRT::device; if (device::is_cuda_available()) return Device::cuda(); if (device::is_vulkan_available()) return Device::vulkan(); if (device::is_metal_available()) return Device::metal(); return Device::cpu(); } } // namespace int main() { const std::filesystem::path dir = std::filesystem::temp_directory_path(); const std::string ckpt = (dir / "first_program.safetensors").string(); const std::string input = (dir / "first_program_input.npy").string(); // ── Stage 1: write the artifacts (a stand-in for a real export) ── // A checkpoint is a NamedTensors: a name -> tensor map. NamedTensors weights; weights.set(kWeight, Tensor::full({kOutputs, kFeatures}, 0.5, DataType::Float32)); weights.set(kBias, Tensor::full({kOutputs}, 0.25, DataType::Float32)); io::save_safetensors(weights, ckpt); io::save_npy(Tensor::ones({kBatch, kFeatures}, DataType::Float32), input); // two rows of ones // ── Stage 2: load, compute, read back ── // Loaders take the target device: the tensors land there directly. const Device dev = pick_device(); const NamedTensors loaded = io::load_safetensors(ckpt, dev); const Tensor x = io::load_npy(input, dev); // One MLP layer: y = relu(x * W^T + b), [2,8] x [4,8]^T -> [2,4]. const Tensor y = ops::relu(ops::linear(x, loaded.get(kWeight), loaded.get(kBias))); // Per-row mean, then back to the host (the reads from part 3). const std::vector pooled = ops::mean(y, {1}).item_as_vec(); std::printf("y = %s\n", y.to_string().c_str()); std::printf("row means = [%.2f, %.2f] (expected 8*1*0.5 + 0.25 = 4.25)\n", pooled[0], pooled[1]); return 0; } ``` ```python title="main.py" import tempfile import clika_runtime as crt import clika_runtime.nn.functional as F def main() -> None: tmp = tempfile.gettempdir() ckpt = f"{tmp}/first_program.safetensors" inp = f"{tmp}/first_program_input.npy" # -- Stage 1: write the artifacts (a stand-in for a real export) -- # A checkpoint is a plain dict: a name -> tensor map. weights = { "mlp.weight": crt.full((4, 8), 0.5), "mlp.bias": crt.full((4,), 0.25), } crt.io.save_safetensors(weights, ckpt) crt.io.save_npy(crt.ones(2, 8), inp) # two rows of ones # -- Stage 2: load, compute, read back -- # Loaders take the target device: the tensors land there directly. dev = crt.Device.gpu() # the part 3 probe loaded = crt.io.load_safetensors(ckpt, device=dev) x = crt.io.load_npy(inp, device=dev) # One MLP layer: y = relu(x * W^T + b), [2,8] x [4,8]^T -> [2,4]. # nn.functional (imported as F by convention) carries the stateless # neural-network operations; the same functions the nn modules call. y = F.relu(F.linear(x, loaded["mlp.weight"], loaded["mlp.bias"])) # Per-row mean, then back to the host (the reads from part 3). pooled = y.mean([1]).to("cpu").numpy() print(f"y = {y}") print(f"row means = [{pooled[0]:.2f}, {pooled[1]:.2f}] " "(expected 8*1*0.5 + 0.25 = 4.25)") if __name__ == "__main__": main() ``` The Kotlin binding is an INFERENCE-level binding ([the two kinds](../../bindings.md)): loading a ready model is its job, and the save-side members of this part are FULL-level work, done from C++ or Python. The compute pieces of this part are typed today (`F.linear`, `F.relu`, `F.mean`, `Tensor.to(device)`); the C++ arm is the pipeline story. On a machine with an NVIDIA GPU (the device line follows what `pick_device()` found): ```text y = Tensor(shape=[2, 4], dtype=Float32, device=CUDA:0, numel=8, data=[4.25, 4.25, 4.25, 4.25, 4.25, 4.25, ...]) row means = [4.25, 4.25] (expected 8*1*0.5 + 0.25 = 4.25) ``` Every element of `y` is `8 * 1 * 0.5 + 0.25 = 4.25`, so both row means print `4.25`, on whatever device the machine offered. ## Model checkpoints are `NamedTensors` A model's weights travel as a `NamedTensors`, a `name -> tensor` map with `set` / `get` / `size` / `for_each`. `io::save_safetensors` / `io::load_safetensors` round-trip it; `io::load_gguf` returns the same map plus the file's metadata; `io::load_npy` reads a single array. The loaders take the **target device**, so weights stream to where they will be used with no separate `.to()` step. That is the placement lesson from part 3, folded into I/O. ## The pipeline shape Stage 2 is the skeleton to keep: **discover -> load onto the device -> compute -> read back**. Scaling it up changes the sizes, not the shape (more names in the checkpoint, a deeper stack of `ops::` calls between load and read). The pieces this series did not need live in the [examples](/examples.md). The `runtime` project wraps this exact shape in sessions, continuous batching and pipelines, and `io`'s later chapters cover GGUF and images. You have a complete, device-portable ClikaRT program. Next: [part 5](05-serve-it.mdx), the same model behind the serving runtime. --- # Your first program Link the bundle, print the version, make your first tensor, and meet the error model. Source: https://docs.clika.io/clikart/getting-started/first-program/your-first-program.md This tutorial has six parts: four build your first ClikaRT program, one serves it, and one deploys it to mobile. Each part is a complete program that compiles and runs on any machine the runtime supports (no accelerator is needed). Core concepts are explained where they first appear. This part links the bundle, prints the version, and makes one tensor. The C++ and Kotlin paths assume the bundle is extracted and `CLIKART_BUNDLE_DIR` points at it (see [Quick install](../installation.md)); the Python wheel carries the runtime itself. Every program here runs compute, so every one of them needs a license credential. `CLIKA_RT_LICENSE` in the environment carries it to the C++ and Kotlin programs and to Python, and `clikart-license-init` stores it once per user account instead; [License the runtime](../../how-to/license-the-runtime.mdx) is the contract. Python has one extra rule, in its tab below. ## The program ```cpp title="main.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; int main() { std::printf("ClikaRT %s\n", ClikaRT::GetVersionInfo().c_str()); // A 2x3 tensor of ones on the CPU (the default device). const Tensor a = Tensor::ones({2, 3}, DataType::Float32); // ops:: are free functions returning the value directly. const Tensor b = ops::add(a, a); std::printf("a + a =\n%s\n", b.to_string().c_str()); try { const Tensor bad = ops::matmul(a, Tensor::ones({5, 7}, DataType::Float32)); } catch (const ClikaRT::Error& e) { std::printf("failed: %s\n", e.what()); // e.status() has the coarse category } return 0; } ``` ```python title="main.py" import clika_runtime as crt def main() -> None: print(f"ClikaRT {crt.version()}") # A 2x3 tensor of ones on the CPU (the default device). Factories take # the shape directly; float32 is the default dtype. a = crt.ones(2, 3) # Operators read as math; every result is a new tensor. b = a + a print(f"a + a =\n{b}") try: bad = a @ crt.ones(5, 7) except RuntimeError as e: print(f"failed: {e}") # the message ends in the machine-readable [code: NAME] if __name__ == "__main__": main() ``` ```kotlin title="Main.kt" import io.clika.runtime.ClikaRt import io.clika.runtime.ClikaRtException import io.clika.runtime.Ops import io.clika.runtime.Tensors fun main() { ClikaRt.load() println("ClikaRT ${ClikaRt.version()}") // A 2x3 tensor of ones on the CPU (the default placement); float32 is // the default dtype. val a = Tensors.ones(longArrayOf(2, 3)) // Every operator is one function of Ops. It returns a fresh tensor and // leaves its operands live; a tensor is AutoCloseable, so release the // ones you named, or scope them with `use`. val b = Ops.add(a, a) println("a + a =\n${b.summary()}") try { Tensors.ones(longArrayOf(5, 7)).use { bad -> Ops.matmul(a, bad) } } catch (e: ClikaRtException) { println("failed: ${e.message}") // e.codeName is the machine channel, e.status the category } b.release() a.release() } ``` ## Set up, build, run The setup is where the languages differ. The CMake side is three lines of substance (the package and one target). ```cmake title="CMakeLists.txt" cmake_minimum_required(VERSION 3.19) project(hello_clikart LANGUAGES CXX) set(CMAKE_CXX_STANDARD 17) set(CMAKE_CXX_STANDARD_REQUIRED ON) find_package(ClikaRT CONFIG REQUIRED) add_executable(hello main.cpp) target_link_libraries(hello PRIVATE ClikaRT::ClikaRT) ``` ```bash cmake -S . -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build ./build/hello ``` The wheel carries the runtime and its backends; the bundle is not needed on this path. [Get ClikaRT](../get-clikart.mdx#python-wheels) has the wheel to download (on Linux, the flavor follows your NVIDIA driver) and the install command. `numpy` is the array bridge for building tensors from Python data (optional by design, and these pages use it). ```bash pip install ./clika_runtime---.whl pip install numpy export CLIKA_RT_LICENSE=CLIKA1-... python main.py ``` `import clika_runtime` loads the runtime, so `CLIKA_RT_LICENSE` has to be set before the import runs. A shell `export` does that. From inside the program, assign `os.environ["CLIKA_RT_LICENSE"]` above the import. The wheel also installs `clikart-license-init` as a console script, which stores the credential once per user account and drops the variable entirely. {/* CERTIFICATION: the Kotlin arm follows the release's Kotlin artifact (clika-runtime-maven-.zip) and the tutorial's gradle project beside the programs, both of the pinned release. */} The Kotlin artifact, `clika-runtime-maven-0.6.4.zip`, extracts to a Maven repository directory that the gradle project names through the `clikaRtMavenRepo` property; it ships the Kotlin API and its JNI bridge. `libClikaRT.so` comes from the bundle, so run the JVM with `-Djava.library.path` pointing at the bundle's `lib` directory. ```kotlin title="build.gradle.kts" // examples/kotlin/tutorial: the Kotlin programs of the getting-started // tutorial, one file per chapter (NN_.kt), built for a desktop JVM // against the release's Kotlin artifact, io.clika:clika-runtime (the // extracted clika-runtime-maven-.zip, named to settings.gradle.kts // by -PclikaRtMavenRepo=). At run time ClikaRt loads two libraries: // libClikaRT.so by name from java.library.path, a runtime distribution's // lib/, and the JNI bridge libclika_rt_jni.so from the artifact's own jar // (its native// entry, extracted at start): // // gradle installDist -PclikaRtMavenRepo= // java -Djava.library.path=/lib -cp 'build/install/tutorial/lib/*' _01_your_first_programKt // // or in one step (CLIKA_RT_LIB_DIR in the environment names /lib too): // // gradle run -Pchapter=01_your_first_program -PclikaRtMavenRepo= -PclikaRtLibDir=/lib // // 06_deploy_to_mobile.kt is the Android chapter's Kotlin half: it compiles // here and has no main to run. plugins { kotlin("jvm") version "2.3.0" application } java { toolchain { languageVersion.set(JavaLanguageVersion.of(17)) } } // One compilation: the chapter files beside this build script. Each chapter's // top-level main lands in its own class, named after the file // (01_your_first_program.kt -> _01_your_first_programKt). sourceSets { main { kotlin.setSrcDirs(listOf(layout.projectDirectory)) kotlin.include("??_*.kt") } } dependencies { // The runtime's Kotlin API, the release's artifact (its version read from the repository directory). implementation("io.clika:clika-runtime:${project.extra["clikaRtVersion"]}") } val chapter = providers.gradleProperty("chapter").orElse("01_your_first_program") val clikaRtLibDir = providers.gradleProperty("clikaRtLibDir") .orElse(providers.environmentVariable("CLIKA_RT_LIB_DIR")) .orElse("") application { mainClass.set(chapter.map { "_${it}Kt" }) } tasks.named("run") { workingDir = projectDir doFirst { if (clikaRtLibDir.get().isEmpty()) { throw GradleException("name the runtime distribution's lib/: -PclikaRtLibDir=/lib or CLIKA_RT_LIB_DIR") } } systemProperty("java.library.path", clikaRtLibDir.get()) } ``` ```bash gradle run -Pchapter=01_your_first_program -PclikaRtMavenRepo="$PWD/clika-runtime-maven-0.6.4" -PclikaRtLibDir="$CLIKART_BUNDLE_DIR/lib" ``` ```text ClikaRT 0.6.4 a + a = Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[2, 2, 2, 2, 2, 2]) failed: ops::matmul: cannot contract a [2, 3] with b [5, 7]: a's last dim (3) must equal b's dim 0 (5) ``` The version line proves the runtime loaded, and the tensor prints its shape, dtype, device and values, every element `2`. The last line is the failure path, demonstrated at the end of the program and explained below. ## One library, one header The umbrella header `ClikaRT/clika_rt.h` includes the entire public API (tensors, operators, devices, I/O, the serving runtime). `ClikaRT::ClikaRT` is the only link target. Nothing else from the bundle enters your build; the accelerator backends are shared libraries the runtime loads on its own at run time ([part 3](03-devices-and-async.mdx)). ## Values out, `ClikaRT::Error` on failure Every fallible call in the public API (factories like `Tensor::ones`, operators like `ops::add`) returns its result directly and raises `ClikaRT::Error` on failure. There are no output parameters and no error codes to check at each call; a shape mismatch or an unavailable device surfaces as an exception where the call was made: ```cpp try { const Tensor bad = ops::matmul(a, Tensor::ones({5, 7}, DataType::Float32)); } catch (const ClikaRT::Error& e) { std::printf("failed: %s\n", e.what()); // e.status() has the coarse category } ``` ```python try: bad = a @ crt.ones(5, 7) except RuntimeError as e: print(f"failed: {e}") # the message ends in the machine-readable [code: NAME] ``` ```kotlin try { Tensors.ones(longArrayOf(5, 7)).use { bad -> Ops.matmul(a, bad) } } catch (e: ClikaRtException) { println("failed: ${e.message}") // e.codeName is the machine channel, e.status the category } ``` Catch it where you can act on it. The programs in this series let a failure terminate the process, the right default for a small batch program. The message names the operation and the values it refused (here the two shapes); the stable, machine-readable channel is the code name, which [Handle errors by code](../../how-to/handle-errors-by-code.mdx) covers. Next: [part 2](02-tensors-and-operators.mdx), tensors and the `ops::` library properly. --- # Get ClikaRT Pick your OS, architecture and accelerator, and get the download, verify and extract commands for the archive the CLIKA platform hands out. Source: https://docs.clika.io/clikart/getting-started/get-clikart.md import GetClikaRT from '@site/src/components/GetClikaRT'; ClikaRT is a set of downloads on the CLIKA platform: one release archive per platform, the Python wheels, the Kotlin artifact and the examples archive, every file with its SHA-256 shown beside it. Take every one you use from the same release. The release the dialog offers is the deployment's own pin, which may differ from the release these pages are written for (`0.6.4`); a download's own `README.md` is the authority for the release it came from, its commands and file names included. Your platform deployment's **Download ClikaRT** dialog (your name in the top bar, or a project's Overview), the platform CLI (`clika-cli runtime-sdk download`) and the MCP tools all hand out the same files; [Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) is the walk through each surface. Pick the platform you are deploying to and copy the commands. Each platform has its own archive, named `ClikaRT__-.tar.xz` (`.zip` on Windows); the tree inside is that platform's complete bundle, the runtime and the Modelverse model library together. The dialog shows the archive's SHA-256 as text with a Copy control; paste it in place of `` below. Then continue with [Quick install](installation.md), which picks up at the extracted bundle. [Complete installation](complete-installation.md) has per-platform depth. ## License credential ClikaRT runs compute under a license, so every process that runs an operator or a model needs the credential CLIKA issued for your deployment. The credential is the `CLIKA1-...` text your project's license shows as its license key, the same text for an online license key and an offline license bundle; an API key (`clika_rk_...`) is not a credential. [Runtime licenses](/platform/concepts/runtime-licenses) is where a credential comes from, and the download dialog shows your project's key beside the archives. Put it in the environment variable `CLIKA_RT_LICENSE`, as the credential text or as the path of a file holding it, before the program starts: ```bash export CLIKA_RT_LICENSE=CLIKA1-... # the credential text export CLIKA_RT_LICENSE=/etc/clika/clikart.license # or the path of a file holding it ``` Or store it once per user account with `clikart-license-init`, which ships in every archive's `bin/` (`clikart-license-init.exe` on Windows) and as a console script of the `clika-runtime` wheel; every later process under that account reads the file when the variable is not set, and the variable wins when both are present: ```bash "$CLIKART_BUNDLE_DIR"/bin/clikart-license-init CLIKA1-... ``` Without a valid credential a call is refused with the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a license past its end date. [License the runtime](../how-to/license-the-runtime.mdx) is the whole contract: where each language and packaging puts the credential, and how a program reads a refusal's code name. ## All combinations The pattern is the same for every platform; only the archive name and the checksum command change. The dialog shows the SHA-256 beside the download button, with the one line that checks the file on your platform: ```bash echo " ClikaRT_linux_x86_64-0.6.4.tar.xz" | sha256sum -c tar -xf ClikaRT_linux_x86_64-0.6.4.tar.xz export CLIKART_BUNDLE_DIR="$PWD/ClikaRT_linux_x86_64-0.6.4" ``` On macOS the check is `shasum -a 256 -c` over the same line. On Windows, in PowerShell: `(Get-FileHash ).Hash -eq ""`, then `tar -xf ` and `$env:CLIKART_BUNDLE_DIR = "$PWD\"`. The platform CLI runs the check itself and refuses a file that does not match. Per platform, the archive and the driver requirements: | OS | Architecture | Archive | Accelerators and drivers | | --- | --- | --- | --- | | Linux | x86_64 | `ClikaRT_linux_x86_64-.tar.xz` | CPU (no driver) · CUDA (NVIDIA driver only) · Vulkan (Vulkan 1.2 or newer) | | Linux | arm64 | `ClikaRT_linux_arm64-.tar.xz` | CPU · CUDA · Vulkan, as above | | Windows | x86_64 | `ClikaRT_windows_x86_64-.zip` | CPU · Vulkan | | Windows | arm64 | `ClikaRT_windows_arm64-.zip` | CPU · Vulkan | | macOS | Apple silicon | `ClikaRT_macos_arm64-.tar.xz` | CPU · Metal (ships with macOS) | | Android | arm64-v8a | `ClikaRT_android_arm64-.tar.xz`, then [Deploy to mobile](first-program/06-deploy-to-mobile.mdx) | CPU · Vulkan | The two Linux archives carry both CUDA images, `lib/libClikaRT_cuda13x.so` and `lib/libClikaRT_cuda12x.so`, and each image is self-contained: the CUDA runtime and cuBLASLt are inside it. A CUDA machine needs its NVIDIA driver and nothing else, no CUDA toolkit install and no `LD_LIBRARY_PATH`; the runtime loads the image the driver serves (a driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image). Every desktop archive verifies the same way after extraction: run `examples/bin/version_00_hello_world` (`.exe` on Windows) from the extracted root; it prints the ClikaRT version. ## Python wheels The `clika-runtime` wheel is the dialog's Python build: pick the platform, the CUDA image (Linux) and the CPython version (3.10 to 3.14), and the download is a zip holding the wheel for that selection, with its SHA-256 shown like the archives'. The wheel carries the runtime and its backends inside the package, and the [Modelverse](/modelverse) model library as `clika_runtime.modelverse`, so the Python path needs no archive and no library paths to set. On Linux (x86_64 and aarch64) the wheel comes in three flavors, told apart by the local version tag; the NVIDIA driver picks the flavor (`nvidia-smi` prints the driver version in its header): | Version | Contents | Pick it when | | --- | --- | --- | | `0.6.4` | every non-CUDA backend, no CUDA image | the machine has no NVIDIA GPU | | `0.6.4+cu13x` | the non-CUDA backends plus the CUDA 13 image | the NVIDIA driver is major version 580 or newer | | `0.6.4+cu12x` | the non-CUDA backends plus the CUDA 12 image | the NVIDIA driver is older than 580 | A CUDA flavor's image is self-contained, as in the archives: the driver is the one requirement, no CUDA toolkit and no `LD_LIBRARY_PATH`. macOS (Apple silicon) and Windows x64 carry no CUDA and ship one unsuffixed wheel each for CPython 3.10 to 3.14; Windows arm64 ships one for CPython 3.11 to 3.14. No wheel declares a requirement on an NVIDIA package. Check the downloaded zip against the SHA-256 the dialog shows, unzip it, and install the wheel inside from the file; `numpy` is the array bridge the tutorial uses, installed separately. ```bash echo " " | sha256sum -c unzip pip install ./clika_runtime-0.6.4+cu13x-cp313-cp313-manylinux_2_28_x86_64.whl pip install numpy python -c "import clika_runtime as crt; print(crt.version())" ``` The last line prints `0.6.4`. The platform CLI's `clika-cli runtime-sdk pip-command` prints one `pip install` line against the deployment's own package index instead, which picks the wheel for the interpreter it runs under. [Use ClikaRT from Python](../how-to/use-clikart-from-python.mdx) is the guide to the wheel's surface. Install these wheels with `pip`. Every wheel is an LZMA-compressed zip, the one compression method that fits the `+cu12x` flavor's CUDA image under the release asset size limit. `pip` reads LZMA on any CPython 3.10 or newer whose build carries the `lzma` module, which the python.org, manylinux, conda and distribution builds do. `uv pip install` does not read LZMA-compressed wheels. ## Modelverse from Python The same wheel carries [Modelverse](/modelverse), the model library, as the `clika_runtime.modelverse` subpackage: model resolution, generation, pipelines and the serving API. There is no second wheel. ```bash python -c "import clika_runtime.modelverse as mv; print(mv.__version__)" ``` The line prints `0.6.4`. Importing the subpackage loads the runtime first, then the Modelverse extension, which finds the runtime library the same wheel installed; a runtime and a model library from different releases refuse to import, naming both versions. The `clikart-cli` command-line program is a console script of the same wheel. ## The Kotlin artifact and the examples Three more downloads sit beside the archives and the wheels in the dialog's listing, taken from the same release: - `clika-runtime-maven-.zip`, the Kotlin artifact: a Maven repository directory holding `io.clika:clika-runtime` as an Android AAR and a desktop JVM jar under one coordinate. Extract it anywhere and name the directory to gradle; on Android the AAR is the whole runtime, and on the desktop JVM the jar runs beside the release archive of the same platform (its `lib/` on `java.library.path`); [Language bindings](../bindings.md) says what it carries and [Package ClikaRT in an Android app](../how-to/package-clikart-in-an-android-app.mdx) how an app ships it. - `release_manifest.json`, the release manifest: one entry per published file with the platform, build type and language it serves, the backends inside it, the files it needs beside it, its digest and size, and the pieces of a split archive in join order; a download page or a script reads the release from it instead of from file names. Its `.sha256` sits beside it like every other file. - `ClikaRT--examples.tar.xz`, the examples archive: complete programs in C++, Python and Kotlin with their recorded output, under one `examples/` directory whose `README.md` and `AGENTS.md` index them. [Additional examples](../examples.md) is the catalog. [System requirements](/system-requirements.md) is the authority on platforms, drivers and hardware. --- # Quick install Install a C++ toolchain, extract the ClikaRT bundle, and prove it works in two commands. Source: https://docs.clika.io/clikart/getting-started/installation.md The fast path from nothing to a running ClikaRT program. For per-platform depth, IDE setup and the full failure catalog, see [Complete installation](complete-installation.md). ## 0. Prerequisites A C++17 compiler and CMake 3.19 or newer: ```bash sudo apt install build-essential cmake # Debian/Ubuntu brew install cmake # macOS; compiler: xcode-select --install winget install Kitware.CMake # Windows; compiler: Visual Studio Build Tools ``` A runtime credential, issued by your platform for the project the program belongs to; step 2 below puts it in place. ## 1. Get the bundle Extract the archive for your platform ([Get ClikaRT](get-clikart.mdx) has the download and checksum commands) and remember where it is: ```bash tar -xf ClikaRT_linux_x86_64-.tar.xz export CLIKART_BUNDLE_DIR="$PWD/ClikaRT_linux_x86_64-" ls "$CLIKART_BUNDLE_DIR" ``` The `ls` is the success check: `bin/`, `cmake/`, `include/`, `lib/` and `examples/` are among the entries. A command-line-only download from the platform's dialog carries `bin/` and `lib/` and no `cmake/`, `include/` or `examples/`; those come with the C/C++ and Everything downloads, and a composed archive's own `README.md` says exactly what it holds. A new terminal loses the `export`; re-run it there, every command below uses it. On Windows, run these in PowerShell (`tar` is built in since Windows 10). ## 2. Place the license credential ClikaRT runs compute under a license. Put the `CLIKA1-...` credential your project's license shows in the environment, or store it once under your user account: ```bash export CLIKA_RT_LICENSE=CLIKA1-... # this shell, and the programs it starts "$CLIKART_BUNDLE_DIR"/bin/clikart-license-init CLIKA1-... # or once, for this user account ``` [License the runtime](../how-to/license-the-runtime.mdx) has the whole contract: what the credential is, where each language and packaging reads it, and what a refused call carries. ## 3. Run a pre-built example The fastest proof the bundle works. Desktop archives ship pre-built example binaries; run one in place. Nothing gets installed, and no GPU or driver is needed: ```bash "$CLIKART_BUNDLE_DIR"/examples/bin/version_00_hello_world ``` ```text ClikaRT 0.6.4 ``` On a platform without pre-built binaries, skip to step 4. On a Jetson, extract the `linux_arm64` archive; its binaries run out of the box there, as on any other arm64 Linux machine. ## 4. Build the version example yourself The same proof through your own toolchain. The three flags, once: `-S` is the source directory, `-B` is the build directory, and `ClikaRT_DIR` tells CMake where the bundle's `cmake/` folder is. ```bash cmake -S "$CLIKART_BUNDLE_DIR/examples/src/version" -B build-version \ -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build-version ./build-version/version_00_hello_world ``` ```text ClikaRT 0.6.4 ``` ## If something failed - `cmake: command not found`: step 0. - `tar: xz: Cannot exec` or `xz: command not found`: install `xz-utils` (Debian/Ubuntu) or `xz`. - `Could not find a package configuration file provided by "ClikaRT"`: `ClikaRT_DIR` is wrong, or the `export` was lost to a new terminal. An error naming a bare `/cmake` path means the variable expanded empty; re-run the `export` from step 1. - The binary is missing after an MSVC build: it lands in a configuration subdirectory, `build-version\Debug\version_00_hello_world.exe`. The install works. [Write your first program](first-program/01-your-first-program.mdx). --- # What to read next Where the documentation goes after the tutorial: how-to guides, the API reference, Modelverse, system requirements. Source: https://docs.clika.io/clikart/getting-started/next-steps.md You installed the bundle and wrote a device-portable program. The rest of the documentation, in a useful reading order: - **[How-to guides](/how-to/index.md)**: problem-oriented recipes past the tutorial, from quantized weights to HTTP serving. [Additional examples](/examples.md) sit inside it: standalone programs in the bundle, one per subsystem; `compute` and `async` deepen what the tutorial introduced, and `io`, `runtime`, `tokenizer` and `http_server` carry it to a served model. - **[API reference](/api/index.md)**: every public namespace, class and function, generated from the headers you compile against. `ClikaRT/clika_rt.h` includes them all. - **[Modelverse](/modelverse)**: the CLIKA model library. Models packaged so ClikaRT loads and runs them as they are, with quantized variants per class of device. - **[System requirements](/system-requirements.md)**: the platform, toolchain, driver and accelerator tables, for when you leave your development machine. --- # ClikaRT at a glance What the ClikaRT runtime is and is not, and the five-minute mental model. Source: https://docs.clika.io/clikart/getting-started/overview.md ClikaRT is CLIKA's inference runtime, a C++ library (with Python and Kotlin bindings) that loads models and runs them on CPUs, GPUs, and other hardware accelerators through one public API. It is not a training framework, and not a bundle of separate CUDA, Vulkan and Metal wrappers. You link one library, include one header, and the same code runs on every backend the runtime ships. ## The mental model Five ideas carry the whole library, in the order you meet them when you build. 1. **Everything arrives in one bundle.** A directory with the public headers, the libraries per platform, a CMake package and the examples. `find_package(ClikaRT CONFIG)` and the target `ClikaRT::ClikaRT` are the entire integration. 2. **Models load from files you already have.** `ClikaRT::io` reads safetensors, GGUF and NumPy checkpoints, plus images and audio, and a model checkpoint arrives as a `name -> tensor` map on the device you name. ```cpp const NamedTensors weights = io::load_safetensors("model.safetensors"); const Tensor w = weights.get("layer.weight"); ``` ```python weights = crt.io.load_safetensors("model.safetensors") w = weights["layer.weight"] ``` The Kotlin binding carries no checkpoint reader: a model reaches an app through the model library (`AutoModel.fromPretrained`), or as an ONNX model or a compiled graph; tensors and operators carry the `Ops` surface, one function per operator. 3. **`Tensor` is the core building block.** Each tensor is assigned to a device (the CPU by default; `.to(device)` moves it), and `ops::` operators run where their inputs live. The CPU backend is always present; CUDA, Vulkan and Metal load at run time where the machine has them. Everything returns values directly and raises `ClikaRT::Error` on failure. ```cpp const Tensor a = Tensor::ones({2, 3}, DataType::Float32); const Tensor b = ops::add(a, a); std::printf("%s\n", b.to_string().c_str()); ``` ```python a = crt.ones(2, 3) b = a + a print(b) ``` ```kotlin val a = Tensors.ones(longArrayOf(2, 3)) val b = a + a // operator extensions return a fresh tensor; operands stay live println(b.summary()) ``` 4. **Execution is asynchronous by nature.** An `ops::` call dispatches work and returns; reads wait for the result, so you never observe unfinished bytes. 5. **Serving is built in.** The serving runtime adds sessions, continuous batching and pipelines, so the model you loaded answers requests. Declare a schema, serve it with a lambda, drive it through an executor: ```cpp rt::FunctionModel model{schema}; model.on_run_once("run", [](rt::PhaseContext& ctx) { ctx.outputs->set("y", ctx.inputs->get("x") * 2.0); }); rt::Executor exec = rt::Executor::create(model); const rt::Response resp = exec.await(exec.enqueue(std::move(req))); ``` The Python lane serves a traced graph: `crt.trace` captures a forward pass as a `ModelGraph`, which runs once finalized, and `crt.runtime` carries the `Executor` and `Pipeline` over it. [Part 5](first-program/05-serve-it.mdx) walks it. The Kotlin binding carries the serving runtime (`FunctionModel`, `Executor`, `Pipeline`) and the model library's server (`Modelverse.serve`), on a phone and on a desktop JVM alike. The [first program series](first-program/01-your-first-program.mdx) turns these into working programs, explaining each where it first appears; the serving runtime has its own [example project](/examples.md). ## Platforms Linux (x86_64, arm64), Android (arm64), Windows (x86_64, arm64) and macOS (Apple silicon), one distribution per platform and architecture, and a distribution works out of the box on every machine of its class: the linux-arm64 dist runs on a Jetson the same way it runs on an arm64 server. [System requirements](/system-requirements.md) has the platform and accelerator tables. --- # How-to guides Problem-oriented recipes. Each guide answers one "how do I X" with a worked example, verified against the bundle. Source: https://docs.clika.io/clikart/how-to.md Practical guides covering common tasks. Each guide answers one concrete "how do I X" with a worked example: real code, built against the bundle, with its real output. Read the [tutorial](../getting-started/first-program/01-your-first-program.mdx) first; the guides assume its ground (tensors, devices, the async model) and go deeper on one problem at a time, in any order. Entries marked as coming are planned and land here as they are written. Every section carries a C++ and a Python arm, each a program that builds and runs against the release with its output recorded beside it, and a Kotlin note where the binding reaches the surface. Where a surface is the C++ tier's alone (the HTTP server, the processor family, a custom kernel), the tab says so and points at the tab that has it. The Kotlin binding is an INFERENCE-level binding ([the two kinds](../bindings.md)): a Kotlin tab shows what runs a ready model, and where a guide reaches past that level the tab names the member it lacks. Three guides are Python-only because their subject exists in the Python package alone: authoring a model in Python, PyTorch interop, and pytrees. The six graph guides (query, edit, optimize and finalize, write a transform, merge and split, and export to ONNX) carry a C++ and a Python arm, the two APIs that read and change a graph's structure. The licensing guide's arms show where each language puts the runtime credential, not a program with an output. ## Weights and data - [Load quantized weights from a GGUF file](load-quantized-weights.mdx): read a block-quantized checkpoint and serve it at its on-disk footprint. - [Load images and audio for inference](load-images-and-audio.mdx): decode files into tensors, resample audio, preprocess with `ops::`. ## Models and graphs - [Run an ONNX model](run-an-onnx-model.mdx): open or build a graph, compile it into a `ModelGraph`, optimize and finalize it, execute by position or by name. - [Preprocess inputs with processors](preprocess-with-processors.mdx): the model's own resize/rescale/normalize recipe, or a log-mel front end, from knobs or its config file. - [Author a model in Python](author-a-model-in-python.mdx): a decoder-only language model as `nn.Module`s from a Hugging Face config and its safetensors, meta init, fused projections, a KV cache and a greedy decode loop, measured against the C++ command line. - [Query a graph](query-a-graph.mdx): a `ModelGraph`'s nodes, values and edges, search, walks and orders, paths and the critical path, dominance, regions and the cheapest cut, and views across edits. - [Edit a graph](edit-a-graph.mdx): change a `ModelGraph` before `finalize()`, from rewiring a value's reads and changing what a node reads to renaming nodes and editing the inputs and outputs, then optimize, finalize and run it. - [Optimize and finalize](optimize-and-finalize.mdx): run the graph optimizer on a `ModelGraph` and read its report, choose its transforms and defaults, place it with `to()`, and `finalize()` it to run. - [Write a transform](write-a-transform.mdx): functions over the graph edits that `optimize()` runs beside the runtime's transforms, operators built with `add_node`, `insert` and `replace`, and rewrite rules over a `Pattern`. - [Merge and split](merge-and-split.mdx): join `ModelGraph`s side by side, as a pipeline or across two devices, cut one into parts at its cheapest cut or by a function, extract the operators between values, and merge the parts back. - [Export a graph to ONNX](export-a-graph-to-onnx.mdx): write a `ModelGraph` as an ONNX model file with `export_onnx`, or build the same model in memory with `OnnxModel::from_graph`; read it back, keep a named dynamic dimension, choose where the weights go and the opset, and meet what the export refuses. ## Text and chat - [Tokenize text and apply a chat template](tokenize-and-chat-templates.mdx): text to token ids and back, byte offsets, the model's own prompt format, batching for a model, streaming decode for a generation loop. ## Serving - [Serve a model over HTTP](serve-over-http.mdx): routes, a JSON inference endpoint, and server-sent events for streaming. ## Execution and memory - [Control asynchronous execution](control-async-execution.mdx): dispatch vs ready, safe host reads, completion callbacks, synchronous and tracing scopes. - [Trace eager code to graphs](trace-eager-code-to-graphs.mdx): capture a function as a `ModelGraph` with `trace`, optimize, finalize and run it, and wrap hot paths in `compile`. - [Wrap existing memory without copying](wrap-existing-memory.mdx): tensors over buffers your application already owns. - [Write a custom operator](write-a-custom-operator.mdx): compose built-ins in an `nn::Module`, or launch your own kernel on the stream's native handle. - [Structure inputs and outputs as pytrees](pytrees.mdx): nested containers of tensors at every boundary (`compile`, `trace`, `eval`, `save`), the registry for your own classes, key paths and serialization. ## Integration and packaging - [Add ClikaRT to an existing CMake project](existing-cmake-project.md): `find_package` against the bundle, or `add_subdirectory`, in a project that already builds. - [Coming from PyTorch or Hugging Face](coming-from-pytorch.mdx): the conventions that differ, each with the one line that bridges it: channels-last and OHWI, integer slots, reads that wait, the accessors, the tokenizer's files, and the generate calls side by side. - [Use ClikaRT from Python](use-clikart-from-python.mdx): the `clika-runtime` wheel, the NumPy boundary, models as `nn.Module`, errors you can branch on. - [Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx): `torch.compile(model, backend="clika")`, zero-copy tensor exchange over DLPack, and `from_torch_module` for an eager module tree. - [Handle errors by code](handle-errors-by-code.mdx): the three channels every failure carries, the stable code name to branch on, the coarse status for policy. - [License the runtime](license-the-runtime.mdx): where the runtime reads the project's credential, one setting per language and packaging, and the code names a refused call carries. - [Package ClikaRT in an Android app](package-clikart-in-an-android-app.mdx): the two artifacts and the libraries copied beside them, the build requirements, the license credential, the thread count and the cache root, a model on the phone, and a server that outlives the screen. ## Tools - Build a command-line model tool (coming): typed argument parsing, progress bars and structured logging in one small tool. ## Complete programs - [Additional examples](/examples.md): the bundle's standalone projects, one per subsystem, each building its topic up chapter by chapter. For every public name, the [API reference](/api/index.md). --- # Author a model in Python Write a Llama-class decoder as clika_runtime.nn modules, load a Hugging Face checkpoint as one resident copy, fuse its projections, decode through a KV cache with a one-token lookahead, and measure it against the C++ command line. Source: https://docs.clika.io/clikart/how-to/author-a-model-in-python.md A model authored in Python over `clika_runtime` runs at the speed of the runtime's operators: Python decides which operator runs next, the kernels do the work, and the loop never reads a value it does not need. This guide walks the wheel's own chapter, `examples/python/howto/author_a_model_in_python/llama/llama_from_scratch.py`, a Llama-class decoder in about six hundred lines that loads a Hugging Face checkpoint, generates greedily and matches the `clikart-cli` command line token for token. Every sample below is an excerpt of that file; the chapter's README holds the command lines and the measured table. The walk is linear: the configuration, the module tree, the load, the fusions, the cache, the attention step, the decode loop, the measurement. [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module basics this page builds on; [Tokenize text and apply a chat template](tokenize-and-chat-templates.mdx) covers the tokenizer the prompt goes through. ## The configuration comes from config.json The architecture is a dataclass whose fields carry the checkpoint's own names, so `config.json` fills it directly. `from_dict` keeps the keys the dataclass declares and drops the rest, and `__post_init__` derives what the file leaves implicit (the key/value head count, the head size). The rotary tables come from one operator, with Llama 3's frequency scaling when the checkpoint declares it. ```python title="llama_from_scratch.py (excerpt)" @dataclasses.dataclass class LlamaConfig: hidden_size: int intermediate_size: int num_hidden_layers: int num_attention_heads: int vocab_size: int num_key_value_heads: int | None = None head_dim: int | None = None rms_norm_eps: float = 1e-5 rope_theta: float = 10000.0 rope_scaling: dict | None = None tie_word_embeddings: bool = False max_position_embeddings: int = 8192 eos_token_id: int | list[int] | None = None def __post_init__(self) -> None: if self.num_key_value_heads is None: self.num_key_value_heads = self.num_attention_heads if self.head_dim is None: self.head_dim = self.hidden_size // self.num_attention_heads @classmethod def from_dict(cls, fields: dict) -> "LlamaConfig": known = inspect.signature(cls).parameters return cls(**{name: value for name, value in fields.items() if name in known}) def rotary_tables(self, max_positions: int, device: crt.Device) -> tuple[crt.Tensor, crt.Tensor]: scaling = self.rope_scaling or {} if self.rope_llama3: return crt.generate_rotary_cache( self.head_dim, max_positions, theta=self.rope_theta, scaling="llama3", scale=float(scaling["factor"]), low_freq_factor=float(scaling["low_freq_factor"]), high_freq_factor=float(scaling["high_freq_factor"]), original_max_pos=int(scaling["original_max_position_embeddings"]), device=device, ) return crt.generate_rotary_cache(self.head_dim, max_positions, theta=self.rope_theta, device=device) ``` ## A module tree with the checkpoint's names `load_state_dict` binds by dotted name, so the tree spells the checkpoint's names: `model.layers.3.self_attn.q_proj.weight` is the attribute path `model.layers[3].self_attn.q_proj` and its `weight`. Every constructor takes the model dtype and declares each layer at it (the chapter resolves it from `--dtype`, else from the checkpoint's embedding table), so a bind adopts a matching checkpoint tensor as it is and no layer sits at the default dtype beside a bfloat16 checkpoint. The attention module owns the four projections and, after the load, one fused group for the first three. ```python title="llama_from_scratch.py (excerpt)" class LlamaAttention(nn.Module): def __init__(self, config: LlamaConfig, dtype: crt.dtype) -> None: super().__init__() hidden, heads, kv_heads, head_dim = (config.hidden_size, config.num_attention_heads, config.num_key_value_heads, config.head_dim) self.num_heads, self.num_kv_heads, self.head_dim = heads, kv_heads, head_dim self.q_proj = nn.Linear(hidden, heads * head_dim, bias=False, dtype=dtype) self.k_proj = nn.Linear(hidden, kv_heads * head_dim, bias=False, dtype=dtype) self.v_proj = nn.Linear(hidden, kv_heads * head_dim, bias=False, dtype=dtype) self.o_proj = nn.Linear(heads * head_dim, hidden, bias=False, dtype=dtype) class LlamaMLP(nn.Module): def __init__(self, config: LlamaConfig, dtype: crt.dtype) -> None: super().__init__() self.gate_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False, dtype=dtype) self.up_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False, dtype=dtype) self.down_proj = nn.Linear(config.intermediate_size, config.hidden_size, bias=False, dtype=dtype) def forward(self, x: crt.Tensor) -> crt.Tensor: return self.down_proj(crt.swiglu(self.gate_up(x))) ``` `LlamaDecoderLayer` holds one attention, one MLP and the two `nn.RMSNorm` weights; `LlamaModel` holds the `nn.Embedding`, an `nn.ModuleList` of layers and the final `norm`; `LlamaForCausalLM` adds `lm_head` and keeps the model dtype as `self.dtype`. The names are the checkpoint's at every level. ## Build storage-free, then adopt the checkpoint The checkpoint comes first: `crt.load` reads the safetensors file into tensors on the device, and a tied checkpoint, which carries no `lm_head.weight`, gets the embedding table bound under the head's name too, so one load adopts one object onto both slots. The model dtype is the given `--dtype`, else the embedding table's own. The tree is then built under `crt.device("meta")`, where every parameter is a shape and a dtype with no bytes; `to_empty` gives the slots their placement, and `load_state_dict(state, strict=True, assign=True)` adopts the checkpoint tensors as the model's parameters: no copy, one resident copy of every weight. ```python title="llama_from_scratch.py (excerpt)" def build_model(config, snapshot, device, max_positions, dtype=None, state=None): if state is None: state = load_state(snapshot, device) state = dict(state) if config.tie_word_embeddings and LM_HEAD_KEY not in state: state[LM_HEAD_KEY] = state[EMBED_KEY] model_dtype = dtype if dtype is not None else state[EMBED_KEY].dtype if dtype is not None: state = cast_state(state, dtype) with crt.device("meta"): model = LlamaForCausalLM(config, model_dtype) model.to_empty(device=device) model.load_state_dict(state, strict=True, assign=True) del state model.tie_weights() model.fuse() model.model.build_rotary(max_positions, device) model.eval() return model ``` `tie_weights` is a check on a tied checkpoint: the head and the embedding must hold one object, and it raises when they do not. After the load, `model.state_dict()` lists the checkpoint's names in the checkpoint's order. A sharded checkpoint loads the same way: `load_state` reads the shards the index file names into one dict. ```python title="check_load.py" model = build_model(config, snapshot, crt.Device.cpu(), max_positions=4096) print(model.lm_head.weight is model.model.embed_tokens.weight) # True print(model.lm_head.weight.dtype) # clika_runtime.bfloat16 print(len(model.state_dict())) # 147 ``` ## Fuse the projections after the load Three projections over the same input are one matmul over the concatenated weights. `nn.fuse_linears` returns the group as an `nn.Linear` whose forward emits the parts' outputs side by side; the parts keep their state-dict names and read their rows back from the fused weight, and the group itself lists no parameters, so holding it as an attribute adds nothing to a save. Fuse once the weights are bound and before the first forward. ```python title="llama_from_scratch.py (excerpt)" def fuse(self) -> None: self.qkv = nn.fuse_linears(self.q_proj, self.k_proj, self.v_proj, name="qkv") # in LlamaMLP def fuse(self) -> None: self.gate_up = nn.fuse_linears(self.gate_proj, self.up_proj, name="gate_up") ``` The forward splits the fused output where the parts meet: ```python title="llama_from_scratch.py (excerpt)" qkv = self.qkv(x) # [tokens, (heads + 2 * kv_heads) * head_dim] q, k, v = crt.split_with_sizes(qkv, [heads * head_dim, kv_heads * head_dim, kv_heads * head_dim], -1) ``` The parts' activations follow one rule: parts that each apply the same pointwise activation keep it, and the group applies it once over the concatenation; parts that apply different ones refuse. A gated `activation=` (`"swiglu"`, `"geglu"`, `"reglu"`, `"situ"`) emits half the concatenated width, the gate and up convention, and needs parts without a bias. Weight-quantized parts (`nn.QLinearWoQ`) join a group too, and share one matmul only under one quantization scheme. A group the runtime cannot fuse (mixed schemes, unequal input widths) serves its parts side by side plus the concatenation, so the output never depends on whether the fusion engaged, only the dispatch count does. `nn.fuse_convs` does the same for convolutions over one input: the group is an `nn.Conv` whose output is the parts' outputs concatenated along the channel axis. Parts share one convolution when the kernel, stride, dilation, padding and its mode, `groups=1`, the dtype, and the presence of a bias agree. A gated activation refuses there, because it halves the channels. ## The KV cache and the step indices `nn.KVCache` owns one key row and one value row per layer, sized once from the configuration. `prepare_step` takes the cumulative query lengths, the past length of each sequence and its slot, and returns the `nn.StepIndices` the attention operator reads: `cu_seqlens_q`, `cu_seqlens_k`, `kvcache_start`, `slot_ids` and `max_seqlen_k`, all device tensors except the last. ```python title="llama_from_scratch.py (excerpt)" def make_cache(config, model, max_positions, device): kv_config = nn.KVCacheConfig( num_layers=config.num_hidden_layers, num_kv_heads=config.num_key_value_heads, head_dim=config.head_dim, kv_dtype=model.model.embed_tokens.weight.dtype, max_seqs=1, max_tokens_per_seq=max_positions, device=device, ) return nn.KVCache.make(kv_config, [nn.KVLayerSpec(preallocate=True) for _ in range(config.num_hidden_layers)]) cache = make_cache(config, model, 4096, device) step = cache.prepare_step([0, len(prompt_ids)], [0], [0]) # the prefill of one sequence ``` Layer `i` reads its rows through `cache.keys(i)` and `cache.values(i)`; the same handles receive the appended keys and values. The program `decode_step.py` beside `check_load.py` makes the two calls a serving loop makes, the prefill of a prompt and one decode step of the token it produced, with nothing else around them. ## The attention step is one operator Rotation at the positions the cache start implies, the append of the new keys and values into the cache row, and the causal grouped-query attention over the whole row are one call, for the prefill and for every decode step alike. The rows are passed as `past_key` / `past_value` and again as `out_present_key` / `out_present_value`, so the append lands in place. ```python title="llama_from_scratch.py (excerpt)" attention, _, _ = crt.group_query_attention_varlen( q, k, v, step.cu_seqlens_q, step.cu_seqlens_k, max_seqlen_k=step.max_seqlen_k, past_key=past_key, past_value=past_value, kvcache_start=step.kvcache_start, rope_cos=rope_cos, rope_sin=rope_sin, is_causal=True, rotary_mode="neox", num_heads=heads, kv_num_heads=kv_heads, out_present_key=past_key, out_present_value=past_value, slot_ids=step.slot_ids, ) return self.o_proj(attention) ``` The operator's contract, as the API reference states it: | Term | Contract | | --- | --- | | query, key, value | hidden-folded `[sum of S, heads * head_dim]`; `num_heads` and `kv_num_heads` drive the head split inside the operator | | `cu_seqlens_q`, `cu_seqlens_k` | the cumulative lengths of the batch's sequences, on the device | | `max_seqlen_q`, `max_seqlen_k` | the longest query and the longest cached sequence of the batch, passed as numbers (`step.max_seqlen_k` above is a host value) | | the continuous cache | `[max_seqs, kv_heads, max_seq, head_dim]` per layer, keys and values as two tensors; `nn.KVCache` owns one pair per layer | | appending in place | the cache row passed as `past_key` / `past_value` and again as `out_present_key` / `out_present_value`; the rotated keys and the values are appended before the attention | | `kvcache_start` | on a continuous cache a rank-1 `[B]` layout selector whose values are not read; each row's write offset is `cu_seqlens_k[b] - q_len[b]` | | `slot_ids` | `[B]` Int32, the cache row each batch row appends to and attends from; absent, batch row `b` uses cache row `b` | | a paged cache | the block pool `[num_blocks, kv_heads, block_size, head_dim]` with a `[B, max_blocks]` Int32 block table; `-1` pads a row's unused tail | | rotary embedding | in the operator, from `rope_cos` / `rope_sin` at the positions the cache implies; `rotary_mode` picks the pairing | | `q_norm_gain`, `k_norm_gain` | a per-head RMS norm applied after the rotation, rank-1 `[head_dim]`, together with `qk_norm_eps` (one without the other refuses); a model that normalizes before the rotation calls `qk_rms_norm` on the projections before this operator and passes no gains | | `head_sink` | `[heads]`, a per-head virtual logit folded into the softmax denominator | | a read-only attend | `attention_over_cache`: the same cache, nothing appended, rotary applied to the query alone | ## A residual add and the norm after it, fused Each block adds a residual and normalizes the sum twice. `crt.add_rms_norm` does both in one operator and returns the pair the next block needs: the normalized stream and the raw residual. The closing call takes the norm weight that follows the block, the next layer's `input_layernorm` or the model's final `norm`, so the first `crt.rms_norm` on the embedding is the only standalone norm in the model. ```python title="llama_from_scratch.py (excerpt)" o = self.self_attn(x, step, past_key, past_value, rope_cos, rope_sin) x1, h1 = crt.add_rms_norm(o, h, None, None, [self.hidden_size], None, self.post_attention_layernorm.weight, None, self.eps) m = self.mlp(x1) x2, h2 = crt.add_rms_norm(m, h1, None, None, [self.hidden_size], None, next_gain, None, self.eps) return x2, h2 ``` The head reads each sequence's last row from the device, `crt.index_select(x, -2, step.cu_seqlens_q[1:] - 1)`, so no shape is read on the host inside the forward. ## Greedy decoding with a one-token lookahead Operators return before their work runs, and a host read waits for the value it needs. The loop uses that: after the prefill, each iteration queues the next step on the device first, with the argmax tensor itself as its input, and reads the current token afterwards. The device is never idle while Python reads a token, and every timing is the host clock at a token read. ```python title="llama_from_scratch.py (excerpt)" cache.reset() prompt = crt.tensor(list(prompt_ids), dtype=crt.int32, device=device) length = len(prompt_ids) logits = model(prompt, cache.prepare_step([0, length], [0], [0]), cache) token = crt.argmax(logits, dims=[-1], index_dtype=crt.int32) # [1] crt.async_eval(token) past = length ids: list[int] = [] for produced in range(max_new_tokens): queued = None if produced + 1 < max_new_tokens: next_logits = model(token, cache.prepare_step([0, 1], [past], [0]), cache) queued = crt.argmax(next_logits, dims=[-1], index_dtype=crt.int32) crt.async_eval(queued) value = int(token.item()) ids.append(value) if value in stops or queued is None: break token = queued past += 1 ``` `crt.async_eval` submits the queued step without waiting; `token.item()` is the one read per token. The stop ids come from `generation_config.json`, with `config.json` as the fallback, the same source the command line reads. ## Run it and measure it The prompt goes through the checkpoint's chat template as one user turn with the assistant priming, rendered at the wall clock (`now_epoch_seconds=int(time.time())`), which is how `clikart-cli prompt` renders it, so the two replies compare byte for byte. ```bash python3 llama_from_scratch.py --model HuggingFaceTB/SmolLM2-135M-Instruct \ --prompt 'Write a long story about a lighthouse keeper who finds a map.' \ --max-new-tokens 64 --device cpu --json reply.json ``` `--json` writes the timings, the memory readings from `clika_runtime.memory_stats(device)` at four points of the run, the rendered prompt and the token ids. The chapter's `measure_decode.py --full` runs the same cell on both sides, alternating, and writes the report: ```bash python3 measure_decode.py --full \ --cli /path/to/clikart-cli --snapshot /path/to/checkpoint --device cpu \ --isl 128 --osl 64 --iters 5 --warmup 1 --repeats 3 \ --prompt 'Write a long story about a lighthouse keeper who finds a map.' --max-new-tokens 64 \ --out ./llama_cpu ``` The chapter README carries the measured table for this cell, the machine it ran on, how each column is defined, and how to read the spread against the difference between the two rows. A number is a claim about one command on one machine: quote it with its command lines and its cell. From here: [Structure inputs and outputs as pytrees](pytrees.mdx) for the containers the compile and save boundaries take, [Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx) for a module tree that starts as `torch.nn.Module`, and [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) for capturing a step as a graph. --- # Coming from PyTorch or Hugging Face The conventions a reader who knows PyTorch and the Hugging Face libraries meets first: channels-last tensors and OHWI weights, integer slots, reads that wait, the JSON and audio accessors, the tokenizer's inputs, and the generate calls side by side. Source: https://docs.clika.io/clikart/how-to/coming-from-pytorch.md {/* CERTIFICATION: every fact on this page is read from the public headers of the pinned release (nn/conv.h, compute/ops.h, json/json.h, io/io.h, tokenizer/tokenizer.h, threading/threading.h) and the wheel's generate_like_transformers README; the samples are excerpts, not a compiled program. */} ClikaRT's Python package reads like PyTorch and its model library loads a checkpoint by the name the Hugging Face hub gives it, so most of what you know carries over. This page lists the places where the runtime's convention differs from the one you expect, each with the one line that bridges it. ## Tensors are channels-last, weights are OHWI Activations are `[N, spatial..., C]` and a convolution weight is `[out_channels, K..., in_channels / groups]` (output-channel-first, channels last). PyTorch stores an activation `[N, C, H, W]` and a convolution weight `[O, C / groups, K, K]`, and every checkpoint exported from it keeps that order, so a weight is brought to OHWI once, at load, with one permute and a contiguous copy: ```cpp const Tensor w_ohwi = ops::contiguous(ops::permute(w_oihw, {0, 2, 3, 1})); // 2-D; {0, 2, 1} for 1-D, {0, 2, 3, 4, 1} for 3-D ``` ```python w_ohwi = w_oihw.permute(0, 2, 3, 1).contiguous() ``` `nn::Conv` and `nn.Conv` declare their geometry in that layout, and [Load images and audio for inference](load-images-and-audio.mdx) decodes a picture straight into `[H, W, C]`. ## A size is an integer at the operator's boundary A split size, an index, a sequence length (`max_seqlen_k` of the attention operators) is an integer slot. The C++ gateway type that takes a number or a tensor accepts a floating-point value too, and refuses it at run time when the slot is integer-typed, so pass an `int` (or an `int64_t`), never a `double` cast from one. ## A read is the wait An operator returns as soon as its work is queued; the kernels run behind it. A host read waits for the value: `item`, `numpy()`, `tolist()`, printing a tensor, `eval`, `to_string()` and `item_as_vec()` in C++. Nothing you read is ever an unfinished value. The consequence for a measurement: time a step at the read of its result, or after `synchronize()` on the producing stream; the return of the dispatch is not the end of the work. [Control asynchronous execution](control-async-execution.mdx) walks the model in full. ## JSON, audio and the tokenizer's files - A JSON document reads an integer through `as_int64()` and a floating-point value through `as_double()`. `operator[]` on a mutable document creates a missing key (it turns a null into an object); `at(key)` and `at(index)` read without creating. - `io::load_audio` returns the samples beside an `AudioInfo` whose `sample_rate`, `channels` and `frames` describe them (`frames / sample_rate` is the duration in seconds). - `Tokenizer::from_huggingface` in C++, `Tokenizer.from_file` in Python and `Tokenizer.fromHuggingface` in Kotlin read a model directory: its `tokenizer.json` with the special ids and the chat template; a SentencePiece model beside its `tokenizer_config.json`; or a `vocab.json` for a model no tokenizer class claims. ## The thread count is read once `CLIKA_RT_NUM_THREADS` sets the CPU worker count and is read before the first compute, so it goes into the environment before the first operator runs: in the shell, or from the program before it touches the runtime. An Android app sets it in its own process before it creates any tensor. ## generate, side by side The model library's Python surface mirrors the `transformers` idioms; the table the wheel's `examples/python/howto/generate_like_transformers/README.md` carries is the whole map. The rows a first program needs: | `transformers` | `clika_runtime.modelverse` | | --- | --- | | `AutoModelForCausalLM.from_pretrained(id)` | `AutoModelForCausalLM.from_pretrained(id, device="cpu")` | | `model.to("cuda")` | `from_pretrained(id, device="cuda:0")`, chosen at load | | `ids = tok(prompt).input_ids; out = model.generate(ids); tok.decode(out)` | `model.generate(prompt)`, which takes text and returns text | | `generate(..., do_sample=False)` | `generate(..., temperature=0.0)` | | `tok.apply_chat_template(messages)` | `model.render(messages)` | | `apply_chat_template` then `generate` | `model.chat(messages)` | | `TextIteratorStreamer` on a second thread | `model.stream_generate(prompt)`, an iterator of text pieces | | `pipeline("text-generation", model=id)` | `pipeline("text-generation", id)` | A model's weights download inside `from_pretrained` into the Hugging Face hub cache the other tooling on the machine shares; `mv.snapshot_download(id)` downloads without loading, and a local directory is a source everywhere a repository id is ([Run fully offline](/modelverse/how-to/run-fully-offline)). ## Where the same idea has another name | You reach for | Here | | --- | --- | | `torch.compile(model)` | `crt.compile(model)`: capture on the first call, replay after ([Trace eager code to graphs](trace-eager-code-to-graphs.mdx)) | | `torch.no_grad()`, `model.eval()` | no gradients exist; `model.eval()` is accepted and changes nothing | | `state_dict()` / `load_state_dict()` | the same names and the same dotted keys; `load_state_dict(state, assign=True)` adopts the checkpoint's tensors as the parameters, one resident copy ([Author a model in Python](author-a-model-in-python.mdx)) | | `torch.device("meta")` | `crt.device("meta")`: shapes and dtypes without bytes, then `to_empty(device=...)` | | DLPack exchange with torch | `clika_runtime.torch` ([Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx)) | --- # Control asynchronous execution Know when dispatched work actually runs, get completion callbacks, and force synchronous or lazy execution with the stream scopes. Source: https://docs.clika.io/clikart/how-to/control-async-execution.md An `ops::` call dispatches work and returns; the kernels run behind it. Where the work lands follows one law: a stream you pass is used as given; an operation placed on a bare `Device` resolves to the ambient placement scope's stream for that device when one is set, otherwise to the calling thread's asynchronous stream; and an operation with no placement of its own follows its inputs. Most programs never notice any of this, because reads wait for the result. This guide is for when you need control anyway: measuring where time goes, reacting the moment a result is ready, stepping op by op while debugging, or building a whole graph before running any of it. Everything below runs on CPU streams, so it behaves the same on any machine; the same rules apply to CUDA, Vulkan and Metal streams. The timing numbers are from one real run and vary with the machine; the ordering they show does not. ## Dispatch is not execution An `ops::` call returns as soon as the work is queued, and the result knows where it queued: `Tensor::stream()` names the stream the operation rode (under the asynchronous default, `stream().is_default()` is false), which is the stream to query and synchronize. `Tensor::status()` names where a result is in that lifecycle, `Stream::query_idle()` asks a stream without blocking, and `Stream::synchronize()` blocks until everything queued has run. Before dispatching anything, `StreamOrDevice(device).resolve()` reads where a bare-Device placement would land. Host reads (`to_string`, `item`, `item_as_vec`) wait on the producing work themselves, so a read is always safe; what you never observe is unfinished bytes. A timing is taken at a host read of the result, or after `synchronize()` on the producing stream: the return of the dispatch is not the end of the work, and a result nobody reads is not evidence of when it ran. ```cpp title="dispatch_vs_ready.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Stream; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; namespace { double ms_since(std::chrono::steady_clock::time_point t0) { return std::chrono::duration( std::chrono::steady_clock::now() - t0).count(); } } // namespace int main() { // The inputs settle on a stream of their own. The chain below carries no // placement, so it rides the calling thread's asynchronous lane, and its // result names that stream: x.stream() is what to query and synchronize. Stream worker = Stream::create(Device::cpu()); const Tensor a = ops::full({2048, 2048}, 0.001, DataType::Float32, worker); const Tensor w = ops::full({2048, 2048}, 0.001, DataType::Float32, worker); worker.synchronize(); // settle the inputs so only the chain is measured const auto t0 = std::chrono::steady_clock::now(); Tensor x = a; // 48 dependent 2048 x 2048 products, about 0.8 TFLOP: the lane is still // at work when the query after the dispatch returns, on any machine. for (int i = 0; i < 48; ++i) x = ops::matmul(x, w); const double dispatch_ms = ms_since(t0); Stream lane = x.stream(); // the stream the chain actually ran on std::printf("dispatch returned after %.1f ms, stream idle: %s\n", dispatch_ms, lane.query_idle() ? "yes" : "no"); lane.synchronize(); std::printf("ready after %.1f ms, stream idle: %s\n", ms_since(t0), lane.query_idle() ? "yes" : "no"); // A host read needs none of the above; it waits on its producer alone. // (item() reads a single-element tensor, so reduce first.) std::printf("max(x) = %g\n", ops::amax(x).item()); return 0; } ``` Every operator returns at once and its work rides the calling thread's lane on the input's device; a read waits for the value it needs. The contract, in full: 1. These settle (they block until the value exists): `numpy()`, `bytes()`, `item()`, `tolist()`, `repr` and `print`, `bool()`, `__array__`, `__dlpack__`, `crt.save`, `crt.eval(*trees)` and `crt.synchronize(target)`; a truthful `shape` / `numel()` / `nbytes` settles only when the producer's shape depends on data. 2. These never settle: `dtype`, `device`, `ndim`, `stride()`, `is_contiguous()`. 3. The scopes are thread-local context managers that also work as decorators: `crt.synchronous()`, `crt.tracing()`, `crt.eager()`, `crt.meta_init()` (the same as `crt.device("meta")`), `crt.device(d)` and `crt.stream(s)`. 4. `crt.async_eval(*trees)` submits the work and returns without waiting, the one-token-lookahead idiom of a decode loop. 5. A data read under `crt.tracing()` or of a storage-free tensor raises `crt.ClikaRTError`; `numpy()` on a device tensor raises `TypeError` (move it to the cpu first); `bool()` of a tensor with more than one element raises `RuntimeError`. ```python title="settle.py" import numpy as np import clika_runtime as crt x = np.ones((64, 64), dtype=np.float32) t = crt.tensor(x) y = t for _ in range(50): y = crt.add(crt.mul(y, 1.01), 0.01) # fifty operators, dispatched without waiting print(y.dtype) # clika_runtime.float32 (metadata never settles) print(float(y.numpy()[0, 0])) # 2.289262533187866 (the first read returns the settled value) tree = {"a": crt.exp(t), "b": [crt.sum(t), None, "text"], "c": (crt.relu(t),)} crt.eval(tree, crt.abs(t)) # settles every tensor leaf of every tree, ignores the rest z = crt.add(crt.mul(t, 2.0), 1.0) crt.async_eval(z) # submitted, not waited for print(float(z.numpy()[0, 0])) # 3.0 (3.0, read later) borrowed = crt.from_numpy(x) # a zero-copy view over the array borrowed.add_(1.0) crt.synchronize() # the calling thread's lane; the array carries the write after this print(x[0, 0]) # 2.0 (2.0) ``` A raw read of a borrowed source array after an in-place operator settles first, as the last four lines do; a fresh `numpy()` call settles on its own. ```text dispatch returned after 23.5 ms, stream idle: no ready after 24.2 ms, stream idle: yes max(x) = 0.00176683 ``` The dispatch and the ready walls sit close together on a CPU stream, and the gap between them is the point: dispatch is still not execution (the stream is not idle when the loop returns), but memory on an asynchronous CPU stream follows execution, so a dependent chain is issued one operation ahead of the one executing and the dispatch loop is paced by the work itself. On a CUDA stream the guarantee is the enqueue rather than the run, so the same loop returns in well under a millisecond while the device works behind it. An op placed on a bare `Device` overlaps with the caller; code that needs it synchronous opts in explicitly, with `Stream::default_stream(device)` as the placement or a `SynchronousStreamScope` around the region. ## React the moment a result is ready `Tensor::on_complete` registers a callback on a result; it fires the moment the producing kernel finishes, on ClikaRT's callback pool while the producer is still running, or inline at registration when the result has already settled. No polling and no blocked thread, the push-style inverse of `synchronize()`. The tensor's storage is kept alive for the callback, which receives it by const reference. ```cpp title="on_complete.cpp" #include #include #include #include #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Stream; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; namespace { double ms_since(std::chrono::steady_clock::time_point t0) { return std::chrono::duration( std::chrono::steady_clock::now() - t0).count(); } } // namespace int main() { Stream worker = Stream::create(Device::cpu()); const Tensor a = ops::full({1024, 1024}, 0.001, DataType::Float32, worker); const Tensor w = ops::full({1024, 1024}, 0.001, DataType::Float32, worker); worker.synchronize(); std::atomic fired{false}; const auto t0 = std::chrono::steady_clock::now(); Tensor x = a; for (int i = 0; i < 10; ++i) x = ops::matmul(x, w); std::printf("main: dispatch done at %.1f ms, registering the callback and doing other work\n", ms_since(t0)); // Fires the moment the producing kernel finishes: on ClikaRT's callback // pool while the producer is still running, inline at registration when // the result has already settled. The storage stays alive for the call. x.on_complete([&](const Tensor& result) { std::printf("callback: fired after %.1f ms, x = %s\n", ms_since(t0), result.to_string().c_str()); fired = true; }); while (!fired) std::this_thread::yield(); return 0; } ``` The Python package carries no completion callback; the push-style `on_complete` is the C++ arm's. The Python shape of the same idea is `crt.async_eval` to submit without waiting, and a read at the point of use: hand the tensor to whoever consumes it, and that consumer's `numpy()` or `item()` is the wait, on its own thread if the dispatching thread must not block. ```text main: dispatch done at 10.8 ms, registering the callback and doing other work callback: fired after 11.8 ms, x = Tensor(shape=[1024, 1024], dtype=Float32, device=CPU, numel=1048576, data=[0.001268, 0.001268, 0.001268, 0.001268, 0.001268, 0.001268, ...]) ``` Use it to hand results to a queue, complete a request, or chain host-side work without dedicating a thread to waiting. Keep callbacks short; they share the callback pool. ## Step synchronously while debugging `SynchronousStreamScope` is RAII: inside it, every op the calling thread dispatches to the stream completes before the call returns, so the program state after each line is exactly what the line computed. Deterministic and slow, which is the right trade while hunting a numeric bug or stepping in a debugger. Open it on a quiescent stream, and pipelining resumes when the scope closes. It is also the region-sized opt-in for code that needs synchronous bare-Device behavior. ```cpp title="sync_scope.cpp" #include #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Stream; using ClikaRT::SynchronousStreamScope; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; int main() { Stream worker = Stream::create(Device::cpu()); // One 4096 x 4096 product, about 137 GFLOP: still at work when the // query after the dispatch returns, on any machine. const Tensor a = ops::full({4096, 4096}, 0.001, DataType::Float32, worker); const Tensor w = ops::full({4096, 4096}, 0.001, DataType::Float32, worker); worker.synchronize(); // An op with no placement rides the calling thread's asynchronous lane; // its result names that stream, which is the one to query. Tensor x = ops::matmul(a, w); // async: dispatched, likely still running std::printf("no scope : idle after dispatch: %s\n", x.stream().query_idle() ? "yes" : "no"); x.stream().synchronize(); { SynchronousStreamScope scope(worker); x = ops::matmul(a, w); // completes before this line returns std::printf("sync scope : idle after dispatch: %s\n", x.stream().query_idle() ? "yes" : "no"); } return 0; } ``` `crt.synchronous()` is the region form: inside it every operator the calling thread dispatches completes before the call returns, so a borrowed source array carries an in-place write the moment the line finishes. It works as a decorator too. ```python title="sync_scope.py" import numpy as np import clika_runtime as crt source = np.zeros((8, 8), dtype=np.float32) borrowed = crt.from_numpy(source) with crt.synchronous(): borrowed.add_(3.0) print(source[0, 0]) # 3.0 (3.0, written before the line returned) result = crt.mul(borrowed, 2.0) print(float(result.numpy()[0, 0])) # 6.0 (6.0) @crt.synchronous() def step(t: crt.Tensor) -> crt.Tensor: return crt.relu(t) print(float(step(borrowed).numpy()[0, 0])) # 3.0 (3.0) ``` ```text no scope : idle after dispatch: no sync scope : idle after dispatch: yes ``` ## Build the whole graph first, run it once `TracingScope` flips the calling thread the other way, to lazy: ops return `Unscheduled` placeholder tensors carrying lineage and no kernel runs. One `synchronize()` (or any read) materializes the graph leaf-first. Use it to declare a computation in full before spending anything, or to hand the runtime the widest possible scheduling view. The same idea with a reusable artifact is [tracing eager code to a graph](trace-eager-code-to-graphs.mdx): capture a function once as a `ModelGraph` and run it repeatedly, instead of scoping one thread's dispatches. ```cpp title="tracing_scope.cpp" #include #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Tensor; using TensorStatus = ClikaRT::Tensor::Status; using ClikaRT::TracingScope; namespace ops = ClikaRT::ops; namespace { const char* name(TensorStatus s) { switch (s) { case TensorStatus::Unscheduled: return "Unscheduled"; case TensorStatus::Evaluated: return "Evaluated"; case TensorStatus::Available: return "Available"; } return "?"; } } // namespace int main() { const Tensor a = ops::ones({4, 4}, DataType::Float32, Device::cpu()); const Tensor b = ops::ones({4, 4}, DataType::Float32, Device::cpu()); Tensor m, s; { TracingScope trace; m = ops::matmul(a, b); // no kernel runs s = ops::add(m, a); std::printf("traced : m=%s s=%s\n", name(m.status()), name(s.status())); s.synchronize(); // materialize the graph, leaf-first } // Outside the scope, ops run eagerly again: a one-element view of the // realized result and its value readback. std::printf("realized: m=%s s=%s, s[0] = %g (4 ones dot ones + 1 = 5)\n", name(m.status()), name(s.status()), ops::select(s.reshape({-1}), 0, 0).item()); return 0; } ``` `crt.tracing()` flips the calling thread to lazy: operators return unscheduled tensors that carry their shape and dtype and no value, a data read inside the region raises `crt.ClikaRTError`, and `crt.eval` materializes the whole chain leaf-first. `crt.eager()` restores the default inside a tracing region for the one operator that must run at once. ```python title="tracing_scope.py" import numpy as np import clika_runtime as crt t = crt.tensor(np.ones((8, 8), dtype=np.float32)) with crt.tracing(): y = crt.add(crt.mul(t, 2.0), 1.0) print(y.shape, y.dtype) # clika_runtime.Size([8, 8]) clika_runtime.float32 ((8, 8) clika_runtime.float32; no kernel has run) try: y.numpy() except crt.ClikaRTError: print("a value read under tracing raises") # printed: the read raised with crt.eager(): at_once = crt.mul(t, 2.0) print(float(at_once.numpy()[0, 0])) # 2.0 (2.0, run at once) crt.eval(y) # materializes the traced chain print(float(y.numpy()[0, 0])) # 3.0 (3.0) ``` ```text traced : m=Unscheduled s=Unscheduled realized: m=Evaluated s=Evaluated, s[0] = 5 (4 ones dot ones + 1 = 5) ``` The four tools compose into one rule of thumb: leave the asynchronous default alone for throughput, read results and let the reads wait, reach for `on_complete` when a thread should not wait, and reserve the two scopes for debugging (synchronous) and up-front graph building (tracing). The bundle's `async` example walks each in its own chapter, including safe-reads patterns this guide leaves implicit. --- # Edit a graph Change a ModelGraph before finalize(): rewire the reads of a value, change what a node reads, rename a node, and edit the graph's inputs and outputs. A refused edit changes nothing; then optimize, finalize and run the edited graph. Source: https://docs.clika.io/clikart/how-to/edit-a-graph.md {/* Every block is a program under examples//howto/edit_a_graph/: the first block of each tab is get_a_graph whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output. */} A `ModelGraph` that you trace or compile can be changed before you call `finalize()`. You can move the reads of a value to another value, remove or bypass a node, change what a node reads, rename a node, and add, remove or rename the graph's inputs and outputs. Each edit checks everything before it changes anything, so a refused edit leaves the graph as it was and says what it refused. The edits are available from C++ and Python, under the same names. ## A graph to edit A trace returns the graph as built, with every operator as written and nothing optimized or finalized, so it is ready to edit. The model below computes `y = Relu(x) + Neg(x)`: `x` feeds a Relu and a Neg, and one Add reads both. Every example on this page uses it. `label` names a node by its input name or its operator, `labels` lists a graph's nodes, and `producers` lists the nodes a node reads, one per input port. Edits run before `finalize()`. A finalized graph serves `run()` and refuses every edit with the code name `FAILED_PRECONDITION`; to edit it, trace or compile the model again. ```cpp title="get_a_graph.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Error; using ClikaRT::Tensor; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::Node; using ClikaRT::graph::NodeKind; using ClikaRT::graph::OpCode; using ClikaRT::graph::Value; namespace ops = ClikaRT::ops; namespace { // y = Relu(x) + Neg(x): x feeds a Relu and a Neg, both read by one Add. std::vector model(const std::vector& inputs) { const Tensor rectified = ops::relu(inputs[0]); const Tensor negated = ops::neg(inputs[0]); return {ops::add(rectified, negated)}; } // A node's label: an input's name, an operator's code. std::string label(const Node& node) { return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code())); } // The labels of `nodes`, separated by spaces. std::string labels(const std::vector& nodes) { std::string out; for (const Node& node : nodes) out += (out.empty() ? "" : " ") + label(node); return out; } // The labels of the nodes that produce what `node` reads, one per input port. std::string producers(const Node& node) { std::string out; for (const Value& value : node.inputs()) out += (out.empty() ? "" : " ") + label(*value.producer()); return out; } // A trace returns the graph as built: every operator as written, nothing optimized or finalized. ModelGraph trace_model() { const std::vector signature = {{"x", DataType::Float32, {2, 3}}}; const std::vector outputs = {"y"}; return ClikaRT::graph::trace(model, signature, "edit", outputs); } } // namespace int main() { const ModelGraph graph = trace_model(); const Node add = graph.find_nodes(OpCode::Add).front(); std::printf("%s | %s\n", labels(graph.nodes()).c_str(), producers(add).c_str()); // x Relu Neg Add | Relu Neg std::printf("%s %s\n", graph.output_names().front().c_str(), graph.is_finalized() ? "true" : "false"); // y false // Edits run before finalize(): a finalized graph refuses every one of them. ModelGraph done = trace_model(); done.finalize(); try { done.rename_output("y", "total"); } catch (const Error& error) { std::printf("%s\n", error.code_name().c_str()); // FAILED_PRECONDITION } return 0; } ``` ```python title="get_a_graph.py" import clika_runtime as crt from clika_runtime.graph import NodeKind, OpCode def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]: x = inputs[0] return [crt.relu(x) + crt.neg(x)] # y = Relu(x) + Neg(x) def label(node: crt.graph.Node) -> str: return node.name if node.kind == NodeKind.Input else node.op_code.name def labels(graph: crt.graph.ModelGraph) -> list[str]: return [label(node) for node in graph.nodes()] def producers(node: crt.graph.Node) -> list[str]: return [label(value.producer()) for value in node.inputs] # A trace returns the graph as built: every operator as written, nothing optimized or finalized. graph = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph (add,) = graph.find_nodes(OpCode.Add) print(labels(graph), producers(add)) # ['x', 'Relu', 'Neg', 'Add'] ['Relu', 'Neg'] print(graph.output_names(), graph.is_finalized()) # ['y'] False # Edits run before finalize(): a finalized graph refuses every one of them. done = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph done.finalize() try: done.rename_output("y", "total") except crt.InvalidArgumentError as error: print(error.code_name) # FAILED_PRECONDITION ``` Edit a graph when a model needs a change its source does not make: an operator to remove before deployment, a value to expose or feed, or a name a caller binds. Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, `label`, `labels`, `producers` and the trace). ## A refused edit changes nothing Every edit checks its arguments and the graph's rules before it changes anything, so a refused edit leaves the nodes, the inputs, the outputs and the constants as they were. Its message names the call and what it refused. The code name is `INVALID_ARGUMENT` for an argument the edit cannot take, and `FAILED_PRECONDITION` for a finalized graph. In C++ the edit throws a `ClikaRT::Error` (a `Result` carries it in a build without exceptions); in Python it raises `InvalidArgumentError`, whose `code_name` says which. Here `rename_node` refuses an input and names the call that renames one. ```cpp title="refused_edit.cpp" int main() { ModelGraph graph = trace_model(); const std::string before = labels(graph.nodes()); try { graph.rename_node(*graph.node("x"), "features"); // an input is renamed with rename_input() } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // INVALID_ARGUMENT | rename_node: 'x' is a graph input; rename it with rename_input() } std::printf("%s %s\n", labels(graph.nodes()) == before ? "true" : "false", graph.input_names().front().c_str()); // true x return 0; } ``` ```python title="refused_edit.py" before = labels(graph) try: graph.rename_node(graph.node("x"), "features") # an input is renamed with rename_input() except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | rename_node: 'x' is a graph input; rename it with rename_input() print(labels(graph) == before, graph.input_names()) # True ['x'] ``` Branch on the code name and report the message; [Handle errors by code](handle-errors-by-code.mdx) covers the channels every failure carries. ## Rewire the reads of a value `replace_all_uses_with(old_value, replacement)` moves every read of a value to another one: each operator input it feeds, and each graph output it returns, which keeps its name. `remove_node(node)` removes an operator that nothing reads, and refuses one that is still read, naming the reader to rewire or remove first. `bypass_node(node)` hands a node's readers the value on its input (the first input and output, unless you name the ports) and removes the node. A value takes another's place only with the same dtype and dims. ```cpp title="rewire.cpp" int main() { ModelGraph graph = trace_model(); const Node relu = graph.find_nodes(OpCode::Relu).front(); const Node neg = graph.find_nodes(OpCode::Neg).front(); const Node add = graph.find_nodes(OpCode::Add).front(); graph.replace_all_uses_with(*neg.output(0), *graph.node("x")->output(0)); // the Add reads x where it read Neg(x) graph.remove_node(neg); // nothing reads the Neg now std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu Add graph.bypass_node(relu); // the Add reads the Relu's input, x, and the Relu goes std::printf("%s | %s\n", labels(graph.nodes()).c_str(), producers(add).c_str()); // x Add | x x return 0; } ``` ```python title="rewire.py" (relu,) = graph.find_nodes(OpCode.Relu) (neg,) = graph.find_nodes(OpCode.Neg) (add,) = graph.find_nodes(OpCode.Add) graph.replace_all_uses_with(neg.output(0), graph.node("x").output(0)) # the Add reads x where it read Neg(x) graph.remove_node(neg) # nothing reads the Neg now print(labels(graph)) # ['x', 'Relu', 'Add'] graph.bypass_node(relu) # the Add reads the Relu's input, x, and the Relu goes print(labels(graph), producers(add)) # ['x', 'Add'] ['x', 'x'] ``` Rewire a graph to remove an operator a deployment does not need, or to let readers take a value the graph already computes. ## Change what a node reads `set_input(node, port, value)` gives one input port another value; the port must read a value already. `set_constant_input(node, port, tensor)` binds a new constant holding the tensor's bytes to the port: the graph keeps a handle to the bytes, with no copy, and places them with the graph at `finalize()`. `add_constant(tensor)` makes a constant that nothing reads yet, for `set_input` or `replace_all_uses_with` to wire in, and `finalize()` drops a constant that nothing reads by then. A weight an operator holds itself, such as a compiled MatMul's, stays as it is, and an edit that would replace it refuses. ```cpp title="node_inputs.cpp" int main() { ModelGraph graph = trace_model(); const Node relu = graph.find_nodes(OpCode::Relu).front(); const Node neg = graph.find_nodes(OpCode::Neg).front(); const Node add = graph.find_nodes(OpCode::Add).front(); graph.set_input(add, 1, *relu.output(0)); // port 1 reads Relu(x) in place of Neg(x) graph.remove_node(neg); // which nothing reads now std::printf("%s\n", producers(add).c_str()); // Relu Relu // Port 1 reads a constant holding the tensor's bytes. graph.set_constant_input(add, 1, Tensor::full({2, 3}, 1.0, DataType::Float32)); std::printf("%s %zu\n", add.input(1)->is_constant() ? "true" : "false", graph.constants().size()); // true 1 const Value half = graph.add_constant(Tensor::full({2, 3}, 0.5, DataType::Float32)); // nothing reads it yet graph.set_input(add, 1, half); // port 1 reads it, and the ones, read by nothing, go std::printf("%s %zu\n", *add.input(1) == half ? "true" : "false", graph.constants().size()); // true 1 return 0; } ``` ```python title="node_inputs.py" (relu,) = graph.find_nodes(OpCode.Relu) (neg,) = graph.find_nodes(OpCode.Neg) (add,) = graph.find_nodes(OpCode.Add) graph.set_input(add, 1, relu.output(0)) # port 1 reads Relu(x) in place of Neg(x) graph.remove_node(neg) # which nothing reads now print(producers(add)) # ['Relu', 'Relu'] graph.set_constant_input(add, 1, crt.ones(2, 3)) # port 1 reads a constant holding the tensor's bytes print(add.input(1).is_constant(), len(graph.constants())) # True 1 half = graph.add_constant(crt.full((2, 3), 0.5)) # a constant nothing reads yet graph.set_input(add, 1, half) # port 1 reads it, and the ones, read by nothing, go print(add.input(1) == half, len(graph.constants())) # True 1 ``` Change a node's inputs to feed a constant of your own into the graph, or to point an operator at another producer. ## Rename a node `rename_node(node, name)` renames an operator, and its outputs keep their names; an input is renamed with `rename_input` instead. A view of the old name is gone afterwards, so look the node up again with `node(name)`. ```cpp title="rename.cpp" int main() { ModelGraph graph = trace_model(); graph.rename_node(graph.find_nodes(OpCode::Add).front(), "total"); // views of the old name are gone afterwards const Node total = *graph.node("total"); // so find the node again by its new one std::printf("%s | %s | %s\n", label(total).c_str(), producers(total).c_str(), graph.output_names().front().c_str()); // Add | Relu Neg | y return 0; } ``` ```python title="rename.py" (add,) = graph.find_nodes(OpCode.Add) graph.rename_node(add, "total") # views of the old name are gone afterwards total = graph.node("total") # so find the node again by its new one print(label(total), producers(total), graph.output_names()) # Add ['Relu', 'Neg'] ['y'] ``` Rename nodes to give them the names your own tools and reports use. ## Edit the graph's inputs and outputs `add_input(spec)` adds a graph input with the spec's name, dtype and dims (a dynamic dim takes its size at `run()`) and returns its input node, whose value the other edits take. `add_output(value, name)` returns a value under a new output name, and `remove_output(name)` stops returning one, while the node that computes it stays. `rename_input(name, new_name)` and `rename_output(name, new_name)` rename an input and an output. Input and output names stay unique, and a name a KV cache layer binds stays as it is. ```cpp title="graph_io.cpp" int main() { ModelGraph graph = trace_model(); const Node relu = graph.find_nodes(OpCode::Relu).front(); const Node neg = graph.find_nodes(OpCode::Neg).front(); const Node add = graph.find_nodes(OpCode::Add).front(); const Node bias = graph.add_input({"bias", DataType::Float32, {2, 3}}); // a new input, bound at run() graph.set_input(add, 1, *bias.output(0)); // y = Relu(x) + bias graph.remove_node(neg); graph.add_output(*relu.output(0), "rectified"); // Relu(x) is returned too graph.rename_input("x", "features"); graph.rename_output("y", "total"); const std::vector ins = graph.input_names(); const std::vector outs = graph.output_names(); std::printf("%s %s | %s %s\n", ins[0].c_str(), ins[1].c_str(), outs[0].c_str(), outs[1].c_str()); // features bias | total rectified graph.remove_output("rectified"); // the Relu that computed it stays: the Add reads it std::printf("%s | %s\n", graph.output_names().front().c_str(), producers(add).c_str()); // total | Relu bias return 0; } ``` ```python title="graph_io.py" (relu,) = graph.find_nodes(OpCode.Relu) (neg,) = graph.find_nodes(OpCode.Neg) (add,) = graph.find_nodes(OpCode.Add) bias = graph.add_input(crt.TensorSpec("bias", crt.float32, [2, 3])) # a new input, bound at run() graph.set_input(add, 1, bias.output(0)) # y = Relu(x) + bias graph.remove_node(neg) graph.add_output(relu.output(0), "rectified") # Relu(x) is returned too graph.rename_input("x", "features") graph.rename_output("y", "total") print(graph.input_names(), graph.output_names()) # ['features', 'bias'] ['total', 'rectified'] graph.remove_output("rectified") # the Relu that computed it stays: the Add reads it print(graph.output_names(), producers(add)) # ['total'] ['Relu', 'bias'] ``` Edit the inputs and outputs to expose an intermediate value, to feed a value from outside the graph, or to match the names a caller binds. ## Optimize, finalize and run the edited graph After the edits, `optimize()` runs the graph optimizer when you want it, and `finalize()` makes the graph runnable, after which it refuses edits. `run()` binds the inputs in `input_names()` order, the new input included. Each program checks the result against a reference it computes itself; the Python program also imports numpy as `np` for that. ```cpp title="after_the_edits.cpp" int main() { ModelGraph graph = trace_model(); const Node neg = graph.find_nodes(OpCode::Neg).front(); const Node add = graph.find_nodes(OpCode::Add).front(); const Node bias = graph.add_input({"bias", DataType::Float32, {2, 3}}); graph.set_input(add, 1, *bias.output(0)); // y = Relu(x) + bias graph.remove_node(neg); graph.optimize(); // optional: the graph optimizer runs over the edited graph graph.finalize(); // the graph runs from here on, and refuses edits const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F}; const std::vector results = graph.run( {Tensor::from_data(x.data(), {2, 3}, DataType::Float32), Tensor::full({2, 3}, 0.5, DataType::Float32)}); const std::vector y = results.front().reshape({-1}).item_as_vec(); bool matches = y.size() == x.size(); for (std::size_t i = 0; matches && i < x.size(); ++i) { matches = y[i] == (x[i] > 0.0F ? x[i] : 0.0F) + 0.5F; // the reference, Relu(x) + 0.5, by hand } for (std::size_t i = 0; i < y.size(); ++i) std::printf("%s%g", i == 0 ? "" : " ", y[i]); std::printf("\n%s\n", matches ? "true" : "false"); // 1.5 0.5 3.5 0.5 5.5 0.5 // true return 0; } ``` ```python title="after_the_edits.py" (neg,) = graph.find_nodes(OpCode.Neg) (add,) = graph.find_nodes(OpCode.Add) bias = graph.add_input(crt.TensorSpec("bias", crt.float32, [2, 3])) graph.set_input(add, 1, bias.output(0)) # y = Relu(x) + bias graph.remove_node(neg) graph.optimize() # optional: the graph optimizer runs over the edited graph graph.finalize() # the graph runs from here on, and refuses edits x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32) b = np.full((2, 3), 0.5, np.float32) (y,) = graph.run([crt.tensor(x), crt.tensor(b)]) print(y.numpy().tolist()) # [[1.5, 0.5, 3.5], [0.5, 5.5, 0.5]] print(np.array_equal(y.numpy(), np.maximum(x, 0) + b)) # True (numpy states the reference) ``` Finalize once the graph has every edit it needs, since a finalized graph takes none. ## Views across edits A `Node`, `Value` or `Edge` view taken before an edit finds its node, value or edge again by its key and keeps answering, and it refuses once an edit removed or renamed what it names, as [Views across edits](query-a-graph.mdx#views-across-edits) on the Query a graph page shows. --- # Add ClikaRT to an existing CMake project Link the bundle into a project that already builds, with find_package or add_subdirectory, and the one linker flag serving-runtime consumers need. Source: https://docs.clika.io/clikart/how-to/existing-cmake-project.md Your application already builds, and you want it to call ClikaRT. The integration is two lines in the CMake file you already have: declare the package, link one target. This guide adds ClikaRT to an existing app both ways (the prebuilt bundle via `find_package`, the bundle source tree via `add_subdirectory`), then covers the two things integrations trip on: where the libraries are found at run time, and the RTTI flag the serving runtime requires. ClikaRT needs C++17 or later; if your project predates `set(CMAKE_CXX_STANDARD 17)`, add it. ## Link the bundle with find_package Say the existing project is a telemetry service with one executable: ```cmake title="CMakeLists.txt (before)" cmake_minimum_required(VERSION 3.19) project(telemetry CXX) set(CMAKE_CXX_STANDARD 17) add_executable(telemetry src/main.cpp src/collect.cpp) ``` Two lines make ClikaRT available to it. `find_package(ClikaRT CONFIG)` loads the package from the extracted bundle, and the imported target `ClikaRT::ClikaRT` carries the include paths, the libraries, and their link order: ```cmake title="CMakeLists.txt (after)" cmake_minimum_required(VERSION 3.19) project(telemetry CXX) set(CMAKE_CXX_STANDARD 17) find_package(ClikaRT CONFIG REQUIRED) add_executable(telemetry src/main.cpp src/collect.cpp) target_link_libraries(telemetry PRIVATE ClikaRT::ClikaRT) ``` CMake finds the package through `ClikaRT_DIR`, pointing at the bundle's `cmake/` directory: ```bash cmake -S . -B build -DClikaRT_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build ``` `-DClikaRT_DIR` is a cache variable, so it is remembered after the first configure. To avoid passing it at all, add the bundle to `CMAKE_PREFIX_PATH` (in a toolchain file, a CMake preset, or the environment) and `find_package` finds it there. A wrong or unset path fails the configure with `Could not find a package configuration file provided by "ClikaRT"`; the fix is always the same, point `ClikaRT_DIR` at `/cmake`. ## Where the libraries are found at run time The package stamps each consumer binary with an rpath to the bundle's platform `lib/` directory, so a freshly built binary runs in place with no library paths to set, `LD_LIBRARY_PATH` included. That rpath names the bundle's absolute path on the build machine. For binaries that ship to other machines, copy the libraries your binary needs from the bundle's `lib/` next to it (or into your package's lib directory) and set your own relative rpath, for example `$ORIGIN/../lib`; the bundle's prebuilt examples ship exactly that way. ## Serving-runtime consumers mirror the library's -fno-rtti Code that only uses tensors, `ops::`, `io::`, tokenizers, or the HTTP pieces builds with your project's existing flags. Code that uses the serving runtime (`runtime::FunctionModel`, `runtime::Model`, the executor) must be compiled with `-fno-rtti`, matching how the library builds. With RTTI left on, the target fails at link with: ```text undefined reference to `typeinfo for ClikaRT::runtime::Model' ``` Scope the flag to the targets that touch `runtime::`: ```cmake target_compile_options(telemetry PRIVATE -fno-rtti) ``` The bundle's own examples set the same flag; it is the one compile-option requirement in the integration. ## Check the bundle into your tree with add_subdirectory For a monorepo that keeps the bundle in the repository (or fetches it into the tree), the bundle root is also a CMake subproject. It defines the same `ClikaRT::ClikaRT` target with the same options, so consumers cannot tell the difference: ```cmake add_subdirectory(third_party/clikart) add_executable(telemetry src/main.cpp src/collect.cpp) target_link_libraries(telemetry PRIVATE ClikaRT::ClikaRT) ``` No `ClikaRT_DIR` and no configure-time flag are involved; the path in the tree is the whole wiring. Prefer `find_package` when the bundle lives outside the repository (a shared install, a CI cache), `add_subdirectory` when it lives inside. ## Verify A two-line call in your existing code proves the link end to end: ```cpp title="src/main.cpp (excerpt as a standalone check)" #include #include "ClikaRT/clika_rt.h" int main() { std::printf("ClikaRT %s\n", ClikaRT::GetVersionInfo().c_str()); std::printf("cuda=%d vulkan=%d\n", ClikaRT::device::is_cuda_available(), ClikaRT::device::is_vulkan_available()); return 0; } ``` ```text ClikaRT 0.6.4 cuda=1 vulkan=1 ``` The availability flags reflect the machine, not the build: the same binary prints different backends on different hardware, which is the point. From here, [the tutorial](../getting-started/first-program/01-your-first-program.mdx) covers the API the newly linked target now reaches, and [Quick install](../getting-started/installation.md) has the bundle layout. --- # Export a graph to ONNX Write a ModelGraph as an ONNX model file, or build the same model in memory: read it back through OnnxModel::open and compile, keep a named dynamic dimension, choose where the weights go and the opset, and meet what the export refuses. Source: https://docs.clika.io/clikart/how-to/export-a-graph-to-onnx.md {/* Every block is a program under examples//howto/export_a_graph_to_onnx/: the first block of each tab is get_the_graph whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output. */} `ModelGraph::export_onnx` writes a graph as a runnable ONNX model file, and `io::OnnxModel::from_graph` builds the same model in memory without writing anything. The model keeps the graph's input and output names, dtypes and dimensions. A dynamic dimension keeps its name, so two inputs that share one still share it when the model is read back, and the graph's weights become the model's initializers. `io::OnnxModel::open` and `compile()` read the file back into a graph that returns the same values. The calls are available from C++ and Python, under the same names. ## The graph to export The examples on this page export one elementwise model, `y = Relu(x * w + b)`, over an input `x` of shape [batch, 37] whose first dimension is dynamic and named `batch`. The model captures `w` and `b`, two vectors of 37 values, so the traced graph holds them as its two weights. `input(batch)` builds an input whose row `i` holds `i + 1` in every column, and `row_sums` prints the sum of each output row, so two graphs that return the same values print the same line. Every value on this page is a multiple of 1/8, which float32 holds exactly. `print_specs` prints an input or an output with a dynamic dimension under its name, and the other helpers name a temporary directory for the files a program writes, list its files and print a refusal. In Python the programs import `clika_runtime` as `crt`, and `ModelGraph` and `Transform` from `clika_runtime.graph`. There `rows(batch)` builds the input, `slots` prints the `(name, dtype, dims)` tuples that `inputs()` and `outputs()` return, with a dynamic dimension as -1, and a value's `spec().dim_names` holds the dimension names. A trace returns the graph as built, with every operator as written and nothing optimized or finalized, and `export_onnx` writes it in that state. This program prints the graph's input and output, then runs it once after `finalize()`. Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, the model and the helpers). ```cpp title="get_the_graph.cpp" #include #include #include #include #include #include #include using ClikaRT::DataType; using ClikaRT::Error; using ClikaRT::Tensor; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::OnnxExportOptions; using ClikaRT::io::OnnxModel; using ClikaRT::spec::TensorSpec; namespace fs = std::filesystem; namespace ops = ClikaRT::ops; namespace { constexpr std::int64_t kWidth = 37; // the length of every row // y = Relu(x * w + b) over x [batch, 37], whose first dim is dynamic and named batch. The model captures // w and b, two [37] vectors, so the graph holds them as its two weights: w[j] = (j - 18) / 8, from -2.25 // to 2.25, and b = 0.5. A trace returns the graph as built: every operator as written, nothing optimized // or finalized. ModelGraph trace_model() { const Tensor w = ops::div(ops::sub(ops::arange(0, kWidth, 1, DataType::Float32), 18), 8); const Tensor b = Tensor::full({kWidth}, 0.5, DataType::Float32); const auto model = [w, b](const std::vector& inputs) -> std::vector { return {ops::relu(ops::add(ops::mul(inputs[0], w), b))}; }; const std::vector signature = { TensorSpec("x", DataType::Float32, {TensorSpec::kDynamicDim, kWidth}, false, {"batch", ""})}; const std::vector outputs = {"y"}; return ClikaRT::graph::trace(model, signature, "elementwise", outputs); } // x [batch, 37], whose row i holds i + 1 in every column. Tensor input(std::int64_t batch) { return ops::cumsum(Tensor::ones({batch, kWidth}, DataType::Float32), 0); } // The sum of each row of y, separated by spaces. std::string row_sums(const Tensor& y) { std::string out; for (const float value : ops::sum(y, {1}).item_as_vec()) { char text[32]; std::snprintf(text, sizeof(text), "%g", static_cast(value)); out += (out.empty() ? "" : " ") + std::string(text); } return out; } // Each spec on its own line after `label`, as " []", a dynamic dim written as its name. void print_specs(const char* label, const std::vector& specs) { for (const TensorSpec& spec : specs) { std::string dims; for (std::size_t d = 0; d < spec.dims.size(); ++d) { const bool named = d < spec.dim_names.size() && !spec.dim_names[d].empty(); dims += (dims.empty() ? "" : ", ") + (named ? spec.dim_names[d] : std::to_string(spec.dims[d])); } std::printf("%s %s %s [%s]\n", label, spec.name.c_str(), ClikaRT::dtype::data_type_name(spec.dtype), dims.c_str()); } } // An empty directory for the files a program writes, under the system's temporary directory. fs::path fresh_directory(const std::string& name) { const fs::path dir = fs::temp_directory_path() / ("export_a_graph_to_onnx_" + name); fs::remove_all(dir); fs::create_directories(dir); return dir; } // The names of the files in dir, sorted and separated by spaces, or "none". std::string files_in(const fs::path& dir) { std::vector names; for (const fs::directory_entry& entry : fs::directory_iterator(dir)) { names.push_back(entry.path().filename().string()); } std::sort(names.begin(), names.end()); std::string out; for (const std::string& name : names) out += (out.empty() ? "" : " ") + name; return out.empty() ? "none" : out; } // A refusal as " | ". void print_refusal(const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } } // namespace int main() { ModelGraph graph = trace_model(); print_specs("input", graph.inputs()); print_specs("output", graph.outputs()); // input x Float32 [batch, 37] // output y Float32 [batch, 37] graph.finalize(); // a finalized graph runs std::printf("%s\n", row_sums(graph.run({input(2)}).front()).c_str()); // 31.625 52.5 return 0; } ``` ```python title="get_the_graph.py" import tempfile from pathlib import Path import clika_runtime as crt from clika_runtime.graph import ModelGraph, Transform WIDTH = 37 # the length of every row def trace_model() -> ModelGraph: # y = relu(x * w + b) over x [batch, 37], whose first dim is dynamic and named batch. The model captures # w and b, two [37] vectors, so the graph holds them as its two weights: w[j] = (j - 18) / 8, from -2.25 # to 2.25, and b = 0.5. A trace returns the graph as built: every operator as written, nothing optimized # or finalized. w = crt.div(crt.sub(crt.arange(0, WIDTH, dtype=crt.float32), 18), 8) b = crt.full((WIDTH,), 0.5, dtype=crt.float32) signature = [crt.TensorSpec("x", crt.float32, [-1, WIDTH], dim_names=["batch", ""])] return crt.trace(lambda inputs: [crt.relu(crt.add(crt.mul(inputs[0], w), b))], signature, output_names=["y"]).graph def rows(batch: int) -> crt.Tensor: # x [batch, 37], whose row i holds i + 1 in every column. return crt.cumsum(crt.ones(batch, WIDTH), 0) def row_sums(y: crt.Tensor) -> str: # The sum of each row of y, separated by spaces. return " ".join(f"{value:g}" for value in crt.sum(y, [1]).tolist()) def slots(listed: list[tuple[str, object, tuple[int, ...]]]) -> str: # (name, dtype, dims) slots as "name [dims]", separated by " | "; a dynamic dim reads -1. return " | ".join(f"{name} {list(dims)}" for name, _, dims in listed) def files_in(folder: Path) -> list[str]: # The names of the files in folder, sorted. return sorted(path.name for path in folder.iterdir()) def refusal(error: crt.ClikaRTError) -> str: # A refusal as " | ". return f"{error.code_name} | {error}" graph = trace_model() print(slots(graph.inputs()), "->", slots(graph.outputs())) # x [-1, 37] -> y [-1, 37] print(graph.node("x").output(0).spec().dim_names) # the dynamic dim keeps its name # ['batch', ''] graph.finalize() # a finalized graph runs print(row_sums(graph.run([rows(2)])[0])) # 31.625 52.5 ``` ## Export the graph and read it back `graph.export_onnx(path)` writes the model to `path`. `OnnxModel::open(path)` reads the file, and `compile()` turns it into a `ModelGraph` again, as built, so `finalize()` readies it to run ([Run an ONNX model](run-an-onnx-model.mdx) covers that side). The read-back graph has the same inputs and outputs, and its dynamic dimension carries the name `batch`, so it runs at any batch size. At batch 3 and at batch 5 the two graphs return the same sums. In Python, `export_onnx` takes the path as a `str` or any `os.PathLike`, such as a `pathlib.Path`, and its options as keyword arguments. `OnnxModel.open` takes the path as a `str`. ```cpp title="export_and_read_back.cpp" int main() { const fs::path dir = fresh_directory("export_and_read_back"); const std::string path = (dir / "elementwise.onnx").string(); ModelGraph graph = trace_model(); graph.export_onnx(path); // writes the graph as traced // OnnxModel::open reads the file, and compile() makes it a graph again: the same inputs and outputs, // with the dynamic dim under its name. ModelGraph back = OnnxModel::open(path).compile(); print_specs("input", back.inputs()); print_specs("output", back.outputs()); // input x Float32 [batch, 37] // output y Float32 [batch, 37] graph.finalize(); back.finalize(); for (const std::int64_t batch : {3, 5}) { // the dynamic dim takes any size const Tensor x = input(batch); std::printf("%s | %s\n", row_sums(graph.run({x}).front()).c_str(), row_sums(back.run({x}).front()).c_str()); } // 31.625 52.5 73.75 | 31.625 52.5 73.75 // 31.625 52.5 73.75 95 116.375 | 31.625 52.5 73.75 95 116.375 fs::remove_all(dir); return 0; } ``` ```python title="export_and_read_back.py" with tempfile.TemporaryDirectory() as tmp: path = Path(tmp) / "elementwise.onnx" graph = trace_model() graph.export_onnx(path) # writes the graph as traced; a str or any os.PathLike names the file # OnnxModel.open reads the file, and compile() makes it a graph again: the same inputs and outputs, with # the dynamic dim under its name. back = crt.io.OnnxModel.open(str(path)).compile() print(slots(back.inputs()), "->", slots(back.outputs())) # x [-1, 37] -> y [-1, 37] print(back.node("x").output(0).spec().dim_names) # ['batch', ''] graph.finalize() back.finalize() for batch in (3, 5): # the dynamic dim takes any size x = rows(batch) print(row_sums(graph.run([x])[0]), "|", row_sums(back.run([x])[0])) # 31.625 52.5 73.75 | 31.625 52.5 73.75 # 31.625 52.5 73.75 95 116.375 | 31.625 52.5 73.75 95 116.375 ``` ## Build the model in memory `OnnxModel::from_graph(graph)` builds, in memory, the model `export_onnx` writes, and writes nothing. The result is an `OnnxModel` like one that `open` returns: inspect its inputs, outputs and initializers, edit it, `optimize()` it, `compile()` it, or `save()` it when you want the file. It reads the graph's weights only when it is saved or compiled, so it holds no copy of them, and it keeps them alive for as long as it lives. Here the model reports the two weights as its initializers, and the graph it compiles to returns the same sums. In Python, `OnnxModel.from_graph` is a static method with the same keyword options, and `num_initializers` is a property. ```cpp title="in_memory.cpp" int main() { ModelGraph graph = trace_model(); // OnnxModel::from_graph builds the model export_onnx writes, in memory: nothing is written. OnnxModel model = OnnxModel::from_graph(graph); print_specs("input", model.inputs()); print_specs("output", model.outputs()); std::printf("%zu initializers\n", model.num_initializers()); // input x Float32 [batch, 37] // output y Float32 [batch, 37] // 2 initializers // It compiles like any other model, and save() writes it when you want the file. ModelGraph back = model.compile(); graph.finalize(); back.finalize(); const Tensor x = input(2); std::printf("%s | %s\n", row_sums(graph.run({x}).front()).c_str(), row_sums(back.run({x}).front()).c_str()); // 31.625 52.5 | 31.625 52.5 const fs::path dir = fresh_directory("in_memory"); model.save((dir / "elementwise.onnx").string()); std::printf("%s\n", files_in(dir).c_str()); // elementwise.onnx fs::remove_all(dir); return 0; } ``` ```python title="in_memory.py" graph = trace_model() # OnnxModel.from_graph builds the model export_onnx writes, in memory: nothing is written. model = crt.io.OnnxModel.from_graph(graph) print(slots(model.inputs()), "->", slots(model.outputs())) print(model.num_initializers, "initializers") # x [-1, 37] -> y [-1, 37] # 2 initializers # It compiles like any other model, and save() writes it when you want the file. back = model.compile() graph.finalize() back.finalize() x = rows(2) print(row_sums(graph.run([x])[0]), "|", row_sums(back.run([x])[0])) # 31.625 52.5 | 31.625 52.5 with tempfile.TemporaryDirectory() as tmp: model.save(str(Path(tmp) / "elementwise.onnx")) print(files_in(Path(tmp))) # ['elementwise.onnx'] ``` ## Choose where the weights go `OnnxExportOptions::save_as_external_data` decides where the weights go when the model is written. `true` puts them in a `_data` file beside the model, `false` keeps them inside the model file, and unset, the default, keeps them inside until they pass 2 GiB. `OnnxModel::open` finds the data file beside the model. Here the data file holds the two weights, 296 bytes, and the model read from it returns the same sums. In Python the option is `save_as_external_data=True`, `False` or `None`. ```cpp title="where_the_weights_go.cpp" int main() { const fs::path dir = fresh_directory("where_the_weights_go"); ModelGraph graph = trace_model(); // save_as_external_data = true writes the weights to a _data file beside the model. OnnxExportOptions beside; beside.save_as_external_data = true; graph.export_onnx((dir / "beside.onnx").string(), beside); // false keeps them in the model file, and so does the default until they pass 2 GiB. OnnxExportOptions inside; inside.save_as_external_data = false; graph.export_onnx((dir / "inside.onnx").string(), inside); graph.export_onnx((dir / "default.onnx").string()); std::printf("%s\n", files_in(dir).c_str()); // beside.onnx beside_data default.onnx inside.onnx std::printf("beside_data: %llu bytes\n", static_cast(fs::file_size(dir / "beside_data"))); // beside_data: 296 bytes: the two [37] float32 weights // OnnxModel::open finds the data file beside the model. ModelGraph back = OnnxModel::open((dir / "beside.onnx").string()).compile(); back.finalize(); std::printf("%s\n", row_sums(back.run({input(2)}).front()).c_str()); // 31.625 52.5 fs::remove_all(dir); return 0; } ``` ```python title="where_the_weights_go.py" with tempfile.TemporaryDirectory() as tmp: folder = Path(tmp) graph = trace_model() # save_as_external_data=True writes the weights to a _data file beside the model. graph.export_onnx(folder / "beside.onnx", save_as_external_data=True) # False keeps them in the model file, and so does None, the default, until they pass 2 GiB. graph.export_onnx(folder / "inside.onnx", save_as_external_data=False) graph.export_onnx(folder / "default.onnx") print(files_in(folder)) # ['beside.onnx', 'beside_data', 'default.onnx', 'inside.onnx'] print("beside_data:", (folder / "beside_data").stat().st_size, "bytes") # beside_data: 296 bytes: the two [37] float32 weights # OnnxModel.open finds the data file beside the model. back = crt.io.OnnxModel.open(str(folder / "beside.onnx")).compile() back.finalize() print(row_sums(back.run([rows(2)])[0])) # 31.625 52.5 ``` ## Set the opset and the other options `opset_version` is the `ai.onnx` opset the model imports, and 0 takes the library's default. `contrib_ops` lets the model use the `com.microsoft` operators, such as GroupQueryAttention and MatMulNBits, and imports that domain only when an operator uses it. `optimize` simplifies the written model, composing pairs of transposes that undo each other and dropping the nodes no output reads. `producer_name` is the producer the model records, and empty takes the library's. `OnnxModel::from_graph` takes the same options, and an operator the chosen opset cannot state is refused, naming the node. In Python the options are the keyword arguments of the same names. ```cpp title="options.cpp" int main() { const fs::path dir = fresh_directory("options"); const std::string path = (dir / "opset17.onnx").string(); ModelGraph graph = trace_model(); OnnxExportOptions options; options.opset_version = 17; // the ai.onnx opset the model imports; 0 takes the library's options.producer_name = "my-exporter"; // the producer the model records; empty takes the library's graph.export_onnx(path, options); std::printf("opset %lld\n", static_cast(OnnxModel::open(path).opset())); // opset 17 // contrib_ops = false keeps the model to the ai.onnx operators, and optimize = false writes it as // lowered, without simplifying it. OnnxExportOptions plain; plain.contrib_ops = false; plain.optimize = false; ModelGraph back = OnnxModel::from_graph(graph, plain).compile(); back.finalize(); std::printf("%s\n", row_sums(back.run({input(2)}).front()).c_str()); // 31.625 52.5 fs::remove_all(dir); return 0; } ``` ```python title="options.py" graph = trace_model() with tempfile.TemporaryDirectory() as tmp: path = Path(tmp) / "opset17.onnx" # opset_version is the ai.onnx opset the model imports (0 takes the library's), and producer_name the # producer the model records (empty takes the library's). graph.export_onnx(path, opset_version=17, producer_name="my-exporter") print("opset", crt.io.OnnxModel.open(str(path)).opset) # opset 17 # contrib_ops=False keeps the model to the ai.onnx operators, and optimize=False writes it as lowered, # without simplifying it. back = crt.io.OnnxModel.from_graph(graph, contrib_ops=False, optimize=False).compile() back.finalize() print(row_sums(back.run([rows(2)])[0])) # 31.625 52.5 ``` ## A graph that rewrites a buffer in place An ONNX model is a function of its inputs, so it cannot state a buffer that a graph changes in place. A trace of a step that writes a captured cache through an in-place operation surfaces the written cache as an extra output, `state_out_0`, and `export_onnx` refuses that graph with `NOT_IMPLEMENTED_FOR_PARAM`, naming the output that carries the update. A graph given a KV cache by `attach_kv_cache()` is refused the same way. Export the graph before `attach_kv_cache()`, or trace the step with the cache as an input and its update as an output. That graph exports, and its model reads the cache in and writes the update out. In Python the refusal raises `crt.UnsupportedError`. ```cpp title="state_buffer.cpp" int main() { const fs::path dir = fresh_directory("state_buffer"); const std::string path = (dir / "step.onnx").string(); // A step that adds x into a cache it captures, in place. The trace surfaces the written cache as the // output state_out_0, and every run rewrites the cache's buffer, which ONNX cannot state. const Tensor cache = Tensor::zeros({2, kWidth}, DataType::Float32); const auto in_place = [cache](const std::vector& inputs) -> std::vector { Tensor written = cache; // a handle on the live buffer ops::add_(written, inputs[0]); return {ops::relu(written)}; }; const std::vector x = {TensorSpec("x", DataType::Float32, {2, kWidth})}; const std::vector y = {"y"}; ModelGraph writes = ClikaRT::graph::trace(in_place, x, "step", y); print_specs("output", writes.outputs()); // output y Float32 [2, 37] // output state_out_0 Float32 [2, 37] try { writes.export_onnx(path); } catch (const Error& error) { print_refusal(error); } // NOT_IMPLEMENTED_FOR_PARAM | export_onnx: the graph rewrites a state buffer (its update is the output // 'state_out_0') in place, which ONNX cannot state; export the graph before attach_kv_cache, or trace // the step with the cache as an input and its update as an output std::printf("files: %s\n", files_in(dir).c_str()); // files: none // The same step with the cache as an input and its update as an output exports. const auto explicit_cache = [](const std::vector& inputs) -> std::vector { const Tensor updated = ops::add(inputs[1], inputs[0]); // cache + x return {ops::relu(updated), updated}; }; const std::vector x_and_cache = {TensorSpec("x", DataType::Float32, {2, kWidth}), TensorSpec("cache", DataType::Float32, {2, kWidth})}; const std::vector y_and_cache = {"y", "cache_out"}; ModelGraph step = ClikaRT::graph::trace(explicit_cache, x_and_cache, "step", y_and_cache); step.export_onnx(path); const OnnxModel model = OnnxModel::open(path); print_specs("input", model.inputs()); print_specs("output", model.outputs()); // input x Float32 [2, 37] // input cache Float32 [2, 37] // output y Float32 [2, 37] // output cache_out Float32 [2, 37] fs::remove_all(dir); return 0; } ``` ```python title="state_buffer.py" cache = crt.zeros(2, WIDTH) def in_place(inputs: list[crt.Tensor]) -> list[crt.Tensor]: # A step that adds x into the cache it captures, in place. The trace surfaces the written cache as the # output state_out_0, and every run rewrites the cache's buffer, which ONNX cannot state. crt.add_(cache, inputs[0]) return [crt.relu(cache)] def explicit_cache(inputs: list[crt.Tensor]) -> list[crt.Tensor]: # The same step with the cache as an input and its update as an output. updated = crt.add(inputs[1], inputs[0]) # cache + x return [crt.relu(updated), updated] with tempfile.TemporaryDirectory() as tmp: path = Path(tmp) / "step.onnx" writes = crt.trace(in_place, [crt.TensorSpec("x", crt.float32, [2, WIDTH])], output_names=["y"]).graph print(writes.output_names()) # ['y', 'state_out_0'] try: writes.export_onnx(path) except crt.UnsupportedError as error: print(refusal(error)) # NOT_IMPLEMENTED_FOR_PARAM | export_onnx: the graph rewrites a state buffer (its update is the output # 'state_out_0') in place, which ONNX cannot state; export the graph before attach_kv_cache, or trace # the step with the cache as an input and its update as an output print(files_in(Path(tmp))) # [] signature = [crt.TensorSpec("x", crt.float32, [2, WIDTH]), crt.TensorSpec("cache", crt.float32, [2, WIDTH])] step = crt.trace(explicit_cache, signature, output_names=["y", "cache_out"]).graph step.export_onnx(path) model = crt.io.OnnxModel.open(str(path)) print(slots(model.inputs()), "->", slots(model.outputs())) # x [2, 37] | cache [2, 37] -> y [2, 37] | cache_out [2, 37] ``` ## What the export refuses A refusal raises `ClikaRT::Error`, whose `code_name()` names the code, and writes nothing. Besides a buffer written in place, the export refuses: - an operator with no ONNX form, naming the node and the operator (`NOT_IMPLEMENTED`). A custom operator is one by nature, since its kernel is your own code ([Write a custom operator](write-a-custom-operator.mdx) builds one); - an opset the library does not know (`INVALID_ARGUMENT`); - a call while `optimize()`, `finalize()`, `attach_kv_cache()`, `to()` or an edit runs on the graph (`FAILED_PRECONDITION`, naming the call). This program makes one from inside a transform that `optimize()` runs ([Write a transform](write-a-transform.mdx)). `OnnxModel::from_graph` refuses the same way. In Python a refusal raises a subclass of `crt.ClikaRTError` whose `code_name` names the code: `crt.UnsupportedError` for a code that starts with `NOT_IMPLEMENTED`, and `crt.InvalidArgumentError` for `INVALID_ARGUMENT` and `FAILED_PRECONDITION`. A path of another type raises `TypeError`. Python has no surface for writing a custom kernel, so the custom operator's refusal is the C++ tab's, and the Python tab shows the others. ```cpp title="refusals.cpp" // A custom operator, y = 2x: its own shape rule and its own kernel, run on the host. class TimesTwo final : public ClikaRT::nn::Module { public: static std::shared_ptr make() { return std::shared_ptr(new TimesTwo()); } Tensor forward(const Tensor& x) const { return this->dispatch({&x, 1})[0]; } std::vector output_shapes( ClikaRT::Span inputs) const override { return {ClikaRT::FakeTensor(inputs[0].shape(), inputs[0].dtype(), inputs[0].stream())}; } void compute(ClikaRT::Span inputs, ClikaRT::Span outputs) const override { const float* x = static_cast(inputs[0].const_data_ptr()); float* y = static_cast(outputs[0].mutable_data_ptr()); const std::int64_t n = inputs[0].numel(); for (std::int64_t i = 0; i < n; ++i) y[i] = 2.0F * x[i]; } private: TimesTwo() = default; }; int main() { const fs::path dir = fresh_directory("refusals"); const std::string path = (dir / "refused.onnx").string(); // An operator with no ONNX form is refused by name. A custom operator is one by nature: its kernel is // your own code. const std::shared_ptr twice = TimesTwo::make(); const auto doubled = [twice](const std::vector& inputs) -> std::vector { return {twice->forward(inputs[0])}; }; const std::vector x = {TensorSpec("x", DataType::Float32, {2, kWidth})}; const std::vector y = {"y"}; ModelGraph custom = ClikaRT::graph::trace(doubled, x, "doubled", y); try { custom.export_onnx(path); } catch (const Error& error) { print_refusal(error); } // NOT_IMPLEMENTED | export_onnx: node 'Custom_1' (Custom) has no ONNX form: it is a user-defined operator // An opset the library does not know. ModelGraph graph = trace_model(); OnnxExportOptions unknown; unknown.opset_version = 99; try { graph.export_onnx(path, unknown); } catch (const Error& error) { print_refusal(error); } // INVALID_ARGUMENT | no compatible ONNX IR version for the given opset version // A call while another call holds the graph: here export_onnx from inside one of optimize()'s // transforms. const auto export_inside = [&path](ModelGraph& held) -> ClikaRT::Result { try { held.export_onnx(path); } catch (const Error& error) { print_refusal(error); } return {}; }; ClikaRT::graph::OptimizeOptions only_this; only_this.transforms = std::vector{ ClikaRT::graph::Transform::from_function("export_inside", export_inside)}; graph.optimize(only_this); // FAILED_PRECONDITION | export_onnx: optimize() is running on this graph; call export_onnx() after it // returns std::printf("files: %s\n", files_in(dir).c_str()); // files: none: a refusal writes nothing fs::remove_all(dir); return 0; } ``` ```python title="refusals.py" graph = trace_model() with tempfile.TemporaryDirectory() as tmp: path = Path(tmp) / "refused.onnx" # An opset the library does not know. try: graph.export_onnx(path, opset_version=99) except crt.InvalidArgumentError as error: print(refusal(error)) # INVALID_ARGUMENT | no compatible ONNX IR version for the given opset version # A call while another call holds the graph: here export_onnx from inside one of optimize()'s # transforms. def export_inside(held: ModelGraph) -> None: try: held.export_onnx(path) except crt.InvalidArgumentError as error: print(refusal(error)) graph.optimize([Transform("export_inside", export_inside)]) # FAILED_PRECONDITION | export_onnx: optimize() is running on this graph; call export_onnx() after it # returns # A path of another type. try: graph.export_onnx(3) except TypeError as error: print(error) # export_onnx(): path is a str or an os.PathLike, not int print(files_in(Path(tmp))) # []: a refusal writes nothing ``` --- # Handle errors by code Read a failure's three channels, branch on the stable code name, and switch policy on the coarse status, in every language. Source: https://docs.clika.io/clikart/how-to/handle-errors-by-code.md Your program needs to react to a runtime failure without parsing prose. Every ClikaRT failure carries the same three channels, in every language: a **coarse status** (the policy switch: retry, reject, fail), a **human message** (diagnostics for a log line), and a **stable code name** (the fine-grained machine channel, e.g. `TIMED_OUT`, `BACKEND_NOT_LOADED`). The rule this page exists for: branch on the code name, never on the message text. A failure you can act on (a bad shape, an option out of range, a missing file) reads as plain text in every build and names the operation and the values it refused. A message of the form `E` means a fault inside the runtime: its value is specific to the build that produced it, so it is nothing to record, compare or branch on; report it verbatim with the runtime version. The code name is machine-readable in every build. ## Read the three channels The program provokes one failure (a shape-mismatched matmul) and reads everything the error carries. ```cpp title="three_channels.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Error; using ClikaRT::Status; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; int main() { const Tensor a = Tensor::ones({3, 4}, DataType::Float32); const Tensor bad = Tensor::ones({5, 5}, DataType::Float32); try { const Tensor y = ops::matmul(a, bad); } catch (const Error& e) { // The three channels of every runtime failure. std::printf("status: %d\n", static_cast(e.status())); std::printf("message: %s\n", e.what()); std::printf("code name: %s\n", e.code_name().c_str()); // Branch on the code NAME: stable in every build, and the only // one of the three to compare against. A failure you can act on, // like this one, states the shapes in plain text in every build; a // message of the form E means a defect inside the runtime, // and its value is specific to the build that produced it. if (e.code_name() == "INVALID_ARGUMENT") std::printf("-> reject this request, keep serving\n"); // The coarse status is the POLICY switch. switch (e.status()) { case Status::Unavailable: /* shed under load: retry later */ break; case Status::InvalidArgument: /* caller bug: answer 400 */ break; default: /* fail loudly */ break; } } return 0; } ``` ```python title="three_channels.py" import numpy as np import clika_runtime as crt def code_name_of(err: BaseException) -> str: """The machine-readable code name a runtime error carries ('' if none): every ClikaRT exception exposes it as ``code_name``.""" return getattr(err, "code_name", "") def main() -> None: a = crt.tensor(np.ones((3, 4), dtype=np.float32)) bad = crt.tensor(np.ones((5, 5), dtype=np.float32)) try: a @ bad except RuntimeError as e: # Python raises a typed exception per status: str(e) is the message # alone, and the stable code name rides the exception's code_name. print(f"message: {e}") print(f"code name: {code_name_of(e)}") # Branch on the code NAME: stable in every build, and the only # one of the three to compare against. A failure you can act on, # like this one, states the shapes in plain text in every build; a # message of the form E means a defect inside the runtime, # and its value is specific to the build that produced it. if code_name_of(e) == "INVALID_ARGUMENT": print("-> reject this request, keep serving") if __name__ == "__main__": main() ``` ```text status: 1 message: ops::matmul: cannot contract a [3, 4] with b [5, 5]: a's last dim (4) must equal b's dim 0 (5) code name: INVALID_ARGUMENT -> reject this request, keep serving ``` Status `1` is `InvalidArgument`, the value of the C++ `Status` enum, whose numbering is append-only. The message names the operation and the two shapes it refused, in a release build as in a debug build. Python prints the status by name rather than by number (`status: InvalidArgument`) and is otherwise identical: the exception carries `.status`, its text and `.code_name`, and its class is the status, so a caller can switch on the type instead of reading the name. ## Branch on the code name, never the message The message exists for a human reading a log. An argument mistake reads as a sentence that names the operation and the values, in every build, but its wording can change between versions, so it is not a contract. A message of the form `E` is a runtime fault code: report it verbatim together with the runtime version; it is not an identifier to record, compare or branch on. The code name is the contract: a stable, append-only name for the fine-grained status the failure originated with, readable in every build. An empty code name is meaningful too: the failure did not originate inside the runtime (the C++ `Error` from your own code, a client-side load failure), so there is nothing machine-readable to branch on. ## A placement failure names its reason One message with a fixed shape: placing work on a device whose backend did not come up fails with the reason spelled out, `cannot allocate on device : `. The `` names what actually happened: the backend library is not shipped beside the runtime, it failed to load (with the loader's cause), or it loaded and brought up no device. The code name stays the branch point; the message is now worth logging as is. On a machine without the Metal backend: ```cpp title="backend_reason.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Error; using ClikaRT::Tensor; int main() { try { const Tensor t = Tensor::ones({2, 2}, DataType::Float32, Device::metal()); std::printf("placed: %s\n", t.to_string().c_str()); } catch (const Error& e) { std::printf("status: %d\n", static_cast(e.status())); std::printf("message: %s\n", e.what()); std::printf("code name: %s\n", e.code_name().c_str()); } return 0; } ``` ```text status: 4 message: cannot allocate on device Metal:0: The metal backend could not load: not shipped in this install. The metal backend is not part of this ClikaRT build. code name: BACKEND_NOT_LOADED ``` The same reason text reaches every binding through its language's error channel, and `clikart-cli` prints it when `--device` names a backend that cannot come up. {/* CERTIFICATION: both captures are real output of the program shown, run under the pinned release's clika_runtime cp313 manylinux wheel (sha256 checked against the digest published with it) in a fresh venv on linux x86_64, with no credential in the environment and no per-user license file present. The refusal texts are the runtime's own and are quoted unchanged. */} ## A licensing refusal carries `LICENSE_FAILED` ClikaRT runs compute under a license ([License the runtime](license-the-runtime.mdx) is where the credential goes). A call refused for licensing carries the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a credential past its end date, and the runtime writes one line to the console naming the refusal. On a process started with no credential at all: ```text [ClikaRT] [error] clikart license invalid — no ClikaRT license is configured for this process ``` ```python import numpy as np import clika_runtime as crt t = crt.tensor(np.ones((2, 3), dtype=np.float32)) try: (t + t).numpy() except crt.ClikaRTError as e: print(type(e).__name__, e.status, e.code_name, sep=" | ") print(e) ``` ```text InternalError | Internal | LICENSE_FAILED no ClikaRT license is configured for this process — a license key comes from the Licenses page of your project on the CLIKA platform; set CLIKA_RT_LICENSE to the key or to a file holding it, or install it with the clikart-license-init tool ``` The message names where the credential comes from and both places it can go, so it is worth logging as is. Licensing is where this page's rule earns its keep: the coarse status reads `Internal`, whose row in the table below sends a reader to file a report, and an unset variable is not something to file. Branch on `LICENSE_FAILED` before you consult the status. Making a tensor is not an operator, so a tensor still builds without a credential; the refusal arrives at the first operator. A process that has to survive a missing credential checks for one before it starts computing, or catches the refusal at its first operator. ## Switch policy on the coarse status The code name says what happened; the coarse status says what KIND of thing happened, which is usually all a policy needs: | Status | Meaning | Usual reaction | | --- | --- | --- | | `InvalidArgument` | the caller's request cannot be right | reject it (a 400 in [the serving guide](serve-over-http.mdx)) | | `NotFound` | a named thing does not exist | reject or fall back (a 404) | | `Unsupported` | this build or backend cannot do it | fail fast at startup, not per request | | `Internal` | the runtime's own invariant broke | log everything, file it | | `Unavailable` | transient: the work was shed under load | retry later (a 503) | ## Without exceptions C++ callers that would rather not pay for exceptions read the same three channels off a `Result`: `ok()`, `status()`, `message()`, `code_name()`. `CLIKART_TRY(...)` captures a throwing call as a `Result` where a failure is an expected outcome; [the serving guide](serve-over-http.mdx) uses it for exactly that on its 400 paths. The serving-side `nn::KVCache` is a worked example of the `Result` surface: its per-layer accessors (`keys_impl`, `values_impl`, `conv_state_impl`, `recurrent_state_impl`) return a failed `Result` for a layer index outside `[0, num_layers())` instead of crashing, and `num_layers()` reports the count without allocating: ```cpp title="kv_bounds.cpp" #include #include using ClikaRT::Result; using ClikaRT::Tensor; namespace nn = ClikaRT::nn; int main() { // The smallest real cache: two full-attention layers, one slot. nn::KVCacheConfig config; config.num_layers = 2; config.num_kv_heads = 1; config.head_dim = 4; config.max_seqs = 1; config.max_tokens_per_seq = 8; const nn::KVLayerSpec specs[2] = {}; // AttentionKV rows by default const nn::KVCache cache = nn::KVCache::make_impl(config, specs).value_or_throw(); // The per-layer accessors return a failed Result for a layer outside // [0, num_layers()); no exception, no crash. std::printf("layers: %d\n", cache.num_layers()); const Result in_range = cache.keys_impl(1); const Result out_range = cache.keys_impl(2); std::printf("keys(1): ok=%s\n", in_range.ok() ? "true" : "false"); std::printf("keys(2): ok=%s, message: %s, code name: %s\n", out_range.ok() ? "true" : "false", out_range.message().c_str(), out_range.code_name().c_str()); return 0; } ``` ```text layers: 2 keys(1): ok=true keys(2): ok=false, message: kv cache layer 2 is out of range: the cache has 2 layer(s), code name: INVALID_ARGUMENT ``` [Tutorial part 1](../getting-started/first-program/01-your-first-program.mdx) introduces the error model this page operationalizes; the python examples' errors chapter walks the same contract with an oracle per claim. --- # License the runtime Where ClikaRT and Modelverse read the credential your platform issues, per language and packaging, and what a refused call looks like. Source: https://docs.clika.io/clikart/how-to/license-the-runtime.md Every ClikaRT runtime needs a credential your platform issued for the project it ships under, and Modelverse, which runs on the same runtime, needs the same one. [Runtime licenses](/platform/concepts/runtime-licenses) is the concept page: why the credential exists, [the two kinds](/platform/concepts/runtime-licenses#the-two-kinds) (an online key, an offline bundle), the entitlements it carries, and [its lifecycle](/platform/concepts/runtime-licenses#lifecycle). This page is the runtime side of that loop: where the runtime looks for the credential, how to put it there from each language and packaging, and what a refused call carries. ## Where the runtime reads the credential The runtime looks in three places, in this order, and uses the first one it finds: 1. The environment variable `CLIKA_RT_LICENSE`, holding either the credential text itself or the path of a file that holds it. 2. The environment variable `CLIKA_LICENSE`, in the same two forms. 3. The per-user file `~/.clika/runtime/license` (on Windows `%USERPROFILE%\.clika\runtime\license`), which `clikart-license-init ` writes once; the tool also takes the credential from a file with `--from-file `, and prompts for it on standard input when given neither, which keeps it out of your shell history. It ships in the bundle's `bin/` and as a console script of the `clika-runtime` Python wheel. The tool writes the file readable by its owner alone and prints the file's path, the credential's kind (an online license's carrier bundle, which the platform validates on first use, or a signed offline bundle with its project and its expiry) and the permissions it set. It refuses to replace a file that is already there, so a second run asks for `--force` rather than overwriting what a machine is running on. The credential is the `CLIKA1-...` license key or license bundle your project's Licenses page reveals. A bare `clika_rk_...` API key is not a runtime credential and is refused everywhere, the license-init tool included. ```bash export CLIKA_RT_LICENSE=/path/to/credential # a file holding the credential export CLIKA_RT_LICENSE="$(cat /path/to/credential)" # or the credential text itself "$CLIKART_BUNDLE_DIR"/bin/clikart-license-init # or the per-user file, once per user ``` Put it in place before the program starts. ## Set it from each language A C++ program carries no code for the credential. The runtime reads the process environment, so set the variable in the shell, the service unit or the launcher that starts the program, or run `clikart-license-init` once on the machine and rely on the per-user file. ```bash export CLIKA_RT_LICENSE=/path/to/credential ./my_program ``` The `clika-runtime` wheel loads the runtime at import, so the variable is set before `import clika_runtime`: in the shell that starts Python, or in `os.environ` at the top of the program. ```python title="licensed.py" import os os.environ["CLIKA_RT_LICENSE"] = "/path/to/credential" # before the import: the wheel loads the runtime here import clika_runtime as crt x = crt.tensor([1.0, 2.0, 3.0]) ``` The wheel also installs `clikart-license-init` as a console script. After `clikart-license-init ` has written the per-user file, the program needs no variable at all. On Android there is no home directory for a per-user file and no `bin/` for the tool, so the credential arrives through the loader itself. `ClikaRtAndroid.load` takes it as its `license` argument; an explicit value wins over an ambient variable, and `null`, the default, leaves the environment alone. ```kotlin title="LicensedApp.kt" import android.app.Application import io.clika.runtime.ClikaRtAndroid class LicensedApp : Application() { override fun onCreate() { super.onCreate() // readCredential() is the app's own: it returns the text the platform // revealed, from wherever the app keeps its secrets. ClikaRtAndroid.load(this, license = readCredential()) } } ``` `ClikaRtAndroid.load` also hands the runtime the app's Context, which is how the license identifies the device from inside an app (the identity Android scopes to the app). A command the app launches instead, such as the runtime's own command line, has no Context: set `XDG_CACHE_HOME` (or `HOME`) to a directory of the app in the child's environment, and the runtime keeps a per-installation identity there. A credential with a hardware seat budget refuses a process that can identify neither the device nor its installation. On a desktop JVM, `ClikaRt.load()` needs nothing extra: the runtime reads the process environment, as in the C++ tab. ## What a refusal looks like A call the runtime may not serve fails through the three channels every runtime failure carries ([Handle errors by code](handle-errors-by-code.mdx)): the coarse status, a message, and the stable code name. The code name is `LICENSE_EXPIRED` when the credential's lifetime has run out and `LICENSE_FAILED` for every other refusal. It reads the same in every language, so branch on it, never on the message. A credential that lacks a grant refuses the call that needs it, with the code the [entitlements](/platform/concepts/runtime-licenses#entitlements) section lists. The coarse status on a licensing refusal reads `Internal`, whose usual policy is to log everything and file a report, which is the wrong reaction to an unset variable. Read the code name first; [Handle errors by code](handle-errors-by-code.mdx) shows the captured refusal and the branch. ## Online and offline An online key needs a route from the runtime to the platform; an offline bundle is verified locally against Clika's signature and needs none. Which to issue, and how each is revoked, is the concept page's [Lifecycle](/platform/concepts/runtime-licenses#lifecycle) section. ## Containers and CI Pass the credential through the environment of the service, never through the image. A `docker run -e`, an orchestrator secret exposed as an environment variable, or a CI job's secret variable all reach the runtime as the variable above; a credential written into an image layer travels with every copy of that image. Keep it out of logs the same way: log a refusal by its code name, and never echo the variable's value. ```bash docker run --rm -e CLIKA_RT_LICENSE="$CLIKA_RT_LICENSE" my-image ``` Issuing, revealing, rotating and revoking a credential is platform work: [Runtime licenses](/platform/concepts/runtime-licenses) is the concept and [the licensing commands](/platform/cli/licensing) drive it from the command line. [Quick install](../getting-started/installation.md) lists the credential among ClikaRT's prerequisites, and [Modelverse's quick install](/modelverse/getting-started/installation) among its own. `clikart-cli` runs on this runtime and reads the credential from the same places. --- # Load images and audio for inference Decode image files into [H, W, C] tensors and audio into resampled waveforms, then make them model-ready with ops. Source: https://docs.clika.io/clikart/how-to/load-images-and-audio.md Vision and audio models consume tensors, and `ClikaRT::io` gets you there from the files themselves: `load_image` decodes the common encoded formats (JPEG, PNG, GIF, BMP, TGA, PSD, HDR, PNM, PIC) into an `[H, W, C]` tensor, and `load_audio` decodes WAV, FLAC, MP3 and OGG (Vorbis) into a `Float32` waveform, resampling and remixing to the layout you name. Both also take an in-memory byte span, so the same calls serve an upload handler as well as a file path. A file outside those sets is refused by name (`Status::Unsupported`, code `NOT_IMPLEMENTED`): `image decode: photo.webp: WebP is not supported; this build decodes JPEG, PNG, ...`; a damaged file, or a file of the other family, is refused as `InvalidArgument` with the file and the container it carries named, so a batch with one bad input says which one. The write direction ships too: `encode_audio` and `save_audio` turn a waveform back into a WAV, and `save_image` writes PNG by default or JPEG, BMP and TGA when the path's extension asks for them ([Write audio back](#write-audio-back)). The runs below use a small RGB PNG (`input.png`, 8x6, a color gradient) and a half-second stereo WAV (`clip.wav`, 8 kHz, a 440 Hz tone on the left channel); any image or audio file of yours works the same. ## Decode an image and make it model-ready `peek_image` reads only the header, the cheap way to validate and route before decoding. `load_image` decodes to `[H, W, C]` with 8-bit pixels; `requested_channels` forces a channel count (3 collapses an alpha channel or expands grayscale, so one call normalizes mixed inputs). The rest of the preprocessing every vision model wants (float, scale, `CHW`, batch dimension) is three ops. ```cpp title="image_to_input.cpp" int main() { const std::string input = input_png(); const io::ImageInfo info = io::peek_image(input); std::printf("header: %lldx%lld, %lld channel(s)\n", static_cast(info.width), static_cast(info.height), static_cast(info.channels)); const Tensor img = io::load_image(input, /*requested_channels=*/3); std::printf("decoded: %s\n", img.to_string().c_str()); // Model-ready: float in [0,1], channels first, leading batch dim. const Tensor x = ops::unsqueeze( ops::permute(img.to(DataType::Float32) * (1.0 / 255.0), {2, 0, 1}), 0); std::printf("input: %s\n", x.to_string().c_str()); std::printf("mean pixel value: %.4f\n", ops::mean(x).item()); return 0; } ``` The image and audio decoders are part of the C++ API today; the C++ tab shows both. From python, a decoded image or waveform enters as a numpy array through `crt.tensor(array)`; the processor module then carries the resize/normalize and feature steps. ```text header: 8x6, 3 channel(s) decoded: Tensor(shape=[6, 8, 3], dtype=UInt8, device=CPU, numel=144, data=[0, 0, 128, 32, 0, 128, ...]) input: Tensor(shape=[1, 3, 6, 8], dtype=Float32, device=CPU, numel=144, data=[0, 0.1255, 0.251, 0.3765, 0.502, 0.6275, ...]) mean pixel value: 0.4444 ``` The decoded tensor is ordinary compute-engine currency from that point on: `.to(device)` moves it, and a model's forward takes it as-is. For bytes already in memory (an HTTP upload, an asset in an archive), pass a `Span` instead of the path; the `http_server` example's `06_image_compute` chapter runs exactly that flow behind an upload endpoint. ## Decode audio at the model's sample rate Audio checkpoints are trained at a fixed sample rate and channel count, and `load_audio` meets them at the file: `target_sample_rate` resamples and `target_channels` remixes during the decode, so the tensor that comes out is already the model's input layout. The result is an `AudioData`: `samples` (`Float32`, `[frames]` mono or `[frames, channels]`) plus the effective `AudioInfo`. ```cpp title="audio_to_input.cpp" int main() { const std::string clip = clip_wav(); // Whatever the file's native layout, ask for 16 kHz mono. const io::AudioData audio = io::load_audio(clip, /*target_sample_rate=*/16000, /*target_channels=*/1); std::printf("decoded: %d Hz, %d channel(s), %lld frames (%.2f s)\n", audio.info.sample_rate, audio.info.channels, static_cast(audio.info.frames), static_cast(audio.info.frames) / audio.info.sample_rate); std::printf("samples: %s\n", audio.samples.to_string().c_str()); // Loudness check: RMS over the waveform. const Tensor rms = ops::sqrt(ops::mean(audio.samples * audio.samples)); std::printf("rms = %.4f\n", rms.item()); return 0; } ``` This step is part of the C++ API today; the C++ tab shows it. ```text decoded: 16000 Hz, 1 channel(s), 8000 frames (0.50 s) samples: Tensor(shape=[8000], dtype=Float32, device=CPU, numel=8000, data=[0.0008917, 0.02581, 0.06115, 0.09333, 0.1175, 0.1378, ...]) rms = 0.1295 ``` Passing 0 for either target keeps the file's native rate or channel count, and `peek_audio`'s `AudioInfo` answers "what would this decode produce" without decoding. From here the waveform feeds a feature front end (`ops::` has the FFT-adjacent reductions and windowed views a spectrogram needs) or goes straight into a model that consumes raw samples. ## Write audio back The write direction is two calls: `encode_audio(samples, sample_rate)` returns the PCM16 WAV bytes of a waveform, and `save_audio(samples, sample_rate, path)` writes the same bytes to a file. A `Float16`, `BFloat16`, `Float32` or `Float64` waveform converts to 16-bit PCM; `Int16` is written as is. `[frames]` is mono and `[frames, channels]` is interleaved multi-channel, the same layouts the decoder produces, so a model's synthesized speech goes to disk with one call and an HTTP handler returns the encoded bytes without touching a file. The path's extension is the format request: `.wav` (or no extension) writes; a name asking for a container the writer does not produce (`reply.mp3`, `reply.ogg`) is refused by name with `Status::Unsupported` and nothing is written, so WAV bytes never land under a foreign suffix. ```cpp title="audio_write.cpp" #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace io = ClikaRT::io; namespace ops = ClikaRT::ops; int main() { // Half a second of a 440 Hz tone at 16 kHz: t = frame / rate, // s = 0.5 * sin(2 pi f t), a Float32 waveform in [-1, 1]. const int rate = 16000; const Tensor t = ops::arange(0.0, 8000.0, 1.0, DataType::Float32) * (1.0 / rate); const Tensor tone = ops::sin(t * (2.0 * 3.14159265358979 * 440.0)) * 0.5; // encode_audio returns the PCM16 WAV bytes; save_audio writes the same // bytes to a file. const std::vector wav_bytes = io::encode_audio(tone, rate); std::printf("encoded: %zu bytes of RIFF/WAVE\n", wav_bytes.size()); io::save_audio(tone, rate, "tone.wav"); // Round trip: decode what was written and check the layout survived. const io::AudioData back = io::load_audio("tone.wav", rate, /*target_channels=*/1); std::printf("reloaded: %d Hz, %d channel(s), %lld frames\n", back.info.sample_rate, back.info.channels, static_cast(back.info.frames)); const Tensor rms = ops::sqrt(ops::mean(back.samples * back.samples)); std::printf("rms = %.4f\n", rms.item()); return 0; } ``` ```python title="audio_write.py" import math import clika_runtime as crt # Half a second of a 440 Hz tone at 16 kHz, a Float32 waveform in [-1, 1]. rate = 16000 t = crt.arange(0.0, 8000.0, 1.0) * (1.0 / rate) tone = crt.sin(t * (2.0 * math.pi * 440.0)) * 0.5 # encode_audio returns the PCM16 WAV bytes; save_audio writes them to a file. wav_bytes = crt.io.encode_audio(tone, rate) print(f"encoded: {len(wav_bytes)} bytes of RIFF/WAVE") crt.io.save_audio(tone, rate, "tone.wav") # Round trip: decode what was written and check the layout survived. # load_audio returns (samples, sample_rate, channels, frames). samples, sample_rate, channels, frames = crt.io.load_audio("tone.wav", rate, 1) print(f"reloaded: {sample_rate} Hz, {channels} channel(s), {frames} frames") rms = crt.sqrt((samples * samples).mean()) print(f"rms = {rms.item():.4f}") ``` ```text encoded: 16044 bytes of RIFF/WAVE reloaded: 16000 Hz, 1 channel(s), 8000 frames rms = 0.3535 ``` The byte count is the 8000 frames as 16-bit PCM plus the 44-byte RIFF header, and the round-trip RMS is the tone's own (0.5 amplitude over root two). Decoding gives you pixels and samples; the model-specific half (resize, normalize, log-mel) is [the processors guide](preprocess-with-processors.mdx). Weights travel the same road in [the GGUF guide](load-quantized-weights.mdx); the bundle's `io` example walks NumPy and safetensors round-trips in its earlier chapters. --- # Load quantized weights from a GGUF file Read a block-quantized checkpoint, keep the weights packed, and serve them through QLinearWoQ at their on-disk footprint. Source: https://docs.clika.io/clikart/how-to/load-quantized-weights.md You have a block-quantized GGUF checkpoint and want to run it without inflating it to dense floats. ClikaRT loads the file as-is: `io::load_gguf` returns the tensors plus the file's metadata, quantized entries keep their packed bytes, and `nn::QLinearWoQ` runs the matmul off the packed form. The checkpoint's on-disk footprint is its in-memory footprint. The programs below use [SmolLM2-135M-Instruct](https://huggingface.co/bartowski/SmolLM2-135M-Instruct-GGUF) (Apache-2.0, 145 MB in Q8_0), small enough to download in a minute. Any `.gguf` file works; nothing here depends on the architecture. ```bash curl -LO "https://huggingface.co/bartowski/SmolLM2-135M-Instruct-GGUF/resolve/main/SmolLM2-135M-Instruct-Q8_0.gguf" ``` Each program is a complete `main.cpp`; build them like any bundle consumer ([tutorial part 1](../getting-started/first-program/01-your-first-program.mdx) has the four-line CMake project). ## Load the file and read its metadata `io::load_gguf` returns a `GgufModel`: the weights as a `NamedTensors` map and the file's metadata as one `Json` object. The loader takes the target device (CPU by default), and the tensor payloads are mmap-backed, so loading is cheap. ```cpp title="inspect_metadata.cpp" int main() { const io::GgufModel model = io::load_gguf(howto::gguf_fixture()); std::printf("tensors %zu\n", model.tensors.size()); std::printf("metadata %zu keys\n", model.metadata.size()); for (const char* key : {"general.architecture", "general.size_label", "llama.block_count", "llama.embedding_length"}) { std::printf(" %-24s = %s\n", key, model.metadata.at(key).dump(0).c_str()); } return 0; } ``` ```python title="inspect_metadata.py" def main() -> None: tensors, metadata_json = crt.io.load_gguf(gguf_path()) metadata = json.loads(metadata_json) print(f"tensors {len(tensors)}") print(f"metadata {len(metadata)} keys") for key in ("general.architecture", "general.size_label", "llama.block_count", "llama.embedding_length"): print(f" {key:<24} = {json.dumps(metadata[key])}") if __name__ == "__main__": main() ``` ```text tensors 272 metadata 37 keys general.architecture = "llama" general.size_label = "135M" llama.block_count = 30 llama.embedding_length = 576 ``` The metadata carries everything the file knows about itself: architecture, hyperparameters, tokenizer configuration (`tokenizer.chat_template` included; [the chat-template guide](tokenize-and-chat-templates.mdx) picks that up). `at(key)` raises `ClikaRT::Error` on a missing key; probe with `contains` when a key is optional. ## What a quantized weight is in memory A quantized entry rides a plain `Tensor` whose element data is the packed block stream, exactly as it sits in the file. `is_quantized()` separates those from the dense entries (norms and embeddings stay floating point in most files). `quantized_view` names what the payload is: the packed bytes, the scheme, and the logical element shape they encode. ```cpp title="inspect_weights.cpp" int main() { const io::GgufModel model = io::load_gguf(howto::gguf_fixture()); int quantized = 0, dense = 0; model.tensors.for_each([&](std::string_view, const Tensor& t) { if (t.is_quantized()) ++quantized; else ++dense; }); std::printf("%d quantized, %d dense\n", quantized, dense); const QTensor q = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight")); std::printf("scheme %s\n", q.scheme.c_str()); std::printf("logical [%lld, %lld]\n", static_cast(q.logical_shape[0]), static_cast(q.logical_shape[1])); std::printf("payload %s\n", q.payload.to_string().c_str()); const Tensor norm = model.tensors.get("blk.0.attn_norm.weight"); std::printf("dense %s\n", norm.to_string().c_str()); return 0; } ``` ```python title="inspect_weights.py" def main() -> None: tensors, _ = crt.io.load_gguf(gguf_path()) quantized = sum(1 for t in tensors.values() if t.is_quantized) print(f"{quantized} quantized, {len(tensors) - quantized} dense") q = crt.quantized_view(tensors["blk.0.ffn_up.weight"]) print(f"scheme {q.scheme}") print(f"logical [{q.logical_shape[0]}, {q.logical_shape[1]}]") print(f"payload {q.payload}") print(f"dense {tensors['blk.0.attn_norm.weight']}") if __name__ == "__main__": main() ``` ```text 211 quantized, 61 dense scheme GGUF_Q8_0 logical [576, 1536] payload Tensor(shape=[1536, 612], dtype=UInt8, device=CPU, numel=940032, quantized=true, mmap=checkpoint:SmolLM2-135M-Instruct-Q8_0.gguf+33749344, data=[232, 27, 234, 24, 66, 230, ...]) dense Tensor(shape=[576], dtype=Float32, device=CPU, numel=576, mmap=checkpoint:SmolLM2-135M-Instruct-Q8_0.gguf+31866976, data=[0.01398, 0.0238, -0.01978, -0.03027, -0.01965, -0.03516, ...]) ``` Two facts to keep. The **logical shape is `[in_features, out_features]`**, the row-contiguous dimension first; that is the orientation every quantized consumer below expects. The **payload is UInt8 `[rows, row_bytes]`**: for Q8_0, each row of 576 elements packs into blocks of 32 (one fp16 scale plus 32 int8 codes each), 34 bytes per block. ## Serve it packed with QLinearWoQ `nn::QLinearWoQ` is a `Linear` over a quantized weight. The weight stays packed for the module's lifetime; the first `forward` reshapes the payload once into the backend's kernel layout, and every later call runs the quantized-weight matmul off that. Nothing is ever materialized dense. The lifecycle is the same as every weight-bearing `nn` module: `make(in, out)` declares the slots, `set_weights` binds the loaded tensor, `forward` runs. ```cpp title="serve_packed.cpp" int main() { const io::GgufModel model = io::load_gguf(howto::gguf_fixture()); const QTensor w = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight")); const std::int64_t in = w.logical_shape[0], out = w.logical_shape[1]; const std::shared_ptr ffn_up = QLinearWoQ::make(in, out); ffn_up->set_weights(w); const Tensor x = Tensor::full({1, in}, 0.01, DataType::Float32); const Tensor y = ffn_up->forward(x); // first call packs, later calls reuse std::printf("y = %s\n", y.to_string().c_str()); return 0; } ``` ```python title="run_quantized.py" def main() -> None: tensors, _ = crt.io.load_gguf(gguf_path()) q = crt.quantized_view(tensors["blk.0.ffn_up.weight"]) in_f, out_f = q.logical_shape # Dequantize to the scheme's float target and run dense math on it. # Serving the matmul off the PACKED form (no dense copy at rest) is # done through the quantized modules of the C++ API; the C++ tab # shows it. w = crt.dequantize(q) x = crt.full((1, in_f), 0.01) y = F.linear(x, w) print(f"y = {y}") if __name__ == "__main__": main() ``` ```text y = Tensor(shape=[1, 1536], dtype=Float32, device=CPU, numel=1536, data=[0.03397, -0.02917, 0.0793, 0.03382, -0.02785, 0.005078, ...]) ``` One orientation trap: `QLinearWoQ` consumes the `[in, out]` logical shape that `load_gguf` produces, while dense `Linear` takes the HuggingFace `[out, in]` layout. Bind the GGUF entry as-is; do not transpose. Moving the module (`ffn_up->to(...)`) re-packs the weight on the target device, and a dtype move is refused: the weight stays quantized at rest. ## Inspect a weight by dequantizing `ops::dequantize` decodes a packed weight into a dense tensor, `[out, in]` row-major. It is the inspection and tooling path, not the serving path; use it to eyeball values or to check a conversion. The program decodes the same weight, checks the packed forward against the dense one, and prints what staying packed saves. ```cpp title="check_dense.cpp" int main() { const io::GgufModel model = io::load_gguf(howto::gguf_fixture()); const QTensor w = ClikaRT::quantized_view(model.tensors.get("blk.0.ffn_up.weight")); const std::int64_t in = w.logical_shape[0], out = w.logical_shape[1]; const Tensor dense = ops::dequantize(w); // [out, in], the scheme's float target std::printf("dense = %s\n", dense.to_string().c_str()); // The packed path and the dense path compute the same values. const std::shared_ptr packed = QLinearWoQ::make(in, out); packed->set_weights(w); const Tensor x = Tensor::full({1, in}, 0.01, DataType::Float32); const Tensor y = packed->forward(x); const Tensor yr = ops::linear(x, dense); std::printf("max |packed - dense| over %lld outputs: %g\n", static_cast(y.numel()), ops::amax(ops::abs(ops::sub(y, yr))).item()); std::printf("bytes: dense %zu, packed %zu (%.2fx)\n", dense.nbytes(), w.payload.nbytes(), static_cast(dense.nbytes()) / static_cast(w.payload.nbytes())); return 0; } ``` ```python title="check_footprint.py" def main() -> None: tensors, _ = crt.io.load_gguf(gguf_path()) q = crt.quantized_view(tensors["blk.0.ffn_up.weight"]) dense = crt.dequantize(q) # [out, in], the scheme's float target print(f"dense = {dense}") # The packed payload is the checkpoint's own byte stream; inflating it # to dense floats shows what staying packed saves. (The packed-vs-dense # VALUE check runs where the packed matmul lives: the C++ tab.) dense_bytes = dense.numpy().size * 4 packed_bytes = q.payload.numpy().size print(f"bytes: dense {dense_bytes}, packed {packed_bytes} " f"({dense_bytes / packed_bytes:.2f}x)") if __name__ == "__main__": main() ``` ```text dense = Tensor(shape=[1536, 576], dtype=Float32, device=CPU, numel=884736, data=[-0.08493, 0.09265, 0.2548, -0.1004, 0.1699, -0.193, ...]) max |packed - dense| over 1536 outputs: 7.18981e-07 bytes: dense 3538944, packed 940032 (3.76x) ``` The packed forward matches the dequantize-then-`ops::linear` reference to float rounding, at 3.76x fewer bytes for Q8_0 (Q4 and Q5 schemes save more). ## When the bytes did not come from GGUF A packed payload from any other source gets the same treatment through `make_quantized(payload, scheme, logical_shape)`: a UInt8 CPU tensor of packed block rows, the scheme's name (`"GGUF_Q8_0"`, `"MXFP4_E8M0"`), and the element shape it encodes, row-contiguous dimension first. Checkpoints that ship MXFP4 as a split pair, 16 nibble-packed code bytes plus one E8M0 scale byte per 32-element group, go through `make_quantized_mxfp4(blocks, scales, logical_shape)`, which weaves the pair into the packed row layout the runtime consumes. Either way the result is the same `QTensor` the programs above served. You can now open any GGUF checkpoint, say what every entry is, and serve its weights at their packed size. The bundle's `io` example (chapter `02_gguf`) inspects arbitrary files from the command line, and the [runtime example](../examples.md) wraps modules like `QLinearWoQ` into sessions and batching for serving. --- # Merge and split Join ModelGraphs into one graph that runs them all, side by side, with the outputs of one feeding the inputs of another, or across two devices; cut a graph into parts at its cheapest cut or by a function, extract the operators between values, and merge the parts back. Source: https://docs.clika.io/clikart/how-to/merge-and-split.md {/* Every block is a program under examples//howto/merge_and_split/: the first block of each tab is get_the_graphs whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output, and it reports across_devices as blocked on a host with no accelerator. */} `merge` joins model graphs into one graph that runs them all. The graphs sit side by side, or the outputs of one feed the inputs of another, on one device or across two. `split` cuts one graph into parts that `merge` joins back, at the cheapest cut through its values or by a function that names each operator's part, and `extract` takes out the operators that compute some values from others. Each call consumes the graphs it is given and returns new ones. The calls are available from C++ and Python, under the same names. ## Graphs to merge and split A trace returns the graph as built, with every operator as written and nothing optimized or finalized, so it is ready to merge or split. The examples on this page use three models over an input of shape [2, 3]. `rectify` computes `h = Relu(x)`, `twice` computes `y = x + x`, and `pool` negates the sum of each row of `Relu(x)`, which pools the [2, 3] input to [2, 1]. `trace_graph` traces a model under the input and output names a section needs. `label` names a node by its input name, or by its operator's code after the part its name starts with, so a merged graph reads `encoder/Relu`. `labels` lists a graph's nodes, and `io` its input names and then its output names. This program runs each graph once. A finalized graph runs, and merge and split refuse it, so the other programs merge and split their graphs before `finalize()`. ```cpp title="get_the_graphs.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Error; using ClikaRT::Tensor; using ClikaRT::graph::Connection; using ClikaRT::graph::GraphPart; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::NameKind; using ClikaRT::graph::Node; using ClikaRT::graph::NodeKind; using ClikaRT::graph::Rename; namespace ops = ClikaRT::ops; namespace { // h = Relu(x). std::vector rectify(const std::vector& inputs) { return {ops::relu(inputs[0])}; } // y = x + x. std::vector twice(const std::vector& inputs) { return {ops::add(inputs[0], inputs[0])}; } // y = -(the sum of each row of Relu(x)): the [2, 3] input pools to [2, 1]. std::vector pool(const std::vector& inputs) { return {ops::neg(ops::sum(ops::relu(inputs[0]), {1}, true))}; } // A trace of `model` over one float32 [2, 3] input, named `input`, with its output named `output`. // A trace returns the graph as built: every operator as written, nothing optimized or finalized. ModelGraph trace_graph(const ClikaRT::graph::TraceFunction& model, const std::string& input, const std::string& output) { const std::vector signature = {{input, DataType::Float32, {2, 3}}}; const std::vector outputs = {output}; return ClikaRT::graph::trace(model, signature, "merge_and_split", outputs); } // A node's label: an input's name, or an operator's code after the part its name starts with. std::string label(const Node& node) { if (node.kind() == NodeKind::Input) return node.name(); const std::string name = node.name(); const std::size_t slash = name.find('/'); const std::string part = slash == std::string::npos ? std::string() : name.substr(0, slash + 1); return part + std::string(ClikaRT::graph::op_code_name(node.op_code())); } // The labels of a graph's nodes, in order, separated by spaces. std::string labels(const ModelGraph& graph) { std::string out; for (const Node& node : graph.nodes()) out += (out.empty() ? "" : " ") + label(node); return out; } // Names separated by spaces. std::string names(const std::vector& list) { std::string out; for (const std::string& name : list) out += (out.empty() ? "" : " ") + name; return out; } // A graph's input names, then its output names. std::string io(const ModelGraph& graph) { return names(graph.input_names()) + " -> " + names(graph.output_names()); } // What a rename names. const char* kind_name(NameKind kind) { switch (kind) { case NameKind::Node: return "node"; case NameKind::Input: return "input"; case NameKind::Output: return "output"; } return "node"; } // The input every program runs, [2, 3]. Tensor input() { const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F}; return Tensor::from_data(x.data(), {2, 3}, DataType::Float32); } // A tensor's values, flattened and separated by spaces. std::string values(const Tensor& tensor) { std::string out; for (const float value : tensor.reshape({-1}).item_as_vec()) { char text[32]; std::snprintf(text, sizeof(text), "%g", static_cast(value)); out += (out.empty() ? "" : " ") + std::string(text); } return out; } } // namespace int main() { ModelGraph encoder = trace_graph(rectify, "x", "h"); ModelGraph head = trace_graph(twice, "h", "y"); ModelGraph pooled = trace_graph(pool, "x", "y"); for (ModelGraph* graph : {&encoder, &head, &pooled}) { std::printf("%s | %s\n", labels(*graph).c_str(), io(*graph).c_str()); graph->finalize(); // a finalized graph runs, and merge and split refuse it std::printf("%s\n", values(graph->run({input()}).front()).c_str()); } // x Relu | x -> h // 1 0 3 0 5 0 // h Add | h -> y // 2 -4 6 -8 10 -12 // x Relu Sum Neg | x -> y // -4 -5 return 0; } ``` ```python title="get_the_graphs.py" from collections.abc import Callable import clika_runtime as crt from clika_runtime.graph import ModelGraph, NameKind, Node, NodeKind Model = Callable[[list[crt.Tensor]], list[crt.Tensor]] def rectify(inputs: list[crt.Tensor]) -> list[crt.Tensor]: return [crt.relu(inputs[0])] # h = Relu(x) def twice(inputs: list[crt.Tensor]) -> list[crt.Tensor]: return [crt.add(inputs[0], inputs[0])] # y = x + x def pool(inputs: list[crt.Tensor]) -> list[crt.Tensor]: # y = -(the sum of each row of Relu(x)): the [2, 3] input pools to [2, 1]. return [crt.neg(crt.sum(crt.relu(inputs[0]), [1], True))] def trace_graph(model: Model, input_name: str, output_name: str) -> ModelGraph: # A trace of `model` over one float32 [2, 3] input, named `input_name`, with its output named `output_name`. # A trace returns the graph as built: every operator as written, nothing optimized or finalized. signature = [crt.TensorSpec(input_name, crt.float32, [2, 3])] return crt.trace(model, signature, output_names=[output_name]).graph def label(node: Node) -> str: # A node's label: an input's name, or an operator's code after the part its name starts with. if node.kind == NodeKind.Input: return node.name part, slash, _ = node.name.partition("/") return (part + slash if slash else "") + node.op_code.name def labels(graph: ModelGraph) -> str: return " ".join(label(node) for node in graph.nodes()) def io(graph: ModelGraph) -> str: # A graph's input names, then its output names. return " ".join(graph.input_names()) + " -> " + " ".join(graph.output_names()) def kind_name(kind: NameKind) -> str: # What a rename names: a node, an input or an output. return kind.name.lower() def x() -> crt.Tensor: # The input every program runs, [2, 3]. return crt.tensor([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], dtype=crt.float32) def values(tensor: crt.Tensor) -> str: # A tensor's values, flattened and separated by spaces. return " ".join(f"{value:g}" for value in crt.to(tensor, "cpu").reshape(-1).tolist()) encoder = trace_graph(rectify, "x", "h") head = trace_graph(twice, "h", "y") pooled = trace_graph(pool, "x", "y") for graph in (encoder, head, pooled): print(labels(graph), "|", io(graph)) graph.finalize() # a finalized graph runs, and merge and split refuse it print(values(graph.run([x()])[0])) # x Relu | x -> h # 1 0 3 0 5 0 # h Add | h -> y # 2 -4 6 -8 10 -12 # x Relu Sum Neg | x -> y # -4 -5 ``` Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, the three models and the helpers). ## Merge graphs side by side `graph::merge(parts)` takes a list of `GraphPart`s, each a graph and the name the merged graph knows it by. Parts that no connection joins sit side by side. The merged graph takes each part's inputs and outputs, parts in the order given, and names every operator of part `p` as `p/`. `MergeOptions::io_names` decides the names of the inputs and outputs. Under `IoNames::PrefixOnCollision`, the default, a name two parts both use becomes `/` and every other name stays as it is. `IoNames::PrefixAlways` prefixes every name, and `IoNames::Keep` keeps every name and refuses a name two parts both use. `MergedGraph::renames` lists each name the merge changed as a `Rename`, which holds its part, whether it names a node, an input or an output, and the name before and after. Here the encoder and the head both read an input named `x`, so the default prefixes the two inputs and keeps the two outputs. In Python, `graph.merge` takes the parts as a dict from name to graph, or as a list of `(name, graph)` pairs, and returns `(merged, renames)`. `io_names` takes the mode's name, such as `"prefix_always"`, and a `Rename`'s fields are `part`, `kind`, `from_` and `to`. ```cpp title="side_by_side.cpp" int main() { // An encoder and a head that both read an input named x. const auto two_parts = [] { std::vector parts; parts.push_back({"encoder", trace_graph(rectify, "x", "h")}); parts.push_back({"head", trace_graph(twice, "x", "y")}); return parts; }; // IoNames::PrefixOnCollision, the default: a name both parts use takes its part's name as a prefix. std::vector parts = two_parts(); const ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts); std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str()); // encoder/x head/x encoder/Relu head/Add | encoder/x head/x -> h y for (const Rename& rename : merged.renames) { // every name the merge changed std::printf("%s %s %s -> %s\n", rename.part.c_str(), kind_name(rename.kind), rename.from.c_str(), rename.to.c_str()); } // encoder input x -> encoder/x // head input x -> head/x // encoder node Relu_1 -> encoder/Relu_1 // head node Add_1 -> head/Add_1 // IoNames::PrefixAlways: every input and output name takes the prefix. ClikaRT::graph::MergeOptions always; always.io_names = ClikaRT::graph::IoNames::PrefixAlways; parts = two_parts(); std::printf("%s\n", io(ClikaRT::graph::merge(parts, {}, always).graph).c_str()); // encoder/x head/x -> encoder/h head/y // IoNames::Keep keeps every name, so it refuses a name both parts use. ClikaRT::graph::MergeOptions keep; keep.io_names = ClikaRT::graph::IoNames::Keep; parts = two_parts(); try { ClikaRT::graph::merge(parts, {}, keep); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } // INVALID_ARGUMENT | merge: parts 'encoder' and 'head' both have an input named 'x', which IoNames::Keep // keeps; merge with IoNames::PrefixOnCollision or PrefixAlways std::printf("%s | %s\n", io(parts[0].graph).c_str(), io(parts[1].graph).c_str()); // x -> h | x -> y: a refused merge leaves every part as it was return 0; } ``` ```python title="side_by_side.py" def two_parts() -> list[tuple[str, ModelGraph]]: # An encoder and a head that both read an input named x. return [("encoder", trace_graph(rectify, "x", "h")), ("head", trace_graph(twice, "x", "y"))] # io_names="prefix_on_collision", the default: a name both parts use takes its part's name as a prefix. merged, renames = crt.graph.merge(two_parts()) print(labels(merged), "|", io(merged)) # encoder/x head/x encoder/Relu head/Add | encoder/x head/x -> h y for rename in renames: # every name the merge changed print(rename.part, kind_name(rename.kind), rename.from_, "->", rename.to) # encoder input x -> encoder/x # head input x -> head/x # encoder node Relu_1 -> encoder/Relu_1 # head node Add_1 -> head/Add_1 # io_names="prefix_always": every input and output name takes the prefix. always, _ = crt.graph.merge(two_parts(), io_names="prefix_always") print(io(always)) # encoder/x head/x -> encoder/h head/y # io_names="keep" keeps every name, so it refuses a name both parts use. parts = two_parts() try: crt.graph.merge(parts, io_names="keep") except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | merge: parts 'encoder' and 'head' both have an input named 'x', which IoNames::Keep # keeps; merge with IoNames::PrefixOnCollision or PrefixAlways print(io(parts[0][1]), "|", io(parts[1][1])) # x -> h | x -> y: a refused merge leaves every part as it was ``` Merge graphs side by side to run several models as one graph, with one `finalize()` and one `run()`. ## Feed one graph into another A `Connection` feeds an output of one part into an input of another. `from` names the part and its output, `to` names the part and its input, and a connected output or input is not one of the merged graph's. Here the encoder's `h` feeds the head's `h`, so the merged graph reads `x`, returns `y`, and computes what running the two graphs one after the other computes. The two ends of a connection have one dtype and one rank, and every size both of them fix is the same. In Python, a connection is a pair of strings, `("encoder.h", "head.h")`. ```cpp title="pipeline.cpp" int main() { std::vector parts; parts.push_back({"encoder", trace_graph(rectify, "x", "h")}); parts.push_back({"head", trace_graph(twice, "h", "y")}); // The encoder's output h feeds the head's input h, and neither is the merged graph's any longer. const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}}; ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections); std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str()); // x encoder/Relu head/Add | x -> y merged.graph.optimize(); // optional: the graph optimizer runs over the merged graph merged.graph.finalize(); // The same input through the two graphs one after the other, for comparison. ModelGraph encoder = trace_graph(rectify, "x", "h"); ModelGraph head = trace_graph(twice, "h", "y"); encoder.finalize(); head.finalize(); std::printf("%s | %s\n", values(merged.graph.run({input()}).front()).c_str(), values(head.run(encoder.run({input()})).front()).c_str()); // 2 0 6 0 10 0 | 2 0 6 0 10 0 return 0; } ``` ```python title="pipeline.py" # The encoder's output h feeds the head's input h, and neither is the merged graph's any longer. parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")} merged, _ = crt.graph.merge(parts, [("encoder.h", "head.h")]) print(labels(merged), "|", io(merged)) # x encoder/Relu head/Add | x -> y merged.optimize() # optional: the graph optimizer runs over the merged graph merged.finalize() # The same input through the two graphs one after the other, for comparison. encoder = trace_graph(rectify, "x", "h") head = trace_graph(twice, "h", "y") encoder.finalize() head.finalize() print(values(merged.run([x()])[0]), "|", values(head.run(encoder.run([x()]))[0])) # 2 0 6 0 10 0 | 2 0 6 0 10 0 ``` Feed one graph into another to serve a model with its preprocessing or its post-processing as one graph. ## Merge a second graph into this one `ModelGraph::merge(second, io_map)` merges `second` into the graph it is called on. Each `io_map` entry feeds one of this graph's outputs into one of `second`'s inputs. This graph keeps every name, and a name of `second` that this graph already uses takes the first free `_`. The call returns those renames, each under the part `second`, and it consumes `second`. Everything else follows `graph::merge`, and a connection between two devices moves its value. Here both graphs are traces of `rectify`, so the second Relu takes a new name. In Python, `first.merge(second, [("h", "x")])` returns the renames. ```cpp title="merge_into.cpp" int main() { ModelGraph first = trace_graph(rectify, "x", "h"); ModelGraph second = trace_graph(rectify, "x", "h"); // first's output h feeds second's input x. first keeps every name, and a name of second that // first already uses takes the first free _. const std::vector renames = first.merge(second, {{"h", "x"}}); std::printf("%s | %s\n", labels(first).c_str(), io(first).c_str()); // x Relu Relu | x -> h for (const Rename& rename : renames) { std::printf("%s %s %s -> %s\n", rename.part.c_str(), kind_name(rename.kind), rename.from.c_str(), rename.to.c_str()); } // second node Relu_1 -> Relu_1_1 std::printf("%zu %zu\n", second.input_names().size(), second.output_names().size()); // 0 0: second is consumed first.finalize(); std::printf("%s\n", values(first.run({input()}).front()).c_str()); // 1 0 3 0 5 0: Relu(Relu(x)) return 0; } ``` ```python title="merge_into.py" first = trace_graph(rectify, "x", "h") second = trace_graph(rectify, "x", "h") # first's output h feeds second's input x. first keeps every name, and a name of second that # first already uses takes the first free _. renames = first.merge(second, [("h", "x")]) print(labels(first), "|", io(first)) # x Relu Relu | x -> h for rename in renames: print(rename.part, kind_name(rename.kind), rename.from_, "->", rename.to) # second node Relu_1 -> Relu_1_1 print(len(second.input_names()), len(second.output_names())) # 0 0: second is consumed first.finalize() print(values(first.run([x()])[0])) # 1 0 3 0 5 0: Relu(Relu(x)) ``` Merge a second graph into a graph to grow it in place while every name it has stays as it is. ## Connect graphs on two devices A part's values sit on the device `to()` sent the part to, and otherwise where they are computed. When a connection's two ends sit on two devices, `CrossDevice::Move`, the default, moves the value onto the input's device, and `CrossDevice::Refuse` refuses the merge, naming the connection and both devices. When the parts do not share one device, the merge places each part on its device before it builds, and the merged graph keeps every part where it was placed. Here the head serves on the accelerator that `Device::gpu()` finds, so the merged graph computes the Relu on the CPU, moves `h` to the accelerator, and returns `y` there. On a host with no accelerator, the program prints `BLOCKED: this section needs an accelerator` and exits with status 3. In Python, the option is `cross_device="refuse"`. ```cpp title="across_devices.cpp" int main() { const ClikaRT::Device accelerator = ClikaRT::Device::gpu(); // the CPU on a host with no accelerator if (accelerator.is_cpu()) { std::printf("BLOCKED: this section needs an accelerator\n"); return 3; } // The encoder computes on the CPU, and the head serves on the accelerator. const auto two_parts = [&accelerator] { std::vector parts; parts.push_back({"encoder", trace_graph(rectify, "x", "h")}); parts.push_back({"head", trace_graph(twice, "h", "y")}); parts[1].graph.to(accelerator); return parts; }; const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}}; // CrossDevice::Move, the default: the merged graph moves h onto the head's device. std::vector parts = two_parts(); ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections); merged.graph.finalize(); const Tensor y = merged.graph.run({input()}).front(); std::printf("%s %s\n", y.device() == accelerator ? "true" : "false", values(y.to(ClikaRT::Device::cpu())).c_str()); // true 2 0 6 0 10 0: y is on the accelerator // CrossDevice::Refuse: a connection between two devices refuses the merge. ClikaRT::graph::MergeOptions refuse; refuse.cross_device = ClikaRT::graph::CrossDevice::Refuse; parts = two_parts(); try { ClikaRT::graph::merge(parts, connections, refuse); } catch (const Error& error) { std::printf("%s\n", error.code_name().c_str()); // INVALID_ARGUMENT, naming the connection and its devices } std::printf("%s | %s\n", io(parts[0].graph).c_str(), io(parts[1].graph).c_str()); // x -> h | h -> y: a refused merge leaves every part as it was return 0; } ``` ```python title="across_devices.py" accelerator = crt.Device.gpu() # the CPU on a host with no accelerator if accelerator.type == "cpu": print("BLOCKED: this section needs an accelerator") raise SystemExit(3) def two_parts() -> dict[str, ModelGraph]: # The encoder computes on the CPU, and the head serves on the accelerator. parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")} parts["head"].to(accelerator) return parts connections = [("encoder.h", "head.h")] # cross_device="move", the default: the merged graph moves h onto the head's device. merged, _ = crt.graph.merge(two_parts(), connections) merged.finalize() y = merged.run([x()])[0] print(y.device == accelerator, values(y)) # True 2 0 6 0 10 0: y is on the accelerator # cross_device="refuse": a connection between two devices refuses the merge. parts = two_parts() try: crt.graph.merge(parts, connections, cross_device="refuse") except crt.InvalidArgumentError as error: print(error.code_name) # INVALID_ARGUMENT, naming the connection and its devices print(io(parts["encoder"]), "|", io(parts["head"])) # x -> h | h -> y: a refused merge leaves every part as it was ``` Connect graphs on two devices to keep one part on the CPU while the model runs on an accelerator. ## Split a graph at its cheapest cut `min_value_cut()` finds the cut through a graph's values that hands the fewest bytes from the nodes before it to the nodes after it, as [The cheapest cut](query-a-graph.mdx#the-cheapest-cut) on the Query a graph page shows. `graph::split(graph, cut)` cuts the graph there into two parts. The part `before` holds the operators that produce the cut values and every operator they depend on, and the part `after` holds every other operator. Connections carry the cut values `after` reads and any graph input both parts read, and a value with no name of its own crosses under the name `_`. Here the cheapest values are the [2, 1] sums in `pool`, 8 bytes against 24 for `x`, so `before` computes the sums and `after` negates them. The split consumes the graph. In Python, `graph.split(graph, cut)` returns `(parts, connections)`, the parts as a dict from name to graph and the connections in the form `graph.merge` takes. ```cpp title="split_at_a_cut.cpp" int main() { ModelGraph pooled = trace_graph(pool, "x", "y"); // The cheapest cut: the values the nodes before it hand to the nodes after it, and their bytes. const ClikaRT::graph::ValueCut cut = pooled.min_value_cut(); std::string crossing; for (const ClikaRT::graph::Value& value : cut.values) { crossing += (crossing.empty() ? "" : " ") + label(*value.producer()); } std::printf("%s %llu\n", crossing.c_str(), static_cast(cut.bytes)); // Sum 8: the [2, 1] float32 sums cross, the cheapest value to hand on ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, cut); for (const GraphPart& part : split.parts) { std::printf("%s: %s | %s\n", part.name.c_str(), labels(part.graph).c_str(), io(part.graph).c_str()); } // before: x Relu Sum | x -> Sum_2_0 // after: Sum_2_0 Neg | Sum_2_0 -> y for (const Connection& connection : split.connections) { std::printf("%s.%s -> %s.%s\n", connection.from.part.c_str(), connection.from.name.c_str(), connection.to.part.c_str(), connection.to.name.c_str()); } // before.Sum_2_0 -> after.Sum_2_0: a value with no name of its own is named _ std::printf("%zu\n", pooled.input_names().size()); // 0: the split consumed the graph return 0; } ``` ```python title="split_at_a_cut.py" pooled = trace_graph(pool, "x", "y") # The cheapest cut: the values the nodes before it hand to the nodes after it, and their bytes. cut = pooled.min_value_cut() print(" ".join(label(value.producer()) for value in cut.values), cut.bytes) # Sum 8: the [2, 1] float32 sums cross, the cheapest value to hand on parts, connections = crt.graph.split(pooled, cut) for name, part in parts.items(): print(f"{name}:", labels(part), "|", io(part)) # before: x Relu Sum | x -> Sum_2_0 # after: Sum_2_0 Neg | Sum_2_0 -> y for source, target in connections: print(source, "->", target) # before.Sum_2_0 -> after.Sum_2_0: a value with no name of its own is named _ print(len(pooled.input_names())) # 0: the split consumed the graph ``` Split a graph at its cheapest cut to run its two halves on two devices, or in two processes, with the least data between them. ## Split a graph by a function `graph::split(graph, partition)` asks a function you write for each operator's part, a name that is not empty and holds no `/` and no `.`. It asks once per operator, in `nodes()` order, and the parts come in an order every connection runs forward in. A value read in a part other than its producer's becomes an output of the producer's part and an input of each part that reads it, under one name. A graph input goes with the first part that reads it, and every later reader receives it through a connection. Here the function names each operator's part by its code. In Python, the function takes a `Node` and returns a `str`. ```cpp title="split_by_a_function.cpp" int main() { ModelGraph pooled = trace_graph(pool, "x", "y"); // Each operator's part, by its code. split asks once per operator, in nodes() order. const auto by_code = [](const Node& node) { if (node.op_code() == ClikaRT::graph::OpCode::Relu) return std::string("rectify"); if (node.op_code() == ClikaRT::graph::OpCode::Sum) return std::string("pool"); return std::string("negate"); }; ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, by_code); for (const GraphPart& part : split.parts) { // in an order every connection runs forward in std::printf("%s: %s | %s\n", part.name.c_str(), labels(part.graph).c_str(), io(part.graph).c_str()); } // rectify: x Relu | x -> Relu_1_0 // pool: Relu_1_0 Sum | Relu_1_0 -> Sum_2_0 // negate: Sum_2_0 Neg | Sum_2_0 -> y for (const Connection& connection : split.connections) { std::printf("%s.%s -> %s.%s\n", connection.from.part.c_str(), connection.from.name.c_str(), connection.to.part.c_str(), connection.to.name.c_str()); } // rectify.Relu_1_0 -> pool.Relu_1_0 // pool.Sum_2_0 -> negate.Sum_2_0 return 0; } ``` ```python title="split_by_a_function.py" def by_code(node: Node) -> str: # Each operator's part, by its code. split asks once per operator, in nodes() order. if node.op_code == crt.graph.OpCode.Relu: return "rectify" if node.op_code == crt.graph.OpCode.Sum: return "pool" return "negate" pooled = trace_graph(pool, "x", "y") parts, connections = crt.graph.split(pooled, by_code) for name, part in parts.items(): # in an order every connection runs forward in print(f"{name}:", labels(part), "|", io(part)) # rectify: x Relu | x -> Relu_1_0 # pool: Relu_1_0 Sum | Relu_1_0 -> Sum_2_0 # negate: Sum_2_0 Neg | Sum_2_0 -> y for source, target in connections: print(source, "->", target) # rectify.Relu_1_0 -> pool.Relu_1_0 # pool.Sum_2_0 -> negate.Sum_2_0 ``` Split a graph by a function to give the operators you choose, by code, by name or by position, a graph of their own. ## Extract the operators between values `graph::extract(graph, inputs, outputs)` returns the operators that compute `outputs` from `inputs` as a graph of their own. A graph input keeps its name, and any other input value becomes an input named after it. An output keeps its graph output name when it has one, and is otherwise named after its value. The extract consumes the graph and frees the weights of the operators it leaves out. Weights stored inside an ONNX file, rather than as external data, share one buffer, which the extracted graph keeps in memory as the graph did. In Python, `graph.extract(graph, inputs, outputs)` takes lists of `Value`s. ```cpp title="extract.cpp" int main() { ModelGraph pooled = trace_graph(pool, "x", "y"); // The operators between x and the sums: the Relu and the Sum, not the Neg after them. const Node sum = pooled.find_nodes(ClikaRT::graph::OpCode::Sum).front(); const std::vector inputs = {*pooled.node("x")->output(0)}; const std::vector outputs = {*sum.output(0)}; ModelGraph pooling = ClikaRT::graph::extract(pooled, inputs, outputs); std::printf("%s | %s\n", labels(pooling).c_str(), io(pooling).c_str()); // x Relu Sum | x -> Sum_2_0: a graph input keeps its name, and an output with none is named after its value std::printf("%zu\n", pooled.input_names().size()); // 0: the extract consumed the graph pooling.finalize(); std::printf("%s\n", values(pooling.run({input()}).front()).c_str()); // 4 5: each row's sum of Relu(x) return 0; } ``` ```python title="extract.py" pooled = trace_graph(pool, "x", "y") # The operators between x and the sums: the Relu and the Sum, not the Neg after them. (sums,) = pooled.find_nodes(crt.graph.OpCode.Sum) pooling = crt.graph.extract(pooled, [pooled.node("x").output(0)], [sums.output(0)]) print(labels(pooling), "|", io(pooling)) # x Relu Sum | x -> Sum_2_0: a graph input keeps its name, and an output with none is named after its value print(len(pooled.input_names())) # 0: the extract consumed the graph pooling.finalize() print(values(pooling.run([x()])[0])) # 4 5: each row's sum of Relu(x) ``` Extract a block of a model to run it alone, to test it, or to serve it on its own. ## Merge the parts back `graph::merge(split.parts, split.connections)` joins the parts of a split back into one graph, with the source's inputs and outputs by name and its operators named `/`. The merged graph computes what the source computes. In Python, `graph.merge(*graph.split(graph, cut))` does the same in one call. ```cpp title="round_trip.cpp" int main() { ModelGraph pooled = trace_graph(pool, "x", "y"); ClikaRT::graph::SplitGraph split = ClikaRT::graph::split(pooled, pooled.min_value_cut()); // The parts merge back into the source's inputs and outputs, by name, with its operators named /. ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(split.parts, split.connections); std::printf("%s | %s\n", labels(merged.graph).c_str(), io(merged.graph).c_str()); // x before/Relu before/Sum after/Neg | x -> y merged.graph.finalize(); ModelGraph traced = trace_graph(pool, "x", "y"); // the same model traced again, for comparison traced.finalize(); std::printf("%s | %s\n", values(merged.graph.run({input()}).front()).c_str(), values(traced.run({input()}).front()).c_str()); // -4 -5 | -4 -5 return 0; } ``` ```python title="round_trip.py" pooled = trace_graph(pool, "x", "y") # split returns the parts and connections merge takes, so the parts merge back into the source's inputs and # outputs, by name, with its operators named /. merged, _ = crt.graph.merge(*crt.graph.split(pooled, pooled.min_value_cut())) print(labels(merged), "|", io(merged)) # x before/Relu before/Sum after/Neg | x -> y merged.finalize() traced = trace_graph(pool, "x", "y") # the same model traced again, for comparison traced.finalize() print(values(merged.run([x()])[0]), "|", values(traced.run([x()])[0])) # -4 -5 | -4 -5 ``` Merge the parts back after you place, optimize or edit them one at a time. ## What the calls consume, and what a refusal leaves A merge consumes every part's graph, and a split or an extract consumes its source. On success those graphs are empty, and a `Node`, `Value` or `Edge` view taken into one of them refuses with the code name `INVALID_ARGUMENT`. Each call checks everything before it changes anything, so a refused call leaves every graph as it was. The one exception is a merge that fixed a dynamic size through a connection before it refused, which leaves that size fixed, the size the merge would give it. The code name is `INVALID_ARGUMENT` for an argument the call cannot take, and `FAILED_PRECONDITION` for a graph it cannot take now, such as a finalized graph. A partition function that throws makes `split` refuse with the function's failure. A thrown `ClikaRT::Error` keeps its status, and any other exception comes back as `INTERNAL`. In Python, `graph.split` raises the function's own exception. [Handle errors by code](handle-errors-by-code.mdx) covers the channels every failure carries. ```cpp title="consumption.cpp" int main() { std::vector parts; parts.push_back({"encoder", trace_graph(rectify, "x", "h")}); parts.push_back({"head", trace_graph(twice, "h", "y")}); const std::vector connections = {{{"encoder", "h"}, {"head", "h"}}}; const Node relu = parts[0].graph.find_nodes(ClikaRT::graph::OpCode::Relu).front(); // a view taken before // A merge consumes every part's graph, and a view into one refuses from then on. const ClikaRT::graph::MergedGraph merged = ClikaRT::graph::merge(parts, connections); std::printf("%s | %zu %zu\n", io(merged.graph).c_str(), parts[0].graph.input_names().size(), parts[1].graph.input_names().size()); // x -> y | 0 0 try { relu.check(); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } // INVALID_ARGUMENT | the node 'Relu_1' is no longer in its graph: that graph no longer exists (a merge, // a split or an extract consumed it, or it was destroyed or assigned over) // A refused merge leaves every part as it was. std::vector again; again.push_back({"encoder", trace_graph(rectify, "x", "h")}); again.push_back({"head", trace_graph(twice, "h", "y")}); const std::vector wrong = {{{"encoder", "h"}, {"head", "z"}}}; try { ClikaRT::graph::merge(again, wrong); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } // INVALID_ARGUMENT | merge: part 'head' has no input 'z' again[0].graph.finalize(); try { ClikaRT::graph::merge(again, connections); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } // FAILED_PRECONDITION | merge: part 'encoder' is finalized; merge before finalize() std::printf("%s | %s\n", io(again[0].graph).c_str(), io(again[1].graph).c_str()); // x -> h | h -> y // A partition function that throws: split refuses with its failure, and the graph stays as it was. ModelGraph pooled = trace_graph(pool, "x", "y"); try { ClikaRT::graph::split(pooled, [](const Node& node) -> std::string { if (node.op_code() == ClikaRT::graph::OpCode::Neg) { throw Error(ClikaRT::Status::InvalidArgument, "no part for Neg"); } return "first"; }); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); } // INVALID_ARGUMENT | split: the partition function failed on the node 'Neg_3': no part for Neg std::printf("%s\n", io(pooled).c_str()); // x -> y return 0; } ``` ```python title="consumption.py" def refuse_neg(node: Node) -> str: # A partition function that raises on the Neg. if node.op_code == crt.graph.OpCode.Neg: raise ValueError("no part for Neg") return "first" parts = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")} connections = [("encoder.h", "head.h")] (relu,) = parts["encoder"].find_nodes(crt.graph.OpCode.Relu) # a view taken before # A merge consumes every part's graph, and a view into one refuses from then on. merged, _ = crt.graph.merge(parts, connections) print(io(merged), "|", len(parts["encoder"].input_names()), len(parts["head"].input_names())) # x -> y | 0 0 try: relu.check() except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | the node 'Relu_1' is no longer in its graph: that graph no longer exists (a merge, # a split or an extract consumed it, or it was destroyed or assigned over) # A refused merge leaves every part as it was. again = {"encoder": trace_graph(rectify, "x", "h"), "head": trace_graph(twice, "h", "y")} try: crt.graph.merge(again, [("encoder.h", "head.z")]) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | merge: part 'head' has no input 'z' again["encoder"].finalize() try: crt.graph.merge(again, connections) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # FAILED_PRECONDITION | merge: part 'encoder' is finalized; merge before finalize() print(io(again["encoder"]), "|", io(again["head"])) # x -> h | h -> y # An exception the partition function raises leaves split as that same exception, and the graph as it was. pooled = trace_graph(pool, "x", "y") try: crt.graph.split(pooled, refuse_neg) except ValueError as error: print(type(error).__name__, "|", error) # ValueError | no part for Neg print(io(pooled)) # x -> y ``` Reuse the graphs a refused call leaves: they are as they were, and a corrected call takes them. --- # Optimize and finalize A traced or compiled ModelGraph comes back as built. Run the graph optimizer with optimize() and read its report, choose its transforms and its defaults, place the graph with to(), and finalize() it to run. Source: https://docs.clika.io/clikart/how-to/optimize-and-finalize.md {/* Every block is a program under examples//howto/optimize_and_finalize/: the first block of each tab is get_a_graph whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output. */} `compile()` and `trace()` return a `ModelGraph` as built, with every operator as the model states it and nothing optimized, placed or finalized. Four calls take it from there. `optimize()` runs the graph optimizer when you want it and reports what it did. `to()` names the device or the stream the graph serves on. `finalize()` places the graph there, packs the weights into their kernel layouts and plans the execution order, and from then on the graph runs. The calls are available from C++ and Python, under the same names. ## A graph as built The model below computes `y = Relu(Neg(Neg(x)))`. The two Negs cancel each other, which gives the graph optimizer one rewrite to find. Every example on this page traces it. `labels` lists a graph's nodes in order, each by its input name or its operator, and `trace_model` traces the model over an `x` of shape [2, 3]. The input every program runs is `input()` in C++ and `x` in Python, and `values` prints a C++ result on one line. A graph that is not finalized refuses `run()` with the code name `FAILED_PRECONDITION`, and the message names the call to make. ```cpp title="get_a_graph.cpp" #include #include #include #include #include using ClikaRT::DataType; using ClikaRT::Error; using ClikaRT::Tensor; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::Node; using ClikaRT::graph::NodeKind; using ClikaRT::graph::OptimizeOptions; using ClikaRT::graph::OptimizeReport; using ClikaRT::graph::Transform; using ClikaRT::graph::TransformReport; namespace ops = ClikaRT::ops; namespace transforms = ClikaRT::graph::transforms; namespace { // y = Relu(Neg(Neg(x))): the two Negs cancel, which the graph optimizer finds. std::vector model(const std::vector& inputs) { return {ops::relu(ops::neg(ops::neg(inputs[0])))}; } // A node's label: an input's name, an operator's code. std::string label(const Node& node) { return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code())); } // The labels of a graph's nodes, in order, separated by spaces. std::string labels(const ModelGraph& graph) { std::string out; for (const Node& node : graph.nodes()) out += (out.empty() ? "" : " ") + label(node); return out; } // A trace returns the graph as built: every operator as written, nothing optimized or finalized. ModelGraph trace_model() { const std::vector signature = {{"x", DataType::Float32, {2, 3}}}; const std::vector outputs = {"y"}; return ClikaRT::graph::trace(model, signature, "lifecycle", outputs); } // The input every program runs, [2, 3]. Tensor input() { const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F}; return Tensor::from_data(x.data(), {2, 3}, DataType::Float32); } // A tensor's values, flattened and separated by spaces. std::string values(const Tensor& tensor) { std::string out; for (const float value : tensor.reshape({-1}).item_as_vec()) { char text[32]; std::snprintf(text, sizeof(text), "%g", static_cast(value)); out += (out.empty() ? "" : " ") + std::string(text); } return out; } } // namespace int main() { const ModelGraph graph = trace_model(); std::printf("%s | %s\n", labels(graph).c_str(), graph.is_finalized() ? "true" : "false"); // x Neg Neg Relu | false // Only a finalized graph runs. try { graph.run({input()}); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // FAILED_PRECONDITION | run: the graph is not finalized; call finalize() first } return 0; } ``` ```python title="get_a_graph.py" import numpy as np import clika_runtime as crt from clika_runtime.graph import DefaultsOptions, NodeKind, transforms def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]: return [crt.relu(crt.neg(crt.neg(inputs[0])))] # y = Relu(Neg(Neg(x))): the two Negs cancel def labels(graph: crt.graph.ModelGraph) -> list[str]: return [node.name if node.kind == NodeKind.Input else node.op_code.name for node in graph.nodes()] # A trace returns the graph as built: every operator as written, nothing optimized or finalized. def trace_model() -> crt.graph.ModelGraph: return crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32) # the input every program runs (numpy as the data entry) graph = trace_model() print(labels(graph), graph.is_finalized()) # ['x', 'Neg', 'Neg', 'Relu'] False # Only a finalized graph runs. try: graph.run([crt.tensor(x)]) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # FAILED_PRECONDITION | run: the graph is not finalized; call finalize() first ``` Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, the helpers, the input and, in Python, the trace). ## Run the graph optimizer `optimize()` runs the graph's default list of transforms. Each iteration runs the whole list once, and the run stops at the first iteration that changes nothing. It returns an `OptimizeReport`. `iterations` counts the iterations that ran, `converged` says the run stopped on its own, and `changed` says it modified the graph. `transforms` holds one `TransformReport` per list entry, in list order, and each row's `applications` counts the iterations in which its transform changed the graph. Here `remove_double_neg` removes both Negs in the first iteration, and the second iteration changes nothing. ```cpp title="optimize.cpp" int main() { ModelGraph graph = trace_model(); const OptimizeReport report = graph.optimize(); std::printf("%s\n", labels(graph).c_str()); // x Relu std::printf("%lld %s %s\n", static_cast(report.iterations), report.converged ? "true" : "false", report.changed ? "true" : "false"); // 2 true true for (const TransformReport& row : report.transforms) { if (row.applications > 0) { std::printf("%s %lld\n", row.transform.name().c_str(), static_cast(row.applications)); // remove_double_neg 1 } } return 0; } ``` ```python title="optimize.py" report = graph.optimize() print(labels(graph)) # ['x', 'Relu'] print(report.iterations, report.converged, report.changed) # 2 True True print([(row.transform.name, row.applications) for row in report.transforms if row.applications]) # [('remove_double_neg', 1)] ``` A view you took before `optimize()` finds its node again by its key, or refuses once the optimizer removed that node, as [Views across edits](query-a-graph.mdx#views-across-edits) on the Query a graph page shows. ## The default list `default_transforms()` returns the list `optimize()` runs on this graph, in the order it runs them, as a starting point to edit and pass back. `transforms::defaults()` (in Python, `transforms.defaults()`) returns the default list for a set of choices without a graph. A graph's own list follows how the graph was built, so the two can differ. For a traced graph with no choices set, they are the same list. Each entry is a `Transform` handle named as its factory is, so `transforms::remove_double_neg()` names `remove_double_neg` (`name()` in C++, the `name` property in Python). Compare and store transforms by name, since the values behind `kind()` move when the runtime's transform list changes. ```cpp title="default_list.cpp" int main() { const ModelGraph graph = trace_model(); const std::vector listed = graph.default_transforms(); const auto holds = [&listed](const Transform& transform) { return std::find(listed.begin(), listed.end(), transform) != listed.end() ? "true" : "false"; }; std::printf("%s\n", listed.front().name().c_str()); // cleanup_graph std::printf("%s %s\n", holds(transforms::remove_double_neg()), holds(transforms::fixate_dyn_shape_queues())); // true false std::printf("%s\n", listed == transforms::defaults() ? "true" : "false"); // true return 0; } ``` ```python title="default_list.py" listed = graph.default_transforms() print(listed[0].name) # cleanup_graph print(transforms.remove_double_neg() in listed, transforms.fixate_dyn_shape_queues() in listed) # True False print(listed == transforms.defaults()) # True ``` ## Run a list of your own `OptimizeOptions::transforms` runs the list you give it instead, in order, each iteration. In Python, pass the list as the first argument of `optimize()`. A list may hold a transform more than once. An empty list runs none of them, though the run still drops the nodes nothing reads. The list is checked before anything runs. A transform placed ahead of one it must follow is refused with the code name `INVALID_ARGUMENT`, and the message names both entries and the order to use. The refused call leaves the graph as it was. ```cpp title="own_list.cpp" int main() { ModelGraph graph = trace_model(); OptimizeOptions options; options.transforms = std::vector{transforms::remove_double_neg()}; const OptimizeReport report = graph.optimize(options); const TransformReport& row = report.transforms.front(); std::printf("%s | %zu %s %lld\n", labels(graph).c_str(), report.transforms.size(), row.transform.name().c_str(), static_cast(row.applications)); // x Relu | 1 remove_double_neg 1 // A list is checked before anything runs: matmul_absorb_transpose must run after remove_double_permute. ModelGraph other = trace_model(); options.transforms = std::vector{transforms::matmul_absorb_transpose(), transforms::remove_double_permute()}; try { other.optimize(options); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // INVALID_ARGUMENT | optimize: transforms[0] (matmul_absorb_transpose) must run after remove_double_permute, which the list places later, at transforms[1]; list remove_double_permute ahead of it } std::printf("%s\n", labels(other).c_str()); // x Neg Neg Relu return 0; } ``` ```python title="own_list.py" report = graph.optimize([transforms.remove_double_neg()]) print(labels(graph), [(row.transform.name, row.applications) for row in report.transforms]) # ['x', 'Relu'] [('remove_double_neg', 1)] # A list is checked before anything runs: matmul_absorb_transpose must run after remove_double_permute. other = trace_model() try: other.optimize([transforms.matmul_absorb_transpose(), transforms.remove_double_permute()]) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | optimize: transforms[0] (matmul_absorb_transpose) must run after remove_double_permute, which the list places later, at transforms[1]; list remove_double_permute ahead of it print(labels(other)) # ['x', 'Neg', 'Neg', 'Relu'] ``` A list can also hold a transform you write, a function that edits the graph and runs beside the runtime's transforms (`Transform::from_function` in C++, `Transform(name, fn)` in Python). [Write a transform](write-a-transform.mdx) covers them. ## Cap the iterations `max_num_iterations` caps one run at that many iterations. It is 100 unless you set it, and a value outside 1 to 2147483647 is refused with `INVALID_ARGUMENT`. A run that reaches the cap still leaves a correct graph. Its report reads `converged` as false, `firing_at_cap` marks the transforms that changed the graph in the last iteration, and the runtime logs a warning that names them. Here the single iteration applies `remove_double_neg`, and the cap ends the run before a second iteration can show that nothing else changes. ```cpp title="iteration_cap.cpp" int main() { ModelGraph graph = trace_model(); OptimizeOptions options; options.max_num_iterations = 1; // one pass over the list const OptimizeReport report = graph.optimize(options); std::printf("%lld %s %s\n", static_cast(report.iterations), report.converged ? "true" : "false", report.changed ? "true" : "false"); // 1 false true for (const TransformReport& row : report.transforms) { if (row.firing_at_cap) std::printf("%s\n", row.transform.name().c_str()); // remove_double_neg } std::printf("%s\n", labels(graph).c_str()); // x Relu options.max_num_iterations = 0; try { graph.optimize(options); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // INVALID_ARGUMENT | optimize: max_num_iterations must be between 1 and 2147483647, got 0 } return 0; } ``` ```python title="iteration_cap.py" report = graph.optimize(max_num_iterations=1) # one pass over the list print(report.iterations, report.converged, report.changed) # 1 False True print([row.transform.name for row in report.transforms if row.firing_at_cap]) # ['remove_double_neg'] print(labels(graph)) # ['x', 'Relu'] try: graph.optimize(max_num_iterations=0) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | optimize: max_num_iterations must be between 1 and 2147483647, got 0 ``` A transform of the runtime's that fails is undone and skipped for the rest of the run, which goes on. Its row reads `failed`, with the status name in `failure_code` and the message in `failure_message`. The graph stays correct, only less optimized. ## Fix the shapes `DefaultsOptions` holds the choices the transforms read, each off by default. With every choice off, the graph keeps its input and output contract and serves every shape it admits. Two choices fix the graph to its shapes instead, collapsing shape arithmetic such as shape cones, masks and reshape templates into constants. Each adds the shape-fixation transforms, `fixate_dyn_shape_queues` among them, to the default list. A traced graph follows `shape_fixation`, which fixes it to the shapes it was traced with. A graph compiled from ONNX follows `specialize`, which fixes it to the input shapes the compile pinned (`CompileOptions::specs` or `sample_inputs`). Without pinned shapes, the arithmetic the model's own static shapes fix collapses either way. The graph-free list takes the choices as stated. Pass the choices to `optimize()` in `OptimizeOptions::defaults` (in Python, the `defaults=` keyword), and to `default_transforms()` to see the list they give. A list of your own that holds a transform these choices switch off is refused, and the message names the choice to set. ```cpp title="defaults.cpp" int main() { const ModelGraph graph = trace_model(); const Transform fixate = transforms::fixate_dyn_shape_queues(); const auto holds = [&fixate](const std::vector& listed) { return std::find(listed.begin(), listed.end(), fixate) != listed.end() ? "true" : "false"; }; // A traced graph follows shape_fixation: its list gains the transforms that fix it to the traced shapes. ClikaRT::graph::DefaultsOptions fixation; fixation.shape_fixation = true; std::printf("%s %s\n", holds(graph.default_transforms()), holds(graph.default_transforms(fixation))); // false true // A graph compiled from ONNX follows specialize; the graph-free list takes the choice as stated. ClikaRT::graph::DefaultsOptions specialize; specialize.specialize = true; std::printf("%s %s\n", holds(transforms::defaults()), holds(transforms::defaults(specialize))); // false true return 0; } ``` ```python title="defaults.py" fixate = transforms.fixate_dyn_shape_queues() # A traced graph follows shape_fixation: its list gains the transforms that fix it to the traced shapes. print(fixate in graph.default_transforms(), fixate in graph.default_transforms(DefaultsOptions(shape_fixation=True))) # False True # A graph compiled from ONNX follows specialize; the graph-free list takes the choice as stated. print(fixate in transforms.defaults(), fixate in transforms.defaults(DefaultsOptions(specialize=True))) # False True ``` ## Channels-last inputs and outputs A model exported channels-first reaches the runtime's channels-last operators through a conversion at each input a spatial operator, such as a convolution, reads. `channels_last_inputs` omits that conversion, so the input is supplied channels-last (`[N, spatial..., C]`) from then on. `channels_last_outputs` does the same for an output that leaves through a conversion. Both change the graph's input and output contract, and `inputs()` and `outputs()` report the layout to supply and the layout delivered. The program builds its ONNX model with a helper outside the block, an `X` of shape [1, 2, 4, 4] that a 1x1 Conv reads, followed by a Relu, and `layout` prints an input's name and dims. `drop_boundary_permutes`, the transform that omits the conversions, runs only under one of the two choices. ```cpp title="channels_last.cpp" int main() { ModelGraph graph = compile_conv_model(); std::printf("%s\n", layout(graph.inputs().front()).c_str()); // X [1, 2, 4, 4] // drop_boundary_permutes runs only under a channels-last choice. OptimizeOptions listed; listed.transforms = std::vector{transforms::drop_boundary_permutes()}; try { graph.optimize(listed); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // INVALID_ARGUMENT | optimize: transforms[0] (drop_boundary_permutes) changes the layout of the graph's inputs and outputs; set DefaultsOptions::channels_last_inputs or channels_last_outputs to run it } OptimizeOptions options; options.defaults.channels_last_inputs = true; // X is supplied channels-last from here on graph.optimize(options); std::printf("%s\n", layout(graph.inputs().front()).c_str()); // X [1, 4, 4, 2] return 0; } ``` ```python title="channels_last.py" graph = compile_conv_model() print(layout(graph)) # X [1, 2, 4, 4] # drop_boundary_permutes runs only under a channels-last choice. try: graph.optimize([transforms.drop_boundary_permutes()]) except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | optimize: transforms[0] (drop_boundary_permutes) changes the layout of the graph's inputs and outputs; set DefaultsOptions::channels_last_inputs or channels_last_outputs to run it graph.optimize(defaults=DefaultsOptions(channels_last_inputs=True)) # X is supplied channels-last from here on print(layout(graph)) # X [1, 4, 4, 2] ``` ## Place the graph with to() `to()` takes a device or a stream. Before `finalize()`, it records where finalizing places the graph and moves nothing, and `device()` reports the new device at once. After `finalize()`, it moves the finalized graph. An operator that changes device re-packs its weights there once, and no copy of a moved weight stays behind. A graph no `to()` moved serves where its operators sit, the CPU for this trace, and a graph compiled from ONNX serves on the device `CompileOptions::weights` names until `to()` names another. In Python, `device` is a property. To compute on a stream of your own, create it and pass it to `to()` once, and every run computes there. On a host with an accelerator, `to(Device::cuda(0))` places the operators and their weights on it; the program stays on the CPU, so it runs on any host. `to()` refuses while a run is in flight, while a KV cache is attached (call it before `attach_kv_cache()`), and while `optimize()`, `finalize()` or an edit runs on the graph. ```cpp title="placement.cpp" int main() { ModelGraph graph = trace_model(); std::printf("%s\n", ClikaRT::device::compute_api_name(graph.device().api)); // CPU: where its operators sit // Before finalize(), to() records where finalize() places the graph; nothing moves yet. graph.to(ClikaRT::Stream::create(ClikaRT::Device::cpu())); std::printf("%s %s\n", ClikaRT::device::compute_api_name(graph.device().api), graph.is_finalized() ? "true" : "false"); // CPU false graph.optimize(); graph.finalize(); // places the operators on that stream and packs their weights there std::printf("%s\n", values(graph.run({input()}).front()).c_str()); // 1 0 3 0 5 0 // After finalize(), to() moves the finalized graph. graph.to(ClikaRT::Device::cpu()); std::printf("%s\n", values(graph.run({input()}).front()).c_str()); // 1 0 3 0 5 0 return 0; } ``` ```python title="placement.py" print(graph.device) # cpu: where its operators sit # Before finalize(), to() records where finalize() places the graph; nothing moves yet. graph.to(crt.Stream.create(crt.Device.cpu())) print(graph.device, graph.is_finalized()) # cpu False graph.optimize() graph.finalize() # places the operators on that stream and packs their weights there (y,) = graph.run([crt.tensor(x)]) print(y.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]] # After finalize(), to() moves the finalized graph. graph.to(crt.Device.cpu()) (again,) = graph.run([crt.tensor(x)]) print(again.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]] ``` ## Finalize and run `finalize()` places the graph on its device, packs the weights into their kernel layouts, absorbs the constants the operators hold and plans the execution order. A second call succeeds and changes nothing. A finalized graph runs, benches and takes a KV cache. It refuses `optimize()` and every edit with the code name `FAILED_PRECONDITION`, and the message says to compile or trace the model again, while `run()` and `to()` keep working. Each program checks the run against a reference it computes itself, the Python one with numpy. ```cpp title="finalize.cpp" int main() { ModelGraph graph = trace_model(); graph.optimize(); graph.finalize(); graph.finalize(); // a second call succeeds and changes nothing std::printf("%s\n", graph.is_finalized() ? "true" : "false"); // true const Tensor y = graph.run({input()}).front(); const std::vector got = y.reshape({-1}).item_as_vec(); const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F}; bool matches = got.size() == x.size(); for (std::size_t i = 0; matches && i < x.size(); ++i) { matches = got[i] == std::max(x[i], 0.0F); // the reference, Relu(x), by hand } std::printf("%s\n", values(y).c_str()); // 1 0 3 0 5 0 std::printf("%s\n", matches ? "true" : "false"); // true // A finalized graph refuses optimize() and every edit; run() and to() keep working. try { graph.optimize(); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // FAILED_PRECONDITION | optimize: the graph is finalized; optimize() runs before finalize(), so compile or trace the model again to optimize it } try { graph.rename_output("y", "out"); } catch (const Error& error) { std::printf("%s\n", error.code_name().c_str()); // FAILED_PRECONDITION } return 0; } ``` ```python title="finalize.py" graph.optimize() graph.finalize() graph.finalize() # a second call succeeds and changes nothing print(graph.is_finalized()) # True (y,) = graph.run([crt.tensor(x)]) print(y.numpy().tolist()) # [[1.0, 0.0, 3.0], [0.0, 5.0, 0.0]] print(np.array_equal(y.numpy(), np.maximum(x, 0))) # True (numpy states the reference) # A finalized graph refuses optimize() and every edit; run() and to() keep working. try: graph.optimize() except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # FAILED_PRECONDITION | optimize: the graph is finalized; optimize() runs before finalize(), so compile or trace the model again to optimize it try: graph.rename_output("y", "out") except crt.InvalidArgumentError as error: print(error.code_name) # FAILED_PRECONDITION ``` Optimize and edit before you finalize. [Edit a graph](edit-a-graph.mdx) covers the edits, and [Query a graph](query-a-graph.mdx) reads the graph at any point in its lifecycle. --- # Package ClikaRT in an Android app What an Android app ships to run the runtime and the model library on the phone: the one artifact and the libraries inside it, the build requirements, the license credential, the thread count, the cache root, and a server that outlives the screen. Source: https://docs.clika.io/clikart/how-to/package-clikart-in-an-android-app.md {/* CERTIFICATION: every fact on this page is read from the Kotlin binding's README and gradle files of the pinned release (compileSdk, minSdk, the plugin versions, the AAR contents, the load entries) and from the release's Kotlin artifact itself (its Maven repository directory); no app on this page is compiled by the docs build. */} The [tutorial's last part](../getting-started/first-program/06-deploy-to-mobile.mdx) pushes a binary to a phone over `adb`. An app does the same thing from inside its own process: it ships the runtime library, loads it through the Kotlin binding, and places the license credential before the first call. This page is the inventory of what the app carries and the few facts that are easy to get wrong the first time. ## The one artifact One Maven artifact, `io.clika:clika-runtime`, carries the runtime and the model library's Kotlin API together. It is published as a Maven repository directory inside `clika-runtime-maven-.zip`, a download of the platform beside the release archives ([Get ClikaRT](../getting-started/get-clikart.mdx)); extract it anywhere and name the directory as a repository, then one dependency line serves both products: ```kotlin repositories { maven { url = uri("/path/to/clika-runtime-maven-") } google() mavenCentral() } dependencies { implementation("io.clika:clika-runtime:") } ``` The artifact is a Kotlin Multiplatform library with two variants; gradle reads its module metadata and picks the one the project builds for: | Variant | Form | Carries | | --- | --- | --- | | Android (`arm64-v8a`) | an AAR | the Kotlin classes of both packages (`io.clika.runtime`, `io.clika.modelverse`), the two JNI bridges, every library of the Android release archive (`libClikaRT.so`, its backend images, `libClikaRT_modelverse.so`) under `jni/arm64-v8a/`, the consumer keep rules, and the archive's third-party notices under `META-INF/licenses/` | | desktop JVM (Linux, Windows, macOS) | a jar | the Kotlin classes of both packages, the two JNI bridges per desktop platform, and the notices; the runtime and model libraries come from the release archive's `lib/` on `java.library.path` | On Android the AAR carries everything: the app copies nothing into `jniLibs`. At run time the binding loads `libClikaRT.so` first, then the runtime bridge; `Modelverse.load` loads `libClikaRT_modelverse.so` and the model library's bridge after it. The binding is built against one release and refuses to load another, naming both versions (`BINDING_VERSION_MISMATCH`). ## The build requirements | Requirement | Value | | --- | --- | | `compileSdk` | 36, the level the binding's modules compile against | | `minSdk` | 28: a device below it installs the AAR but cannot load the runtime library | | Android Gradle plugin | 9.3.2, which needs gradle 9.5 or newer | | Kotlin | 2.3 or newer in the app: the binding's classes are compiled by Kotlin 2.3, and an older compiler refuses their metadata | | JDK | 17 | | NDK | none: an app that takes the published artifact builds nothing native | | ABI | `arm64-v8a` alone | The binding reaches nothing by reflection, so an app's R8 pass needs no rule of its own: the AAR carries the consumer rules for the classes its bridges bind by name. ## Load, and place the credential first An app has no shell to export a variable from and no per-user license file, so the credential is an argument of the load. `ClikaRtAndroid.load(context, license = "CLIKA1-...")` places it for the process and loads the runtime; an app that uses the model library calls `Modelverse.load(context, license)` instead, which runs the whole sequence, the runtime first. A credential given there replaces one already in the process environment; the default, `null`, leaves the environment as it is. The credential's text is never logged. Ship it the way you ship any other secret the app needs at start-up: a build-time field read from a file outside the source tree, or a value the app asks for once and keeps in its own secret store. Without a valid credential the first operator throws a `ClikaRtException` (a `ModelverseException` from the model library) whose `codeName` is `LICENSE_FAILED`, or `LICENSE_EXPIRED` past the end date. Two more facts are placed before the first call. The CPU worker count, `CLIKA_RT_NUM_THREADS`, is read once before the first compute; an app sets it in its own process with `Os.setenv("CLIKA_RT_NUM_THREADS", "", true)` before the load. A phone with a few large cores and several small ones runs best at the large cores' count; measure on the device. The cache root is the app's own cache directory: `ClikaRtAndroid.load` sets `XDG_CACHE_HOME` to `Context.getCacheDir()` unless the app set it itself, and the hub cache the model library downloads into sits under it in the Hugging Face layout, so a snapshot downloaded on another machine and copied there serves with no network. Every call into the runtime runs on a thread of the app's own, never on the main thread: a model load blocks for the load, and `generate` blocks for the reply. After the load, `ClikaRtAndroid.trimOnMemoryPressure(context)` lets the runtime give back its cached memory when the platform reports pressure, and `ClikaRtAndroid.bridgeLogs()` routes the runtime's log lines to logcat. ## A model on the phone `AutoModelForCausalLM.fromPretrained(source, LoadOptions(contextLength = 4096, cacheDir = ...))` takes a hub repository id (the snapshot downloads inside the load), a snapshot directory, or a `.gguf` file. Set the context length on a phone: the key-value cache is sized for the window, and a checkpoint's own window is often tens of thousands of tokens. `chat(messages, config, listener)` returns at once and calls the listener on the model's worker thread, `onToken` per visible piece and `onDone` with the report; the handle's `cancel()` stops the decode at its next token. `generate(prompt = ...)` is the blocking form. `Modelverse.registry()` lists the families the library serves with the keys a checkpoint is matched by, so an app checks a model before it downloads one. The whole surface, kind by kind, is the binding's README. ## A server that outlives the screen `Modelverse.serve(model, ServeOptions(port = 8129))` hosts the library's chat API on the phone (the OpenAI chat route, the model list and the chat page when enabled), on the server's own threads; `start()` returns at once and `baseUrl` is the address another device on the network points its client at. A server started from an `adb shell` dies with the shell and with the phone's next reboot; an app keeps it alive by hosting it in a foreground service with an ongoing notification (the connected-device service type, which Android 14 and newer admits with a network-class permission beside it), so the process survives the activity leaving the screen and starts again with the app. ## Where the runnable proof is The examples archive ([Additional examples](../examples.md)) carries the Android programs. The Kotlin chapters under its `kotlin/clika_rt/` directory are instrumented tests that run the same calls on a connected device (`gradle connectedAndroidTest -PclikaRtMavenRepo=`): the operator surface through `Ops`, and the device and backend facts through `Backends`. Its `android/` directory holds three sample apps over the same artifact, one screen each: `hello` loads the runtime and prints the version, the backends with their devices and one operator's result; `chat` loads a chat model and streams its replies; `serve` hosts a loaded model for the other devices on the network from a foreground service. Every one of them takes the artifact through the `clikaRtMavenRepo` property and copies nothing into `jniLibs`. --- # Preprocess inputs with processors Model-specific image and audio preprocessing: build a resize/rescale/normalize pipeline or a log-mel front end from knobs, or load it from the model's own config. Source: https://docs.clika.io/clikart/how-to/preprocess-with-processors.md Decoding a file into pixels or samples is half the input story; the other half is what the MODEL was trained to receive: resized, rescaled, normalized pixels, or a log-mel spectrogram at a fixed rate. `processor::ImageProcessor` and `processor::AudioProcessor` carry that half. Build one from explicit knobs when you know the recipe, or load it straight from the model's own `preprocessor_config.json` so the recipe cannot drift from the checkpoint. [Load images and audio for inference](load-images-and-audio.mdx) covers the decode step this page starts after; the processors consume the same tensors it produces. ## Build an image pipeline from knobs `from_args` covers the common recipe with no config file: resize, center-crop, rescale, normalize. The example uses a constant image so the result is checkable by hand: every input pixel 127 becomes `(127/255 - 0.5) / 0.5`. ```cpp title="image_knobs.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Tensor; using ClikaRT::processor::ImageProcessor; int main() { const Tensor image = Tensor::full({8, 8, 3}, 127, DataType::UInt8); const float mean[] = {0.5F, 0.5F, 0.5F}; const float std_[] = {0.5F, 0.5F, 0.5F}; ImageProcessor proc = ImageProcessor::from_args( /*resize=*/{{4, 4}}, /*crop=*/{}, /*rescale=*/{}, mean, std_); const Tensor out = proc.process(image); std::printf("features = %s\n", out.to_string().c_str()); return 0; } ``` ```python title="image_knobs.py" import numpy as np import clika_runtime as crt image = crt.tensor(np.full((8, 8, 3), 127, dtype=np.uint8)) proc = crt.processor.ImageProcessor.from_args( resize=(4, 4), mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5] ) out = proc.process(image) print(out.dtype, out.shape) # clika_runtime.float32 clika_runtime.Size([4, 4, 3]) print(out.numpy()[0, 0, 0]) # (127/255 - 0.5) / 0.5 ``` The Kotlin binding carries the processor family (`ImageProcessor`, `AudioProcessor`, `VideoProcessor` and the `Processor` composite); the worked programs on this page are the C++ and Python ones, and [Language bindings](../bindings.md) says what Kotlin carries. ```text features = Tensor(shape=[4, 4, 3], dtype=Float32, device=CPU, numel=48, data=[-0.003922, -0.003922, -0.003922, -0.003922, -0.003922, -0.003922, ...]) ``` Knob semantics worth knowing: `rescale` defaults to 1/255; `mean` and `std` must arrive together (both present turns normalize on, both absent leaves it off); and `process` also takes a file path or an in-memory byte span, folding the decode step in when you have not done it yourself. ## Or load the model's own recipe A checkpoint that ships `preprocessor_config.json` names its exact pipeline. `from_huggingface` reads that file (or the model directory holding it), so preprocessing follows the checkpoint instead of a hand-copied recipe. Unrecognized keys are ignored with a logged warning naming them. ```python proc = crt.processor.ImageProcessor.from_huggingface("SmolVLM-Instruct") features = proc.process("photo.jpg") # decode + the model's own recipe ``` ## The audio front end `AudioProcessor.from_args()` is the standard log-mel front end, 16 kHz and 80 mel bins by default; feed it a waveform and the sample rate it actually has. ```python title="audio_frontend.py" import math import clika_runtime as crt # 1 s of a 440 Hz tone, composed on the runtime: t = n / rate, s = 0.5 sin(2 pi f t). t = crt.arange(0.0, 16000.0, 1.0, dtype=crt.float32) / 16000.0 waveform = 0.5 * crt.sin(2.0 * math.pi * 440.0 * t) audio = crt.processor.AudioProcessor.from_args() feats = audio.process(waveform, sample_rate=16000) print(feats.dtype, list(feats.shape)) # float32, an 80-mel-bin axis inside ``` The frame axis follows the hop length, so its extent depends on the clip; the 80-bin axis is the front end's signature. The same `from_huggingface` path exists here, reading the audio half of a checkpoint's preprocessor config. Where this fits: [decode](load-images-and-audio.mdx) turns files into tensors; processors turn tensors into MODEL inputs; and the tensor a processor returns feeds a forward or a [compiled graph](run-an-onnx-model.mdx) directly. --- # Structure inputs and outputs as pytrees Flatten and rebuild nested containers of tensors, name every leaf by its key path, register your own classes and dataclasses as nodes, serialize a structure, and pass whole trees through compile, trace, eval and save. Source: https://docs.clika.io/clikart/how-to/pytrees.md A pytree is a nested arrangement of Python containers whose ends are the values you care about: a dict of lists of tensors, a tuple of dicts, a dataclass holding both. `clika_runtime.pytree` takes such a tree apart into a flat list of leaves plus a `TreeSpec` that remembers the shape, and puts the two halves back together. Every boundary of the Python surface that accepts several tensors at once (`compile`, `trace`, `eval`, `save` and `load`) accepts a pytree, so a model's inputs and outputs keep the structure the code was written in. The functions carry PyTorch's spellings (`tree_flatten`, `tree_map`, `register_pytree_node`), so code written against `torch.utils._pytree` reads the same here. Two rules differ from a plain container walk and are worth reading first. ## Leaves, and why None is one of them Lists, tuples, dicts, `OrderedDict`, `defaultdict`, `deque`, namedtuples, and every class registered with the node registry are nodes; everything else (numbers, strings, bytes, tensors, objects the registry does not know) is a leaf. Leaves come out depth first, left to right, and a flatten followed by an unflatten reproduces the same container types. ```python title="flatten.py" from clika_runtime import pytree tree = {"a": [1, (2, None)], "b": 3, "c": {"d": (4,), "e": []}} leaves, spec = pytree.tree_flatten(tree) print(leaves) # [1, 2, None, 3, 4] print(spec.num_leaves) # 5 print(spec) # TreeSpec({'a': [*, (*, *)], 'b': *, 'c': {'d': (*,), 'e': []}}) rebuilt = pytree.tree_unflatten(spec, leaves) assert rebuilt == tree and type(rebuilt) is dict ``` `None` is a leaf. A model that takes an optional input (a KV cache that is empty on the first step) keeps the same flat index for every other input whether or not the optional one is present, so a compiled graph's argument slots stay stable. `none_is_leaf=False` turns `None` into a childless node instead, and `is_leaf=lambda x: isinstance(x, list)` stops the descent at every list, which then counts as one leaf. ```python title="leaf_rules.py" from clika_runtime import pytree with_cache = {"input_ids": 1, "cache": (2, 3)} without_cache = {"input_ids": 1, "cache": None} leaves_a, _ = pytree.tree_flatten(with_cache) leaves_b, spec_b = pytree.tree_flatten(without_cache) print(leaves_a[0], leaves_b[0]) # 1 1 print(leaves_b) # [1, None] print(pytree.tree_unflatten(spec_b, [10, None])) # {'input_ids': 10, 'cache': None} leaves, spec = pytree.tree_flatten({"x": 1, "cache": None}, none_is_leaf=False) print(leaves) # [1] print(spec) # TreeSpec({'x': *, 'cache': None}, none_is_leaf=False) ``` ## Map and reduce `tree_map` applies a function to every leaf and rebuilds the same structure; extra trees pair their leaves by position and must share the first tree's structure (a mismatch raises `ValueError`). `tree_map_` runs the function for its side effect and returns the tree it was given; `tree_map_only` filters by type or by predicate. The reductions (`tree_all`, `tree_any`, `tree_reduce`, `tree_sum`, `tree_max`, `tree_min`) read the leaves without rebuilding anything, `tree_leaves` returns the flat list and `tree_iter` yields it lazily. ```python title="map.py" from clika_runtime import pytree print(pytree.tree_map(lambda x: x is None, {"x": 1, "y": None})) # {'x': False, 'y': True} print(pytree.tree_map(lambda x, y: x + y, {"a": 1, "b": 2}, {"a": 10, "b": 20})) # {'a': 11, 'b': 22} mixed = {"a": 1, "b": "text", "c": [2, None, 2.5]} print(pytree.tree_map_only(int, lambda x: x + 1, mixed)) # {'a': 2, 'b': 'text', 'c': [3, None, 2.5]} ``` ## Name every leaf by its key path `tree_flatten_with_path` returns `(path, leaf)` pairs; a path is a tuple of key entries (`MappingKey`, `SequenceKey`, `GetAttrKey`, `DataclassKey`), `keystr` renders it as the indexing expression that reaches the leaf, and `key_get` follows it. Key paths are how the file format names the leaves of a saved tree (below) and how a diagnostic points at one tensor inside a large state. ```python title="paths.py" from typing import NamedTuple from clika_runtime import pytree class Point(NamedTuple): x: object y: object tree = {"a": [1, 2], "p": Point(3, 4)} pairs = pytree.tree_flatten_with_path(tree)[0] print([leaf for _, leaf in pairs]) # [1, 2, 3, 4] print([pytree.keystr(path) for path, _ in pairs]) # ["['a'][0]", "['a'][1]", "['p'].x", "['p'].y"] path = (pytree.MappingKey("a"), pytree.SequenceKey(1)) print(pytree.key_get(tree, path)) # 2 ``` ## Register your own containers A class the registry does not know is a leaf. `register_pytree_node(cls, flatten_fn, unflatten_fn)` opens it: `flatten_fn` returns `(children, context)` or `(children, context, entries)` (the entries name the children for key paths), and `unflatten_fn(children, context)` rebuilds an instance. The argument order is PyTorch's. After registration, every tree function descends into the class and every rebuild produces an instance of it. ```python title="register.py" from clika_runtime import pytree class Point: def __init__(self, x, y): self.x, self.y = x, y pytree.register_pytree_node( Point, lambda p: ((p.x, p.y), None, ("x", "y")), lambda children, _ctx: Point(*children), ) print(pytree.tree_leaves({"p": Point(1, 2), "n": 3})) # [1, 2, 3] mapped = pytree.tree_map(lambda v: v * 10, Point(1, 2)) print(type(mapped).__name__, mapped.x, mapped.y) # Point 10 20 spec = pytree.tree_structure(Point(1, 2)) print(spec.entries(), spec.paths()) # ['x', 'y'] [('x',), ('y',)] ``` A class can carry its own flattening as a method pair and register with the decorator. Note the method order: `__tree_unflatten__(cls, context, children)` takes the context first, while the function form above takes `(children, context)`. ```python title="register_class.py" from clika_runtime import pytree @pytree.register_pytree_node_class class Pair: def __init__(self, a, b): self.a, self.b = a, b def __tree_flatten__(self): return (self.a, self.b), "pair", ("a", "b") @classmethod def __tree_unflatten__(cls, context, children): return cls(*children) leaves, spec = pytree.tree_flatten(Pair(1, [2, 3])) print(leaves, spec.context) # [1, 2, 3] pair rebuilt = pytree.tree_unflatten(spec, [10, 20, 30]) print(rebuilt.a, rebuilt.b) # 10 [20, 30] ``` Dataclasses register by field: the data fields become children, the meta fields ride the context and come back unchanged. ```python title="register_dataclass.py" import dataclasses from clika_runtime import pytree @dataclasses.dataclass class Sample: weight: object bias: object name: str pytree.register_dataclass(Sample, data_fields=["weight", "bias"], meta_fields=["name"]) leaves, spec = pytree.tree_flatten(Sample(1, 2, "s")) print(leaves) # [1, 2] print(pytree.tree_unflatten(spec, [10, 20])) # Sample(weight=10, bias=20, name='s') print([pytree.keystr(p) for p, _ in pytree.tree_flatten_with_path(Sample(1, 2, "s"))[0]]) # ['.weight', '.bias'] ``` Registering a built-in container, an instance instead of a class, or the same class twice raises; `unregister_pytree_node(cls)` makes instances leaves again. A registration meant for one library goes into a namespace so it never changes what other code sees: `register_pytree_node(..., namespace="mylib")` opens the class only for calls that pass `namespace="mylib"` (`tree_leaves(tree, namespace="mylib")`), `is_registered(cls, namespace="mylib")` reports it, and a namespace with no registration of its own falls back to the global one. ## Dictionary order Dicts flatten in insertion order, and the order is part of the structure: `{"b": 1, "a": 2}` and `{"a": 2, "b": 1}` have different specs. Code that wants two dicts with the same keys to share one structure regardless of insertion order wraps the calls in `dict_insertion_ordered(False)`, which flattens by sorted key inside the block; the rebuilt dict still comes back in the original order. ```python title="dict_order.py" from clika_runtime import pytree tree = {"b": 1, "a": 2, "c": 3} leaves, spec = pytree.tree_flatten(tree) print(leaves) # [1, 2, 3] print(list(pytree.tree_unflatten(spec, leaves))) # ['b', 'a', 'c'] with pytree.dict_insertion_ordered(False): print(pytree.tree_leaves({"b": 1, "a": 2})) # [2, 1] print(pytree.tree_structure({"b": 1, "a": 2}) == pytree.tree_structure({"a": 0, "b": 0})) # True ``` ## Treespecs A `TreeSpec` is a value: it compares and hashes by structure, prints as the container shape with a `*` per leaf, and rebuilds a tree from any sequence of the right length (`spec.unflatten(leaves)`; a wrong count raises `ValueError` naming the expected number). `is_prefix` asks whether one structure is the top of another, `compose` grafts an inner structure onto every leaf of an outer one, and `flatten_up_to` flattens a tree only as deep as the spec goes. ```python title="treespec.py" from clika_runtime import pytree spec = pytree.tree_structure({"a": [1, (2, None)], "b": 3}) print(spec) # TreeSpec({'a': [*, (*, *)], 'b': *}) print(spec.num_leaves, spec.num_nodes, spec.num_children) # 4 7 2 print(spec.paths()) # [('a', 0), ('a', 1, 0), ('a', 1, 1), ('b',)] short = pytree.tree_structure([1, 2]) deep = pytree.tree_structure([1, (2, 3)]) print(short.is_prefix(deep), deep.is_prefix(short)) # True False outer = pytree.tree_structure([1, 2]) inner = pytree.tree_structure((1, 2)) print(outer.compose(inner)) # TreeSpec([(*, *), (*, *)]) ``` `treespec_dumps` writes a spec as JSON and `treespec_loads` reads it back, so a structure can travel beside a file or a request. Built-in containers and namedtuples serialize as they are; a registered class needs a `serialized_type_name`, and a context that is not JSON needs the `to_dumpable_context` / `from_dumpable_context` pair at registration. An unnamed custom node or an unknown type name in the document raises `ValueError`. ```python title="serialize.py" import json from clika_runtime import pytree spec = pytree.tree_structure({"a": [1, (2, None)], "b": 3}) text = pytree.treespec_dumps(spec) print(json.loads(text)["version"] == pytree.SERIALIZATION_PROTOCOL) # True loaded = pytree.treespec_loads(text) print(loaded == spec) # True print(loaded.unflatten([1, 2, None, 3])) # {'a': [1, (2, None)], 'b': 3} ``` ## Pytrees at the boundaries `crt.compile` takes a callable whose arguments and return value are pytrees of tensors: it flattens the call's inputs, treats the tensor leaves as graph inputs and every other leaf as a static value that is part of the guard, and rebuilds the output structure on the way out; `dynamic=True` keeps one graph across batch sizes. ```python title="compile_tree.py" import numpy as np import clika_runtime as crt import clika_runtime.nn as nn BATCH, D_IN, D_OUT = 5, 131, 64 rng = np.random.default_rng(0) w = rng.standard_normal((D_IN, D_OUT)).astype(np.float32) * 0.1 b = rng.standard_normal(D_OUT).astype(np.float32) x = rng.standard_normal((BATCH, D_IN)).astype(np.float32) class Head(nn.Module): def __init__(self, w: np.ndarray, b: np.ndarray) -> None: super().__init__() self.weight = nn.Parameter(crt.tensor(w)) self.bias = nn.Parameter(crt.tensor(b)) def forward(self, x: crt.Tensor) -> crt.Tensor: return crt.softmax(crt.relu(crt.add(crt.matmul(x, self.weight), self.bias)), dim=-1) class Wrapped(nn.Module): def __init__(self) -> None: super().__init__() self.head = Head(w, b) def forward(self, batch: dict[str, crt.Tensor]) -> dict[str, crt.Tensor]: p = self.head(batch["x"]) return {"probs": p, "argmax": crt.argmax(p, dims=[1])} compiled = crt.compile(Wrapped(), dynamic=True) out = compiled({"x": crt.tensor(x)}) print(sorted(out)) # ['argmax', 'probs'] again = compiled({"x": crt.tensor(np.tile(x, (2, 1)))}) print(tuple(again["probs"].shape)) # (10, 64) ``` `crt.trace` records a function once over stand-ins and returns a graph that runs by position or by name; the traced function takes its tensor inputs as one list and returns a list, and [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) walks through it. Operators return before their work runs. `crt.eval(*trees)` flattens whatever trees it is given, skips the leaves that are not tensors, and returns once every tensor leaf holds its value; a later read then costs no wait. ```python title="eval_tree.py" import numpy as np import clika_runtime as crt x = np.random.default_rng(0).standard_normal((5, 131)).astype(np.float32) t = crt.tensor(x) tree = {"a": crt.exp(t), "b": [crt.sum(t), None, "text"], "c": (crt.relu(t),)} crt.eval(tree, crt.abs(t)) print(np.allclose(tree["a"].numpy(), np.exp(x.astype(np.float64)), rtol=1e-5, atol=1e-6)) # True ``` `crt.save` writes a tensor, a state dict, or any pytree of tensors as a safetensors file, and `crt.load` reads it back. A state dict (`crt.save(state, path)`, `crt.load(path, device="cpu")`) saves under its own keys in insertion order, so another safetensors reader sees it as it is; a deeper tree saves one entry per leaf named by `keystr` of its key path, with the structure carried in the file's metadata, and `crt.load` rebuilds the same tree. ```python title="save_tree.py" from pathlib import Path import numpy as np import clika_runtime as crt from clika_runtime import pytree rng = np.random.default_rng(0) tree = { "layers": [ {"w": crt.tensor(rng.standard_normal((32, 131)).astype(np.float32)), "b": crt.tensor(np.zeros(32, np.float32))} for _ in range(2) ], "step": crt.tensor(np.array([7], dtype=np.int64)), } crt.save(tree, Path("tree.safetensors")) back = crt.load(Path("tree.safetensors")) print(pytree.tree_structure(back) == pytree.tree_structure(tree)) # True print(sorted(pytree.keystr(kp) for kp, _ in pytree.tree_flatten_with_path(tree)[0])) # ["['layers'][0]['b']", "['layers'][0]['w']", "['layers'][1]['b']", "['layers'][1]['w']", "['step']"] ``` The same structure rule holds for a module's state: `crt.save(module.state_dict(), path)` followed by `other.load_state_dict(crt.load(path))` reproduces the forward, and the file is plain safetensors any other tool reads. [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module surface the leaves belong to. --- # Query a graph Read a ModelGraph as data: its nodes, values and edges, search by name or predicate, walks and topological orders, paths and the critical path, dominance, regions, cones and the cheapest cut, views across edits, and networkx. Source: https://docs.clika.io/clikart/how-to/query-a-graph.md {/* Every block is a program under examples//howto/query_a_graph/: the first block of each tab is get_a_graph whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output. */} A `ModelGraph`, traced from your own code or compiled from a model file, answers questions about its own structure: which operators it holds, what feeds what, which orders it can run in, which paths join two nodes, and where it can be split. The answers are views (`Node`, `Value` and `Edge`) that you read, compare and use as keys, and every query reads the graph as it stands. The graph queries are available from C++ and Python. ## Get a graph A trace records a function once over stand-in tensors and returns the graph as built: every operator stays as written, nothing optimized or finalized, which makes the graph easy to read. The model below has two branches that meet, one tensor read twice and two outputs. Every example on this page uses it, and `label` names a node by its input name or its operator. ```cpp title="get_a_graph.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::Node; using ClikaRT::graph::NodeKind; using ClikaRT::graph::OpCode; namespace ops = ClikaRT::ops; namespace { // Two branches that meet, one tensor read twice, two outputs. std::vector model(const std::vector& inputs) { const Tensor a = ops::relu(inputs[0]); const Tensor b = ops::sigmoid(inputs[0]); const Tensor c = ops::add(a, b); const Tensor d = ops::mul(c, c); // c feeds d twice: two edges, on input ports 0 and 1 const Tensor e = ops::tanh(a); return {ops::sub(d, e), b}; } // A node's label: an input's name, an operator's code. std::string label(const Node& node) { return node.kind() == NodeKind::Input ? node.name() : std::string(ClikaRT::graph::op_code_name(node.op_code())); } // The labels of `nodes`, separated by spaces. std::string labels(const std::vector& nodes) { std::string out; for (const Node& node : nodes) out += (out.empty() ? "" : " ") + label(node); return out; } // The trace returns the graph as recorded: every operator stays as written, // nothing optimized or finalized. ModelGraph trace_model() { const std::vector signature = {{"x", DataType::Float32, {2, 3}}}; return ClikaRT::graph::trace(model, signature, "query"); } } // namespace int main() { const ModelGraph graph = trace_model(); std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu Sigmoid Add Mul Tanh Sub std::printf("%zu\n", graph.find_nodes(OpCode::Mul).front().in_degree()); // 2 return 0; } ``` ```python title="get_a_graph.py" import clika_runtime as crt from clika_runtime.graph import NodeKind, OpCode def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]: x = inputs[0] a, b = crt.relu(x), crt.sigmoid(x) c = a + b d = c * c # c feeds d twice: two edges, on input ports 0 and 1 e = crt.tanh(a) return [d - e, b] def label(node: crt.graph.Node) -> str: return node.name if node.kind == NodeKind.Input else node.op_code.name # The trace returns the graph as recorded: every operator stays as written, # nothing optimized or finalized. graph = crt.trace(model, [crt.TensorSpec("x", crt.float32, [2, 3])]).graph print([label(node) for node in graph.nodes()]) # ['x', 'Relu', 'Sigmoid', 'Add', 'Mul', 'Tanh', 'Sub'] print(graph.find_nodes(OpCode.Mul)[0].in_degree()) # 2 ``` A compiled model answers the same queries: `compile()` returns the same `ModelGraph` type. The program builds a one-operator ONNX file first; any `.onnx` file works in its place. ```cpp title="compile_a_model.cpp" int main() { const ModelGraph graph = OnnxModel::open(relu_model_path()).compile(); std::printf("%s\n", labels(graph.nodes()).c_str()); // x Relu return 0; } ``` ```python title="compile_a_model.py" graph = crt.io.OnnxModel.open(path).compile() print([label(node) for node in graph.nodes()]) # ['x', 'Relu'] ``` Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, `label` and the trace). ## Nodes, values and edges A `Node` is an operator or a graph input. A `Value` is a tensor that a node produces on one output port, and an `Edge` is one read of a value, from the producer's output port to the reader's input port. The graph is a multigraph keyed by those ports: `d = c * c` reads one value over two edges. A node answers its inputs and outputs, `input(port)`, `output(port)`, `in_edges()` and `out_edges()`; a value answers `producer()`, `consumers()` and `uses()`. ```cpp title="nodes_and_edges.cpp" int main() { const ModelGraph graph = trace_model(); const Node mul = graph.find_nodes(OpCode::Mul).front(); const std::vector reads = mul.in_edges(); // one edge per input port, each keyed by its two ports for (const Edge& edge : reads) { std::printf("%s %d -> %s %d\n", label(edge.src()).c_str(), edge.src_port(), label(edge.dst()).c_str(), edge.dst_port()); } // Add 0 -> Mul 0 // Add 0 -> Mul 1 const Value value = *mul.input(0); // the tensor both edges carry const Node add = *value.producer(); std::printf("%s %s %s\n", value == *mul.input(1) ? "true" : "false", label(add).c_str(), labels(value.consumers()).c_str()); // true Add Mul std::printf("%zu %zu\n", graph.edges().size(), graph.edges_between(add, mul).size()); // 9 2 return 0; } ``` ```python title="nodes_and_edges.py" (mul,) = graph.find_nodes(OpCode.Mul) for edge in mul.in_edges(): # one edge per input port, each keyed by its two ports print(label(edge.src()), edge.src_port(), "->", label(edge.dst()), edge.dst_port()) # Add 0 -> Mul 0 # Add 0 -> Mul 1 value = mul.input(0) # the tensor both edges carry print(value == mul.input(1), label(value.producer()), [label(node) for node in value.consumers()]) # True Add ['Mul'] print(len(graph.edges()), len(graph.edges_between(value.producer(), mul))) # 9 2 ``` Read the views when your own code walks the structure, as an exporter or a quantizer does. ## Search by name and by predicate `find_nodes_by_name` matches a regular expression against a node's whole name, or any part of it with the `Anywhere` name match. `find_nodes` takes an `OpCode` or a predicate over a `Node`, and `node(name)` looks up one node. Each list comes back in name order. Operator names carry the operator and a number, so match them with a pattern rather than a literal. ```cpp title="search.cpp" int main() { const ModelGraph graph = trace_model(); const Regex relu = Regex::compile("Relu_[0-9]+"); const Regex sig = Regex::compile("Sig"); std::printf("%s\n", labels(graph.find_nodes_by_name(relu)).c_str()); // Relu std::printf("%s\n", labels(graph.find_nodes_by_name(sig, NameMatch::Anywhere)).c_str()); // Sigmoid std::printf("%s\n", labels(graph.find_nodes([](const Node& node) { return node.in_degree() == 2; })).c_str()); // Add Mul Sub std::printf("%s %s\n", label(*graph.node("x")).c_str(), labels(graph.find_nodes(OpCode::Tanh)).c_str()); // x Tanh return 0; } ``` ```python title="search.py" print([label(node) for node in graph.find_nodes_by_name("Relu_[0-9]+")]) # ['Relu'] print([label(node) for node in graph.find_nodes_by_name("Sig", crt.graph.NameMatch.Anywhere)]) # ['Sigmoid'] print([label(node) for node in graph.find_nodes(lambda node: node.in_degree() == 2)]) # ['Add', 'Mul', 'Sub'] print(label(graph.node("x")), [label(node) for node in graph.find_nodes(OpCode.Tanh)]) # x ['Tanh'] ``` ## Walks, visits, ancestors and descendants `walk(start)` lists the nodes a breadth-first walk reaches, the start first. The options `direction`, `max_depth`, `max_nodes`, `stop_at` and `edge_filter` (the fields of `WalkOptions` in C++, keyword arguments in Python) bound any walk. `descendants(node)` and `ancestors(node)` leave the node itself out and list the nearest first. `visit(start, visitor)` shows a callable each node, and the callable answers the `VisitAction` `Continue`, `Prune` (do not go past this node) or `Stop`. ```cpp title="walks.cpp" int main() { const ModelGraph graph = trace_model(); const Node x = *graph.node("x"); const Node relu = graph.find_nodes(OpCode::Relu).front(); const Node sub = graph.find_nodes(OpCode::Sub).front(); WalkOptions one_step; one_step.max_depth = 1; std::printf("%s\n", labels(graph.walk(x, one_step)).c_str()); // x Relu Sigmoid std::printf("%s\n", labels(graph.descendants(relu)).c_str()); // Add Tanh Mul Sub std::printf("%s\n", labels(graph.ancestors(sub)).c_str()); // Mul Tanh Add Relu Sigmoid x std::vector seen; graph.visit(x, [&seen](const Node& node) { seen.push_back(node); return node.op_code() == OpCode::Add ? VisitAction::Prune : VisitAction::Continue; }); // the walk does not go past the Add, so the Mul is never shown std::printf("%s\n", labels(seen).c_str()); // x Relu Sigmoid Add Tanh Sub return 0; } ``` ```python title="walks.py" x = graph.node("x") (relu,) = graph.find_nodes(OpCode.Relu) (sub,) = graph.find_nodes(OpCode.Sub) print([label(node) for node in graph.walk(x, max_depth=1)]) # ['x', 'Relu', 'Sigmoid'] print([label(node) for node in graph.descendants(relu)]) # ['Add', 'Tanh', 'Mul', 'Sub'] print([label(node) for node in graph.ancestors(sub)]) # ['Mul', 'Tanh', 'Add', 'Relu', 'Sigmoid', 'x'] seen: list[str] = [] def look(node: crt.graph.Node) -> crt.graph.VisitAction: seen.append(label(node)) return crt.graph.VisitAction.Prune if node.op_code == OpCode.Add else crt.graph.VisitAction.Continue graph.visit(x, look) # the walk does not go past the Add, so the Mul is never shown print(seen) # ['x', 'Relu', 'Sigmoid', 'Add', 'Tanh', 'Sub'] ``` Use a walk to collect everything upstream or downstream of a node, and a visit when the decision to go on depends on the node you are at. ## Topological orders, generations and depths `topological_order()` lists every node after the nodes it reads, in the `Deterministic` order by default (`Structural` and `MinMemory` are the other two `TopologicalOrder` values). `topological_generations()` groups the nodes by the longest path that ends at them, and `depths()` gives that length per node. `random_topological_order(seed)` draws one of the valid orders; the programs check the order they draw instead of printing it, since the order depends on the seed. ```cpp title="orders.cpp" int main() { const ModelGraph graph = trace_model(); std::printf("%s\n", labels(graph.topological_order()).c_str()); // x Relu Sigmoid Add Mul Tanh Sub std::string layers; for (const std::vector& layer : graph.topological_generations()) layers += "[" + labels(layer) + "]"; std::printf("%s\n", layers.c_str()); // [x][Relu Sigmoid][Add Tanh][Mul][Sub] std::string depths; for (const auto& [node, depth] : graph.depths()) { depths += (depths.empty() ? "" : " ") + label(node) + "=" + std::to_string(depth); } std::printf("%s\n", depths.c_str()); // x=0 Relu=1 Sigmoid=1 Add=2 Mul=3 Tanh=2 Sub=4 // Any valid order, drawn from the seed: check it rather than print it. std::printf("%s\n", graph.is_topological_order(graph.random_topological_order(7)) ? "true" : "false"); // true return 0; } ``` ```python title="orders.py" print([label(node) for node in graph.topological_order()]) # ['x', 'Relu', 'Sigmoid', 'Add', 'Mul', 'Tanh', 'Sub'] print([[label(node) for node in layer] for layer in graph.topological_generations()]) # [['x'], ['Relu', 'Sigmoid'], ['Add', 'Tanh'], ['Mul'], ['Sub']] print({label(node): depth for node, depth in graph.depths().items()}) # {'x': 0, 'Relu': 1, 'Sigmoid': 1, 'Add': 2, 'Mul': 3, 'Tanh': 2, 'Sub': 4} print(graph.is_topological_order(graph.random_topological_order(7))) # True: any valid order, drawn from the seed ``` A generation's nodes read only earlier generations, so they are the nodes that can run at the same time. ## Paths and path counts A path never repeats a node. `all_simple_paths(src, dst)` lists the paths as nodes, where parallel edges give one path, and `all_simple_edge_paths` lists them as edges, where each parallel edge gives its own. Both search at the call, and Python hands the paths back through an iterator. Without a `limit`, a search that finds more than `kMaxPathsWithoutLimit` paths (`MAX_PATHS_WITHOUT_LIMIT` in Python) fails with an invalid-argument error instead of returning part of the list; `count_paths` counts without listing and has no such cap. `shortest_path` gives the path with the fewest edges and `has_path` answers reachability. The options `cutoff`, `avoid` and `edge_filter` (the fields of `PathOptions` in C++, keyword arguments in Python) bound every path query. ```cpp title="paths.cpp" int main() { const ModelGraph graph = trace_model(); const Node x = *graph.node("x"); const Node sub = graph.find_nodes(OpCode::Sub).front(); // The two edges into the Mul give one path. const std::vector> paths = graph.all_simple_paths(x, sub); for (const std::vector& path : paths) std::printf("%s\n", labels(path).c_str()); // x Relu Add Mul Sub // x Relu Tanh Sub // x Sigmoid Add Mul Sub std::printf("%llu %zu\n", static_cast(graph.count_paths(x, sub)), graph.all_simple_edge_paths(x, sub).size()); // 3 5 std::printf("%s %s\n", labels(graph.shortest_path(x, sub)).c_str(), graph.has_path(sub, x) ? "true" : "false"); // x Relu Tanh Sub false PathOptions first; first.limit = 1; std::printf("%zu %zu\n", graph.all_simple_paths(x, sub, first).size(), ClikaRT::graph::kMaxPathsWithoutLimit); // 1 10000 return 0; } ``` ```python title="paths.py" x = graph.node("x") (sub,) = graph.find_nodes(OpCode.Sub) for path in graph.all_simple_paths(x, sub): # the two edges into the Mul give one path print([label(node) for node in path]) # ['x', 'Relu', 'Add', 'Mul', 'Sub'] # ['x', 'Relu', 'Tanh', 'Sub'] # ['x', 'Sigmoid', 'Add', 'Mul', 'Sub'] print(graph.count_paths(x, sub), len(list(graph.all_simple_edge_paths(x, sub)))) # 3 5 print([label(node) for node in graph.shortest_path(x, sub)], graph.has_path(sub, x)) # ['x', 'Relu', 'Tanh', 'Sub'] False print(len(list(graph.all_simple_paths(x, sub, limit=1))), crt.graph.MAX_PATHS_WITHOUT_LIMIT) # 1 10000 ``` ## The critical path `critical_path(node_cost, edge_cost)` returns the costliest path as a `WeightedPath` carrying its nodes, its edges and its summed cost: the longest chain of work when each node costs its run time. Given an `edge_cost`, `shortest_path` returns the cheapest path in the same form. ```cpp title="critical_path.cpp" int main() { const ModelGraph graph = trace_model(); // A cost per node (and optionally per edge): here the Sigmoid costs 2, every other node 1. const WeightedPath path = graph.critical_path([](const Node& node) { return node.op_code() == OpCode::Sigmoid ? 2.0 : 1.0; }); std::printf("%s %g\n", labels(path.nodes).c_str(), path.cost); // x Sigmoid Add Mul Sub 6 return 0; } ``` ```python title="critical_path.py" # A cost per node (and optionally per edge): here the Sigmoid costs 2, every other node 1. path = graph.critical_path(lambda node: 2.0 if node.op_code == OpCode.Sigmoid else 1.0) print([label(node) for node in path.nodes], path.cost) # ['x', 'Sigmoid', 'Add', 'Mul', 'Sub'] 6.0 ``` ## Structure: cycles, dominance, regions and cones `find_cycle()` is empty for every traced or compiled graph, and `would_create_cycle(src, dst)` checks an edge before an edit adds it. `weakly_connected_components()` groups the nodes that edges join in either direction. `dominators()` maps each node to the last node that every path from the inputs to it crosses, `split_points()` lists the nodes that every path from the inputs to the outputs crosses, and `merge_point(nodes)` finds where the paths from several nodes meet. `region(entry, exit)` is the single-entry, single-exit block between two nodes, `producer_cone` and `consumer_cone` are everything a set of nodes reads or feeds, and `is_convex` tells whether a set can be cut out and replaced as one piece. ```cpp title="structure.cpp" int main() { const ModelGraph graph = trace_model(); std::map node; for (const Node& n : graph.nodes()) node.emplace(label(n), n); std::printf("%zu %s\n", graph.find_cycle().size(), graph.would_create_cycle(node.at("Sub"), node.at("x")) ? "true" : "false"); // 0 true std::printf("%zu\n", graph.weakly_connected_components().size()); // 1 std::string dominators; for (const auto& [n, dominator] : graph.dominators()) { dominators += (dominators.empty() ? "" : " ") + label(n) + ":" + label(dominator); } std::printf("%s\n", dominators.c_str()); // x:x Relu:x Sigmoid:x Add:x Mul:Add Tanh:Relu Sub:x std::printf("%s %s\n", labels(graph.split_points()).c_str(), label(*graph.merge_point({node.at("Add"), node.at("Tanh")})).c_str()); // x Sub std::printf("%s\n", labels(graph.region(node.at("Add"), node.at("Mul"))).c_str()); // Add Mul std::printf("%s\n", labels(graph.producer_cone({node.at("Mul")})).c_str()); // x Relu Sigmoid Add std::printf("%s %s\n", graph.is_convex({node.at("Add"), node.at("Mul")}) ? "true" : "false", graph.is_convex({node.at("Relu"), node.at("Mul")}) ? "true" : "false"); // true false return 0; } ``` ```python title="structure.py" node = {label(n): n for n in graph.nodes()} print(graph.find_cycle(), graph.would_create_cycle(node["Sub"], node["x"])) # [] True print(len(graph.weakly_connected_components())) # 1 print({label(n): label(dominator) for n, dominator in graph.dominators().items()}) # {'x': 'x', 'Relu': 'x', 'Sigmoid': 'x', 'Add': 'x', 'Mul': 'Add', 'Tanh': 'Relu', 'Sub': 'x'} print([label(n) for n in graph.split_points()], label(graph.merge_point([node["Add"], node["Tanh"]]))) # ['x'] Sub print([label(n) for n in graph.region(node["Add"], node["Mul"])]) # ['Add', 'Mul'] print([label(n) for n in graph.producer_cone([node["Mul"]])]) # ['x', 'Relu', 'Sigmoid', 'Add'] print(graph.is_convex([node["Add"], node["Mul"]]), graph.is_convex([node["Relu"], node["Mul"]])) # True False ``` ## The cheapest cut `min_value_cut(before, after)` finds the split of the nodes into two parts that hands the fewest bytes from the first part to the second, and returns a `ValueCut` of the values that cross and their bytes. `before` and `after` pin nodes to either part. Of the cheapest cuts it returns the one nearest the inputs; networkx's `minimum_cut` returns the one nearest the sink instead. Use it to decide where to split a model across two devices or two processes. ```cpp title="value_cut.cpp" int main() { const ModelGraph graph = trace_model(); std::map node; for (const Node& n : graph.nodes()) node.emplace(label(n), n); // The producers of the values that cross the cut, and their bytes. const auto show = [](const ValueCut& cut) { std::string crossing; for (const Value& value : cut.values) crossing += (crossing.empty() ? "" : " ") + label(*value.producer()); std::printf("%s %llu\n", crossing.c_str(), static_cast(cut.bytes)); }; show(graph.min_value_cut()); // x 24: every float32 [2, 3] value holds 24 bytes show(graph.min_value_cut({node.at("Add")}, {node.at("Sub")})); // Relu Sigmoid Add 72 return 0; } ``` ```python title="value_cut.py" node = {label(n): n for n in graph.nodes()} cut = graph.min_value_cut() # every float32 [2, 3] value holds 24 bytes print([label(value.producer()) for value in cut.values], cut.bytes) # ['x'] 24 pinned = graph.min_value_cut(before=[node["Add"]], after=[node["Sub"]]) print([label(value.producer()) for value in pinned.values], pinned.bytes) # ['Relu', 'Sigmoid', 'Add'] 72 ``` ## Views across edits An edit such as `optimize()` changes the graph under the views you hold. A view finds its node, value or edge again by its key (a node's name, a value's producer and port, an edge's two ends and ports) and keeps answering. Once the key is gone, `check()` refuses, naming the key, and so does a query given the view. In C++ the view's other reads answer empty; in Python they raise the same `InvalidArgumentError` as `check()`. Equality and hashing keep working in both, so a map or dictionary keyed by views still finds its entries. ```cpp title="edits.cpp" int main() { const std::vector signature = {{"x", DataType::Float32, {2, 3}}}; ModelGraph twice = ClikaRT::graph::trace( [](const std::vector& inputs) -> std::vector { return {ops::relu(ops::relu(inputs[0]))}; }, signature, "twice"); const Node inner = twice.node("x")->successors().front(); const Node outer = inner.successors().front(); // Views key a map across edits. const std::unordered_map index = {{inner, "inner"}, {outer, "outer"}}; twice.optimize(); // Relu(Relu(x)) is Relu(x): the optimizer removes the outer Relu inner.check(); // the kept Relu still answers std::printf("%s %s\n", label(inner).c_str(), inner.output(0)->is_graph_output() ? "true" : "false"); // Relu true try { outer.check(); // refuses, naming the node, and so does a query given the view } catch (const Error& error) { std::printf("%s\n", error.code_name().c_str()); // INVALID_ARGUMENT } // Every other read of the gone view answers empty. std::printf("%s '%s'\n", index.at(outer).c_str(), outer.name().c_str()); // outer '' return 0; } ``` ```python title="edits.py" twice = crt.trace(lambda inputs: [crt.relu(crt.relu(inputs[0]))], [crt.TensorSpec("x", crt.float32, [2, 3])]).graph (inner,) = twice.node("x").successors() (outer,) = inner.successors() index = {inner: "inner", outer: "outer"} # views key a dict across edits twice.optimize() # Relu(Relu(x)) is Relu(x): the optimizer removes the outer Relu print(inner.check(), label(inner), inner.output(0).is_graph_output()) # None Relu True try: outer.check() # so does every other read of the view, and any query given it except crt.InvalidArgumentError: print("the outer Relu is gone") # the outer Relu is gone print(index[outer], repr(outer).startswith("Node(gone: ")) # outer True ``` ## Hand the graph to networkx `to_networkx()` returns the graph as a `networkx.MultiDiGraph`, for the algorithms networkx provides. Its nodes are the `Node` views, which keep the graph alive, and each edge is keyed `(src_port, dst_port)`. It needs the `networkx` package, which the program imports as `nx`. networkx is a Python library, so this section has a Python program only. From C++, `edges()` and the node views carry the same structure to any graph library. ```python title="to_networkx.py" nx_graph = graph.to_networkx() # a networkx.MultiDiGraph whose nodes are the Node views print(nx_graph.number_of_nodes(), nx_graph.number_of_edges()) # 7 9 print(nx.dag_longest_path_length(nx_graph)) # 4 print(sorted(label(node) for node in nx_graph.successors(graph.node("x")))) # ['Relu', 'Sigmoid'] ``` --- # Run an ONNX model Open an ONNX file (or build one from scratch), compile it into a ModelGraph, optimize and finalize it, and execute it by position or by name. Source: https://docs.clika.io/clikart/how-to/run-an-onnx-model.md {/* CERTIFICATION: the python and C++ arms are verified against the pinned release (tools/tutorial_check.py over the programs the page embeds: build_onnx, inspect_onnx and run_onnx all match their recordings in both languages, on the release's cp313 wheel and its payload). Given no path, the C++ and Python IO-contract programs build the tiny MLP into a temporary file, so they run with no file on disk. */} You have a model as an `.onnx` file and want ClikaRT to run it. Three types carry the whole story. `io::OnnxModel` is the format-level object, the ONNX graph as data: open it, inspect it, edit it, save it. `compile()` is the one crossing into the executable world, where every node parses into its runtime operator, weights bind, and shapes resolve. The result is a `ModelGraph` as built. `optimize()` runs the graph optimizer when you want it, `finalize()` places the graph and packs its weights, and the finalized graph has one `run`. The programs below first open an existing file and read its contract, then build a tiny model from scratch so the page runs with no file on disk, then compile and run it. Any `.onnx` file works in the first section; nothing depends on the architecture. ## Open a model and read its IO contract Opening is cheap and does not compile anything: you get the graph as data. The IO contract (names, dtypes, dims, with dynamic dims reading as named placeholders or -1) is what you need to prepare feeds; print it before anything else when a model is new to you. ```cpp title="inspect_onnx.cpp (the inspection)" int main(int argc, char** argv) { const std::string path = argc > 1 ? argv[1] : tiny_model_path(); OnnxModel model = OnnxModel::open(path); std::printf("nodes: %zu initializers: %zu opset: %lld\n", model.num_nodes(), model.num_initializers(), static_cast(model.opset())); // Hold the returned vectors: a range-for over a temporary would iterate // storage that is gone before the first step. const std::vector inputs = model.inputs(); const std::vector outputs = model.outputs(); for (const auto& s : inputs) std::printf(" input %s\n", s.name.c_str()); for (const auto& s : outputs) std::printf(" output %s\n", s.name.c_str()); return 0; } ``` ```python title="inspect_onnx.py (the inspection)" model = crt.io.OnnxModel.open(sys.argv[1] if len(sys.argv) > 1 else tiny_model_path()) print(f"nodes: {model.num_nodes} initializers: {model.num_initializers} " f"opset: {model.opset}") for name, dtype, dims in model.inputs(): print(f" input {name} {list(dims)}") # a dynamic dim reads as -1 for name, dtype, dims in model.outputs(): print(f" output {name} {list(dims)}") ``` The program opens the path on its command line; with none, `tiny_model_path()` builds the MLP of the next section into a temporary file, so the page runs with no file on disk. ## Build a graph from scratch When there is no file yet (a test, a fixture, a tool that emits ONNX), the same object builds a graph node by node: declare inputs, add initializers (weights enter as ordinary tensors), add nodes by operator type, name the outputs. The tiny MLP here is `y = relu(X W + B)` with a dynamic batch dimension. ```cpp title="build_onnx.cpp (excerpt of the build stage)" OnnxModel build_tiny_mlp() { Dim batch; // one Dim object: every use is the SAME dynamic dimension OnnxModel model = OnnxModel::create("tiny_mlp", /*opset_version=*/21); model.add_input("X", DataType::Float32, {batch, 4}); const std::string w = model.add_initializer( Tensor::full({4, 3}, 0.5, DataType::Float32)); const std::string b = model.add_initializer( Tensor::full({3}, 0.25, DataType::Float32)); const std::vector mm = model.add_node("MatMul", {"X", w}, 1); const std::vector sum = model.add_node("Add", {mm[0], b}, 1); const std::vector y = model.add_node("Relu", {sum[0]}, 1); model.add_output(y[0], DataType::Float32, {batch, 3}); return model; } ``` ```python title="build_onnx.py (the build stage)" def build_tiny_mlp() -> "crt.io.OnnxModel": model = crt.io.OnnxModel.create("tiny_mlp", opset_version=21) model.add_input("X", crt.float32, [-1, 4]) # any value <= 0 is dynamic w = model.add_initializer(crt.tensor(np.full((4, 3), 0.5, dtype=np.float32))) b = model.add_initializer(crt.tensor(np.full((3,), 0.25, dtype=np.float32))) (mm,) = model.add_node("MatMul", ["X", w]) (summed,) = model.add_node("Add", [mm, b]) (out,) = model.add_node("Relu", [summed]) model.add_output(out, crt.float32, [-1, 3]) return model ``` `save(path)` writes the graph; `open(path)` round-trips it structure-intact, and `optimize()` runs the rewrite pipeline to a fixed point (a minimal graph survives unchanged). ## Compile and run `compile()` can fail like any load of real weights and shapes, so it returns through the error contract; branch on the code name. It returns the graph as built. `optimize()` runs the graph optimizer when you want it, and `finalize()` readies the graph to run ([Optimize and finalize](optimize-and-finalize.mdx) covers both). The finalized `ModelGraph` runs positionally (one tensor per `input_names()` entry, in that order) and, in Python, also by name. ```cpp title="run_onnx.cpp (compile and run)" int run(ClikaRT::io::OnnxModel& model) { ClikaRT::Result compiled = CLIKART_TRY(model.compile()); if (!compiled.ok()) { std::printf("compile: FAILED [%s]\n", compiled.code_name().c_str()); return 1; } ModelGraph graph = std::move(compiled.value()); graph.optimize(); // the graph optimizer, when wanted graph.finalize(); // only a finalized graph runs const std::vector x = {1.0F, 2.0F, 3.0F, 4.0F, -1.0F, 0.5F, 2.0F, -2.0F}; std::vector outputs = graph.run( {Tensor::from_data(x.data(), {2, 4}, DataType::Float32)}); std::printf("y = %s\n", outputs[0].to_string().c_str()); return 0; } ``` ```python title="run_onnx.py (compile and run)" graph = model.compile() graph.optimize() # the graph optimizer, when wanted graph.finalize() # only a finalized graph runs print(f"inputs {graph.input_names()} -> outputs {graph.output_names()}") x = crt.tensor(np.array([[1, 2, 3, 4], [-1, 0.5, 2, -2]], dtype=np.float32)) (y,) = graph.run([x]) # positional: input_names() order (y2,) = graph.run({"X": x}) # named: the same result print(y.numpy()) codes = [n.op_code for n in graph.nodes()] # the compiled graph as data assert crt.graph.OpCode.MatMul in codes # The one-call loader: open, compile, optimize and finalize, then call the model with named inputs. model.save("tiny_mlp.onnx") loaded = crt.onnx.load("tiny_mlp.onnx", dynamic_axes={"X": {0: "batch"}}, device="cpu") out = loaded(X=x) # a dict keyed by output name assert list(out) == loaded.output_names ``` `crt.onnx.load` opens, compiles, optimizes and finalizes in one call: `dynamic_axes` names the axes that vary between calls, `device` places the weights, and the returned model takes its inputs by keyword and returns a dict keyed by output name. `optimize=False` returns the graph as built, which runs after `model.graph.finalize()`. `graph.nodes()` reads the compiled graph as data, one `crt.graph.OpCode` per node; [Trace eager code to graphs](trace-eager-code-to-graphs.mdx) walks that surface. With the fixed all-half weights and the 0.25 bias above, the math fits in your head; the C++ program's run over the two-row input prints: ```text y = relu(X W + B): [5.25, 5.25, 5.25] [0, 0, 0] ``` The first row is `(1+2+3+4) * 0.5 + 0.25 = 5.25` per output; the second row's pre-activation is negative in every column, so relu zeroes it. Two pointers from here. A compiled `ModelGraph` is the same type that [tracing eager code](trace-eager-code-to-graphs.mdx) produces, so everything downstream of `compile()` is shared, [optimizing and finalizing](optimize-and-finalize.mdx) included. And a model too big to build by hand arrives as a file: the IO-contract section works unchanged on a checkpoint you downloaded. --- # Serve a model over HTTP Put compute behind HTTP endpoints with the built-in server, answer JSON requests with tensor results, and stream tokens with server-sent events. Source: https://docs.clika.io/clikart/how-to/serve-over-http.md You have working compute and want it behind an HTTP endpoint. `ClikaRT::http` ships a server in the same library: declare routes with lambdas, parse and build bodies with `ClikaRT::Json`, and stream with server-sent events. No web framework enters the ship path. Each program below starts a server, drives it with the built-in HTTP client in the same process, and prints the exchange, so it runs self-contained. To poke a server from outside instead, replace `bind_to_any_port` with `listen_async("0.0.0.0", 8080)` and use curl. A served process runs compute like any other, so give it a license credential where you set the rest of its environment (`CLIKA_RT_LICENSE`, or the per-user file `clikart-license-init` writes; [Get ClikaRT](../getting-started/get-clikart.mdx#license-credential)). Without one, a handler the runtime refuses for licensing reports the code name `LICENSE_FAILED`. ## Start a server and add routes Routes are declared before the server starts: `get`/`post` for fixed paths, `route` with a regex for path parameters (`req.param(0)` is the first capture). Handlers return a `ServerResponse` and run concurrently on the server's I/O pool, so anything they share needs a lock. ```cpp title="routes.cpp" #include #include #include namespace http = ClikaRT::http; using ClikaRT::json::Json; using ClikaRT::Result; int main() { http::HttpServer server = http::HttpServer::create(); server.get("/health", [](http::ServerRequest&) -> Result { Json o = Json::object(); o["status"] = "ok"; o["runtime"] = ClikaRT::GetVersionInfo(); return http::ServerResponse::json(o.dump()); }); server.route(http::Method::Get, R"(/models/(\w+))", [](http::ServerRequest& req) -> Result { Json o = Json::object(); o["model"] = req.param(0); o["loaded"] = false; return http::ServerResponse::json(o.dump()); }); const int port = server.bind_to_any_port("127.0.0.1"); server.wait_until_ready(); const std::string base = "http://127.0.0.1:" + std::to_string(port); std::printf("GET /health -> %s\n", http::get_text(base + "/health").c_str()); std::printf("GET /models/smol -> %s\n", http::get_text(base + "/models/smol").c_str()); server.stop(); return 0; } ``` The Python server lands with the HTTP bindings of the `clika_runtime.http` module; the samples on this page run once the wheel carries it. Routes are decorators or calls: `get` / `post` for fixed paths, `route(method, pattern, pattern=True)` for a regular expression with named captures, read back with `request.param(name)`. A handler returns an `http.Response`, a `str` (sent as `text/plain`) or a `dict` / `list` (sent as JSON). Handlers run on the server's own threads and hold the interpreter lock only while they run Python code. ```python title="routes.py" import http.client import json from clika_runtime import http as crt_http server = crt_http.HttpServer() @server.get("/health") def health(request: crt_http.Request) -> dict[str, str]: return {"status": "ok"} # a dict is sent as JSON @server.route("GET", r"/models/(?P[^/]+)", pattern=True) def model(request: crt_http.Request) -> str: return f"model {request.param('name')}" # a str is sent as text/plain @server.post("/echo") def echo(request: crt_http.Request) -> crt_http.Response: return crt_http.Response.bytes(request.body, "application/octet-stream") with server: # the block ends with server.stop() port = server.bind_to_any_port("127.0.0.1") client = http.client.HTTPConnection("127.0.0.1", port) client.request("GET", "/health") print(json.loads(client.getresponse().read())) # {'status': 'ok'} client.request("GET", "/models/tiny") print(client.getresponse().read().decode()) # model tiny client.request("POST", "/echo", body=b"ping") print(client.getresponse().read()) # b'ping' ``` `server.use(middleware)` wraps every route with `middleware(request, next)`; `server.group(prefix)` scopes routes under a path prefix; `listen(host, port)` blocks the calling thread until `stop()` and `listen_async` serves on the server's own thread. Every one of them releases the interpreter lock while it waits. The Kotlin binding does not carry the serving runtime; the C++ arm is the serving story today. A served model is consumed from Kotlin with any JVM HTTP client, and [Modelverse's serving guide](/modelverse/how-to/serve-openai-compatible) shows the OpenAI-compatible route. ```text GET /health -> {"status":"ok","runtime":"0.6.4"} GET /models/smol -> {"model":"smol","loaded":false} ``` ## A JSON inference endpoint The serving shape every model endpoint repeats: parse the body, validate, build a tensor from the request, compute, and put the result back into JSON. Bad input gets a clean 400 with a reason, not an exception. The compute here is one dense layer with fixed weights, standing in for a loaded model ([the GGUF guide](load-quantized-weights.mdx) is where real weights come from). The Python tab carries the compute half for real: the same scoring model compiled once to a `ModelGraph` and run per request; only the HTTP transport stays C++. ```cpp title="score_endpoint.cpp" #include #include #include #include #include namespace http = ClikaRT::http; namespace ops = ClikaRT::ops; using ClikaRT::DataType; using ClikaRT::json::Json; using ClikaRT::Result; using ClikaRT::Tensor; constexpr std::int64_t kFeatures = 4; int main() { // The "model": y = x * W^T + b, weights fixed for a reproducible page. const Tensor w = Tensor::full({2, kFeatures}, 0.5, DataType::Float32); const Tensor b = Tensor::full({2}, 0.25, DataType::Float32); http::HttpServer server = http::HttpServer::create(); server.post("/score", [&](http::ServerRequest& req) -> Result { Result body = CLIKART_TRY(Json::parse(req.body())); if (!body.ok()) return http::ServerResponse::json(R"({"error":"invalid json"})", 400); Result feats = CLIKART_TRY(body.value().at("features")); if (!feats.ok() || feats.value().size() != kFeatures) { Json err = Json::object(); err["error"] = Json("'features' must hold " + std::to_string(kFeatures) + " numbers"); return http::ServerResponse::json(err.dump(), 400); } std::vector x(kFeatures); for (std::size_t i = 0; i < kFeatures; ++i) x[i] = static_cast(feats.value().at(i).as_double()); const Tensor input = Tensor::from_data(x.data(), {1, kFeatures}, DataType::Float32); const std::vector scores = ops::linear(input, w, b).reshape({-1}).item_as_vec(); Json out = Json::object(); out["scores"] = Json::array(); for (float s : scores) out["scores"].push_back(Json(static_cast(s))); return http::ServerResponse::json(out.dump()); }); const int port = server.bind_to_any_port("127.0.0.1"); server.wait_until_ready(); const std::string base = "http://127.0.0.1:" + std::to_string(port); std::printf("POST /score [1,2,3,4] -> %s\n", http::post_text(base + "/score", R"({"features":[1, 2, 3, 4]})").c_str()); // Client helpers raise ClikaRT::Error on any status >= 400; CLIKART_TRY // captures that as a Result when a failure is an expected outcome. Result bad = CLIKART_TRY(http::post_text(base + "/score", R"({"features":[1]})")); std::printf("POST /score [1] -> %s\n", bad.ok() ? bad.value().c_str() : bad.message().c_str()); server.stop(); return 0; } ``` ```python title="score_compute.py" # The model is compiled to a ModelGraph once at startup and run per request; # the route below is the served endpoint (the ONNX chapters of the python # examples build bigger graphs the same way). import json import numpy as np import clika_runtime as crt from clika_runtime import http as crt_http FEATURES = 4 # The "model": y = x * W^T + b, as a compiled graph with fixed weights. model = crt.io.OnnxModel.create("score", opset_version=21) model.add_input("X", crt.float32, [1, FEATURES]) w = model.add_initializer(crt.tensor(np.full((FEATURES, 2), 0.5, dtype=np.float32))) b = model.add_initializer(crt.tensor(np.full(2, 0.25, dtype=np.float32))) (mm,) = model.add_node("MatMul", ["X", w]) (scores,) = model.add_node("Add", [mm, b]) model.add_output(scores, crt.float32, [1, 2]) graph = model.compile() # every node parses, weights bind, shapes resolve graph.optimize() # the graph optimizer graph.finalize() # only a finalized graph runs def score(body: str) -> tuple[int, str]: """The handler shape: parse, validate, run the graph, answer JSON.""" try: feats = json.loads(body).get("features") except ValueError: return 400, json.dumps({"error": "invalid json"}) if not isinstance(feats, list) or len(feats) != FEATURES: return 400, json.dumps({"error": "'features' must hold 4 numbers"}) x = crt.tensor(np.asarray([feats], dtype=np.float32)) (y,) = graph.run({"X": x}) return 200, json.dumps({"scores": y.numpy().reshape(-1).tolist()}) server = crt_http.HttpServer() @server.post("/score") def score_route(request: crt_http.Request) -> crt_http.Response: status, body = score(request.body.decode()) return crt_http.Response.json(json.loads(body), status=status) print("POST /score [1,2,3,4] ->", *score('{"features":[1, 2, 3, 4]}')) # 200 {"scores": [5.25, 5.25]} print("POST /score [1] ->", *score('{"features":[1]}')) # 400 {"error": "'features' must hold 4 numbers"} ``` A bad body answers a 400 with a reason, a good one a 200 with the scores; the route runs the same `score` the two prints call. The Kotlin binding carries the serving runtime (`FunctionModel`, `Executor`, `Pipeline`) and the HTTP server (`HttpServer`); the worked program on this page is the C++ one, and [Language bindings](../bindings.md) says what Kotlin carries. ```text POST /score [1,2,3,4] -> {"scores":[5.25,5.25]} POST /score [1] -> 400 Bad Request on POST /score ``` The handler captures the weight tensors by reference; they outlive the server. A real model swaps the `ops::linear` line for its forward and nothing else changes shape. On the failure, the in-process client surfaces the status as the `Result`'s message; an external client (curl, a browser) reads the JSON error body the handler wrote. ## Stream results with server-sent events Token-by-token streaming (the transport behind LLM chat responses) is one factory away: `ServerResponse::sse` takes a `next` callback, and the server pulls it until it returns an empty optional, one SSE frame per event. The client here buffers the finite stream and prints the raw frames; a browser or an SSE-aware client consumes them incrementally. ```cpp title="stream_tokens.cpp" #include #include #include #include #include namespace http = ClikaRT::http; using ClikaRT::Result; namespace { constexpr const char* kTokens[] = {"Tensors ", "stream ", "one ", "by ", "one."}; constexpr int kTokenCount = static_cast(sizeof kTokens / sizeof kTokens[0]); } // namespace int main() { http::HttpServer server = http::HttpServer::create(); server.get("/generate", [](http::ServerRequest&) -> Result { auto sent = std::make_shared(0); // per-connection cursor return http::ServerResponse::sse( [sent]() -> Result> { if (*sent >= kTokenCount) return std::optional{}; // close http::ServerSentEvent ev; ev.event = "token"; ev.data = kTokens[(*sent)++]; return std::optional{ev}; }); }); const int port = server.bind_to_any_port("127.0.0.1"); server.wait_until_ready(); const std::string raw = http::get_text("http://127.0.0.1:" + std::to_string(port) + "/generate"); std::printf("raw SSE body:\n%s", raw.c_str()); server.stop(); return 0; } ``` A streamed body pulls its frames from a generator: `crt_http.Response.sse(events)` sends one `event:` / `data:` frame per item, and the server pulls the next item on its own thread as the client reads. The producer is the generation loop; [the tokenizer guide's streaming-decode section](tokenize-and-chat-templates.mdx) turns its ids into exactly the text pieces a stream sends, one non-empty piece per `token` event. ```python title="sse.py" from clika_runtime import http as crt_http server = crt_http.HttpServer() @server.get("/events") def events(request: crt_http.Request) -> crt_http.Response: def frames(): for piece in ("Tensors ", "stream ", "one ", "by ", "one."): yield {"event": "token", "data": piece} # one frame per generated piece return crt_http.Response.sse(frames()) ``` The same route with a decoder behind it yields each non-empty `push` result of the streaming decoder as a frame's `data`. The Kotlin binding carries the serving runtime, the streaming decoder and the HTTP server; the worked program on this page is the C++ one, and [Language bindings](../bindings.md) says what Kotlin carries. ```text raw SSE body: event: token data: Tensors event: token data: stream event: token data: one event: token data: by event: token data: one. ``` Each frame is an `event:` line, a `data:` line, and a blank line; a generation loop replaces the fixed token array with reads from its decoder, and [the tokenizer guide's streaming-decode section](tokenize-and-chat-templates.mdx) is that decoder: each non-empty `push` result is one frame's `data`. The server pulls `next` on an I/O worker, so a slow producer stalls only its own connection. Middleware (logging, auth), static mounts, and an image-upload endpoint that runs compute are in the bundle's `http_server` example, chapters `03` to `06`. Wiring a served endpoint into sessions and continuous batching is the `runtime` example's ground. The finished version of this page's story ships in Modelverse: [an OpenAI-compatible endpoint](/modelverse/how-to/serve-openai-compatible) over these same server pieces, with its `01_serve` example as the smallest complete server ([Additional examples](/modelverse/examples) has it). --- # Tokenize text and apply a chat template Turn text into token ids and back, align tokens to source bytes, render a conversation with the model's own chat template, and batch for a model. Source: https://docs.clika.io/clikart/how-to/tokenize-and-chat-templates.md Your model consumes token ids, and a chat model expects its prompt formatted exactly the way it was trained. `ClikaRT::Tokenizer` covers both: one loader reads a HuggingFace model directory, `encode`/`decode` convert text to ids and back, and the model's own chat template renders conversations. No Python and no external tokenizer library are involved. The programs below use the tokenizer of [SmolLM2-135M-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct), the model from [the GGUF guide](load-quantized-weights.mdx). Two small files are all a tokenizer needs: ```bash mkdir -p SmolLM2-135M-Instruct && cd SmolLM2-135M-Instruct curl -LO "https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/resolve/main/tokenizer.json" curl -LO "https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/resolve/main/tokenizer_config.json" cd .. ``` The bundle also ships a self-contained tokenizer under `examples/src/tokenizer/data/hf_model`, if you would rather not download anything. ## Load a tokenizer and round-trip some text `Tokenizer::from_huggingface` takes the model directory, detects the artifact inside it (`tokenizer.json`, `tokenizer.model`, `tekken.json`, or `vocab.json` plus `merges.txt`), and overlays `tokenizer_config.json` for the special-token ids and the chat template. `Tokenizer::from_file` loads one tokenizer file (or a directory holding one) directly, detecting its format the same way. ```cpp title="roundtrip.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); std::printf("vocab %lld, bos %lld, eos %lld, chat template: %s\n", static_cast(tok.vocab_size()), static_cast(tok.bos_id()), static_cast(tok.eos_id()), tok.has_chat_template() ? "yes" : "no"); const std::string text = "ClikaRT runs the same code on every backend."; const std::vector ids = tok.encode(text); std::printf("encoded %zu tokens:", ids.size()); for (std::int32_t id : ids) std::printf(" %d", id); std::printf("\ndecoded: %s\n", tok.decode(ids).c_str()); return 0; } ``` ```python title="roundtrip.py" def main() -> None: tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir()) print(f"vocab {tok.vocab_size}, bos {tok.bos_id}, eos {tok.eos_id}, " f"chat template: {'yes' if tok.has_chat_template else 'no'}") text = "ClikaRT runs the same code on every backend." ids = tok.encode(text) print(f"encoded {len(ids)} tokens:", *ids) print(f"decoded: {tok.decode(ids)}") if __name__ == "__main__": main() ``` ```text vocab 49152, bos 1, eos 2, chat template: yes encoded 12 tokens: 51 1418 6335 16895 7313 260 1142 2909 335 897 25817 30 decoded: ClikaRT runs the same code on every backend. ``` ## See where each token came from `tokenize` returns one `Token` per piece: the id, the surface string, and the byte span `[begin, end)` in the original text. Slice the original by that span when you need alignment (highlighting, span labeling, streaming cursors); the spans line up exactly, dropped spaces included. ```cpp title="offsets.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); const std::string text = "Quantized weights stay packed."; const std::vector tokens = tok.tokenize(text, /*add_special_tokens=*/false); std::printf(" id [begin,end) source span\n"); for (const Token& t : tokens) { const std::string_view span(text.data() + t.begin, t.end - t.begin); std::printf(" %-6d [%2zu,%2zu) \"%.*s\"\n", t.id, t.begin, t.end, static_cast(span.size()), span.data()); } return 0; } ``` ```python title="offsets.py" def main() -> None: tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir()) text = "Quantized weights stay packed." ids = tok.encode(text, add_special_tokens=False) # The python binding returns ids; id_to_token shows each piece. A # leading 'G-with-breve' marks a token that starts with a space; the # byte-span view is the C++ tokenize surface. print(" id token") for i in ids: print(f" {i:<6} {tok.id_to_token(i)!r}") if __name__ == "__main__": main() ``` ```text id [begin,end) source span 24696 [ 0, 5) "Quant" 1005 [ 5, 9) "ized" 10379 [ 9,17) " weights" 2951 [17,22) " stay" 13448 [22,29) " packed" 30 [29,30) "." ``` ## Render a conversation with the model's chat template A chat model's prompt format (its role markers, turn separators, generation priming) ships with the model as a Jinja2 template in `tokenizer_config.json`, and the loader attached it above. `apply_chat_template` renders a `messages` array the OpenAI-API shape into the exact prompt string; `encode_chat` goes straight to ids. Never hand-build these markers: the template is the model's contract. ```cpp title="chat_template.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); Json messages = Json::array(); Json system = Json::object(); system["role"] = "system"; system["content"] = "You are a concise assistant."; messages.push_back(std::move(system)); Json user = Json::object(); user["role"] = "user"; user["content"] = "What does a tokenizer do?"; messages.push_back(std::move(user)); const std::string prompt = tok.apply_chat_template(messages); std::printf("=== rendered prompt ===\n%s\n=======================\n", prompt.c_str()); // The template writes the prompt's own bos/eos framing, so encode_chat adds // no special tokens on top of it. const std::vector ids = tok.encode_chat(messages, /*add_generation_prompt=*/true); std::printf("encode_chat produced %zu tokens\n", ids.size()); return 0; } ``` ```python title="chat_template.py" def main() -> None: tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir()) # The messages array travels as JSON text at this surface. messages = json.dumps([ {"role": "system", "content": "You are a concise assistant."}, {"role": "user", "content": "What does a tokenizer do?"}, ]) prompt = tok.apply_chat_template(messages) print(f"=== rendered prompt ===\n{prompt}\n=======================") # The template writes the prompt's own bos/eos framing, so encode_chat adds # no special tokens on top of it. ids = tok.encode_chat(messages, add_generation_prompt=True) print(f"encode_chat produced {len(ids)} tokens") if __name__ == "__main__": main() ``` ```text === rendered prompt === <|im_start|>system You are a concise assistant.<|im_end|> <|im_start|>user What does a tokenizer do?<|im_end|> <|im_start|>assistant ======================= encode_chat produced 107 tokens ``` The rendered prompt ends with the assistant-turn priming (`add_generation_prompt` defaults to true), so the model continues as the assistant. For tool calling, extra template variables, or a reproducible clock, pass a `ChatTemplateInputs` instead of the bare messages array; the two-argument form above covers plain conversations. ## Batch for a model Feeding a model takes tensors, not vectors. `encode_batch` with `return_tensors` produces the standard quartet: padded `input_ids` `[B, S]`, an `attention_mask`, per-sequence lengths, and the `cu_seqlens` prefix-sum table. `varlen = true` skips padding entirely and lays the ids out flat, the shape variable-length attention consumes. ```cpp title="batch.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); const std::vector texts = { "Short prompt.", "A somewhat longer prompt that pads the short one.", }; EncodeOptions opts; opts.return_tensors = true; // Decoder-only checkpoints often ship no pad token; designate one (eos is // the usual choice) or the padded encode raises ClikaRT::Error. opts.pad_id = static_cast(tok.eos_id()); const ClikaRT::tokenizer::Encoded batch = tok.encode_batch(texts, opts); std::printf("input_ids %s\n", batch.input_ids->to_string().c_str()); std::printf("attention_mask %s\n", batch.attention_mask->to_string().c_str()); std::printf("seq_lengths %s\n", batch.seq_lengths->to_string().c_str()); opts.varlen = true; const ClikaRT::tokenizer::Encoded flat = tok.encode_batch(texts, opts); std::printf("varlen ids %s\n", flat.input_ids->to_string().c_str()); std::printf("cu_seqlens %s\n", flat.cu_seqlens->to_string().c_str()); return 0; } ``` ```python title="batch.py" def main() -> None: tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir()) texts = [ "Short prompt.", "A somewhat longer prompt that pads the short one.", ] # The python binding returns the ragged ids, one list per text; pad on # the tensor side with the lengths below. The padded quartet (input_ids, # attention_mask, seq_lengths, cu_seqlens) is the C++ and C surface. batch = tok.encode_batch(texts) for row, ids in enumerate(batch.ids): print(f"text {row}: {len(ids):2} ids {ids}") if __name__ == "__main__": main() ``` ```text input_ids Tensor(shape=[2, 10], dtype=Int32, device=CPU, numel=20, data=[20355, 6011, 30, 2, 2, 2, ...]) attention_mask Tensor(shape=[2, 10], dtype=Int32, device=CPU, numel=20, data=[1, 1, 1, 0, 0, 0, ...]) seq_lengths Tensor(shape=[2], dtype=Int32, device=CPU, numel=2, data=[3, 10]) varlen ids Tensor(shape=[13], dtype=Int32, device=CPU, numel=13, data=[20355, 6011, 30, 49, 7932, 2848, ...]) cu_seqlens Tensor(shape=[3], dtype=Int32, device=CPU, numel=3, data=[0, 3, 13]) ``` `EncodeOptions` also carries truncation (`max_length`, `truncation_side`), the padding side (`Left` suits decoder-only batch generation), and a target `device` so the tensors land where the model computes. The bundle's `tokenizer` example walks each of these one chapter at a time, and the `templating` example covers the Jinja2-compatible engine behind `apply_chat_template` on its own. ## Stream the decode of a generation loop A generation loop produces ids one at a time, and `decode(ids)` over the growing list re-decodes everything on every step. `Tokenizer::streaming_decoder` is the incremental form: `push(id)` returns exactly the newly-stable text, and the pieces concatenate to what `decode` would have produced. The catch it handles for you is the UTF-8 boundary: one code point can span tokens, so `push` holds bytes back until they are displayable and returns an empty string meanwhile; you never emit half a character. `finish()` flushes whatever the tail held (a trailing incomplete sequence as-is) and resets the decoder for a fresh stream; a decoder serves one stream at a time. ```cpp title="stream_decode.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); // Stand-in for a generation loop: the ids a real decoder would emit one // at a time (the roundtrip section's sentence, so the ids match). const std::vector ids = tok.encode("ClikaRT runs the same code on every backend."); StreamingDecoder stream = tok.streaming_decoder(); std::string assembled; int emitted = 0; for (std::int32_t id : ids) { // push returns exactly the newly-stable text: empty while a // multi-byte code point is still incomplete, never a torn character. const std::string piece = stream.push(id); if (!piece.empty()) ++emitted; assembled += piece; } assembled += stream.finish(); // flush the tail; the decoder resets std::printf("%zu ids -> %d incremental pieces\n", ids.size(), emitted); std::printf("assembled: %s\n", assembled.c_str()); std::printf("assembled == decode(ids): %s\n", assembled == tok.decode(ids) ? "yes" : "no"); return 0; } ``` ```python title="stream_decode.py" def main() -> None: tok = crt.tokenizer.Tokenizer.from_huggingface(model_dir()) # Stand-in for a generation loop: the ids a real decoder would emit one # at a time (the roundtrip section's sentence, so the ids match). ids = tok.encode("ClikaRT runs the same code on every backend.") stream = tok.streaming_decoder() pieces = [stream.push(i) for i in ids] # "" while a code point is incomplete assembled = "".join(pieces) + stream.finish() # flush; the decoder resets print(f"{len(ids)} ids -> {sum(1 for p in pieces if p)} incremental pieces") print(f"assembled: {assembled}") print(f"assembled == decode(ids): " f"{'yes' if assembled == tok.decode(ids) else 'no'}") if __name__ == "__main__": main() ``` {/* CERTIFICATION: the run below follows the API contract (the pieces concatenate to decode(ids)); the piece count is not verified against a built bundle. */} ```text 32 ids -> 32 incremental pieces assembled: ClikaRT runs the same code on every backend. assembled == decode(ids): yes ``` Every push emitted text here because the sentence is plain ASCII; text with accents, CJK, or emoji is where the empty returns appear, and exactly why the boundary handling exists. The decoder skips special tokens by default (`streaming_decoder(false)` keeps them), and the handle stays valid even after the `Tokenizer` that made it is gone, so a generation worker can own just the decoder. This is the producer half of token streaming: each non-empty piece is one frame for the transport. [The serving guide](serve-over-http.mdx) sends exactly these pieces as `token` events over server-sent events. --- # Trace eager code to graphs Capture an eager function or module as a ModelGraph with trace, inspect it, optimize, finalize and run it, and wrap hot paths in compile for capture-and-replay. Source: https://docs.clika.io/clikart/how-to/trace-eager-code-to-graphs.md Eager code runs op by op. Tracing runs your function ONCE over data-free stand-ins (shapes and dtypes matter, values are never read) and captures the operator graph as a `ModelGraph`, the same type the ONNX loader compiles to. The graph comes back as recorded: `optimize()` runs the graph optimizer when you want it, and `finalize()` readies it for the same `run`. Two rules make a function traceable: - **List in, list out.** A traced callable receives its input tensors as one list and returns its outputs as a list. A bare tensor return does not auto-wrap; return `[y]`. - **No value reads.** The stand-ins carry no data, so reading a value during tracing raises. Shape-driven math is fine. In Python the two rules relax to pytrees: a traced function may take a list, a dict or any nested container of tensors and return one, and the graph keeps the names. No value is read during the capture in either form. ```python title="trace_function.py" import numpy as np import clika_runtime as crt w = crt.tensor(np.full((4, 3), 0.1, dtype=np.float32)) b = crt.tensor(np.zeros(3, dtype=np.float32)) def fn(inputs: list[crt.Tensor]) -> list[crt.Tensor]: return [crt.softmax(crt.relu(crt.add(crt.matmul(inputs[0], w), b)), dim=-1)] x = crt.tensor(np.ones((2, 4), dtype=np.float32)) # Capture: fn runs once over stand-ins; the graph names its inputs and outputs. graph = crt.trace(fn, example_inputs=[x], input_names=["x"], output_names=["p"]) print(graph.input_names(), graph.output_names()) # ['x'] ['p'] # The graph comes back as recorded: optimize it, then finalize it to run. graph.optimize() graph.finalize() # The graph runs like the function did, by position or by name. (positional,) = graph.run([x]) (named,) = graph.run({"x": x}) print(positional.numpy()[0]) # [0.33333334 0.33333334 0.33333334] ``` A module traces the same way, and the traced graph is inspectable node by node: each node carries an `op_code` from `crt.graph.OpCode` and a `kind`. ```python title="trace_module.py" import numpy as np import clika_runtime as crt import clika_runtime.nn as nn class Head(nn.Module): def __init__(self, w: np.ndarray, b: np.ndarray) -> None: super().__init__() self.weight = nn.Parameter(crt.tensor(w)) self.bias = nn.Parameter(crt.tensor(b)) def forward(self, x: crt.Tensor) -> crt.Tensor: return crt.softmax(crt.relu(crt.add(crt.matmul(x, self.weight), self.bias)), dim=-1) rng = np.random.default_rng(0) head = Head(rng.standard_normal((4, 3)).astype(np.float32), np.zeros(3, dtype=np.float32)) x = crt.tensor(np.ones((2, 4), dtype=np.float32)) traced = crt.trace(head, example_inputs=(x,)) graph = traced.graph # the ModelGraph, as recorded codes = {node.op_code for node in graph.nodes()} print(crt.graph.OpCode.MatMul in codes, crt.graph.OpCode.Softmax in codes) # True True print(sum(node.kind == crt.graph.NodeKind.Input for node in graph.nodes())) # 1 ``` As recorded, the graph holds one node per operator the forward calls, the `MatMul`, the `Add`, the `Relu` and the `Softmax`, after its one input node. `optimize()` folds the bias add and the relu into the `MatMul`, which leaves it and the `Softmax`. The same capture surface in C++: `ClikaRT::graph::trace` takes the callable and the example inputs and returns the `ModelGraph` as recorded, and `ClikaRT::compile` wraps a callable in capture-and-replay exactly as below. [Optimize and finalize](optimize-and-finalize.mdx) walks `optimize()`, `to()` and `finalize()` in C++, and the shipped examples include a full tracing walkthrough. ## The compile wrapper `compile` wraps a callable in capture-and-replay: the first call runs eagerly AND captures; later calls replay the graph. A shape change recaptures, invisible to values and visible on the counter. ```python title="compile_function.py" step = crt.compile(fn, [crt.TensorSpec("x", crt.float32, [2, 4])]) print(step.state) # State.Pending (State.Pending: nothing captured yet) (first,) = step([x]) # runs eagerly and records the graph print(step.state) # State.Compiled (State.Compiled: later calls replay) (again,) = step([x]) # replays the captured graph; Python is not called print(step.recapture_count) # 0 (0) graph = step.take_graph() # the captured ModelGraph, for standalone use (replayed,) = graph.run([x]) ``` The pytree form takes a module or a function over dicts and returns dicts; `dynamic=True` keeps one graph across batch sizes instead of recapturing on a shape change. ```python title="compile_module.py" class Wrapped(nn.Module): def __init__(self) -> None: super().__init__() self.head = head def forward(self, batch: dict[str, crt.Tensor]) -> dict[str, crt.Tensor]: p = self.head(batch["x"]) return {"probs": p, "argmax": crt.argmax(p, dims=[1])} compiled = crt.compile(Wrapped(), dynamic=True) out = compiled({"x": x}) print(sorted(out)) # ['argmax', 'probs'] (['argmax', 'probs']) again = compiled({"x": crt.tensor(np.ones((4, 4), dtype=np.float32))}) print(again["probs"].shape) # clika_runtime.Size([4, 3]) ``` `fullgraph=True` turns a fallback into an error: a Python branch on a tensor value cannot be recorded, and `crt.compile(fn, fullgraph=True)` raises `crt.ClikaRTError` at the call instead of serving it eagerly. `step.reset()` re-arms the wrapper, and `step.state` reads `Pending`, `Compiled` or `Fallback`. --- # Use ClikaRT from Python The clika-runtime wheel: NumPy in and out with explicit copy semantics, math that reads as math, modes as strings, models as nn.Module, and errors typed by class. Source: https://docs.clika.io/clikart/how-to/use-clikart-from-python.md The `clika-runtime` wheel puts the runtime behind one import: `import clika_runtime as crt` loads `libClikaRT.so` and its backends from inside the wheel, with no library paths to set. The Python surface follows PyTorch's shapes: dtype objects such as `crt.float32`, a string spelling for every mode argument, `nn.Module` for models, and one exception class per failure kind. NumPy plays three roles: data entry, data exit, and the independent oracle you check results against. Compute runs in the runtime. This guide assumes the wheel is installed; [First steps](../getting-started/index.md) covers getting it. Everything below is one script's worth of ground: the license credential, the NumPy boundary, operator chains, modes as strings, device placement, a model as `nn.Module`, and the error contract. ## The credential goes in before the import The import is what loads the runtime, so `CLIKA_RT_LICENSE` has to hold the credential before `import clika_runtime` runs. Export it in the shell, or assign it above the import: ```python title="license.py" import os os.environ["CLIKA_RT_LICENSE"] = "CLIKA1-..." # the credential text, or the path of a file holding it import clika_runtime as crt ``` The alternative drops the variable: `clikart-license-init `, a console script the wheel installs, stores the credential once under your user account. `clika_runtime.torch` and `clika_runtime.modelverse` follow the same rule, because all three are the one runtime. [License the runtime](license-the-runtime.mdx) is the whole contract; without a valid credential a call raises `crt.ClikaRTError` with `code_name` `LICENSE_FAILED`. ## NumPy in, NumPy out `crt.tensor(array)` copies the array in, `crt.from_numpy(array)` borrows its memory, and `t.numpy()` is the exit. Lists and scalars enter too, at the dtype NumPy would pick for them. ```python title="boundary.py" import numpy as np import clika_runtime as crt a = np.ones((2, 5), dtype=np.float32) t = crt.tensor(a) print(t.shape, t.dtype, t.device) # clika_runtime.Size([2, 5]) clika_runtime.float32 cpu assert t.dtype == crt.float32 # one dtype object per storable dtype assert isinstance(crt.float32, crt.dtype) assert t.device == crt.Device("cpu") a[0, 0] = 999.0 # entry copied: the tensor is unmoved assert t.numpy()[0, 0] == 1.0 borrowed = crt.from_numpy(a) # from_numpy shares the array's memory assert borrowed.numpy()[0, 0] == 999.0 assert crt.tensor([1, 2, 3]).dtype == crt.int64 half = t.to(crt.float16) # narrowing is explicit assert half.dtype == crt.float16 print(half) # tensor([[1., 1., 1., 1., 1.], # [1., 1., 1., 1., 1.]], dtype=clika_runtime.float16) ``` Every NumPy-native dtype enters as itself (the float family, the signed ints, `uint8`, `bool`); a dtype with no tensor twin, `complex64` for example, is refused with a `TypeError` that names the routes out. Payload dtypes NumPy cannot spell, bfloat16 among them, cross through `bytes()` and `Tensor.from_bytes()` instead of the array bridge. A tensor prints as `tensor([...])`: the dtype is named when it is not `float32`, the device when it is not the CPU, and a tensor above a thousand elements is abbreviated to its edge items. A tensor prints as its values, wrapped in `tensor(...)`, with no shape or device header: ```python print(2 * crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3)) + 3) ``` ```text tensor([[ 3., 5., 7.], [ 9., 11., 13.]]) ``` ## Math that reads as math Operators compose the way the expression reads: Python numbers broadcast, `@` is matmul, and method chains mirror the functional forms. Check anything against NumPy; that is what the oracle role means. ```python title="tensor_math.py" import numpy as np import clika_runtime as crt x = crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3)) y = 2.0 * x + 3.0 # scalars broadcast z = (y - 3.0).abs().amax().item() # a method chain down to one float print(z) # 10.0 assert np.allclose((x ** 2).numpy(), x.numpy() ** 2) a = crt.tensor(np.ones((2, 3), dtype=np.float32)) b = crt.tensor(np.ones((3, 2), dtype=np.float32)) print((a @ b).numpy()) # [[3. 3.] # [3. 3.]] ``` ## Modes are strings A mode argument takes its spelling as a string: `approximate="tanh"`, `mode="reflect"`, `rounding_mode="floor"`, `activation="relu"`. An unknown spelling raises a `ValueError` that lists the accepted ones. ```python title="modes.py" import numpy as np import clika_runtime as crt x = crt.tensor(np.linspace(-3.0, 3.0, 7, dtype=np.float32)) tanh_form = crt.gelu(x, approximate="tanh") assert not np.array_equal(tanh_form.numpy(), crt.gelu(x).numpy()) assert np.array_equal(crt.gelu(x).numpy(), crt.gelu(x, approximate="none").numpy()) padded = crt.pad(x, [2, 2], mode="reflect") assert np.array_equal(padded.numpy(), np.pad(x.numpy(), (2, 2), mode="reflect")) try: crt.div(x, 2.0, rounding_mode="ceil") except ValueError as e: print(e) # rounding_mode: 'ceil' is not a RoundingMode; choose one of 'none', 'trunc', 'floor' ``` ## Placement Placement is a constructor argument or a move: `crt.tensor(arr, device=...)` lands data where you say, `.to("cpu")` moves it, and `crt.Device.gpu()` names the machine's accelerator, or the CPU when it has none, so the same script runs everywhere. Each backend has a namespace: `crt.cuda`, `crt.vulkan` and `crt.metal` mirror `crt.accelerator`, with `is_available()`, `device(index)` and `synchronize()`. ```python title="placement.py" import numpy as np import clika_runtime as crt gpu = crt.Device.gpu() t = crt.zeros(2, device=gpu) assert t.device == gpu assert np.array_equal(crt.to(t, "cpu").numpy(), np.zeros(2, dtype=np.float32)) if crt.cuda.is_available(): x = crt.ones(2, 3, device=crt.cuda.device(0)) crt.cuda.synchronize() print(crt.Device("cpu")) # cpu ``` ## A model is an nn.Module Assigning a layer in `__init__` registers it, as in PyTorch: `load_state_dict` binds dotted names, `named_parameters()` enumerates them, `state_dict()` exports the same names back out, and the instance is callable. A `Parameter` is a `Tensor`. ```python title="model.py" import numpy as np import clika_runtime as crt import clika_runtime.nn as nn class TinyMlp(nn.Module): def __init__(self, d_in: int, d_hidden: int, d_out: int) -> None: super().__init__() # The first layer fuses its activation as an epilogue. self.up = nn.Linear(d_in, d_hidden, activation="relu") self.down = nn.Linear(d_hidden, d_out, bias=False) def forward(self, x: crt.Tensor) -> crt.Tensor: return self.down(self.up(x)) model = TinyMlp(4, 8, 2) assert isinstance(model.up.weight, nn.Parameter) and isinstance(model.up.weight, crt.Tensor) result = model.load_state_dict({ "up.weight": crt.tensor(np.full((8, 4), 0.1, dtype=np.float32)), "up.bias": crt.tensor(np.zeros(8, dtype=np.float32)), "down.weight": crt.tensor(np.full((2, 8), 0.1, dtype=np.float32)), }) assert result.missing_keys == [] and result.unexpected_keys == [] y = model(crt.tensor(np.ones((3, 4), dtype=np.float32))) print(y.shape) # clika_runtime.Size([3, 2]) exported = model.state_dict() # the same dotted names back out print(sorted(exported)) # ['down.weight', 'up.bias', 'up.weight'] ``` `load_state_dict` is strict by default: a missing or unexpected key raises a `RuntimeError` naming both sets; `strict=False` returns the report instead, with `missing_keys` and `unexpected_keys`. The `state_dict()` -> fresh `load_state_dict()` round trip reproduces the forward, which is the portable way to hand weights between processes. [Author a model in Python](author-a-model-in-python.mdx) builds a full decoder this way. ## When it fails A runtime failure raises an exception typed by its kind: `crt.InvalidArgumentError` (also a `ValueError`) for a shape, dtype, device or option the call cannot accept, `crt.NotFoundError` (also a `FileNotFoundError`) for a missing file or entry, `crt.UnsupportedError`, `crt.OutOfMemoryError` and `crt.UnavailableError` for their kinds, all under `crt.ClikaRTError`. Every instance carries `.code_name` (the fine code, stable across builds) and `.status` (the coarse class); `str(err)` is the message alone. Branch on the class or the code name, never on the message text: an argument mistake reads as a sentence naming the operation and the values, while an `E` message is an internal fault code specific to the build that produced it (report it verbatim with the runtime version). ```python title="errors.py" import numpy as np import clika_runtime as crt try: a = crt.tensor(np.ones((2, 3), dtype=np.float32)) b = crt.tensor(np.ones((4, 5), dtype=np.float32)) _ = a @ b # shape mismatch except crt.InvalidArgumentError as e: assert isinstance(e, ValueError) print(f"failed with code {e.code_name!r}") # failed with code 'INVALID_ARGUMENT' try: crt.load("absent.safetensors") except crt.NotFoundError as e: assert isinstance(e, FileNotFoundError) assert e.code_name == "NO_SUCHFILE" ``` From here, the rest of the Python surface follows the same grammar: [GGUF and quantized weights](load-quantized-weights.mdx) and [tokenizers](tokenize-and-chat-templates.mdx) have Python arms on their pages, [ONNX models](run-an-onnx-model.mdx) compile and run, [tracing](trace-eager-code-to-graphs.mdx) turns eager functions into graphs, [pytrees](pytrees.mdx) carry structured inputs and outputs across those boundaries, and [PyTorch interop](use-clikart-with-pytorch.mdx) covers the torch backend and tensor exchange. The wheel's own example programs double as a smoke suite for an installed wheel. --- # Use ClikaRT with PyTorch Run a torch.compile model on the runtime with backend=\"clika\", move tensors across the two libraries over DLPack without a copy, and convert an eager torch.nn.Module tree into runtime layers. Source: https://docs.clika.io/clikart/how-to/use-clikart-with-pytorch.md `clika_runtime.torch` is the PyTorch side of the wheel: a `torch.compile` backend that lowers the captured graph onto the runtime's operators, a tensor exchange over DLPack that shares memory instead of copying, and a converter from `torch.nn.Module` trees to `clika_runtime.nn` layers. It ships with the `torch` extra: ```bash pip install "clika-runtime[torch]" ``` Importing `clika_runtime` on its own never imports torch; only `clika_runtime.torch` reaches for it, at the first use, and it raises an `ImportError` naming the extra when torch is absent. A process that never touches the interop package pays nothing for it. ## Compile a model onto the runtime Registration puts the backend under the name `"clika"`; from then on `torch.compile(model, backend="clika")` sends every captured graph to the runtime, and the compiled module returns torch tensors like any other backend. Registering twice is a no-op; registering under a name torch already owns (`"inductor"`) raises. ```python title="compile.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime.torch as crt_torch crt_torch.register_torch_backend() print("clika" in torch._dynamo.list_backends()) # True torch.manual_seed(2) model = torch.nn.Sequential(torch.nn.Linear(131, 64), torch.nn.ReLU(), torch.nn.Linear(64, 131)).eval() x = torch.randn(5, 131) with torch.no_grad(): expected = model(x) got = torch.compile(model, backend="clika")(x) print(got.shape) # torch.Size([5, 131]) print(torch.allclose(got, expected, rtol=1e-3, atol=1e-3)) # True ``` The wheel also declares the backend as a `torch_dynamo_backends` entry point, so an installed distribution serves `torch.compile(model, backend="clika")` with no `clika_runtime.torch` import in the calling code: torch resolves the name through `importlib.metadata.entry_points(group="torch_dynamo_backends")`, where `clika` loads `clika_runtime.torch.clika_backend`. Passing the function itself works too: `torch.compile(model, backend=crt_torch.clika_backend)`. `torch_backend(name)` registers under a name of your choosing for the length of a `with` block and restores torch's backend registry on exit, byte for byte: ```python title="scoped_backend.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime.torch as crt_torch torch._dynamo.list_backends(None) # torch imports its own backends lazily with crt_torch.torch_backend("clika-scoped") as backend: print("clika-scoped" in torch._dynamo.list_backends()) # True print("clika-scoped" in torch._dynamo.list_backends()) # False ``` ### What the backend does with a graph `torch.compile` hands the backend an FX graph: placeholders for the inputs, one node per operation, an output node. The backend lowers every node onto a runtime operator once and records the result; every later call converts the inputs, replays the recorded runtime graph, and converts the outputs back. Static shapes specialize the graph, as they do for any dynamo backend. A graph break in the function (a `print`, a Python branch on a value) splits the capture into several graphs, each lowered on its own; `captured_graphs()` lists what the backend has lowered in this process and `clear_captured_graphs()` empties the list. `clika_runtime.torch.compile(fn, **torch_compile_kwargs)` is `torch.compile(fn, backend=clika_backend, ...)` with one addition: an operator the runtime cannot lower raises `LoweringError` out of the compiled call, and one error names every unsupported operator in the graph rather than the first. ```python title="graph_breaks.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime.torch as crt_torch def fn(a: torch.Tensor, b: torch.Tensor) -> torch.Tensor: h = torch.relu(a @ b) print("break") # a graph break: two graphs; prints twice (the recording run, then the replay) return torch.tanh(h).sum(dim=-1) a, b = torch.randn(7, 131), torch.randn(131, 29) got = crt_torch.compile(fn)(a, b) print(torch.allclose(got, fn(a, b), rtol=1e-3, atol=1e-3)) # True print(len(crt_torch.captured_graphs())) # 2 def unsupported(a: torch.Tensor) -> torch.Tensor: return torch.special.zeta(a, a) + torch.special.bessel_j0(a) try: crt_torch.compile(unsupported, fullgraph=True)(torch.rand(3, 131) + 2.0) except crt_torch.LoweringError as error: print(error) # 2 unsupported torch operation(s) in the graph: # torch._C._special.special_zeta (node: special_zeta) # torch._C._special.special_bessel_j0 (node: special_bessel_j0) ``` The compiled path follows torch's dtype: a float32 model compares with its eager result to within 1e-3 relative and absolute, the tolerance the interop test suite holds every compiled model to. ## Tensors across the two libraries `from_torch` and `to_torch` exchange tensors through the DLPack protocol. The result shares the source's memory (a write through either side is visible from the other), the source buffer stays alive for the result's lifetime, and nothing is copied: on the CPU, and on a CUDA device both libraries address. The dtypes that cross are the ones both sides spell in the protocol: bool, the four unsigned and four signed integer widths, float16, bfloat16, float32 and float64; a torch float8 tensor or a runtime sub-byte tensor is refused with a `TypeError` naming the dtype and the cast that gets it across. ```python title="exchange.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime as crt import clika_runtime.torch as crt_torch source = torch.randn(5, 131) runtime_tensor = crt_torch.from_torch(source) print(runtime_tensor.dtype == crt.float32, tuple(runtime_tensor.shape)) # True (5, 131) print(runtime_tensor.numpy().ctypes.data == source.data_ptr()) # True: one buffer source[0, 0] = 7.0 print(float(runtime_tensor.numpy()[0, 0])) # 7.0 view = crt_torch.to_torch(runtime_tensor) print(view.data_ptr() == source.data_ptr()) # True: the round trip never copies view[1, 1] = -3.0 print(float(source[1, 1])) # -3.0 print(crt_torch.to_clika_dtype(torch.bfloat16) == crt.bfloat16) # True print(crt_torch.to_torch_dtype(crt.bfloat16) == torch.bfloat16) # True ``` Two things the exchange does on purpose. A torch tensor that is not contiguous is made contiguous first (one copy on the torch side), so the runtime tensor shares that dense buffer and later writes through the strided source do not reach it. And the exchange never moves a tensor: `device=` on either function names where the result must live, and a device other than the tensor's own raises `ValueError` naming both; move the tensor first (`tensor.to("cuda:0")`, `clika_runtime.to(tensor, "cpu")`) and convert the result. ```python title="refusals.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime as crt import clika_runtime.torch as crt_torch try: crt_torch.from_torch(torch.ones(4, 131).to(torch.float8_e4m3fn)) except TypeError as error: print("float8_e4m3fn" in str(error)) # True try: crt_torch.from_torch(torch.ones(3), device="cuda:0") except ValueError as error: print("cpu" in str(error) and "cuda:0" in str(error)) # True strided = torch.randn(5, 262)[:, ::2] dense = crt_torch.from_torch(strided) print(dense.is_contiguous(), tuple(dense.shape)) # True (5, 131) ``` On a machine where torch and the runtime both see a CUDA device (`torch.cuda.is_available()` and `crt.is_available("cuda")`), `from_torch(torch.randn(5, 131, device="cuda"))` lands on the runtime's `cuda:0` without a copy, and `to_torch` of that tensor reads the same device address. `from_torch_state_dict(model.state_dict())` converts a checkpoint one tensor at a time and keeps tied weights tied: two entries over the same storage, offset, shape and strides come back as one runtime tensor under both names. ## Convert a module tree `from_torch_module(module)` builds the `clika_runtime.nn` twin of a `torch.nn.Module` and loads its weights. Leaves convert by class (`Linear`, `Conv1d/2d/3d`, `ConvTranspose1d/2d/3d`, `Embedding`, `LayerNorm`, `RMSNorm`, the stateless activations, `Dropout`, `Identity`, `Flatten`) and the containers (`Sequential`, `ModuleList`, `ModuleDict`) keep their shape. The runtime layers are declared without storage first, then the weights bind through `load_state_dict(assign=True)` one tensor at a time, so each weight exists once: the converted module's parameters carry the same names, in the same order, at the same byte count as the torch module's. ```python title="convert.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime as crt import clika_runtime.torch as crt_torch torch.manual_seed(1) model = torch.nn.Sequential( torch.nn.Linear(131, 64), torch.nn.GELU(), torch.nn.LayerNorm(64), torch.nn.Linear(64, 16) ).eval() converted = crt_torch.from_torch_module(model) print(isinstance(converted, crt.nn.Module)) # True print([name for name, _ in converted.named_parameters()] == [name for name, _ in model.named_parameters()]) # True torch_bytes = sum(p.numel() * p.element_size() for p in model.parameters()) runtime_bytes = sum(p.nbytes for p in converted.parameters()) print(runtime_bytes == torch_bytes) # True: one copy of every weight, at its dtype x = torch.randn(5, 131) with torch.no_grad(): expected = model(x) got = crt_torch.to_torch(converted(crt_torch.from_torch(x))) print(torch.allclose(got, expected, rtol=1e-3, atol=1e-3)) # True ``` Tied weights stay one resident copy. A head whose `weight` is the embedding table binds the same runtime tensor into both slots, the converted `state_dict()` lists both keys, and `parameters()` counts the table once, as torch does: ```python title="tied.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime.torch as crt_torch torch.manual_seed(3) embed = torch.nn.Embedding(50, 64) head = torch.nn.Linear(64, 50, bias=False) head.weight = embed.weight model = torch.nn.Sequential(embed, torch.nn.LayerNorm(64), head).eval() converted = crt_torch.from_torch_module(model) print(list(converted.state_dict()) == list(model.state_dict())) # True distinct_torch = sum(p.numel() * p.element_size() for p in {id(p): p for p in model.parameters()}.values()) distinct_runtime = sum(p.nbytes for p in {id(p): p for p in converted.parameters()}.values()) print(distinct_runtime == distinct_torch) # True: the tied table is one resident copy ``` Convolution weights are permuted on the way in, from torch's channels-first layout (`OIHW`; `IOHW` for a transposed convolution) to the runtime's channels-last one (`OHWI`; `IHWO`), so a converted convolution consumes the runtime's channels-last activations directly. `device=` lands every converted weight on that device (each tensor crosses on the CPU and moves once on the runtime side; the torch module itself is never moved), and `dtype=` casts the floating-point weights after the move. A leaf class the converter does not know raises `NotImplementedError` naming it; `strict=False` keeps such a module as a structure-only container whose children, parameters and buffers are converted under their names and whose own forward is left to the compiled path. `register_module_converter(name)` adds a converter for a class of your own. ```python title="strict.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime as crt import clika_runtime.torch as crt_torch class Odd(torch.nn.Module): def forward(self, x: torch.Tensor) -> torch.Tensor: return x.flip(0) try: crt_torch.from_torch_module(torch.nn.Sequential(torch.nn.Linear(131, 8), Odd())) except NotImplementedError as error: print("Odd" in str(error)) # True lenient = crt_torch.from_torch_module(torch.nn.Sequential(torch.nn.Linear(131, 8), Odd()), strict=False) print(isinstance(lenient, crt.nn.Module)) # True ``` ## torch containers as pytrees Importing `clika_runtime.torch` registers `torch.Size` (and the FX immutable list and dict) with the pytree registry, so a shape inside a nested input flattens to its integers and rebuilds as a `torch.Size`; `register_torch_pytree_nodes()` performs the same registration explicitly. With `transformers` installed, `register_hf_pytree_nodes()` registers its `ModelOutput` classes and caches, so a model output flattens to the fields that are set. ```python title="size_pytree.py" try: import torch except ImportError: print("BLOCKED: torch is not installed (pip install \"clika-runtime[torch]\")") raise SystemExit(3) import clika_runtime.torch as crt_torch from clika_runtime import pytree crt_torch.register_torch_pytree_nodes() size = torch.Size([2, 3, 131]) leaves, spec = pytree.tree_flatten({"shape": size, "n": 1}) print(leaves) # [2, 3, 131, 1] rebuilt = pytree.tree_unflatten(spec, leaves) print(isinstance(rebuilt["shape"], torch.Size), rebuilt["shape"] == size) # True True ``` [Structure inputs and outputs as pytrees](pytrees.mdx) covers the registry these nodes join, and [Use ClikaRT from Python](use-clikart-from-python.mdx) covers the tensor and module surface the converted model lands on. --- # Wrap existing memory without copying Tensor::from_blob over buffers your application already owns: the borrow and adopt contracts, strided views, and what happens on a device move. Source: https://docs.clika.io/clikart/how-to/wrap-existing-memory.md Your data already sits in memory that some other part of the program owns: a decoder's output buffer, an arena, a mapped file, another library's array. `Tensor::from_blob` wraps that memory as a tensor with no copy; the tensor's data pointer is your pointer. What needs deciding is ownership, and the API makes the two contracts explicit: - **No deleter passed: borrowed.** You keep ownership. The buffer must stay alive, and its layout unchanged, for as long as the tensor or any view of it is in use. Writes through the buffer are visible through the tensor and the other way around; it is the same memory. - **Deleter passed: adopted.** The tensor takes ownership and calls `deleter(data)` once, when the last reference drops. The Python variant borrows a numpy array's memory. ## Borrow a buffer and compute on it ```cpp title="borrow.cpp" #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; int main() { std::vector buf(8, 1.0F); // memory the application owns const Tensor view = Tensor::from_blob(buf.data(), {8}, DataType::Float32); std::printf("shares memory: %s\n", view.const_data_ptr() == buf.data() ? "yes (no copy)" : "no"); std::printf("sum = %g\n", ops::sum(view).item()); buf[0] = 100.0F; // write through the buffer... std::printf("sum after buf[0] = 100: %g\n", ops::sum(view).item()); return 0; } ``` ```python title="main.py" import numpy as np import clika_runtime as crt def main() -> None: # Borrow: from_numpy wraps the array's own memory, no copy. The array # and the tensor see the same bytes, mutation aliases both ways, and # the tensor keeps the array alive. The array must be writable and # contiguous; the refusals name the fix. buf = np.zeros(4, dtype=np.float32) view = crt.from_numpy(buf) buf[0] = 100.0 print(f"sum after buf[0] = 100: {view.sum().item():g}") # The other direction: numpy() on a CPU tensor is a zero-copy view of # the tensor's memory; a device tensor asks you to move it first # (t.to('cpu').numpy()). t = crt.ones(2, 2) arr = t.numpy() print(f"shared bytes: {arr.sum():g}") # A non-contiguous array does not borrow; the error names the fix. try: crt.from_numpy(np.zeros((4, 4), dtype=np.float32)[:, ::2]) except TypeError as e: print(f"non-contiguous refused: {e}") if __name__ == "__main__": main() ``` `crt.tensor(array)` stays the copying entry when an independent tensor is wanted. ```text shares memory: yes (no copy) sum = 8 sum after buf[0] = 100: 107 ``` The borrow contract in one sentence: the runtime never reuses or overwrites borrowed memory, and in exchange you guarantee it outlives every tensor that sees it. A `vector` that reallocates (or a stack buffer that goes out of scope) under a live view is the bug this contract exists to name. ## Hand ownership over with a deleter When the producer wants to fire and forget, pass a deleter. The tensor (and every tensor computed from it) keeps the buffer alive; the deleter runs exactly once, when the last reference drops, and it runs in your runtime: an exception it throws never crosses the library boundary. It also frees the runtime to reuse the buffer as scratch, which the borrow contract forbids. ```cpp title="adopt.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; namespace { constexpr std::int64_t kCount = 8; // elements in the caller-owned buffer } // namespace int main() { float* buf = static_cast(std::malloc(kCount * sizeof(float))); for (std::int64_t i = 0; i < kCount; ++i) buf[i] = static_cast(i); { const Tensor adopted = Tensor::from_blob( buf, {kCount}, DataType::Float32, {}, [](void* p) { std::printf("deleter: buffer released\n"); std::free(p); }); std::printf("mean = %g\n", ops::mean(adopted).item()); std::printf("leaving the tensor's scope...\n"); } std::printf("scope closed\n"); return 0; } ``` Handing ownership over with a deleter is part of the C++ API today; the C++ tab shows it. The python borrow keeps the ARRAY as the owner: `crt.from_numpy(array)` holds the array alive for the tensor's lifetime, so no deleter changes hands. ```text mean = 3.5 leaving the tensor's scope... deleter: buffer released scope closed ``` The deleter must not throw (a throw is swallowed). Adoption is the right contract at module boundaries: the producer allocates, the consumer wraps and forgets the allocation ever existed. ## Wrap non-contiguous memory with strides The strided overload views memory that is not laid out contiguously, without rearranging a byte. Strides are in elements, one per dimension. A worked case: cropping a region of interest out of a pitched image buffer, the layout every camera API and GPU readback hands you (rows padded to a pitch wider than the image). ```cpp title="strided_roi.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::Tensor; namespace ops = ClikaRT::ops; namespace { constexpr std::int64_t kRows = 4, kPitch = 6; // the whole image, row-major constexpr std::int64_t kRoiRow = 1, kRoiCol = 2; // where the region starts constexpr std::int64_t kRoiRows = 2, kRoiCols = 3; // its extent } // namespace int main() { // A 4x6 single-channel image, row-major. buf[r][c] = r*10 + c. std::vector buf(kRows * kPitch); for (std::int64_t r = 0; r < kRows; ++r) for (std::int64_t c = 0; c < kPitch; ++c) buf[r * kPitch + c] = static_cast(r * 10 + c); // The 2x3 region starting at row 1, column 2: shape {2, 3}, and the // ORIGINAL row pitch as the row stride. No pixel is copied. const Tensor roi = Tensor::from_blob(buf.data() + kRoiRow * kPitch + kRoiCol, {kRoiRows, kRoiCols}, {kPitch, 1}, DataType::Float32); std::printf("roi = %s\n", roi.to_string().c_str()); std::printf("sum = %g (12+13+14+22+23+24 = 108)\n", ops::sum(roi).item()); return 0; } ``` Wrapping strided memory with explicit strides is part of the C++ API today; the C++ tab shows it. `crt.from_numpy` takes a C-contiguous array; a strided source enters through `np.ascontiguousarray` (the refusal names it). ```text roi = Tensor(shape=[2, 3], dtype=Float32, device=CPU, numel=6, data=[12, 13, 14, 22, 23, 24]) sum = 108 (12+13+14+22+23+24 = 108) ``` The same shape covers any pitched or tiled layout: a submatrix of a row-major matrix, a plane in a planar image, a batch entry inside a larger allocation. `ops::contiguous` materializes an owned compact copy when a consumer needs one. ## Device moves and pinned memory `.to(device)` on a wrapped tensor behaves like on any other: on a discrete accelerator the move is a real transfer to device memory (the wrap saved the host-side copy, not the transfer), while unified-memory hardware moves for free. Two related notes. `from_data` is the copying cousin: it copies your bytes into an owned tensor so the source's lifetime stops mattering; take it when the buffer is short-lived and the tensor is not. And `from_blob`'s `pinned_for` parameter is tag-only: it asserts pages you already page-locked for a device, letting transfers take the pinned path; it cannot pin memory for you. ## Give the pool's idle reserve back The runtime side of the memory story: the pool keeps memory it handed out and got back (`MemoryStats::cached_bytes`), so the next allocation is cheap. After a model unloads, or before a second model must fit beside the first, that idle reserve is memory the device (on a shared-memory part, the host) cannot use for anything else. `device::release_cached_memory(device)` returns it to the driver or the OS: deferred reservations drain, parked buffers retire and empty blocks release, the same three steps the runtime takes before it reports out-of-memory. Memory still in use, or whose last use has not completed on the device, is never touched; a later call can release more once that work retires. Automatic placement calls it on a failed accelerator before the CPU fallback loads. ```cpp title="release_cached.cpp" #include #include using ClikaRT::DataType; using ClikaRT::Device; using ClikaRT::Tensor; static void report(const char* when) { const ClikaRT::device::MemoryStats s = ClikaRT::device::memory_stats(Device::cpu()); std::printf("%-14s active %8.1f MiB, cached %8.1f MiB\n", when, s.active_bytes / 1048576.0, s.cached_bytes / 1048576.0); } int main() { { // 256 MiB of Float32 work: the pool reserves real memory for it. Tensor big = Tensor::zeros({64, 1024, 1024}, DataType::Float32); big.synchronize(); report("in use:"); } // The tensor is gone, but the pool keeps its bytes idle for reuse. report("dropped:"); // Hand the idle reserve back to the system. Memory still in use, or // whose last use has not completed, is never touched. ClikaRT::device::release_cached_memory(Device::cpu()); report("released:"); return 0; } ``` ```text in use: active 256.0 MiB, cached 0.0 MiB dropped: active 0.0 MiB, cached 256.0 MiB released: active 0.0 MiB, cached 0.0 MiB ``` The bundle's `compute` example covers device movement and this wrap in its `03_data_movement` and `06_zero_copy` chapters; [the custom-operator guide](write-a-custom-operator.mdx) uses `from_blob` to hand a hand-written kernel's output back to the runtime. --- # Write a custom operator Subclass nn::Module two ways: compose built-ins into a reusable block, or run your own kernel through the runtime with output_shapes and compute. Source: https://docs.clika.io/clikart/how-to/write-a-custom-operator.md The built-in `ops::` library does not have to be the end of the line. `nn::Module` is both the parameter container (PyTorch-style dotted names, so HuggingFace checkpoints bind by name) and the custom-op entry point: a leaf module that overrides `output_shapes()` and `compute()` runs its own kernel through the runtime, eager and tracing alike. This guide builds one of each. ## Compose built-ins into a module A composite module subclasses `nn::Module`, registers its children in the constructor, and defines its own `forward` composing `ops::` and child modules. Registration is what buys the dotted names: the block below enumerates `up.weight` and `down.weight`, which is exactly how a checkpoint refers to them. ```cpp title="mlp_block.cpp" #include #include #include #include using ClikaRT::DataType; using ClikaRT::nn::Linear; using ClikaRT::NamedTensors; using ClikaRT::Tensor; namespace nn = ClikaRT::nn; namespace ops = ClikaRT::ops; // A residual MLP block: y = down(relu(up(x))) + x. class MlpBlock final : public nn::Module { public: static std::shared_ptr make(std::int64_t dim, std::int64_t hidden) { std::shared_ptr m(new MlpBlock()); m->up_ = Linear::make(dim, hidden); m->down_ = Linear::make(hidden, dim); m->register_module("up", m->up_); m->register_module("down", m->down_); return m; } Tensor forward(const Tensor& x) const { Tensor h = this->up_->forward(x); ops::relu_(h); // fresh from the projection: nothing else sees it h = this->down_->forward(h); ops::add_(h, x); // the residual add writes the fresh down output; x stays a read return h; } private: MlpBlock() = default; std::shared_ptr up_; std::shared_ptr down_; }; int main() { const std::shared_ptr block = MlpBlock::make(4, 8); std::printf("parameters (dotted, checkpoint-shaped):\n"); for (const auto& [name, t] : block->named_parameters()) std::printf(" %-12s fake=%d\n", name.c_str(), static_cast(t.is_fake())); // Bind weights by name, the way a checkpoint would. NamedTensors ckpt; ckpt.set("up.weight", Tensor::full({8, 4}, 0.1, DataType::Float32)); ckpt.set("down.weight", Tensor::full({4, 8}, 0.1, DataType::Float32)); block->load_state_dict(ckpt); const Tensor x = Tensor::ones({1, 4}, DataType::Float32); std::printf("y = %s\n", block->forward(x).to_string().c_str()); return 0; } ``` ```python title="mlp_block.py" import clika_runtime as crt import clika_runtime.nn as nn class MlpBlock(nn.Module): """A residual MLP block: y = down(relu(up(x))) + x. A model is a Module subclass: assigning a module REGISTERS it: load_state_dict, named_parameters and repr all see `up` and `down` automatically. The first linear fuses its ReLU as an activation epilogue, named by its string spelling; without the epilogue the same block writes F.relu(self.up(x)) with `import clika_runtime.nn.functional as F`. Bias defaults to on, so these layers opt out explicitly.""" def __init__(self, dim: int, hidden: int) -> None: super().__init__() self.up = nn.Linear(dim, hidden, bias=False, activation="relu") self.down = nn.Linear(hidden, dim, bias=False) def forward(self, x: crt.Tensor) -> crt.Tensor: return self.down(self.up(x)) + x def main() -> None: block = MlpBlock(4, 8) print(block) # the module tree, registered names included print("parameters (dotted, checkpoint-shaped):") for name, t in block.named_parameters(): print(f" {name}") # Weights arrive as a plain dict; dotted keys route to the registered # submodules. block.load_state_dict({ "up.weight": crt.full((8, 4), 0.1), "down.weight": crt.full((4, 8), 0.1), }) y = block(crt.ones(1, 4)) # calling the module runs forward print(f"y = {y}") if __name__ == "__main__": main() ``` ```text parameters (dotted, checkpoint-shaped): up.weight fake=1 down.weight fake=1 y = Tensor(shape=[1, 4], dtype=Float32, device=CPU, numel=4, data=[1.32, 1.32, 1.32, 1.32]) ``` The math checks out by hand: `up` maps ones to 0.4 per unit, relu passes it, `down` sums 8 of them times 0.1 to 0.32, and the residual adds the input back, 1.32. Python prints a tensor as its values alone, so its last line reads `y = tensor([[1.3200, 1.3200, 1.3200, 1.3200]])`. `Linear::make` declares storage-free slots, so the parameters read as fake until `load_state_dict` binds them; `block->to(device)` moves the whole tree. ## Run your own kernel through the runtime A leaf custom op overrides two virtuals. `output_shapes()` is the shape rule: it reads the inputs' metadata (`FakeTensor`: symbolic dims, dtype, placement) and returns the outputs' metadata. `compute()` is the kernel, with one accessor pair per residence: for a plain host loop over CPU-resident tensors, inputs read through `const_data_ptr()` and the pre-allocated outputs write through `mutable_data_ptr()` (both wait for the data and refuse a device tensor); a kernel on the op's device reads `device_const_data_ptr()` and writes `device_mutable_data_ptr()`, with no host wait and the pointer ordered on the op's stream. An exception thrown inside `compute()` surfaces as a typed error in the caller's runtime instead of crossing the library boundary. `forward` hands both to `dispatch()`, which runs the op through the runtime: eager mode calls `compute()` now, a tracing scope records the op from the shape rule alone. A leaf that dispatches must be owned by a `shared_ptr`. ```cpp title="softclip_op.cpp" #include #include #include #include #include #include using ClikaRT::DataType; using ClikaRT::FakeTensor; using ClikaRT::Tensor; namespace nn = ClikaRT::nn; // Elementwise soft clip: y = x / (1 + |x|). class SoftClip final : public nn::Module { public: static std::shared_ptr make() { return std::shared_ptr(new SoftClip()); } Tensor forward(const Tensor& x) const { return dispatch({&x, 1})[0]; } // Shape rule: one output, same shape and dtype as the input. std::vector output_shapes( ClikaRT::Span inputs) const override { return {FakeTensor(inputs[0].shape(), inputs[0].dtype(), inputs[0].stream())}; } // Kernel: runs on the op's stream; here the CPU path, a plain loop over // host pointers. A device kernel (a CUDA arm, say) reads // `device_const_data_ptr()`, writes `device_mutable_data_ptr()`, and // launches on `this->stream().native_handle()`; the host accessors below // wait for the data and refuse a device tensor. void compute(ClikaRT::Span inputs, ClikaRT::Span outputs) const override { const Tensor& in = inputs[0]; const float* x = static_cast(in.const_data_ptr()); float* y = static_cast(outputs[0].mutable_data_ptr()); const std::int64_t n = in.numel(); for (std::int64_t i = 0; i < n; ++i) y[i] = x[i] / (1.0F + std::fabs(x[i])); } private: SoftClip() = default; }; int main() { const std::shared_ptr clip = SoftClip::make(); std::vector v = {-9.0F, -1.0F, 0.0F, 1.0F, 9.0F}; const Tensor x = Tensor::from_data(v.data(), {5}, DataType::Float32); std::printf("y = %s\n", clip->forward(x).to_string().c_str()); return 0; } ``` ```python title="softclip_op.py" import clika_runtime as crt import clika_runtime.nn as nn class SoftClip(nn.Module): """Elementwise soft clip: y = x / (1 + |x|), composed from the built-in operations; a custom module needs nothing beyond a forward. Writing a custom KERNEL (your own shape rule and compute) is done through the C++ API; the C++ tab walks it.""" def forward(self, x: crt.Tensor) -> crt.Tensor: return x / (1 + x.abs()) def main() -> None: clip = SoftClip() x = crt.tensor([-9.0, -1.0, 0.0, 1.0, 9.0]) print(f"y = {clip(x)}") if __name__ == "__main__": main() ``` ```text y = Tensor(shape=[5], dtype=Float32, device=CPU, numel=5, data=[-0.9, -0.5, 0, 0.5, 0.9]) ``` The runtime allocated the output, placed it beside the input, and ran the kernel; the op composes with everything else (`ops::` calls before and after, `on_complete`, the scopes) because it went through `dispatch` like a built-in. Python's line reads `y = tensor([-0.9000, -0.5000, 0.0000, 0.5000, 0.9000])`: the same values in its own tensor form. ## Launching on an accelerator `compute()` runs on whatever device its tensors live on. Inside it, `this->stream()` is the op's stream: `stream().device()` says where you are, and `stream().native_handle()` is the backend's queue as an opaque pointer, `cudaStream_t` on CUDA (`nullptr` on the CPU backend, which has no device queue). The launch pattern, from the `nn/module.h` contract: ```text Stream s = this->stream(); // the op's stream auto cu = static_cast(s.native_handle()); // the CUDA queue const float* q = static_cast(inputs[0].device_const_data_ptr()); float* o = static_cast(outputs[0].device_mutable_data_ptr()); my_kernel<<>>(q, ..., o); // your kernel ``` The device accessors return the device address with no host wait, ordered on the op's stream, so the launch above is valid as written; the host accessors (`const_data_ptr()`, `mutable_data_ptr()`) are for CPU-resident tensors and would refuse here. Size the launch from `get_device_properties(stream().device())`, and take stream-ordered scratch from `stream().allocate(shape, dtype)`, which recycles safely on that stream only. The handle is owned by the runtime: never destroy it, and keep the `Stream` alive while using it. The bundle's `flash_attention` example is the complete worked case, a hand-written CUDA attention kernel dispatched through this exact interface. A rule of thumb for choosing the shape: compose built-ins when the math decomposes into `ops::` (the runtime already fuses and places them); write a leaf when you have a kernel the library does not, and keep its `output_shapes` honest, because tracing trusts it without running `compute`. --- # Write a transform Write graph transforms of your own: functions over the ModelGraph edits that optimize() runs beside the runtime's transforms, operators built with add_node, insert and replace, and rewrite rules that replace each occurrence of a Pattern. Source: https://docs.clika.io/clikart/how-to/write-a-transform.md {/* Every block is a program under examples//howto/write_a_transform/: the first block of each tab is get_a_graph whole, and every later block is its program's docs region. tools/tutorial_check.py runs each program against its recorded output. */} A transform rewrites a `ModelGraph` inside `optimize()`. The runtime ships its own transforms, and you can write more. A transform you write is a function over the graph edits that [Edit a graph](edit-a-graph.mdx) covers, and `optimize()` runs it beside the runtime's transforms in one list, at its place, once per iteration. Three more calls build operators inside a transform: `add_node` adds one operator, and `insert` and `replace` add the operators that a function over the public ops records. A rewrite rule is a transform that finds each occurrence of a `Pattern` and replaces it. The calls are available from C++ and Python, under the same names. ## A graph to transform The model below computes `y = Relu(Relu(x)) + Neg(x)`, a Relu over a Relu and a Neg that the Add reads, and the examples on this page rewrite both. `trace_model` traces a function over an `x` of shape [2, 3], the model unless it is given another, and returns the graph as built. Each program prints a graph as the formula it returns. `formula` names a value by its operator and the values that operator reads, a graph input by its name and a constant as `c`, and `formula_of` gives the formula of the graph's output, which the node names and their order do not change. `rows` lists a report's rows as each transform's name and applications, and in C++ `optimize_with` runs a list through `OptimizeOptions::transforms`. `drop_repeated_relu` is the first transform on this page. It bypasses each Relu that reads a Relu with `bypass_node`, so that Relu's readers read the one before it. The program runs `remove_redundant_relu`, one of the runtime's transforms, in a list of its own, as [Optimize and finalize](optimize-and-finalize.mdx#run-a-list-of-your-own) shows. It leaves one Relu where the model has two, and its row counts one application. ```cpp title="get_a_graph.cpp" #include #include #include #include #include #include #include #include #include using ClikaRT::Error; using ClikaRT::Result; using ClikaRT::Tensor; using ClikaRT::graph::ModelGraph; using ClikaRT::graph::Node; using ClikaRT::graph::NodeKind; using ClikaRT::graph::OpCode; using ClikaRT::graph::OptimizeOptions; using ClikaRT::graph::OptimizeReport; using ClikaRT::graph::Pattern; using ClikaRT::graph::RewriteMatch; using ClikaRT::graph::Transform; using ClikaRT::graph::TransformReport; using ClikaRT::graph::Value; namespace ops = ClikaRT::ops; namespace transforms = ClikaRT::graph::transforms; namespace { // y = Relu(Relu(x)) + Neg(x): a Relu over a Relu, and a Neg that the Add reads. std::vector model(const std::vector& inputs) { const Tensor rectified = ops::relu(ops::relu(inputs[0])); const Tensor negated = ops::neg(inputs[0]); return {ops::add(rectified, negated)}; } // The formula a value computes: its operator over the values it reads, a graph input by its name, and a // constant as c. std::string formula(const Value& value) { const std::optional producer = value.producer(); if (!producer.has_value()) return "c"; if (producer->kind() == NodeKind::Input) return producer->name(); std::string reads; for (const Value& read : producer->inputs()) reads += (reads.empty() ? "" : ", ") + formula(read); return std::string(ClikaRT::graph::op_code_name(producer->op_code())) + "(" + reads + ")"; } // The formula of the value the graph returns. std::string formula_of(const ModelGraph& graph) { for (const Node& node : graph.nodes()) { for (const Value& value : node.outputs()) { if (value.is_graph_output()) return formula(value); } } return ""; } // A trace returns the graph as built: every operator as written, nothing optimized or finalized. ModelGraph trace_model(const ClikaRT::graph::TraceFunction& fn = model) { const std::vector signature = {{"x", ClikaRT::DataType::Float32, {2, 3}}}; const std::vector outputs = {"y"}; return ClikaRT::graph::trace(fn, signature, "transform", outputs); } // optimize() over `list`, in order. OptimizeReport optimize_with(ModelGraph& graph, std::vector list) { OptimizeOptions options; options.transforms = std::move(list); return graph.optimize(options); } // Each row of an optimize() report, as "name applications", separated by commas. std::string rows(const OptimizeReport& report) { std::string out; for (const TransformReport& row : report.transforms) { out += (out.empty() ? "" : ", ") + row.transform.name() + " " + std::to_string(row.applications); } return out; } // A Relu that reads a Relu is bypassed: its readers read the first Relu instead. Result drop_repeated_relu(ModelGraph& graph) { for (const Node& relu : graph.find_nodes(OpCode::Relu)) { const std::optional producer = relu.input(0)->producer(); if (producer.has_value() && producer->op_code() == OpCode::Relu) graph.bypass_node(relu); } return {}; } } // namespace int main() { ModelGraph graph = trace_model(); std::printf("%s\n", formula_of(graph).c_str()); // Add(Relu(Relu(x)), Neg(x)) // optimize() runs a list of transforms. This one of the runtime's removes the repeated Relu. const OptimizeReport report = optimize_with(graph, {transforms::remove_redundant_relu()}); std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str()); // Add(Relu(x), Neg(x)) | remove_redundant_relu 1 return 0; } ``` ```python title="get_a_graph.py" from collections.abc import Callable import clika_runtime as crt from clika_runtime.graph import NodeKind, OpCode, Pattern, RewriteMatch, Transform, transforms def model(inputs: list[crt.Tensor]) -> list[crt.Tensor]: x = inputs[0] return [crt.relu(crt.relu(x)) + crt.neg(x)] # y = Relu(Relu(x)) + Neg(x) # The formula a value computes: its operator over the values it reads, a graph input by its name, and a # constant as c. def formula(value: crt.graph.Value) -> str: producer = value.producer() if producer is None: return "c" if producer.kind == NodeKind.Input: return producer.name return f"{producer.op_code.name}({', '.join(formula(read) for read in producer.inputs)})" # The formula of the value the graph returns. def formula_of(graph: crt.graph.ModelGraph) -> str: (returned,) = [value for node in graph.nodes() for value in node.outputs if value.is_graph_output()] return formula(returned) # A trace returns the graph as built: every operator as written, nothing optimized or finalized. def trace_model(fn: Callable[[list[crt.Tensor]], list[crt.Tensor]] = model) -> crt.graph.ModelGraph: return crt.trace(fn, [crt.TensorSpec("x", crt.float32, [2, 3])], output_names=["y"]).graph # Each row of an optimize() report, as (name, applications). def rows(report: crt.graph.OptimizeReport) -> list[tuple[str, int]]: return [(row.transform.name, row.applications) for row in report.transforms] # A Relu that reads a Relu is bypassed: its readers read the first Relu instead. def drop_repeated_relu(graph: crt.graph.ModelGraph) -> None: for relu in graph.find_nodes(OpCode.Relu): producer = relu.input(0).producer() if producer is not None and producer.op_code == OpCode.Relu: graph.bypass_node(relu) graph = trace_model() print(formula_of(graph)) # Add(Relu(Relu(x)), Neg(x)) # optimize() runs a list of transforms. This one of the runtime's removes the repeated Relu. report = graph.optimize([transforms.remove_redundant_relu()]) print(formula_of(graph), rows(report)) # Add(Relu(x), Neg(x)) [('remove_redundant_relu', 1)] ``` Write a transform when a rewrite should run inside `optimize()`, beside the runtime's transforms and at every iteration, so it sees what their rewrites expose and they see what it changes. Each program on this page is complete and runs on its own. From here on, a block shows the part of its program that follows the opening lines the first block shows (the includes or imports, `model`, the helpers, `drop_repeated_relu` and, in Python, the trace). ## Write a transform `Transform::from_function(name, fn)` in C++ and `Transform(name, fn)` in Python make a transform from a function. `fn` gets the graph that `optimize()` is running and edits it with the graph edits. In C++ it returns a `Result`, and a failed one ends the run. In Python it returns None. The `@transforms.transform` decorator makes a transform from the function it decorates, named after the function, and `@transforms.transform(name, runs_after=...)` sets the name and the order. The name is how the transform reads in the report and in errors, so it may be neither empty nor the name of one of the runtime's transforms. `optimize()` runs each entry of the list at its place, once per iteration, and stops at the first iteration that changes nothing. Each entry has one row in the report, and its `applications` counts the iterations in which the transform changed the graph. A transform changes the graph when one of its edits does, so `count_nodes` below, which only reads the graph, counts none, though it runs in both iterations and records the node count each time. A handle you made equals itself and the copies the report holds. ```cpp title="a_transform.cpp" int main() { ModelGraph graph = trace_model(); // A transform you write: a name, and the function optimize() calls at its place in the list. const Transform drop = Transform::from_function("drop_repeated_relu", drop_repeated_relu); // One that reads the graph and changes nothing: it records the node count each time it runs. std::vector counts; const Transform count_nodes = Transform::from_function("count_nodes", [&counts](ModelGraph& g) -> Result { counts.push_back(g.nodes().size()); return {}; }); const OptimizeReport report = optimize_with(graph, {count_nodes, drop}); std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str()); // Add(Relu(x), Neg(x)) | count_nodes 0, drop_repeated_relu 1 std::printf("%lld | %zu %zu\n", static_cast(report.iterations), counts.at(0), counts.at(1)); // 2 | 5 4: count_nodes ran in both iterations std::printf("%s\n", report.transforms.at(1).transform == drop ? "true" : "false"); // true return 0; } ``` ```python title="a_transform.py" drop = Transform("drop_repeated_relu", drop_repeated_relu) # a name, and the function optimize() calls counts: list[int] = [] @transforms.transform # a transform named after the function it decorates def count_nodes(g: crt.graph.ModelGraph) -> None: counts.append(len(g.nodes())) # it reads the graph and changes nothing report = graph.optimize([count_nodes, drop]) print(formula_of(graph), rows(report)) # Add(Relu(x), Neg(x)) [('count_nodes', 0), ('drop_repeated_relu', 1)] print(report.iterations, counts) # 2 [5, 4]: count_nodes ran in both iterations print(report.transforms[1].transform == drop) # True ``` Write a transform as a function when the rewrite reads more of the graph than a pattern states, or edits parts of it that a rule does not, such as the graph's inputs and outputs. ## Order the transforms `runs_after` lists the transforms, the runtime's or yours, that a transform must follow in any list that holds both. It is the third argument of `Transform::from_function` and the `runs_after=` keyword of `Transform` and of the decorator. `optimize()` checks the list before anything runs. A list that places a transform ahead of one it must follow is refused with the code name `INVALID_ARGUMENT`, and the message names both entries and the order to use. A transform the list does not hold asks nothing. Here `drop_repeated_relu` follows `remove_double_neg`, which can expose a Relu over a Relu. ```cpp title="order.cpp" int main() { ModelGraph graph = trace_model(); // remove_double_neg can expose a Relu over a Relu (Relu(Neg(Neg(Relu(x))))), so drop_repeated_relu runs after it. const Transform drop = Transform::from_function("drop_repeated_relu", drop_repeated_relu, {transforms::remove_double_neg()}); try { optimize_with(graph, {drop, transforms::remove_double_neg()}); // a list is checked before anything runs } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // INVALID_ARGUMENT | optimize: transforms[0] (drop_repeated_relu) must run after remove_double_neg, which the list places later, at transforms[1]; list remove_double_neg ahead of it } const OptimizeReport report = optimize_with(graph, {transforms::remove_double_neg(), drop}); std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str()); // Add(Relu(x), Neg(x)) | remove_double_neg 0, drop_repeated_relu 1 return 0; } ``` ```python title="order.py" # remove_double_neg can expose a Relu over a Relu (Relu(Neg(Neg(Relu(x))))), so drop_repeated_relu runs after it. drop = Transform("drop_repeated_relu", drop_repeated_relu, runs_after=[transforms.remove_double_neg()]) try: graph.optimize([drop, transforms.remove_double_neg()]) # a list is checked before anything runs except crt.InvalidArgumentError as error: print(error.code_name, "|", error) # INVALID_ARGUMENT | optimize: transforms[0] (drop_repeated_relu) must run after remove_double_neg, which the list places later, at transforms[1]; list remove_double_neg ahead of it report = graph.optimize([transforms.remove_double_neg(), drop]) print(formula_of(graph), rows(report)) # Add(Relu(x), Neg(x)) [('remove_double_neg', 0), ('drop_repeated_relu', 1)] ``` Declare the order when a transform depends on another one's result, so that no list can run them the other way. ## What a transform may not do Inside a transform, its graph serves the graph edits and every query. It refuses `optimize()`, `finalize()`, `to()` and `attach_kv_cache()` with the code name `FAILED_PRECONDITION`, since `optimize()` is still working on it. The graph's inputs and outputs change only through their own edits (`add_input`, `add_output`, `remove_output`, `rename_input` and `rename_output`), and every other edit keeps them as they are. The graph also refuses an edit from another thread, and moving from the graph or assigning over it leaves both graphs as they were. A transform that fails ends the run, and the graph is as `optimize()` found it, with the changes of every transform in the run undone. In C++ a transform fails by returning a failed `Result` or by throwing. `optimize()` raises that failure with its status and its code name, and the message starts with the entry that failed, as in `optimize: transforms[1] (no_neg) failed:`. In Python, `optimize()` raises the exception the transform raised, as that same object. The finalize refusal comes back as `finalize()` raised it, and an exception class of your own comes back as itself. ```cpp title="refusals.cpp" int main() { ModelGraph graph = trace_model(); // Inside a transform, its own graph refuses finalize(): optimize() is still working on the graph. const Transform finalizes = Transform::from_function("finalizes", [](ModelGraph& g) -> Result { g.finalize(); return {}; }); try { optimize_with(graph, {finalizes}); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // FAILED_PRECONDITION | optimize: transforms[0] (finalizes) failed: finalize: the transform 'finalizes' is running on this graph inside optimize(); call finalize() after optimize() returns } // A failure ends the run: optimize() raises it with its status and code, and every change of the run is undone. const Transform no_neg = Transform::from_function("no_neg", [](ModelGraph& g) -> Result { if (g.find_nodes(OpCode::Neg).empty()) return {}; return Result(ClikaRT::Status::Unsupported, "the graph still computes a Neg", "NEG_NOT_SERVED"); }); try { optimize_with(graph, {Transform::from_function("drop_repeated_relu", drop_repeated_relu), no_neg}); } catch (const Error& error) { std::printf("%s | %s\n", error.code_name().c_str(), error.what()); // NEG_NOT_SERVED | optimize: transforms[1] (no_neg) failed: the graph still computes a Neg } std::printf("%s\n", formula_of(graph).c_str()); // Add(Relu(Relu(x)), Neg(x)): the bypassed Relu is back return 0; } ``` ```python title="refusals.py" # Inside a transform, its own graph refuses finalize(): optimize() is still working on the graph. @transforms.transform def finalizes(g: crt.graph.ModelGraph) -> None: g.finalize() try: graph.optimize([finalizes]) except crt.InvalidArgumentError as error: # the refusal finalize() raised, as it raised it print(error.code_name, "|", error) # FAILED_PRECONDITION | finalize: the transform 'finalizes' is running on this graph inside optimize(); call finalize() after optimize() returns class NegNotServed(Exception): """The target the graph is built for runs no Neg.""" # An exception a transform raises ends the run: optimize() raises that same exception, and every change of the # run is undone. @transforms.transform def no_neg(g: crt.graph.ModelGraph) -> None: if g.find_nodes(OpCode.Neg): raise NegNotServed("the graph still computes a Neg") try: graph.optimize([Transform("drop_repeated_relu", drop_repeated_relu), no_neg]) except NegNotServed as error: print(type(error).__name__, "|", error) # NegNotServed | the graph still computes a Neg print(formula_of(graph)) # Add(Relu(Relu(x)), Neg(x)): the Relu that drop_repeated_relu bypassed is back ``` Fail a transform when it meets a graph it cannot handle: the run ends, and the graph stays as `optimize()` found it. ## Build one operator with add_node `add_node(code, inputs, attributes)` adds an operator of `code`, built the way its public op builds it, so the op's own checks apply. `inputs` holds its values, one per input port in the op's order, with `std::nullopt` (None in Python) for an optional operand left out. `attributes` holds its settings under the names a node's attributes report (`Node::attributes()` in C++, `node.attributes` in Python). A setting left out takes the op's default, an integer also serves a number, and an enumerated setting takes its enumerator's name. Nothing reads the new operator's outputs until an edit wires them in, and `finalize()` drops a node that nothing reads by then. A code that no single public op call builds, such as an operator with several outputs or a quantized one, is refused, and `insert` builds it from the public ops that compute it. Here `fold_neg_into_sub` turns `a + Neg(b)` into `a - b`. The Sub takes the Add's own settings, `alpha` and `activation`, which the two ops name alike. `replace_all_uses_with` moves the Add's reads to the Sub, the graph's output among them, and the Add and the Neg go. ```cpp title="add_node.cpp" int main() { ModelGraph graph = trace_model(); // a + Neg(b) is a - b: a Sub with the Add's own settings takes the Add's place, and the Neg goes once nothing // reads it. const Transform fold = Transform::from_function("fold_neg_into_sub", [](ModelGraph& g) -> Result { for (const Node& add : g.find_nodes(OpCode::Add)) { const std::optional neg = add.input(1)->producer(); if (!neg.has_value() || neg->op_code() != OpCode::Neg) continue; // The Sub reads a and b, with the Add's settings (alpha, activation) under their own names. const Node sub = g.add_node(OpCode::Sub, {add.input(0), neg->input(0)}, add.attributes()); g.replace_all_uses_with(*add.output(0), *sub.output(0)); g.remove_node(add); if (neg->out_degree() == 0) g.remove_node(*neg); } return {}; }); const OptimizeReport report = optimize_with(graph, {fold}); std::printf("%s | %s\n", formula_of(graph).c_str(), rows(report).c_str()); // Sub(Relu(Relu(x)), x) | fold_neg_into_sub 1 return 0; } ``` ```python title="add_node.py" # a + Neg(b) is a - b: a Sub with the Add's own settings takes the Add's place, and the Neg goes once nothing # reads it. @transforms.transform def fold_neg_into_sub(g: crt.graph.ModelGraph) -> None: for add in g.find_nodes(OpCode.Add): neg = add.input(1).producer() if neg is None or neg.op_code != OpCode.Neg: continue sub = g.add_node(OpCode.Sub, [add.input(0), neg.input(0)], add.attributes) # alpha and activation g.replace_all_uses_with(add.output(0), sub.output(0)) g.remove_node(add) if neg.out_degree() == 0: g.remove_node(neg) report = graph.optimize([fold_neg_into_sub]) print(formula_of(graph), rows(report)) # Sub(Relu(Relu(x)), x) [('fold_neg_into_sub', 1)] ``` Build with `add_node` when the new operator is one call of a public op, with settings you read off the graph or choose yourself. ## Splice in a function with insert and replace `insert(fn, inputs)` traces `fn`, a function over the public ops, over one tensor per input, each with that value's dtype, device and shape, and adds the operators it records, reading the values given. In C++ `fn` returns a `std::vector`, and in Python a Tensor or a sequence of Tensors. `insert` returns the values `fn` returns, which nothing reads until an edit wires them in, as `add_node`'s outputs are wired. `replace(old_outputs, fn, inputs)` is `insert` followed by the wiring and the removal: every read of `old_outputs[i]` reads new value `i`, the graph's outputs included, and the operators that computed the old outputs from the inputs go, with every producer only they read. An operator an input is computed from stays. `fn` computes values, so it writes no input in place and does not edit the graph. Here two transforms put one Relu in place of a Relu over a Relu. `with_insert` wires the new Relu in with `replace_all_uses_with` and removes the pair itself, and `with_replace` does all of that in one call. Each rewrites one pair per call and returns, since its edits remove nodes that its loop still holds, and the next iteration finds the next pair. ```cpp title="insert_replace.cpp" int main() { // Relu(Relu(v)) is Relu(v): a function over the public ops. const ClikaRT::graph::TraceFunction relu_of = [](const std::vector& in) { return std::vector{ops::relu(in[0])}; }; // insert() adds the operators relu_of records over the values given, and returns their values, which nothing // reads until an edit wires them in. const Transform with_insert = Transform::from_function("with_insert", [&relu_of](ModelGraph& g) -> Result { for (const Node& outer : g.find_nodes(OpCode::Relu)) { const std::optional inner = outer.input(0)->producer(); if (!inner.has_value() || inner->op_code() != OpCode::Relu) continue; const std::vector made = g.insert(relu_of, {*inner->input(0)}); g.replace_all_uses_with(*outer.output(0), made.at(0)); g.remove_node(outer); g.remove_node(*inner); return {}; // one pair per call: the next iteration finds the next one } return {}; }); // replace() does all of that in one call: the operators that compute the old output from the inputs go, the // outer Relu and the inner one only it reads. const Transform with_replace = Transform::from_function("with_replace", [&relu_of](ModelGraph& g) -> Result { for (const Node& outer : g.find_nodes(OpCode::Relu)) { const std::optional inner = outer.input(0)->producer(); if (!inner.has_value() || inner->op_code() != OpCode::Relu) continue; g.replace({*outer.output(0)}, relu_of, {*inner->input(0)}); return {}; } return {}; }); ModelGraph graph = trace_model(); ModelGraph other = trace_model(); const std::string inserted = rows(optimize_with(graph, {with_insert})); const std::string replaced = rows(optimize_with(other, {with_replace})); std::printf("%s | %s\n", formula_of(graph).c_str(), inserted.c_str()); // Add(Relu(x), Neg(x)) | with_insert 1 std::printf("%s | %s\n", formula_of(other).c_str(), replaced.c_str()); // Add(Relu(x), Neg(x)) | with_replace 1 return 0; } ``` ```python title="insert_replace.py" def relu_of(tensors: list[crt.Tensor]) -> crt.Tensor: return crt.relu(tensors[0]) # Relu(Relu(v)) is Relu(v): a function over the public ops # insert() adds the operators relu_of records over the values given, and returns their values, which nothing # reads until an edit wires them in. @transforms.transform def with_insert(g: crt.graph.ModelGraph) -> None: for outer in g.find_nodes(OpCode.Relu): inner = outer.input(0).producer() if inner is not None and inner.op_code == OpCode.Relu: (made,) = g.insert(relu_of, [inner.input(0)]) g.replace_all_uses_with(outer.output(0), made) g.remove_node(outer) g.remove_node(inner) return # one pair per call: the next iteration finds the next one # replace() does all of that in one call: the operators that compute the old output from the inputs go, the # outer Relu and the inner one only it reads. @transforms.transform def with_replace(g: crt.graph.ModelGraph) -> None: for outer in g.find_nodes(OpCode.Relu): inner = outer.input(0).producer() if inner is not None and inner.op_code == OpCode.Relu: g.replace([outer.output(0)], relu_of, [inner.input(0)]) return other = trace_model() inserted = rows(graph.optimize([with_insert])) replaced = rows(other.optimize([with_replace])) print(formula_of(graph), inserted) # Add(Relu(x), Neg(x)) [('with_insert', 1)] print(formula_of(other), replaced) # Add(Relu(x), Neg(x)) [('with_replace', 1)] ``` Use `replace` when the new values are a computation over values the graph already has, and `insert` when you wire the result in yourself. ## Write a rule A rule rewrites the occurrences of a `Pattern`, which names the operators to find, one pattern node each, and the edges between them. `Pattern::chain` (in Python, `Pattern.chain` or `Pattern([...])`) builds a straight line, and `add_node` and `add_edge` build any other shape, as the next section does. `transforms::rewrite(name, pattern, replacement, condition)` in C++ and `transforms.rewrite(name, pattern, replacement, condition)` in Python make the rule, a transform like any other. Each iteration searches the graph for the pattern once and takes the occurrences in the order the search returns them. The rule replaces each occurrence its condition accepts, every one when there is no condition, with the values `replacement(match, tensors)` computes from one traced tensor per value the occurrence reads, traced as `replace` traces its function. Every occurrence is rewritten in the same iteration, so the rule's row counts that iteration once, and the next iteration searches again, so an occurrence a replacement creates is rewritten then. In Python, the `@transforms.rule(pattern, condition=..., name=..., runs_after=...)` decorator makes a rule from the replacement it decorates, named after the function unless given a name. The match is a `RewriteMatch`. `nodes` holds the matched node for each pattern node, in the order the nodes were added, with `std::nullopt` (None in Python) for an optional node the occurrence lacks. `inputs` holds the values the occurrence reads from outside it, each once, and the replacement gets one tensor per entry, in that order. `outputs` holds the values it makes that a node outside it reads or the graph returns, and the replacement returns one new value per entry. Here the replacement also records what it reads of each occurrence. ```cpp title="rules.cpp" int main() { // y = Relu(Relu(x)) + Relu(Relu(Neg(x))): two Relus over a Relu. ModelGraph pairs = trace_model([](const std::vector& in) { const Tensor first = ops::relu(ops::relu(in[0])); const Tensor second = ops::relu(ops::relu(ops::neg(in[0]))); return std::vector{ops::add(first, second)}; }); std::printf("%s\n", formula_of(pairs).c_str()); // Add(Relu(Relu(x)), Relu(Relu(Neg(x)))) std::vector seen; // what the replacement reads of each occurrence const OpCode pair[] = {OpCode::Relu, OpCode::Relu}; const Transform collapse_relu = transforms::rewrite( "collapse_relu", Pattern::chain(pair), [&seen](const RewriteMatch& match, const std::vector& in) { seen.push_back(std::to_string(match.nodes.size()) + " nodes, reads " + formula(match.inputs.at(0)) + ", " + std::to_string(match.outputs.size()) + " output"); return std::vector{ops::relu(in[0])}; // one Relu over the value the pair reads }); const std::string report = rows(optimize_with(pairs, {collapse_relu})); std::sort(seen.begin(), seen.end()); for (const std::string& occurrence : seen) std::printf("%s\n", occurrence.c_str()); // 2 nodes, reads Neg(x), 1 output // 2 nodes, reads x, 1 output std::printf("%s | %s\n", formula_of(pairs).c_str(), report.c_str()); // Add(Relu(x), Relu(Neg(x))) | collapse_relu 1 return 0; } ``` ```python title="rules.py" # y = Relu(Relu(x)) + Relu(Relu(Neg(x))): two Relus over a Relu. def two_pairs(inputs: list[crt.Tensor]) -> list[crt.Tensor]: x = inputs[0] return [crt.relu(crt.relu(x)) + crt.relu(crt.relu(crt.neg(x)))] seen: list[tuple[list[str], list[str], int]] = [] # what the replacement reads of each occurrence def one_relu(match: RewriteMatch, tensors: list[crt.Tensor]) -> crt.Tensor: nodes = [node.op_code.name for node in match.nodes] seen.append((nodes, [formula(value) for value in match.inputs], len(match.outputs))) return crt.relu(tensors[0]) # one Relu over the value the pair reads collapse_relu = transforms.rewrite("collapse_relu", Pattern.chain([OpCode.Relu, OpCode.Relu]), one_relu) pairs = trace_model(two_pairs) print(formula_of(pairs)) # Add(Relu(Relu(x)), Relu(Relu(Neg(x)))) report = pairs.optimize([collapse_relu]) print(sorted(seen)) # [(['Relu', 'Relu'], ['Neg(x)'], 1), (['Relu', 'Relu'], ['x'], 1)] print(formula_of(pairs), rows(report)) # Add(Relu(x), Relu(Neg(x))) [('collapse_relu', 1)] ``` Write a rule when the rewrite is local: a fixed arrangement of operators, a condition on the match, and a replacement computed from the values it reads. ## The occurrences a rule skips A rule skips an occurrence, and its condition never sees it, in three cases: an earlier replacement in the same iteration removed one of its nodes, the occurrence is not convex, or nothing outside it reads a value it makes. An occurrence is not convex when a path leaves it and comes back, a path no replacement can keep, and `ModelGraph::is_convex` answers the same question. Every other occurrence goes to `condition(match)`, which decides whether to replace it, by the truth of its result in Python. With no condition, the rule replaces every occurrence. Which occurrence an earlier replacement removes depends on the order the search returns them, so no program here shows that case. The program below builds its pattern node by node, the Add as the root and the Neg feeding its second operand through `add_edge`, and counts the condition's calls. The condition accepts an Add with its default settings whose value is the one the occurrence hands on. On the page's graph the condition is asked once and the fold runs. On `n = Neg(x); y = Relu(n) + n`, the path from the Neg through the Relu leaves the occurrence and comes back, so the rule skips it and the condition is never asked. ```cpp title="skipped_occurrences.cpp" int main() { // Add(a, Neg(b)), built node by node: the Add is the root, and the Neg feeds its second operand. Pattern pattern; const Pattern::NodeId add = pattern.add_node(OpCode::Add); const Pattern::NodeId neg = pattern.add_node(OpCode::Neg); pattern.add_edge(neg, add, 0, 1); // the Neg's output 0 feeds the Add's input 1 int asked = 0; // how many occurrences the condition is asked about // a + Neg(b) is a - b for an Add with its default settings, when its value is the one the occurrence hands on. const auto plain_sum = [&asked](const RewriteMatch& match) { ++asked; const Node& total = *match.nodes[0]; const std::optional alpha = total.attribute("alpha"); const double* scale = alpha.has_value() ? std::get_if(&alpha->value) : nullptr; return match.outputs.size() == 1 && scale != nullptr && *scale == 1.0 && total.fused_activation() == ops::Activation::Identity; }; const Transform fold_neg_into_sub = transforms::rewrite( "fold_neg_into_sub", std::move(pattern), [](const RewriteMatch&, const std::vector& in) { return std::vector{ops::sub(in[0], in[1])}; // match.inputs holds a, then b }, plain_sum); ModelGraph graph = trace_model(); std::string report = rows(optimize_with(graph, {fold_neg_into_sub})); std::printf("%s | %s | asked %d\n", formula_of(graph).c_str(), report.c_str(), asked); // Sub(Relu(Relu(x)), x) | fold_neg_into_sub 1 | asked 1 // n = Neg(x), y = Relu(n) + n: the path from the Neg through the Relu leaves the occurrence and comes back in. asked = 0; ModelGraph looped = trace_model([](const std::vector& in) { const Tensor n = ops::neg(in[0]); return std::vector{ops::add(ops::relu(n), n)}; }); report = rows(optimize_with(looped, {fold_neg_into_sub})); std::printf("%s | %s | asked %d\n", formula_of(looped).c_str(), report.c_str(), asked); // Add(Relu(Neg(x)), Neg(x)) | fold_neg_into_sub 0 | asked 0 return 0; } ``` ```python title="skipped_occurrences.py" # Add(a, Neg(b)), built node by node: the Add is the root, and the Neg feeds its second operand. pattern = Pattern() add = pattern.add_node(OpCode.Add) neg = pattern.add_node(OpCode.Neg) pattern.add_edge(neg, add, to_port=1) asked: list[str] = [] # the Add of each occurrence the condition is asked about # a + Neg(b) is a - b for an Add with its default settings, when its value is the one the occurrence hands on. def plain_sum(match: RewriteMatch) -> bool: total = match.nodes[0] asked.append(total.name) return (len(match.outputs) == 1 and total.attribute("alpha") == 1 and total.attribute("activation") == "Identity") @transforms.rule(pattern, condition=plain_sum) def fold_neg_into_sub(match: RewriteMatch, tensors: list[crt.Tensor]) -> crt.Tensor: return tensors[0] - tensors[1] # match.inputs holds a, then b report = graph.optimize([fold_neg_into_sub]) print(formula_of(graph), rows(report), len(asked)) # Sub(Relu(Relu(x)), x) [('fold_neg_into_sub', 1)] 1 # n = Neg(x), y = Relu(n) + n: the path from the Neg through the Relu leaves the occurrence and comes back in. def looped(inputs: list[crt.Tensor]) -> list[crt.Tensor]: n = crt.neg(inputs[0]) return [crt.relu(n) + n] asked.clear() loop = trace_model(looped) report = loop.optimize([fold_neg_into_sub]) print(formula_of(loop), rows(report), len(asked)) # Add(Relu(Neg(x)), Neg(x)) [('fold_neg_into_sub', 0)] 0 ``` Write a condition for each fact the replacement relies on, such as a matched node's settings, since the rule itself checks only the three cases above. ## The same API as the runtime's A rule you write with this API can give the same graph as one of the runtime's own transforms. `clamp_to_relu` turns a Clamp with a zero floor and no ceiling into a Relu, and the program writes it as a rule over one Clamp node. The rule's condition reads the Clamp's settings: a floor held as the literal zero, no ceiling, and one input, since a bound given as a tensor is an input of its own. Both runs turn the first Clamp into a Relu and keep the one with a ceiling, so the two graphs return the same formula. Their node names differ, since `replace` names the new Relu afresh where the runtime's transform keeps the Clamp's name. ```cpp title="same_api.cpp" int main() { // y = Clamp(x, 0) + Clamp(x, 0, 6): a zero floor alone, then a floor and a ceiling. const ClikaRT::graph::TraceFunction clamps = [](const std::vector& in) { const Tensor floored = ops::clamp(in[0], 0.0); const Tensor bounded = ops::clamp(in[0], 0.0, 6.0); return std::vector{ops::add(floored, bounded)}; }; // A Clamp with a zero floor held as its literal, and no ceiling, is a Relu. const auto zero_floor = [](const RewriteMatch& match) { const std::optional lower = match.nodes[0]->attribute("min"); const std::optional upper = match.nodes[0]->attribute("max"); const ClikaRT::Scalar* bound = lower.has_value() ? std::get_if(&lower->value) : nullptr; const bool zero = bound != nullptr && ((bound->is_float() && bound->float_value() == 0.0) || (bound->is_int() && bound->int_value() == 0)); const bool no_ceiling = !upper.has_value() || std::holds_alternative(upper->value); return match.inputs.size() == 1 && zero && no_ceiling; }; Pattern one_clamp; one_clamp.add_node(OpCode::Clamp); const Transform clamp_to_relu_rule = transforms::rewrite( "clamp_to_relu_rule", std::move(one_clamp), [](const RewriteMatch&, const std::vector& in) { return std::vector{ops::relu(in[0])}; }, zero_floor); ModelGraph by_runtime = trace_model(clamps); ModelGraph by_rule = trace_model(clamps); const std::string runtime_rows = rows(optimize_with(by_runtime, {transforms::clamp_to_relu()})); const std::string rule_rows = rows(optimize_with(by_rule, {clamp_to_relu_rule})); std::printf("%s | %s\n", formula_of(by_runtime).c_str(), runtime_rows.c_str()); // Add(Relu(x), Clamp(x)) | clamp_to_relu 1 std::printf("%s | %s\n", formula_of(by_rule).c_str(), rule_rows.c_str()); // Add(Relu(x), Clamp(x)) | clamp_to_relu_rule 1 std::printf("%s\n", formula_of(by_rule) == formula_of(by_runtime) ? "true" : "false"); // true return 0; } ``` ```python title="same_api.py" # y = Clamp(x, 0) + Clamp(x, 0, 6): a zero floor alone, then a floor and a ceiling. def clamps(inputs: list[crt.Tensor]) -> list[crt.Tensor]: x = inputs[0] return [crt.clamp(x, 0.0) + crt.clamp(x, 0.0, 6.0)] # A Clamp with a zero floor held as its literal, and no ceiling, is a Relu. def zero_floor(match: RewriteMatch) -> bool: settings = match.nodes[0].attributes return len(match.inputs) == 1 and settings.get("min") == 0 and settings.get("max") is None clamp_to_relu_rule = transforms.rewrite("clamp_to_relu_rule", Pattern([OpCode.Clamp]), lambda match, tensors: crt.relu(tensors[0]), zero_floor) by_runtime, by_rule = trace_model(clamps), trace_model(clamps) runtime_rows = rows(by_runtime.optimize([transforms.clamp_to_relu()])) rule_rows = rows(by_rule.optimize([clamp_to_relu_rule])) print(formula_of(by_runtime), runtime_rows) # Add(Relu(x), Clamp(x)) [('clamp_to_relu', 1)] print(formula_of(by_rule), rule_rows) # Add(Relu(x), Clamp(x)) [('clamp_to_relu_rule', 1)] print(formula_of(by_rule) == formula_of(by_runtime)) # True ``` Write your own version of one of the runtime's transforms when you need a variant of it, such as a condition of your own, and compare the graphs the two give. ## Finalize and run the transformed graph After the transforms, `finalize()` makes the graph runnable, and `run()` computes with it, as for any graph. A list of your own can start from the default list: `default_transforms()` returns it, and a transform added at its end runs after the runtime's. Each program checks the result against a reference it computes itself, and the Python program also imports numpy as `np` for that. ```cpp title="after_the_transforms.cpp" int main() { ModelGraph graph = trace_model(); // The runtime's default list, with a transform of your own at its end. std::vector list = graph.default_transforms(); list.push_back(Transform::from_function("drop_repeated_relu", drop_repeated_relu)); optimize_with(graph, list); graph.finalize(); // the graph runs from here on, and refuses edits const std::vector x = {1.0F, -2.0F, 3.0F, -4.0F, 5.0F, -6.0F}; const std::vector results = graph.run({Tensor::from_data(x.data(), {2, 3}, ClikaRT::DataType::Float32)}); const std::vector y = results.front().reshape({-1}).item_as_vec(); bool matches = y.size() == x.size(); for (std::size_t i = 0; matches && i < x.size(); ++i) { matches = y[i] == (x[i] > 0.0F ? x[i] : 0.0F) - x[i]; // the reference, Relu(x) - x, by hand } for (std::size_t i = 0; i < y.size(); ++i) std::printf("%s%g", i == 0 ? "" : " ", static_cast(y[i])); std::printf("\n%s\n", matches ? "true" : "false"); // 0 2 0 4 0 6 // true return 0; } ``` ```python title="after_the_transforms.py" # The runtime's default list, with a transform of your own at its end. graph.optimize(graph.default_transforms() + [Transform("drop_repeated_relu", drop_repeated_relu)]) graph.finalize() # the graph runs from here on, and refuses edits x = np.array([[1.0, -2.0, 3.0], [-4.0, 5.0, -6.0]], np.float32) # numpy as the data entry (y,) = graph.run([crt.tensor(x)]) print(y.numpy().tolist()) # [[0.0, 2.0, 0.0], [4.0, 0.0, 6.0]] (numpy as the data exit) print(np.array_equal(y.numpy(), np.maximum(x, 0) - x)) # True (numpy states the reference) ``` Check a graph your transforms rewrote against a reference like this one before you serve it. [Optimize and finalize](optimize-and-finalize.mdx) covers the runtime's transforms and their defaults, [Edit a graph](edit-a-graph.mdx) the edits a transform makes, and [Query a graph](query-a-graph.mdx) the reads a transform or a condition can make. --- # System requirements Platforms, accelerators, host software, memory and storage for running ClikaRT. Source: https://docs.clika.io/clikart/system-requirements.md ClikaRT ships one distribution per platform. Every distribution includes the CPU backend; accelerator backends are included where the platform has them and load on demand at run time. ## Platforms | Platform | Architecture | Backends | | --- | --- | --- | | Linux | x86_64 | CPU, CUDA, Vulkan | | Linux | arm64 | CPU, CUDA, Vulkan | | Android | arm64-v8a | CPU, Vulkan | | Windows | x86_64 | CPU, Vulkan (CUDA on request) | | Windows | arm64 | CPU, Vulkan | | macOS | Apple silicon | CPU, Metal | ## Host software | Requirement | Detail | | --- | --- | | C++ toolchain | A C++17 compiler for your own code; the library itself has no compiler requirement | | CMake | 3.19 or newer | | Linux glibc | 2.28 or newer (manylinux_2_28 compatible) | | macOS | 14 (Sonoma) or newer, on Apple silicon | | NVIDIA driver | The driver alone: the runtime's CUDA images carry the CUDA runtime and cuBLASLt inside them, so no CUDA toolkit install and no `LD_LIBRARY_PATH`. A driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image; the Linux archives carry both, and the Python wheel comes in one flavor per image | | Vulkan | A driver supporting Vulkan 1.2 or newer | | Python | CPython 3.10 to 3.14 for the `clika-runtime` wheel, which carries the runtime and `clika_runtime.modelverse`, on Linux, macOS and Windows x64; CPython 3.11 to 3.14 on Windows arm64. The wheel is an LZMA-compressed zip, which `pip` reads and `uv pip` does not | | License credential | The `CLIKA1-...` credential your project's license carries, in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes. Every process that runs an operator needs one ([License credential](getting-started/get-clikart.mdx#license-credential)) | ## Accelerators | Backend | Hardware | | --- | --- | | CUDA | NVIDIA GPUs from compute capability 7.0 (Volta) upward, Jetson included | | Vulkan | Desktop and mobile GPUs with a conformant Vulkan driver (NVIDIA, AMD, Intel, Qualcomm Adreno) | | Metal | Apple silicon | ## Memory A running model needs memory in proportion to the size of its weights file. The minimums below were measured on devices running the platform's performance benchmark (up to 8 concurrent requests, prompts up to about 2K tokens), with models whose weights are up to about 2.5 GB. The KV cache drives the peak, so fewer concurrent requests or a shorter context need less, and more requests, a longer context or a larger model can need more than these rules give. | Backend | Platforms | Minimum memory | | --- | --- | --- | | CPU | Linux x86_64, Linux arm64, Windows x86_64, macOS | 2 GB of RAM plus about 3.5 times the model's file size | | CUDA on a discrete NVIDIA GPU | Linux x86_64 | 1 GB of system RAM plus about 0.5 times the model's file size, in addition to GPU memory for the model | | CUDA on Jetson | Linux arm64 | 2 GB plus about 4 times the model's file size, in the memory the CPU and GPU share | | Metal on Apple silicon | macOS | no rule: the one measurement, on a 32 GB machine, read about 11 to 14 GB for models from about 0.9 to 2.4 GB, nearly the same whatever the model's size, which does not give a minimum; a rule needs the benchmark on a 16 GB and an 8 GB Mac | | Vulkan on an integrated GPU | Linux x86_64, Windows x86_64 | no rule: the host's resident-memory counter leaves out the GPU driver's own allocations, so a figure read that way understates the need; a rule needs a GPU-memory counter beside it | The file size is the size of the model's weights on disk as the platform lists the model. The measurements used models stored as bf16 safetensors; a quantized file of the same model is smaller on disk, and these rules were not measured for it. ## Storage ### Engine package The engine package is the ClikaRT build the platform delivers to a device for a benchmark or a model deployment. The platform delivers it to Linux, macOS and Windows x86_64 devices; Windows arm64 devices receive none. Where it can, the platform sends only the part of the engine the device uses. The device keeps the download in its cache next to the extracted tree, so plan for both together. They live under the device agent's staging directory: `/var/lib/clika-runtime-agent/staging` for a system-wide agent on Linux and macOS, `%ProgramData%\Clika\DeviceAgent\staging` on Windows, otherwise `~/.clika-rt/staging`. | Device | Download | Installed | Both together | | --- | --- | --- | --- | | Linux x86_64, no GPU | the dialog's figure | about 320 MB | the download plus about 320 MB | | Linux x86_64, NVIDIA GPU, driver 580 or newer | the dialog's figure | about 1.6 GB | the download plus about 1.6 GB | | Linux x86_64, NVIDIA GPU, older driver | the dialog's figure | about 3.0 GB | the download plus about 3.0 GB | | Linux x86_64, another GPU (AMD, Intel): the whole package | about 1.96 GB | about 4.1 GB | about 6.1 GB | | Linux arm64, no GPU | the dialog's figure | about 210 MB | the download plus about 210 MB | | Linux arm64, with the Vulkan backend | the dialog's figure | about 350 MB | the download plus about 350 MB | | Linux arm64, NVIDIA GPU, driver 580 or newer | the dialog's figure | about 1.8 GB | the download plus about 1.8 GB | | Linux arm64, NVIDIA GPU, older driver | the dialog's figure | about 3.6 GB | the download plus about 3.6 GB | | Linux arm64, the whole package | about 2.41 GB | about 5.1 GB | about 7.5 GB | | Windows x86_64, no GPU | the dialog's figure | about 220 MB | the download plus about 220 MB | | Windows x86_64, with a GPU: the whole package | about 159 MB | about 370 MB | about 530 MB | | Windows arm64, the whole package | about 154 MB | about 330 MB | about 480 MB | | macOS, Apple silicon: the whole package | about 31 MB | about 170 MB | about 200 MB | | Android arm64, the whole package | about 64 MB | about 310 MB | about 380 MB | The installed figures are the release archive's tree, summed, less the backend images a row does not carry: on Linux x86_64 the CUDA 13 image is about 1.2 GB, the CUDA 12 image about 2.5 GB and the Vulkan backend about 150 MB of the 4.1 GB; on Linux arm64 about 1.4 GB, 3.3 GB and 150 MB of the 5.1 GB; on Windows the Vulkan backend is about 150 MB of the tree; on macOS the Metal backend about 40 MB. A whole package's download is the release archive itself, whose size the platform's dialog and `clika-cli runtime-sdk list` show beside it; a composed download (one backend's worth of the tree) is smaller, and the dialog shows its size in the same place, so a row that reads "the dialog's figure" takes it from there. The first time the platform prepares a part of the engine for one kind of device, that one job or deployment receives the whole package for the platform instead (on Linux x86_64, the "another GPU" row). Each model deployment keeps its own extracted copy, and after an engine update the device also keeps the previous version's download for a while. ### Models Required storage grows with the models you pull, by each model's file size. Benchmarks keep their models in a cache of up to 20 GB by default and remove the oldest first. A model deployment keeps its model until the deployment is deleted with its model removed. A model used by both is stored twice. ## Platform software The device agent takes about 25 MB of disk. `clika-cli` takes about 15 MB of disk and about 70 MB of RAM while it runs. --- # ClikaRT CLI ClikaRT CLI is the command line interface that infers, serves and benchmarks popular models, built on top of Modelverse, the CLIKA model library, and running them on ClikaRT. Source: https://docs.clika.io/modelverse.md ClikaRT CLI is a command line interface that lets you infer, serve or benchmark popular models, one command each. It is built on top of Modelverse, the CLIKA model library: models packaged so that [ClikaRT](/clikart) loads and runs them as they are, a catalog of registered model families covering language, vision, audio and multimodal models, and a C++ library when you want the same machinery inside your own application. The executable is `clikart-cli`: one command takes a model name to generated text, one more serves it over HTTP, one more benchmarks it on the device it runs on. ## Why ClikaRT CLI For the AI developer. The checkpoints you already use are the input. ClikaRT CLI resolves a Hugging Face Hub repo id, a pasted Hugging Face URL or a local directory to model files, matches them to a registered family, and runs them; ONNX exports, GGUF quantizations and safetensors checkpoints all load through ClikaRT unchanged. `clikart-cli prompt "..."` is a working generation before you have written any code, and every knob you expect (sampling, system prompt, context length, KV cache mode) is a flag. For the backend engineer. A model becomes an OpenAI-compatible endpoint in one command. `clikart-cli serve` hosts `/v1/chat/completions` with streaming, plus embeddings, transcription and the other engine routes a model family provides, and existing OpenAI clients point at it by changing one base URL. `clikart-cli` is built for scripts: stdout carries only the payload, diagnostics go to stderr, and the exit codes follow a fixed four-value contract. No Python runs anywhere. For the embedded developer. A model family ships quantized variants, and you pick the one that fits the device. A GGUF repo with ten quantizations is a selector away (`/:Q6_K`), the option table with file sizes prints before anything downloads, and the same model runs wherever ClikaRT runs, from a workstation GPU to a phone. For the defense, healthcare and finance developer. Nothing here requires a network at run time. Fetch a model on a connected machine, move the directory, and point the CLI at it; a local directory is a first-class model source, and an offline flag makes any network touch an error instead of a surprise. The install is one archive extracted into one directory, runtime and Modelverse together, and the models arrive the same way. For the business. Every model in the catalog has a known license: the catalog records who published each family and under what terms. You do not have to vet checkpoints from unknown sources. Models come straight from Hugging Face by repository id, so the models your team already uses work as-is. Quantized variants run the same model on cheaper hardware, which lowers serving cost. And you do not have to build inference for the popular models yourself. ClikaRT CLI already runs them. ## What you get - **The catalog.** Dozens of registered model families, from Llama, Qwen and Gemma through Whisper, CLIP, DETR and Depth Anything. Each family declares its input and output modalities, the checkpoint formats it matches, and the commands it can run. `clikart-cli list` prints it. - **The executable.** `clikart-cli` inspects (`info`), downloads (`fetch`) and runs models. Which commands a model supports (prompt, serve, transcribe, embed, bench and more) depends on its family, discovered per model. - **The server.** An OpenAI-compatible HTTP server with streaming chat completions, a built-in web chat page and a health probe, plus per-modality routes for transcription, embeddings, depth, detection and segmentation. - **The library.** The surface behind all of it, in C++ with Python and Kotlin bindings: fetch a snapshot, load a runnable model, build a serving pipeline, or mount your own engine on the server. For C++, one `find_package(Modelverse CONFIG)` integrates it; the bindings arrive through their package managers. - **The packaging.** One archive per platform, the ClikaRT runtime and Modelverse inside it, extracted and run in place. A manifest pins the exact runtime each build linked against, so a mismatched pair refuses with a readable error. No installer and no downloads at run time. ## Where to go next - [Getting started](getting-started/index.md): install the archive and run your first model. - [How-to guides](how-to/index.md): problem-oriented recipes, from quantization selection to offline deployment. - [Model requirements](model-requirements.md): memory and device figures per model variant. - [Additional examples](examples.md): the example programs, one per subsystem. - [ClikaRT](/clikart): the runtime underneath, with its own tutorial and API reference. --- # Additional examples The Modelverse example programs in the release's examples archive, standalone CMake projects against the installed package, from the catalog probe to vision, OCR, reranking and vision-language chat. Source: https://docs.clika.io/modelverse/examples.md The Modelverse examples are standalone `find_package(Modelverse CONFIG)` projects on the public API only, one shared `README.md` walk-through beside them, and together they cover the library surface the [tutorial](getting-started/first-model/01-pick-a-model.md) meets through the `clikart-cli` executable. They live in the release's examples archive, `ClikaRT--examples.tar.xz` ([Get ClikaRT](/clikart/getting-started/get-clikart) names the download), under `cpp/modelverse/`; the release archive's own `examples/src` carries the runtime's examples ([the ClikaRT catalog](/clikart/examples)). Building them needs the release archive for your platform, extracted, with `CLIKART_BUNDLE_DIR` naming its directory: the one `cmake/` directory there carries both products' packages. Read top to bottom; each row assumes a little of the ones above it. | Example | What it shows | | --- | --- | | `init_model` | The registered catalog and identity resolution: list the families, resolve a model's identity from a local directory or a Hugging Face Hub repo id. The snapshot fetches configs and companions only, so identity costs no weight download. | | `00_generate` | The minimal end-to-end path: explicit snapshot, registry match, generative pipeline, one prompt, text. The program [part 4 of the tutorial](getting-started/first-model/04-use-it-from-code.mdx) builds. | | `01_serve` | The ServeAPI in process: implement the `ChatEngine` interface, mount it on `ServeApi`, and round-trip one `/v1/chat/completions` request through the built-in HTTP client. Runs with no arguments and no checkpoint, so it is also the fastest server smoke test. | | `02_custom_node` | Your own `ClikaRT::runtime::Model` node composed with the library's serving nodes in one pipeline: a synthesized two-layer llama generates and a short custom node post-processes the reply, with the request surface unchanged. Offline, runs with no arguments. The worked version against a real model is [Add your own node to a model pipeline](how-to/add-a-pipeline-node.md). | | `03_conversational_pipeline` | A conversational AI as one pipeline: wav bytes in, the reply waveform out, through speech to text, the chat template, a chat model and text to speech across eight nodes; real models, so it fetches on first run and takes a question wav plus the voice-reference wav the speech model requires. The guide is [A conversational AI as one pipeline](how-to/conversational-ai-pipeline.md). | | `04_classify_zero_shot` | Zero-shot image classification: the labels you name scored against a picture by cosine in a dual-tower checkpoint's shared space (SigLIP); the picture drawn by the program; no arguments. | | `05_detect_objects` | Object detection over the checkpoint's own label set (DETR): boxes in the picture's pixels above a confidence; no arguments. | | `06_detect_by_phrase` | Open-vocabulary detection (OWL-ViT): the prompt's phrases encoded once, a box per thing found, labeled with its phrase; no arguments. | | `07_estimate_depth` | Monocular depth (Depth Anything V2): one value per pixel at the picture's size, metric or relative; no arguments. | | `08_read_text` | Text recognition (GLM-OCR): a page of block letters drawn with the operators, read back as text through the registry's text-recognition door; no arguments. | | `09_rerank` | Reranking (Qwen3-Reranker): one relevance score per document against a query, printed best first; no arguments. | | `10_describe_a_picture` | Vision-language chat (Qwen3-VL): one question about a picture, the picture riding the user turn; greedy, no arguments. | A program that runs with no arguments and prints the same text on every run sits beside its recorded output (`.out`), the text the release printed; a program that takes a checkpoint or a clip, or prints a port, is built and not recorded. ## Build and run From the extracted `examples/` directory: ```bash cmake -S cpp/modelverse -B build-examples \ -DModelverse_DIR="$CLIKART_BUNDLE_DIR/cmake" cmake --build build-examples ``` ```bash build-examples/init_model # catalog only, offline build-examples/init_model -r openai/whisper-large-v3-turbo # + identity from Hugging Face build-examples/00_generate Qwen/Qwen2.5-0.5B-Instruct "Once upon a time, in a port town by a cold sea," build-examples/01_serve build-examples/02_custom_node # offline, no arguments build-examples/05_detect_objects # the vision, OCR, reranking and chat programs: no arguments, the checkpoint from the local cache ``` `01_serve` prints its bound port, answers one request against itself, and exits `PASS`, which makes it the one to run first when checking a new machine: ```text serving on 127.0.0.1:41627 HTTP 200 choice: echo: Say hello. 01_serve: PASS ``` Every program runs compute, so it needs the license credential in `CLIKA_RT_LICENSE` or in the per-user file `clikart-license-init` writes; without one a call is refused with the code name `LICENSE_FAILED`. Where models come from, cache placement and the offline path are the same for the examples as for `clikart-cli`; the examples' shared `README.md` restates them next to the code. --- # First steps New to Modelverse? Start here. What the model library is, how to install it, and your first model from catalog to served endpoint. Source: https://docs.clika.io/modelverse/getting-started.md New to Modelverse? This section is where to start. It gives enough orientation to hold the product in your head, an install you can verify in minutes, and a first model that goes from the catalog to a served endpoint. Read it in order: 1. **[Modelverse at a glance](overview.mdx)**: what the model library is and is not, and the five-minute mental model. 2. **[Get Modelverse](get-modelverse.md)**: pick your platform, get the download and verify commands. 3. **[Quick install](installation.md)**: extract the archive and prove it works in two commands. 4. **Tutorial series**: four parts, each a complete step. [Pick a model](first-model/01-pick-a-model.md) (the catalog, identity, what a download would cost), [fetch and prompt](first-model/02-fetch-and-prompt.md) (weights on disk, first generated text), [serve and chat](first-model/03-chat-and-serve.md) (an OpenAI-compatible endpoint with a built-in chat page), and [use it from your code](first-model/04-use-it-from-code.mdx) (the same model inside your own program). 5. **Adding your own model**: the author-side series, three parts on the smallest real model (an embedding model: text in, vector out). [A model from scratch](own-model/01-a-model-from-scratch.md) (a checkpoint is a directory you can write by hand), [join the catalog](own-model/02-join-the-catalog.md) (register a family of your own), and [make it embed](own-model/03-make-it-embed.md) (the smallest real backend, the standard commands in two lines). Read it when you bring your own architecture; the first series does not depend on it. 6. **[What to read next](next-steps.md)**: where to go once it runs. ## How the Modelverse docs are layered - **This section** orients: condensed, in reading order, concepts explained where they first appear. - **[How-to guides](../how-to/index.md)**: problem-oriented recipes, one per "how do I X", with [additional examples](../examples.md) as the end-to-end reading inside it. - **[ClikaRT](/clikart)** documents the runtime underneath, including the API reference the C++ surface builds on and the [system requirements](/clikart/system-requirements) Modelverse inherits. --- # Serve and chat Host the model behind an OpenAI-compatible HTTP endpoint, hold a conversation on the built-in chat page, and call it with curl. Source: https://docs.clika.io/modelverse/getting-started/first-model/chat-and-serve.md The model generates on demand; this part keeps it loaded. `serve` hosts the model behind an OpenAI-compatible server with a built-in chat page. One command, no configuration files: ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct serve ``` ```text [2026-10-04 07:56:43.933] [modelverse] [info] serving on http://127.0.0.1:8000 (model=Qwen2.5-0.5B-Instruct); open it in a browser for the chat page [2026-10-04 07:56:43.933] [modelverse] [info] endpoints: POST /v1/chat/completions /v1/messages /v1/messages/count_tokens /v1/audio/transcriptions; GET /v1/models /health /props / (web UI) /dashboard ``` The defaults bind `127.0.0.1:8000`; `--host 0.0.0.0` opens it to the network and `--port` moves it. Three things are now running: - **The API.** `POST /v1/chat/completions`, streaming (server-sent events) and non-streaming, in the OpenAI request and response shape. - **A web chat page.** `http://127.0.0.1:8000/` serves a built-in chat UI from inside the binary (`--no-web-ui` disables it). - **A health probe.** `GET /health` answers `{"status":"ok"}`, for load balancers and scripts. ## Chat in the browser The chat page is the conversation surface: a scrolling transcript with an input box, the assistant reply streaming token by token as the pipeline decodes it. The conversation carries its history, so follow-up questions see earlier turns. The same conversation is available to any OpenAI client through the API, where the system message is the first `messages` entry and the sampling knobs (`temperature`, `top_p`) ride each request. One templated turn from the terminal stays `prompt`, part 2's command. `prompt`, `serve` and `bench` are the text-generation surface, and there is no terminal `chat` verb: `clikart-cli chat` is refused by name, `error: family 'qwen' (text->text) provides no chat (it provides: prompt, serve, bench, mm_bench)`. ## Call it over HTTP From a second terminal: ```bash curl -s http://127.0.0.1:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "messages": [{"role": "user", "content": "Say hello in French."}], "max_tokens": 32 }' ``` ```text {"id":"chatcmpl-1791068204-0","object":"chat.completion","created":1791068204,"model":"Qwen2.5-0.5B-Instruct","choices":[{"index":0,"message":{"role":"assistant","content":"Bonjour! C'est un plaisir de vous rencontrer. Comment puis-je vous aider aujourd'hui ?"},"finish_reason":"stop"}],"usage":{"prompt_tokens":34,"completion_tokens":20,"total_tokens":54,"prompt_tokens_details":{"cached_tokens":0}}} ``` The `id`, `created` and `usage` values are your run's own. Add `"stream": true` and the response arrives as server-sent events, one delta per chunk, the way OpenAI clients expect. An existing client needs one change: ```python from openai import OpenAI client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused") reply = client.chat.completions.create( model="qwen", messages=[{"role": "user", "content": "Say hello in French."}], ) print(reply.choices[0].message.content) ``` The endpoint surface is bigger than chat: a Whisper model's `serve` answers `POST /v1/audio/transcriptions`, an embedding model's answers `POST /v1/embeddings`, and so on per family. [Serve a model over the OpenAI and Anthropic APIs](../../how-to/serve-openai-compatible.md) has the full route table and the operational flags. Next: [part 4](04-use-it-from-code.mdx), the same model inside your own program. --- # Fetch and prompt Download the model snapshot, generate your first text, and control sampling from the command line. Source: https://docs.clika.io/modelverse/getting-started/first-model/fetch-and-prompt.md Part 1 established what the model is; this part downloads it and makes it generate. At the end you have the weights on disk in a directory you control and a repeatable one-shot generation command. Generation is compute, so the credential from [Quick install](../installation.md#2-place-the-license-credential) has to be in place before the `prompt` command below. `fetch` downloads files and needs none. ## Fetch the snapshot `fetch` downloads a source's files (companions first, then the weights, one progress bar per file) and prints exactly one thing on stdout: the local path the download landed at. ```bash clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct ``` ```text /home/you/.cache/huggingface/hub/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775 ``` Progress, timing and the file count print to stderr, so the path is safe to capture: ```bash MODEL_DIR=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct) ``` By default snapshots land in the standard Hugging Face Hub cache, shared with other Hugging Face tooling on the machine; `--cache-dir models` keeps them in a directory of your choosing instead. Fetching a source that is already cached verifies and returns immediately, so scripts can fetch unconditionally. A gated repo needs a token. The CLI reads it the way the other Hugging Face tooling does: `HF_TOKEN` in the environment, else the token a Hub login stored (`hf auth login` writes it to `$HF_HOME/token`; `HF_TOKEN_PATH` names another file). There is no token flag, deliberately, so a token can never land in shell history. Qwen2.5 0.5B Instruct is not gated, so its fetch needs no token; for a gated repo, accept its license once on the repo's Hub page, provide a token, and the fetch proceeds (without them it refuses by name: `'/' is gated or private on https://huggingface.co; set $HF_TOKEN or log in to the hub`). Two related commands you already have: `fetch --dry` prints the total of what the fetch downloads without downloading it (`info --dry` adds the file listing behind that total), and a local directory used as a source skips fetching entirely. One refusal worth meeting on purpose: `fetch` takes checkpoints in the formats from part 1 (safetensors, torch containers, GGUF) and refuses anything else before a byte of weights moves. An ONNX-only export, for example: ```text $ clikart-cli fetch onnx-community/Qwen2-0.5B-Instruct-ONNX error: 'onnx-community/Qwen2-0.5B-Instruct-ONNX' ships its weights only in formats the loaders do not read: ONNX (onnx/model.onnx, onnx/model_bnb4.onnx, onnx/model_fp16.onnx and 5 more). The loaders read safetensors, GGUF and the torch container ``` ## The first generation `prompt` runs one templated generation of the positional message: the model's own chat template wraps your text, the pipeline decodes, and the generated text is the stdout payload. The template is why a question gets an answer here: a raw, untemplated completion would CONTINUE the question's shape instead (part 4 shows that mode from the library). Prompt answers one shot; a conversation runs against the served model in part 3. ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct prompt "The capital of France is" ``` ```text Paris. ``` The first run loads the weights (a progress bar on stderr); repeat runs on a warm cache start in seconds. The same invocation works with `"$MODEL_DIR"` in place of the repo id, which is the fully offline form. ## The knobs are flags Sampling and context are controlled per invocation. The ones you will use first: ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct prompt "Name three rivers." \ --max-new-tokens 32 \ --temperature 0.2 \ --seed 7 ``` ```text Three rivers that I can name for you are the Yangtze River, the Yellow River, and the Pearl River. Each of these rivers is significant in Chinese ``` - `--max-new-tokens` caps the generation length; `--max-seq` caps the whole context. - `--temperature` shapes sampling; `--greedy` disables it for the single most likely continuation, and `--seed` makes a sampled run repeatable. - A conversation's knobs move into the request once the model is served: the system message is the first `messages` entry, and `temperature`/`top_p` ride each OpenAI request ([part 3](03-chat-and-serve.md)). - `--device cuda` places the model explicitly; the default picks the best available device, and `clikart-cli devices` shows the candidates. `--device auto` names that default: the same pick, falling back down the accelerator order (CUDA, TPU, Metal, Vulkan) when a load fails, and the CPU last. `GET /props` on a served model reports which device it landed on. A vision-capable model takes media the same way. With the flagship multimodal family the message and the image travel together: ```bash clikart-cli google/gemma-4-E2B-it prompt "What is on the sign?" --image photo.jpg ``` The root commands' options also have a file form: `generate-template` writes a JSON file of every root option at its default, `--template ` loads one, and explicit flags win over it (a family's own verbs, like `prompt` and `serve`, carry their options as flags only). [Script clikart-cli](../../how-to/script-the-cli.md) covers that workflow. Next: [part 3](03-chat-and-serve.md), the model behind an HTTP endpoint, with a conversation on its built-in chat page. --- # Pick a model Read the catalog, resolve a model's identity, and see what a download would cost, all before any weights move. Source: https://docs.clika.io/modelverse/getting-started/first-model/pick-a-model.md This tutorial takes one model from the catalog to a served endpoint in four parts, each a complete session, and ends with the same model running inside a C++ program. Core concepts are explained where they first appear. This part picks the model and learns everything about it without downloading a single weight. It assumes the install directory exists and `clikart-cli` is on your `PATH` (see [Quick install](../installation.md)). The model is Qwen2.5 0.5B Instruct, small enough to run on any machine in the [system requirements](/clikart/system-requirements); every command works the same with any other model in the catalog. ## The catalog knows the families `list` prints every registered model family: its modalities, the commands it provides, and whether it is runnable on this build. No network is involved; the catalog is compiled into the `clikart-cli` executable. ```bash clikart-cli list ``` ```text registered model families (47) bert runnable * Input Modalities: text * Output Modalities: embedding * Valid Combos: text -> embedding * Model Types: bert * Commands: embed, similar, bench, serve * Web UI: generic page * Vendor: Google * License page: https://github.com/google-research/bert/blob/master/LICENSE ------------------------------------------ ... ``` A family is the model architecture Modelverse knows how to run; a model you fetch is a checkpoint of that family. The `Commands` row is the contract for parts 2 and 3: whatever it lists is what that model can do. ## A source names a model Everything model-specific starts from a source: a Hugging Face Hub repo id (`/`), a pasted Hugging Face URL, or a local directory. Modelverse runs checkpoints in the formats model publishers ship on the Hub: safetensors, torch containers (`pytorch_model.bin`), and GGUF. (ONNX exports are a different lane: the ClikaRT runtime runs ONNX models directly; its how-to guide covers that path.) `info` resolves a source's identity: ```bash clikart-cli info Qwen/Qwen2.5-0.5B-Instruct ``` ```text Qwen/Qwen2.5-0.5B-Instruct family=qwen model_type=qwen2 architecture=Qwen2ForCausalLM variant: dense components: tokenizer=yes image=no audio=no video=no license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT model: Qwen/Qwen2.5-0.5B-Instruct provider: Qwen source: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct license: apache-2.0 https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/blob/main/LICENSE ``` Identity resolution reads configuration files only. `clikart-cli` downloads a few KB of JSON, matches it against the registered families, and reports what it found; weights do not move. This is deliberate: you can interrogate a 70B model from a laptop. The last three lines are the licensing report: the family's license page, the checkpoint's provider and source, and the checkpoint's own license with whether the repository is gated. Read them before you ship what the model produces. ## `--dry` shows what a fetch would cost Add `--dry` and `info` also lists the files a fetch downloads, with sizes, split into companions (configs, tokenizer) and weights, and says how much of the repository it leaves behind (other weight formats, files the fetch never takes): ```bash clikart-cli info Qwen/Qwen2.5-0.5B-Instruct --dry ``` ```text Qwen/Qwen2.5-0.5B-Instruct family=qwen model_type=qwen2 architecture=Qwen2ForCausalLM variant: dense components: tokenizer=yes image=no audio=no video=no license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT model: Qwen/Qwen2.5-0.5B-Instruct provider: Qwen source: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct license: apache-2.0 https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/blob/main/LICENSE companions: config.json 659 B generation_config.json 242 B merges.txt 1.6 MiB tokenizer.json 6.7 MiB tokenizer_config.json 7.1 KiB vocab.json 2.6 MiB weights: model.safetensors 942 MiB total: 953 MiB (weights 942 MiB, 7 files) repository: 953 MiB in 10 files; the rest is not fetched (other weight formats, files the fetch never takes) (dry; nothing downloaded) ``` Nothing here is Qwen-specific. `info google/gemma-4-E2B-it` reports `variant: dense + multimodal`; `info openai/whisper-large-v3-turbo` reports an audio component. Some repos ship several weight options to choose between; [Run a specific GGUF quantization](../../how-to/run-a-gguf-quantization.md) covers picking one. ## The model's own commands `clikart-cli` has four commands that work without naming a model (`list`, `devices`, `info` and `fetch`; `generate-template`, their file-form helper, rides beside them). Every other command runs on a model you name: pass the source and `--help`, and `clikart-cli` resolves its family and prints that family's own command surface: ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct --help ``` ```text Qwen/Qwen2.5-0.5B-Instruct family=qwen text->text usage: clikart-cli Qwen/Qwen2.5-0.5B-Instruct [options] commands the 'qwen' family provides for this model options: -h, --help show this help and exit --verbose, -v debug diagnostics (the effective-options report) --quiet, -q warnings and errors only --no-color plain output (also honored: $NO_COLOR, TERM=dumb, non-tty) commands: prompt one templated generation of the positional user message serve run the http server (OpenAI and Anthropic APIs) bench the family-owned benchmark flow mm_bench the multi-modality benchmark flow (media axes over the serving path) ``` Asking a model for a command its family does not provide refuses precisely and names what it does provide, with exit code 3 ([Script clikart-cli](../../how-to/script-the-cli.md) has the full exit contract). Next: [part 2](02-fetch-and-prompt.md), the weights arrive and the model speaks. --- # Use it from your code Link the Modelverse library and run the same model inside your own program: snapshot, load, pipeline, text. Source: https://docs.clika.io/modelverse/getting-started/first-model/use-it-from-code.md {/* CERTIFICATION: the C++ sample compiles and runs against the pinned release bundle (linux_x86_64, CPU); the output block is that run's capture. load.max_seq mirrors the CLI's 4096 default (the library otherwise keeps the checkpoint's full window and sizes the KV cache from it). The Kotlin arm is a staged binding surface and carries its own in-tab marker. */} Everything the `clikart-cli` executable did in parts 1 through 3 is a library call. This part writes the smallest program that does what `prompt` does: resolve a snapshot, load the runnable model, and generate. ## The project For C++, the install directory you already have is also the SDK (the headers, the library, and the ClikaRT runtime beside them): the platform's one archive for your platform carries the runtime and Modelverse together ([Get Modelverse](../get-modelverse.md)), so there is no second download. A two-file CMake project is the whole setup. The C++17 toolchain from the [ClikaRT quick install](/clikart/getting-started/installation) prerequisites is the one requirement. ```cmake title="CMakeLists.txt" cmake_minimum_required(VERSION 3.19) project(hello_modelverse LANGUAGES CXX) set(CMAKE_CXX_STANDARD 17) set(CMAKE_CXX_STANDARD_REQUIRED ON) find_package(Modelverse CONFIG REQUIRED) add_executable(hello main.cpp) target_link_libraries(hello PRIVATE Modelverse::modelverse) # Copy the Modelverse library and the ClikaRT runtime next to the binary, # so the program runs from the build directory as-is. modelverse_stage_runtime(TARGET hello) ``` `find_package(Modelverse CONFIG)` wires everything: it defines the one link target (`Modelverse::modelverse`), and finds the ClikaRT runtime installed in the same root (a `ClikaRT::ClikaRT` you already provide is respected instead). Python arrives in the `clika-runtime` wheel, which carries the model library as `clika_runtime.modelverse`; Kotlin arrives in the one Maven artifact, `io.clika:clika-runtime`, which carries the model library as `io.clika.modelverse` (the artifact's zip, extracted, is the Maven repository the build names). Both are downloads of the platform ([Get ClikaRT](/clikart/getting-started/get-clikart)): ```bash pip install ./clika_runtime-0.6.4-cp313-cp313-manylinux_2_28_x86_64.whl # Python: the wheel, the model library inside # Kotlin/Gradle: implementation("io.clika:clika-runtime:0.6.4") ``` ## The program ```cpp title="main.cpp" #include #include "clika_modelverse/hub/hub.h" #include "clika_modelverse/registry/model_registry.h" #include "clika_modelverse/runtime/serving_pipeline.h" #include "clika_modelverse/runtime/text_nodes.h" namespace mv = clika_modelverse; namespace mvr = clika_modelverse::runtime; int main() { // 1. Model files, exactly as `fetch` gets them. An already-fetched or // local directory passes through untouched. mv::hub::SnapshotOptions snap; snap.cache_dir = "models"; // beside the program; empty = the shared Hugging Face cache const mv::hub::SnapshotResult snapped = mv::hub::snapshot( "Qwen/Qwen2.5-0.5B-Instruct", snap); // 2. The registry matches the snapshot to its family and returns the // runnable model, on the device you name (CPU by default). mv::LoadOptions load; load.max_seq = 4096; // the CLI's default context cap; 0 keeps the // checkpoint's full window and sizes the KV cache from it mv::GenerativeModel model = mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load); // 3. A serving pipeline around it: tokenizer -> decoder -> detokenizer, // the same assembly `prompt` and `serve` run on. mv::generation::GenerationConfig gen = model.defaults(); gen.max_new_tokens = 64; gen.stop = {"."}; // raw completion: stop at the first sentence end mvr::PipelineOptions opts; mvr::GenerativePipeline pipe = mvr::build_generative_pipeline(model, gen, opts); // 4. One request through it. ClikaRT::runtime::Request req; req.inputs.set("prompt", mvr::string_to_byte_tensor( "Once upon a time, in a port town by a cold sea,")); const auto sid = pipe.pipeline->enqueue(std::move(req)); const ClikaRT::runtime::Response resp = pipe.pipeline->await(sid); std::printf("%s\n", mvr::byte_tensor_to_string(resp.outputs.get("text")).c_str()); return 0; } ``` ```python title="main.py" import clika_runtime.modelverse as mv # 1. and 2. Model files and the runnable model in one call: the snapshot # downloads into the hub cache (a local directory passes through untouched), # the registry matches it to its family and loads it on the device named. model = mv.AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", device="cpu") # 3. and 4. The generation: text in, text out; the keyword arguments are the # decode knobs over the checkpoint's own defaults. print(model.generate("Once upon a time, in a port town by a cold sea,", max_new_tokens=64, stop=["."])) ``` The model served in part 3 answers any OpenAI client as well (`clikart-cli serve`, then a client pointed at `http://127.0.0.1:8000/v1`), which is the path for a program that must not load the model itself. {/* CERTIFICATION: the Kotlin arm shows the Android binding's shape as the pinned release's Kotlin package declares it (Modelverse.load, AutoModelForCausalLM.fromPretrained, LoadOptions, generate); it is not compiled by the docs build. */} ```kotlin title="Chat.kt" import io.clika.modelverse.AutoModelForCausalLM import io.clika.modelverse.LoadOptions import io.clika.modelverse.Modelverse // On a worker thread of the app, never the main thread. fun reply(context: android.content.Context, license: String): String { // 0. The runtime, the model library and the bridge, once per process, // with the license credential placed before any model runs. Modelverse.load(context, license) // 1. and 2. Model files and the runnable model in one call: a hub // repository id downloads into the app's cache (a snapshot directory // or a .gguf file passes through untouched), and the registry loads // it on the CPU. The context length sizes the key-value cache for a phone. val model = AutoModelForCausalLM.fromPretrained( "Qwen/Qwen2.5-0.5B-Instruct", LoadOptions(contextLength = 4096), ) // 3. and 4. The generation: text in, text out, blocking for the reply; // chat(messages, config, listener) is the streaming form. return model.use { it.generate(prompt = "Once upon a time, in a port town by a cold sea,") } } ``` Build and run the C++ project: ```bash cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake" cmake --build build ./build/hello ``` ```text there lived a young woman named Akira. ``` The raw pipeline CONTINUES text; there is no chat template in the loop, which is the visible difference from `prompt` in part 2 (templated, answers a question). That is why this program feeds it a story opener and stops at the first sentence end: give a raw completion a question and it rambles on in the question's shape. A program that answers a question renders the conversation through the model's chat template first, as the executable does; the next section is that step. ## Answer a question A chat model's prompt format ships with the model as a template, and the tokenizer renders a `messages` array in the OpenAI API's shape into the exact prompt string; [Tokenize and chat templates](/clikart/how-to/tokenize-and-chat-templates#render-a-conversation-with-the-models-chat-template) is the contract. Never hand-build the role markers. `Tokenizer::apply_chat_template` renders the messages. In the program above, the rendered string takes the story opener's place as the `prompt` input, and the `.` stop goes, so the pipeline answers as the assistant. ```cpp title="chat_template.cpp" int main() { Tokenizer tok = Tokenizer::from_huggingface(model_dir()); Json messages = Json::array(); Json system = Json::object(); system["role"] = "system"; system["content"] = "You are a concise assistant."; messages.push_back(std::move(system)); Json user = Json::object(); user["role"] = "user"; user["content"] = "What does a tokenizer do?"; messages.push_back(std::move(user)); const std::string prompt = tok.apply_chat_template(messages); std::printf("=== rendered prompt ===\n%s\n=======================\n", prompt.c_str()); // The template writes the prompt's own bos/eos framing, so encode_chat adds // no special tokens on top of it. const std::vector ids = tok.encode_chat(messages, /*add_generation_prompt=*/true); std::printf("encode_chat produced %zu tokens\n", ids.size()); return 0; } ``` `chat` renders and generates in one call, and `render` is the rendered prompt alone; the conversation is yours to carry, so append the reply and the next call sees the whole history. ```python title="chat_messages.py" def main() -> None: model = mv.AutoModelForCausalLM.from_pretrained(resolve_model(), device="cpu", offline=True) messages = [ {"role": "system", "content": "You answer in one short sentence."}, {"role": "user", "content": "Why is the sky blue?"}, ] # A checkpoint without a chat template reads a plain prompt instead, so a # program that may meet either one says which it got. if not model.has_chat_template: print("BLOCKED: this checkpoint ships no chat template") sys.exit(3) # render is the prompt text the template produced, the same string # apply_chat_template returns on the transformers side. print(f"=== rendered prompt ===\n{model.render(messages)}\n=======================") reply = model.chat(messages, max_new_tokens=32, temperature=0.0) print(f"=== reply ===\n{reply}\n=============") # The conversation is the caller's: append the reply and the next call # carries the whole history. messages.append({"role": "assistant", "content": reply}) followup = model.chat(messages + [{"role": "user", "content": "In fewer words?"}], max_new_tokens=32, temperature=0.0) print(f"=== follow-up ===\n{followup}\n=================") assert reply.strip(), "the first reply is empty" assert followup.strip(), "the follow-up reply is empty" assert messages[-1]["role"] == "assistant", "the history did not gain the assistant turn" if __name__ == "__main__": main() ``` ## What the four steps are The program is the executable's anatomy laid bare, and each step is independently useful: - **The snapshot call** resolves any source (a Hugging Face Hub repo id, URL, or local directory) to a directory of model files. Point it at a directory you shipped with your application and no network code is ever built in. - **The registry load** is the catalog from part 1 as a function: identity match, factory, weights onto the device. Where the weights load is `LoadOptions::where`, a `Stream` or a `Device`: a Device (or nothing; the `device` field then decides) loads through a fresh stream of the model's own, a Stream loads through yours, behind whatever it already carries. The load synchronizes the weights, so their stream stops mattering once it returns; each inference call then runs on the stream its input tensors arrive on (or the calling thread's stream when they carry none), which is what lets two models fed from two streams overlap. The CLI's `--device` fills the same option. - **The pipeline build** assembles the serving pipeline. It is a [ClikaRT serving-runtime](/clikart) pipeline underneath, so sessions, continuous batching and the async model documented there apply unchanged. - **The request** is the generation loop. `serve` from part 3 is this loop behind HTTP; the library's `ServeApi` lets you mount the same engines, or your own, in-process. Failures follow the runtime's error model per language: C++ returns values directly and raises `ClikaRT::Error`, Python raises `clika_runtime.ClikaRTError` with the same code name, and Kotlin throws `ModelverseException` with its `codeName`; a served model reports a failure as the HTTP status the client raises. You have taken a model from the catalog to your own binary. [What to read next](../next-steps.md). --- # Get Modelverse Pick your platform's ClikaRT archive; Modelverse ships inside it, with the download and verify commands. Source: https://docs.clika.io/modelverse/getting-started/get-modelverse.md Modelverse ships inside the ClikaRT release archives: one archive per platform, carrying the ClikaRT runtime and Modelverse together (the `clikart-cli` executable, the library, the public headers, and the CMake packages for both products). Installing ClikaRT installs Modelverse; there is no separate Modelverse download and no second install step. The archive comes from your CLIKA Platform deployment: [Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) covers the web dialog, the platform CLI and the MCP tools that hand it out, with the license key the runtime needs beside it. The commands start from that archive and the checksum shown with it. The archive comes from the CLIKA platform's **Download ClikaRT** dialog, the platform CLI (`clika-cli runtime-sdk download`) or the MCP tools ([Download the ClikaRT SDK](/platform/how-to/download-the-clikart-sdk) walks each surface); the dialog shows the archive's SHA-256 beside its download button. The release the dialog offers is the deployment's own pin, which may differ from the release these pages are written for (`0.6.4`); a download's own `README.md` is the authority for the release it came from, its commands and file names included. Take the one file for your platform, check it against that digest, and extract: ```bash echo " ClikaRT_linux_x86_64-0.6.4.tar.xz" | sha256sum -c tar -xf ClikaRT_linux_x86_64-0.6.4.tar.xz export MODELVERSE_INSTALL_DIR="$PWD/ClikaRT_linux_x86_64-0.6.4" ``` `MODELVERSE_INSTALL_DIR` is the extracted directory, the same one the ClikaRT documentation calls `CLIKART_BUNDLE_DIR`; one directory carries both products. (macOS checks with `shasum -a 256 -c` over the same line. Windows, in PowerShell: `(Get-FileHash ).Hash -eq ""`, then `tar -xf `.) ## License credential Running a model needs the license credential CLIKA issued for your deployment, because Modelverse runs on the ClikaRT runtime and the runtime runs compute under a license. The credential is the `CLIKA1-...` text your project's license shows as its license key ([Runtime licenses](/platform/concepts/runtime-licenses) is where it comes from). Set it for the process, as the credential text or as the path of a file holding it: ```bash export CLIKA_RT_LICENSE=CLIKA1-... "$MODELVERSE_INSTALL_DIR"/bin/clikart-cli prompt "Hello" --device cpu ``` `bin/clikart-license-init ` stores it once under your user account instead, and the variable wins when both are present. Without a valid credential the model commands are refused with the code name `LICENSE_FAILED`, or `LICENSE_EXPIRED` for a license past its end date. From Python, `clika_runtime.modelverse` takes the credential from the process's environment the same way. [License the runtime](/clikart/how-to/license-the-runtime) is the whole contract for both products. Per platform, the archive to pick and the check that it runs on your machine: | OS | Architecture | Archive | Verify with | Notes | | --- | --- | --- | --- | --- | | Linux | x86_64 | `ClikaRT_linux_x86_64-.tar.xz` | `bin/clikart-cli devices` | CPU always; CUDA with the NVIDIA driver alone, Vulkan with a Vulkan 1.2 driver | | Linux | arm64 | `ClikaRT_linux_arm64-.tar.xz` | `bin/clikart-cli devices` | CPU · CUDA · Vulkan, as above | | Windows | x86_64 | `ClikaRT_windows_x86_64-.zip` | `bin\clikart-cli.exe devices` | CPU · Vulkan | | Windows | arm64 | `ClikaRT_windows_arm64-.zip` | `bin\clikart-cli.exe devices` | CPU · Vulkan | | macOS | Apple silicon | `ClikaRT_macos_arm64-.tar.xz` | `bin/clikart-cli devices` | CPU · Metal (ships with macOS) | | Android | arm64-v8a | `ClikaRT_android_arm64-.tar.xz` | push the directory to the device, run over `adb shell` | CPU · Vulkan | The Linux archives carry both CUDA images, and each image is self-contained: the CUDA runtime and cuBLASLt are inside it, so a CUDA machine needs its NVIDIA driver and nothing else, no CUDA toolkit install and no `LD_LIBRARY_PATH`. The runtime loads the image the driver serves (a driver of major version 580 or newer serves the CUDA 13 image, an older driver the CUDA 12 image). From Python, Modelverse ships inside the runtime's `clika-runtime` wheel as `clika_runtime.modelverse` ([Get ClikaRT](/clikart/getting-started/get-clikart#python-wheels) picks the wheel for your machine), so one install gives both: ```bash pip install ./clika_runtime-0.6.4-cp313-cp313-manylinux_2_28_x86_64.whl python -c "import clika_runtime.modelverse as mv; print(mv.__version__)" ``` The subpackage loads the runtime before its own extension, and a runtime and a model library from different releases refuse to import, naming both versions. The `clikart-cli` executable comes from the same wheel, which installs it as a console script, and from the archive's `bin/`. The C++ library ships in the archive only. Every archive carries `dist.json` at its top level, which describes the platform build itself (the target it serves, the backends inside, the library layout), and `VERSION`, the release it belongs to. The runtime and the model library in one archive are built together; a model library paired with a runtime of another release is refused with a readable error naming both versions, never undefined behavior. Accelerators, drivers and hardware are the runtime's: the [ClikaRT system requirements](/clikart/system-requirements) are the authority. Then continue with [Quick install](installation.md), which picks up at the extracted directory. --- # Quick install Extract the release archive and prove Modelverse works in two commands, no toolchain required. Source: https://docs.clika.io/modelverse/getting-started/installation.md The fast path from nothing to a running `clikart-cli` executable. It needs no compiler and no runtime dependencies; everything below runs the extracted directory in place. ## 0. Prerequisites A terminal and `tar` with xz support (Linux, macOS and Windows 10+ have it out of the box); the archive carries the ClikaRT runtime and Modelverse together, so there is no toolchain to install. Running a model needs a credential your platform issued for the project; [License the runtime](/clikart/how-to/license-the-runtime) says where `clikart-cli` reads it. (A C++17 toolchain becomes relevant only in [part 4 of the tutorial](first-model/04-use-it-from-code.mdx), when you embed Modelverse in your own program.) Running a model needs memory and disk that grow with the model's file size; the [ClikaRT system requirements](/clikart/system-requirements#memory) give the minimum memory per backend, measured at the platform's benchmark settings. ## 1. Extract the archive One archive per platform ([Get Modelverse](get-modelverse.md) has the download and checksum commands), one directory out of it, and that directory is the whole install: ```bash tar -xf ClikaRT_linux_x86_64-.tar.xz export MODELVERSE_INSTALL_DIR="$PWD/ClikaRT_linux_x86_64-" ls "$MODELVERSE_INSTALL_DIR" ``` The `ls` is the success check: `bin/`, `cmake/`, `include/`, `lib/`, `dist.json` and `VERSION` are there, and `README-Modelverse.md` sits beside the runtime's own `README.md`. The runtime's libraries and the Modelverse library share `lib/`, which is exactly where `bin/clikart-cli` looks for them. A new terminal loses the `export`; re-run it there. ## 2. Place the license credential Modelverse runs on the ClikaRT runtime, and the runtime runs compute under a license. Put the `CLIKA1-...` credential your project's license shows in the environment, or store it once under your user account: ```bash export CLIKA_RT_LICENSE=CLIKA1-... # this shell, and the programs it starts "$MODELVERSE_INSTALL_DIR"/bin/clikart-license-init CLIKA1-... # or once, for this user account ``` [License the runtime](/clikart/how-to/license-the-runtime) has the whole contract. ## 3. Ask it about this machine The fastest proof the install works: it loads the runtime, probes the compute backends, and prints what it found. Nothing gets installed and no network is touched. ```bash "$MODELVERSE_INSTALL_DIR"/bin/clikart-cli devices ``` ```text == compute devices on this machine == CPU available CUDA available Vulkan available device name driver runtime memory_gb sm_count compute_capability warp max_threads_per_block shared_mem_kb tensor_cores fp16 bf16 fp8_e4m3 status CPU:0 AMD - - 61.95013 6 - 1 0 0 no yes yes no ok CUDA:0 NVIDIA GeForce RTX 4080 13.2 13.3 15.569031 76 8.9 32 1024 48 yes yes yes yes ok Vulkan:0 NVIDIA GeForce RTX 4080 595.71.05 1.3.0 15.9921875 0 - 32 1024 0 yes yes yes no ok ``` Your table shows your hardware; `CPU available` alone is a pass. A backend listed as `compiled, not loaded` was built into this dist but found no usable device or driver on this machine. The `status` column reads `ok` for a device this build can run work on; a device that is present but unusable under this build names the reason there instead of failing at first use. ## 4. Print the catalog The second command proves the model registry is intact, still with no network: ```bash "$MODELVERSE_INSTALL_DIR"/bin/clikart-cli list ``` ```text registered model families (52 runnable; 10 hidden, --all shows all) chatterbox runnable * Input Modalities: text * Output Modalities: audio * Valid Combos: text -> audio * Commands: speak, serve, bench * Web UI: generic page * Vendor: Resemble AI * License page: https://github.com/resemble-ai/chatterbox/blob/master/LICENSE ----------------------------------------- ... ``` The list is long; `--full` adds each family's description. For convenience, put `bin/` on your `PATH`; the commands in the rest of the documentation assume `clikart-cli` resolves. ```bash export PATH="$MODELVERSE_INSTALL_DIR/bin:$PATH" ``` ## If something failed - `tar: xz: Cannot exec` or `xz: command not found`: install `xz-utils` (Debian/Ubuntu) or `xz`. - `No such file or directory` on the binary: the archive does not match this machine; `cat "$MODELVERSE_INSTALL_DIR/dist.json"` names the platform it was built for. - A shared-library error on start: the binary finds its libraries in `lib/` relative to itself, so run it inside the extracted directory; copying `bin/clikart-cli` out alone breaks that lookup. - A runtime-mismatch error on start: the model library is built against the runtime of its own release and refuses another, naming both versions; a directory hand-mixed from two releases is refused. Extract one archive, unmodified, and it clears. The install works. [Pick a model](first-model/01-pick-a-model.md). --- # What to read next Where the documentation goes after the tutorial: how-to guides, model requirements, the examples, and the ClikaRT docs underneath. Source: https://docs.clika.io/modelverse/getting-started/next-steps.md You installed the archive, fetched a model, served it, and ran it from C++. The rest of the documentation, in a useful reading order: - **[How-to guides](../how-to/index.md)**: problem-oriented recipes past the tutorial, from picking a GGUF quantization to running with no network at all. [Additional examples](../examples.md) sit inside it: the standalone example programs, from the catalog probe to an in-process server. - **[Your own models and nodes](../how-to/register-your-own-family.md)**: the how-to group for bringing your own architecture, from the full registration contract to [composing your own pipeline node](../how-to/add-a-pipeline-node.md). - **[Model requirements](../model-requirements.md)**: the memory and device figures per model variant, for sizing a deployment before you fetch anything. - **[ClikaRT](/clikart)**: the runtime underneath. Its tutorial explains tensors, devices and the async model the pipelines run on; its API reference covers every public name your C++ program touches; its [system requirements](/clikart/system-requirements) are the platform authority. --- # Modelverse at a glance What the Modelverse model library is and is not, and the five-minute mental model. Source: https://docs.clika.io/modelverse/getting-started/overview.md Modelverse is the CLIKA model library, built on the [ClikaRT](/clikart) runtime. It is a catalog of model families, the `clikart-cli` executable, which inspects, fetches and runs any model in it, and the library behind both (C++ first, with Python and Kotlin bindings over the same surface). It is not a training framework, not a conversion pipeline, and not a hosted service. Models arrive as the checkpoint files their authors published; Modelverse knows which family they belong to and what that family can do. ## The mental model Five ideas carry the whole product, in the order you meet them. 1. **The catalog is a registry of model families.** A family (Llama, Qwen, Whisper, CLIP, ...) declares its input and output modalities, the checkpoints it matches, and the commands it provides. The catalog is built into `clikart-cli`; listing it needs no network. ```bash clikart-cli list ``` 2. **A model is a source you name.** A Hugging Face Hub repo id, a pasted Hugging Face URL, or a local directory all name a model. Identity resolution reads config files only, so asking what something is never downloads weights. ```bash clikart-cli info Qwen/Qwen2.5-0.5B-Instruct clikart-cli info ./my-model-dir ``` A GGUF repo that ships several quantizations takes a selector, `/:Q6_K`; with several options and no selection, `clikart-cli` refuses and prints the option table instead of guessing. 3. **The `clikart-cli` executable has four commands that work without naming a model: `list`, `devices`, `info`, and `fetch`.** Every other command runs on a model you name: `clikart-cli `. Which commands a model supports depends on what kind of model it is: a text-generation model has `prompt`, `serve`, and `bench`; a speech model has `transcribe`. `clikart-cli --help` prints the commands a model supports: ```bash clikart-cli --help # that model's commands clikart-cli prompt "The capital of France is" ``` 4. **Serving is one of those commands.** `serve` hosts the model behind an OpenAI-compatible HTTP server, with streaming chat completions, a built-in web chat page, and per-modality routes (transcription, embeddings, depth and more). Existing OpenAI clients connect by changing their base URL. ```bash clikart-cli serve --port 8000 ``` 5. **Everything `clikart-cli` does, the library does.** The executable is a thin layer over the `clika_modelverse` library: a snapshot call fetches, the registry loads a runnable model, a serving pipeline generates. Your application makes the same two calls in its own language: {/* CERTIFICATION: the Python arm's two lines were run against the pinned release's clika_modelverse wheel and loaded a GenerativeModel. The C++ arm and the CLI are the other shipped surfaces. The Kotlin arm shows the Modelverse Kotlin binding's shape; its artifact ships with the releases that carry it. */} ```cpp mv::hub::SnapshotResult snapped = mv::hub::snapshot("Qwen/Qwen2.5-0.5B-Instruct", opts); mv::GenerativeModel model = mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load); ``` ```python snapped = mv.hub.snapshot("Qwen/Qwen2.5-0.5B-Instruct", cache_dir="models") model = mv.ModelRegistry.builtin().load_generative(snapped.local_dir) ``` ```kotlin val snapped = Hub.snapshot("Qwen/Qwen2.5-0.5B-Instruct", cacheDir = "models") val model = ModelRegistry.builtin().loadGenerative(snapped.localDir) ``` The [tutorial series](first-model/01-pick-a-model.md) turns these into working sessions, one idea per part. ## Platforms Modelverse ships inside the ClikaRT release archive, one archive per platform: Linux (x86_64, arm64), Windows (x86_64, arm64), macOS and Android, and from Python it ships inside the `clika-runtime` wheel as `clika_runtime.modelverse`. Extract the archive and run in place; the manifest inside pins the exact runtime this build linked against, so a mismatched pair refuses with a readable error. The [ClikaRT system requirements](/clikart/system-requirements) apply unchanged; Modelverse adds no requirements of its own. --- # A model from scratch Write an embedding model's checkpoint by hand (config, tokenizer, random weights), load it through the registry, and compare two texts, all offline. Source: https://docs.clika.io/modelverse/getting-started/own-model/a-model-from-scratch.md This series is for adding your own model, as opposed to running the catalog's. It walks the same ground the [first tutorial](../first-model/01-pick-a-model.md) covered from the consumer side, now from the author's side, in three parts with a running result each. The worked model is an embedding model on purpose: text goes in, one vector comes out, and there is no machinery beyond that idea. This part demystifies the checkpoint itself: a model is a directory of files, and you can write one by hand. The program below synthesizes a tiny embedding checkpoint (an 8-word vocabulary, hidden size 32, a gemma3-shaped text encoder), loads it through the registry, and compares three texts. Nothing downloads and the weights are random; random weights still embed, they are only bad at it, and that is enough to see every file a model needs and where each one enters. Same two-file CMake project as [part 4 of the first tutorial](../first-model/04-use-it-from-code.mdx); only `main.cpp` changes. ## The files an embedding model needs - `config.json` names the architecture (`model_type`, `architectures`) and its geometry (layers, heads, hidden size). Identity resolution reads exactly this. - `tokenizer.json` turns text into token ids; the toy one below is a word-level vocabulary of eight entries. - The weights (`model.safetensors` here) are tensors under the names the architecture expects. - An embedding checkpoint additionally ships its module chain, the sentence-transformers layout: `modules.json` lists the stages (transformer, pooling, dense heads, normalize), and each configurable stage carries its own small directory. ## The program ```cpp title="main.cpp" #include #include #include #include #include #include #include #include "clika_modelverse/modules/pooling.h" #include "clika_modelverse/registry/model_registry.h" namespace fs = std::filesystem; namespace mv = clika_modelverse; using ClikaRT::DataType; using ClikaRT::Tensor; // The whole identity: model_type and architectures are what the registry // matches; the rest is geometry the loader shapes the encoder from. constexpr const char* kConfigJson = R"({ "model_type": "gemma3_text", "architectures": ["Gemma3TextModel"], "hidden_size": 32, "num_hidden_layers": 4, "num_attention_heads": 4, "num_key_value_heads": 2, "head_dim": 8, "vocab_size": 8, "intermediate_size": 64, "max_position_embeddings": 128, "rms_norm_eps": 1e-6, "rope_theta": 1000000.0, "rope_local_base_freq": 10000.0, "sliding_window": 16, "sliding_window_pattern": 2, "query_pre_attn_scalar": 8, "use_bidirectional_attention": true })"; // A word-level tokenizer: eight words, one id each. constexpr const char* kTokenizerJson = R"({ "version": "1.0", "pre_tokenizer": {"type": "WhitespaceSplit"}, "model": {"type": "WordLevel", "vocab": {"a": 0, "b": 1, "c": 2, "d": 3, "e": 4, "f": 5, "g": 6, "[eos]": 7}, "unk_token": "a"}, "decoder": null, "added_tokens": [{"id": 7, "content": "[eos]", "special": true, "normalized": false}] })"; // The module chain: transformer -> mean pooling -> two dense heads -> normalize. constexpr const char* kModulesJson = R"([ {"idx": 0, "name": "0", "path": "", "type": "sentence_transformers.models.Transformer"}, {"idx": 1, "name": "1", "path": "1_Pooling", "type": "sentence_transformers.models.Pooling"}, {"idx": 2, "name": "2", "path": "2_Dense", "type": "sentence_transformers.models.Dense"}, {"idx": 3, "name": "3", "path": "3_Dense", "type": "sentence_transformers.models.Dense"}, {"idx": 4, "name": "4", "path": "4_Normalize", "type": "sentence_transformers.models.Normalize"} ])"; constexpr const char* kPoolingJson = R"({ "word_embedding_dimension": 32, "pooling_mode_cls_token": false, "pooling_mode_mean_tokens": true, "pooling_mode_max_tokens": false, "pooling_mode_mean_sqrt_len_tokens": false, "pooling_mode_weightedmean_tokens": false, "pooling_mode_lasttoken": false })"; Tensor randn(std::mt19937& rng, std::vector shape) { std::int64_t n = 1; for (const std::int64_t d : shape) n *= d; std::normal_distribution dist(0.0f, 0.2f); std::vector host(static_cast(n)); for (float& v : host) v = dist(rng); return Tensor::from_data(host.data(), shape, DataType::Float32); } // One dense head: its own directory with a config and one weight. void write_dense(std::mt19937& rng, const fs::path& dir, std::int64_t in, std::int64_t out) { fs::create_directories(dir); std::ofstream(dir / "config.json") << R"({"in_features": )" << in << R"(, "out_features": )" << out << R"(, "bias": false, "activation_function": "torch.nn.modules.linear.Identity"})"; ClikaRT::NamedTensors sd; sd.set("linear.weight", randn(rng, {out, in})); ClikaRT::io::save_safetensors(sd, (dir / "model.safetensors").string()); } // Random weights under the names the gemma3-shaped encoder expects (bare // keys, no prefix); the fixed seed keeps every run identical. fs::path write_snapshot() { const fs::path dir = fs::temp_directory_path() / "my_first_model"; fs::create_directories(dir); std::ofstream(dir / "config.json") << kConfigJson; std::ofstream(dir / "tokenizer.json") << kTokenizerJson; std::ofstream(dir / "modules.json") << kModulesJson; fs::create_directories(dir / "1_Pooling"); std::ofstream(dir / "1_Pooling" / "config.json") << kPoolingJson; constexpr std::int64_t kVocab = 8, kHidden = 32, kLayers = 4, kHeads = 4; constexpr std::int64_t kKvHeads = 2, kHeadDim = 8, kFfn = 64; constexpr std::int64_t kDenseMid = 48, kDim = 16; std::mt19937 rng(20260831); ClikaRT::NamedTensors sd; const auto put = [&](const std::string& name, std::vector shape) { sd.set(name, randn(rng, std::move(shape))); }; put("embed_tokens.weight", {kVocab, kHidden}); put("norm.weight", {kHidden}); for (std::int64_t l = 0; l < kLayers; ++l) { const std::string p = "layers." + std::to_string(l) + "."; put(p + "self_attn.q_proj.weight", {kHeads * kHeadDim, kHidden}); put(p + "self_attn.k_proj.weight", {kKvHeads * kHeadDim, kHidden}); put(p + "self_attn.v_proj.weight", {kKvHeads * kHeadDim, kHidden}); put(p + "self_attn.o_proj.weight", {kHidden, kHeads * kHeadDim}); put(p + "self_attn.q_norm.weight", {kHeadDim}); put(p + "self_attn.k_norm.weight", {kHeadDim}); put(p + "mlp.gate_proj.weight", {kFfn, kHidden}); put(p + "mlp.up_proj.weight", {kFfn, kHidden}); put(p + "mlp.down_proj.weight", {kHidden, kFfn}); put(p + "input_layernorm.weight", {kHidden}); put(p + "post_attention_layernorm.weight", {kHidden}); put(p + "pre_feedforward_layernorm.weight", {kHidden}); put(p + "post_feedforward_layernorm.weight", {kHidden}); } ClikaRT::io::save_safetensors(sd, (dir / "model.safetensors").string()); write_dense(rng, dir / "2_Dense", kHidden, kDenseMid); write_dense(rng, dir / "3_Dense", kDenseMid, kDim); return dir; } int main() { // 1. A checkpoint is a directory; this one is yours, written just now. const fs::path dir = write_snapshot(); std::printf("wrote %s\n", dir.string().c_str()); // 2. The registry reads config.json, matches architecture // "Gemma3TextModel" to the gemma-embedding family, and returns the // runnable model, exactly as for a fetched checkpoint. mv::EmbeddingModel model = mv::ModelRegistry::builtin().load_embedding(dir.string(), {}); // 3. Embed three texts: one [3, 16] Float32 tensor, one row per text // (the module chain pools, projects to 16 dims, and normalizes). const std::string texts[] = {"a b c", "a b d", "f g"}; const Tensor rows = model.embed(texts); std::printf("embeddings: %s\n", rows.to_string().c_str()); // 4. Compare them: the cosine matrix is unit rows times their own // transpose. The explicit normalize keeps the demo self-contained // (the chain's Normalize stage already produced unit rows). const Tensor unit = mv::modules::l2_normalize_rows(rows); const Tensor sim = ClikaRT::ops::matmul(unit, unit, {}, std::nullopt, false, /*transpose_b=*/true); const std::vector s = ClikaRT::ops::reshape(sim, {9}).item_as_vec(); std::printf("close pair (a b c ~ a b d): %.3f\n", s[1]); std::printf("far pair (a b c ~ f g): %.3f\n", s[2]); return 0; } ``` ```bash cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake" cmake --build build ./build/hello ``` ```text [2026-09-22 00:30:20.206] [modelverse] [info] hub: '/tmp/my_first_model' already on disk; loading from /tmp/my_first_model [2026-09-22 00:30:20.208] [gemma-embedding] [info] embedding encoder on CPU:-1 (dim 16) wrote /tmp/my_first_model embeddings: Tensor(shape=[3, 16], dtype=Float32, device=CPU, numel=48, data=[0.03671, -0.04182, -0.2858, 0.1808, 0.1262, 0.399, ...]) close pair (a b c ~ a b d): 0.400 far pair (a b c ~ f g): -0.324 ``` Random weights, so the numbers mean little; what matters is the shape of what happened. A directory you wrote from scratch went through the same registry, loader and embedding surface as a fetched checkpoint would, and three texts became three vectors you can compare, because a checkpoint is nothing more than these files. ## What the registry did with it `config.json`'s `model_type` and `architectures` are the identity keys; the architecture `Gemma3TextModel` matched the built-in `gemma-embedding` family, and that family's loader shaped the encoder from the geometry fields and bound your `model.safetensors` names to it. The tokenizer file turned each text into ids before the encoder saw them, and `modules.json` told the loader what follows the encoder: mean pooling, the two dense projections, and the final normalize. Delete the directory and nothing else remembers it. Your own architecture will not say `"architectures": ["Gemma3TextModel"]`, and then the match fails; that refusal, and fixing it by registering a family of your own, is [part 2](02-join-the-catalog.md). --- # Your architecture joins the catalog Rename the toy model's architecture so nothing matches it, then register a family of your own and watch the same directory resolve. Source: https://docs.clika.io/modelverse/getting-started/own-model/join-the-catalog.md Part 1 rode the built-in `gemma-embedding` family. Your real architecture has its own name, and the catalog does not know it; this part makes the failure visible and then fixes it the way every built-in family fixes it, by self-registration. ## Break the match first In part 1's `kConfigJson`, rename the identity keys the way a Hugging Face checkpoint of your own architecture would name them: ```json "model_type": "my_model", "architectures": ["MyModelModel"], ``` Rebuild and run, and the registry refuses precisely; the raised `ClikaRT::Error` carries: ```text wrote /tmp/my_first_model no model template registered for model_type='my_model' / architecture='MyModelModel' ``` Nothing else changed. The files are fine; the catalog has no entry whose identity keys match them. ## Register the family A family is one translation unit that pushes a `ModelRegistration` into the shared registry at static-initialization time, and a TU compiled into your own program registers exactly like a built-in one. Add this file to the project and add it to `add_executable`: ```cpp title="my_model_family.cpp" #include "clika_modelverse/models/registration.h" namespace { using namespace clika_modelverse; // The identity template: what a resolved checkpoint of this family IS. // The registry fills architecture, model_type and variant from config.json. class MyModel final : public Model { public: std::string_view family() const noexcept override { return "my-model"; } }; constexpr std::string_view kModelTypes[] = {"my_model"}; constexpr std::string_view kArchitectures[] = {"MyModelModel"}; // Constructing the registrar is the whole hookup; there is no list to edit. const ModelFamilyRegistrar kRegistrar{ [] { ModelRegistration registration{}; registration.metadata.family = "my-model"; registration.metadata.vendor = "my-org"; registration.metadata.description = "the tutorial's own model family"; registration.metadata.input_combos = kTextOnlyCombos; registration.metadata.output_combos = kTextOnlyCombos; registration.hf_model_types = kModelTypes; registration.hf_architectures = kArchitectures; registration.factory = make_model(); return registration; }(), }; } // namespace ``` To watch it resolve, have `main` ask for identity instead of embeddings for a moment: ```cpp const std::unique_ptr matched = mv::ModelRegistry::builtin().open(dir.string()); std::printf("%s resolved: family=%s model_type=%s\n", dir.string().c_str(), std::string(matched->family()).c_str(), std::string(matched->model_type()).c_str()); ``` ```text wrote /tmp/my_first_model /tmp/my_first_model resolved: family=my-model model_type=my_model ``` ## What you registered, and what you did not The registration so far is identity-only: `metadata` (the family's name, its publisher, its declared modalities), the match keys, and the `factory` that builds the identity object. That is enough for the directory to resolve, for the family to appear in a catalog listing, and for every runnable command to refuse with a precise message naming what the family provides (nothing yet). An incomplete declaration (no family name, no modalities, no identity factory) is refused at process start, not papered over downstream. Running is a separate concern on the same registration, and that split is deliberate: identity must stay cheap (no weights) and total (every checkpoint of yours resolves), while running is opt-in per capability. Wiring it is [part 3](03-make-it-embed.md); the field-by-field reference for everything a registration can carry is [Register your own model family](../../how-to/register-your-own-family.md). --- # Make it embed Give the my-model family the smallest real embedding backend, opt into the similar and serve commands, and compare texts through your own family. Source: https://docs.clika.io/modelverse/getting-started/own-model/make-it-embed.md Part 2's family resolves but cannot run. For an embedding family the whole running contract is one small interface, `EmbeddingBackend` (`clika_modelverse/models/base/embedding_model.h`): token ids in, embedding rows out, plus the width, the pooling declaration and the device. This part implements the smallest real backend, wires the factory, and opts into the standard commands. ## The smallest real backend The toy backend embeds each text as the mean of its tokens' embedding-table rows, L2-normalized. That is a real embedding model (the bag-of-words baseline retrieval systems start from), and it needs exactly one weight: ```cpp title="my_model_backend.cpp" #include "clika_modelverse/models/base/embedding_model.h" namespace { using namespace clika_modelverse; class MyModelBackend final : public EmbeddingBackend { public: MyModelBackend(ClikaRT::Tensor table, ClikaRT::Device device) : table_(std::move(table)), device_(device) {} ClikaRT::Tensor embed_ids( const std::vector>& sequences) const override { std::vector rows; rows.reserve(sequences.size()); for (const std::vector& ids : sequences) { const ClikaRT::Tensor picked = ClikaRT::ops::index_select(table_, 0, ids); rows.push_back(ClikaRT::ops::mean(picked, /*dim=*/0)); } return ClikaRT::ops::l2_normalize(ClikaRT::ops::stack(rows), /*dim=*/-1); } std::int64_t embedding_dim() const override { return table_.shape()[1]; } Pooling pooling() const override { return Pooling::Mean; } bool normalizes() const override { return true; } ClikaRT::Device device() const override { return device_; } private: ClikaRT::Tensor table_; ClikaRT::Device device_; }; } // namespace ``` A real encoder runs its layers between the lookup and the pooling; where those layers go, and how a checkpoint's names bind onto them, is the [full contract's](../../how-to/register-your-own-family.md) territory. The interface does not change with the depth: however sophisticated the encoder, it enters the family as this same `EmbeddingBackend`. One optional declaration to know: `max_concurrent_sessions()` says how many requests one handle's forward may run at once, and its default of 1 is the safe answer for a backend that keeps per-request state; a forward that is a pure function of its inputs and weights can declare itself unbounded. ## The registration delta The factory reads the snapshot, builds the backend from the one weight it uses, and wraps it with the snapshot's tokenizer; the registration gains the factory and the command surface: ```cpp ClikaRT::Result build_my_model(const std::string& dir, const Model& identity, const LoadOptions& options) { auto weights = ClikaRT::io::load_safetensors(dir + "/model.safetensors"); auto table = weights.get("embeddings.word_embeddings.weight").to(options.device); auto backend = std::make_unique(std::move(table), options.device); auto tokenizer = ClikaRT::tokenizer::from_huggingface(dir); return EmbeddingModel(std::move(backend), std::move(tokenizer), std::string(identity.family()), std::string(identity.model_type())); } // The command surface, composed from the same public builders every // embedding family uses: `similar` and `serve`, two lines. FamilyApp build_my_model_cli(const FamilyCliContext& context) { FamilyApp fam{batteries::cli::make_family_app( context, "commands the 'my-model' family provides"), {}}; batteries::cli::similar_command(fam, context); batteries::cli::serve_command(fam, context, batteries::serving::generic_serve); return fam; } ``` ```cpp registration.cli = CliSurface{build_my_model_cli, my_model_cli_verbs()}; registration.factory = make_model(); registration.build_embedding = build_my_model; ``` ## Run it Part 1's `main` works again unchanged (it never named a family, only a directory), now through your own: ```text wrote /tmp/my_first_model embeddings: Tensor(shape=[3, 32], dtype=Float32, device=CPU:0, numel=96, ...) close pair (a b c ~ a b d): 0.667 far pair (a b c ~ f g): -0.041 ``` The close pair scores closer than in part 1, and honestly so: mean-pooled bags of shared words ARE similar, which is exactly what this backend measures. And because the commands came from the shared battery, your family now serves them like any catalog family: `similar` compares texts from the command line, and `serve` answers `POST /v1/embeddings` and `POST /v1/similarity` ([the route table](../../how-to/serve-openai-compatible.md)). ## Where the full contract lives Everything real that the toy skipped is registration fields on the same struct, documented in [Register your own model family](../../how-to/register-your-own-family.md) and field-by-field in `clika_modelverse/models/registration.h`: encoders with layers and their weight binding, the generative, speech and reranking factories with their GGUF twins, `probe_snapshot` for repos without a `config.json`, and custom verb surfaces. For custom logic AROUND a catalog model rather than a model of your own, [Add your own node to a model pipeline](../../how-to/add-a-pipeline-node.md) is the shorter road. --- # How-to guides Problem-oriented recipes. Each guide answers one "how do I X" with a worked session and its real output. Source: https://docs.clika.io/modelverse/how-to.md Practical guides covering common tasks. Each guide answers one concrete "how do I X" with a worked session: real commands or real code, with the output they produce. Read the [tutorial](../getting-started/first-model/01-pick-a-model.md) first; the guides assume its ground (sources, fetching, the family-owned commands) and go deeper on one problem at a time, in any order. ## Models and weights - [Run a specific GGUF quantization](run-a-gguf-quantization.md): pick one weight option of a repo that ships many, by tag, glob or exact file. - [Run fully offline](run-fully-offline.md): fetch on a connected machine, move the directory, and make any network touch an error. ## Running and serving - [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md): the full route table, streaming, operational flags, and pointing existing clients at it. - [Transcribe audio](transcribe-audio.md): speech-to-text from the command line and over HTTP. - [Translate text](translate-text.md): neural machine translation from the command line and over HTTP, with the model's own language codes. - [Speak text](speak-text.md): text-to-speech with voice references and the synthesis knobs. - [Estimate depth, segment and detect](depth-segment-detect.md): the three image verbs, one artifact per input, open-vocabulary detection included. - [Benchmark a model on this machine](bench-a-model.md): the `bench` flow, its sweep axes and how to read the structured report. - [Embed, compare and rerank](embed-and-rerank.md): dense vectors, similarity matrices and cross-encoder reranking, for texts and images. ## Scripting and integration - [Script clikart-cli](script-the-cli.md): the exit-code contract, the stdout/stderr split, option templates and shell completion. - [Add Modelverse to an existing CMake project](existing-cmake-project.md): `find_package` against the installed directory, staging the runtime beside your binary, and picking the dist on cross builds. ## Your own models and nodes The custom-model track at its full depth; the gentle guided walk is the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md). - [Register your own model family](register-your-own-family.md): what a snapshot directory must contain, identity matching, and the registration that makes your architecture runnable. - [Add your own node to a model pipeline](add-a-pipeline-node.md): compose custom pre- or post-processing with the zoo's serving nodes, on the same request surface. - [A conversational AI as one pipeline](conversational-ai-pipeline.md): speech-to-text, a chat model and text-to-speech composed into a single pipeline, wav bytes in and the spoken answer out. ## Complete programs - [Additional examples](../examples.md): the standalone example programs, one per subsystem. Sizing questions (which variant fits which device) live in [Model requirements](../model-requirements.md). For every public C++ name, the ClikaRT API reference on [/clikart](/clikart). --- # Add your own node to a model pipeline Compose custom pre- or post-processing with the zoo's serving nodes in one pipeline, on the same request surface. Source: https://docs.clika.io/modelverse/how-to/add-a-pipeline-node.md Your application needs logic the model does not have: redaction, templating, routing, scoring, any transform of what goes in or comes out. In Modelverse that logic is not a wrapper around the pipeline; it is a node inside it. The zoo's serving nodes and your code meet on one contract, `ClikaRT::runtime::Model`, and `ClikaRT::runtime::Pipeline::create` composes any mix of them into one pipeline with one request/response surface. A node implements three members: `schema()` declares its input and output tensors by name, `phases()` declares its execution phases, and `Phase_RunOnce` does the work. The zoo's tokenizer, generative decoder and detokenizer implement exactly the same three, which is why yours can stand beside them without an adapter layer. ## The program Same two-file project as [part 4 of the tutorial](../getting-started/first-model/04-use-it-from-code.mdx); only `main.cpp` changes. The custom node here uppercases the reply, standing in for any post-processing you own: ```cpp title="main.cpp" #include #include #include #include #include #include #include #include "clika_modelverse/generation/generation_config.h" #include "clika_modelverse/hub/hub.h" #include "clika_modelverse/registry/model_registry.h" #include "clika_modelverse/runtime/decoder_node.h" #include "clika_modelverse/runtime/pipeline_io.h" #include "clika_modelverse/runtime/text_nodes.h" namespace mv = clika_modelverse; namespace mvr = clika_modelverse::runtime; namespace crt = ClikaRT::runtime; using ClikaRT::DataType; // The detokenizer's reply routes here under this name instead of going // straight out. constexpr const char* kDraftText = "draft_text"; // Everything a user writes: declare I/O, declare a phase, do the work. class ShoutNode final : public crt::Model { public: ShoutNode() { schema_.inputs.push_back(ClikaRT::spec::TensorSpec{ kDraftText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim}, /*optional=*/false}); schema_.outputs.push_back(ClikaRT::spec::TensorSpec{ mvr::kText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim}, /*optional=*/false}); phases_.push_back(crt::PhaseSpec{"shout", crt::PhaseKind::RunOnce}); } const crt::ModelSchema& schema() const override { return schema_; } ClikaRT::Span phases() const override { return phases_; } void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override { std::string text = mvr::byte_tensor_to_string(ctx.inputs->get(kDraftText)); std::transform(text.begin(), text.end(), text.begin(), [](unsigned char c) { return static_cast(std::toupper(c)); }); ctx.outputs->set(mvr::kText, mvr::string_to_byte_tensor(text)); } private: crt::ModelSchema schema_; std::vector phases_; }; int main() { // The zoo half, exactly as in the tutorial: snapshot, registry, model. mv::hub::SnapshotOptions snap; snap.cache_dir = "models"; const mv::hub::SnapshotResult snapped = mv::hub::snapshot( "Qwen/Qwen2.5-0.5B-Instruct", snap); mv::LoadOptions load; load.max_seq = 4096; // the CLI's default context cap, as in part 4 mv::GenerativeModel model = mv::ModelRegistry::builtin().load_generative(snapped.local_dir, load); mv::generation::GenerationConfig gen = model.defaults(); gen.max_new_tokens = 24; gen.stop = {"."}; // raw completion: one sentence is the demo // The zoo's three serving nodes, constructed directly. mvr::TokenizerNode tok(model.tokenizer(), /*add_special_tokens=*/false); mvr::GenerativeDecoderModel dec(model.provider(), gen, /*max_active=*/1, mv::KVCacheMode::Continuous, &model.tokenizer(), /*prefill_chunk_tokens=*/0); mvr::DetokenizerNode detok(model.tokenizer()); ShoutNode shout; // yours // One pipeline over all four. Edges wire by name; the one rename (the // detokenizer's `text` becomes `draft_text`) routes the reply through // the custom node instead of straight out. std::vector steps(4); steps[0].name = "tokenize"; steps[0].model = &tok; steps[1].name = "decode"; steps[1].model = &dec; steps[2].name = "detokenize"; steps[2].model = &detok; steps[2].output_map.emplace_back(mvr::kText, kDraftText); steps[3].name = "shout"; steps[3].model = &shout; // The pipeline's public request surface: what callers set and read. crt::ModelSchema ext; ext.inputs.push_back(ClikaRT::spec::TensorSpec{ mvr::kPrompt, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim}, /*optional=*/false}); ext.outputs.push_back(ClikaRT::spec::TensorSpec{ mvr::kText, DataType::UInt8, {ClikaRT::spec::TensorSpec::kDynamicDim}, /*optional=*/false}); crt::Pipeline pipeline = crt::Pipeline::create(std::move(steps), std::move(ext)); // The same Request/Response surface every pipeline speaks. crt::Request req; req.inputs.set(mvr::kPrompt, mvr::string_to_byte_tensor( "Once upon a time, in a port town by a cold sea,")); const crt::SessionId sid = pipeline.enqueue(std::move(req)); const crt::Response resp = pipeline.await(sid); std::printf("%s\n", mvr::byte_tensor_to_string(resp.outputs.get(mvr::kText)).c_str()); pipeline.shutdown(); return 0; } ``` ```text THERE LIVED A YOUNG GIRL NAMED KAITO. ``` One build flag matters here: the custom node derives from the serving runtime's `Model`, so its translation unit compiles with `-fno-rtti`, matching how the library builds; without it the link fails on `typeinfo for ClikaRT::runtime::Model`. The [ClikaRT integration guide](/clikart/how-to/existing-cmake-project) covers the flag and which consumers need it. ## What the composition rests on - **Edges wire by tensor name.** Each step's outputs feed the next step's same-named inputs; `output_map` renames one edge where names must differ. The single rename above is the whole routing change: without it the detokenizer's `text` would leave the pipeline directly. - **The external schema is the caller's contract.** Only what it declares is settable and readable from outside; everything between the nodes stays internal. Requests and responses are exactly the ones part 4 used, so a caller cannot tell a customized pipeline from a stock one. - **Placement in the chain is yours.** A node before the tokenizer transforms the prompt (templating, redaction); a node after the detokenizer transforms the reply, as here; a fully custom model runs beside a zoo model in the same pipeline. The `02_custom_node` program in [Additional examples](../examples.md) is this guide's offline twin: it synthesizes a toy checkpoint at run time, so the composition runs in moments with no download and no arguments. When HTTP is the goal, mount an engine on the built-in server instead ([Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md)); a pipeline node changes what a model computes, an engine changes what the server serves. And when the model itself is yours rather than the catalog's, that is the other half of the custom track: the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md), with [Register your own model family](register-your-own-family.md) as its full contract. --- # Benchmark a model on this machine The bench flow: serving-shaped sweep cells, the structured report, and how to read its columns before committing a deployment. Source: https://docs.clika.io/modelverse/how-to/bench-a-model.md Whether a model meets your latency and throughput targets is a property of this machine, this quantization and this serving shape, and `bench` measures exactly that. Every runnable text family provides it, the flow drives the same serving pipeline `prompt` and `serve` run on, and the result is a structured table, never ad-hoc prints. ## The default sweep ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct bench --device cuda ``` The default sweep runs three serving-shaped cells (balanced `128:128`, prefill-heavy `2048:128`, decode-heavy `128:1024`) over the default concurrency ladder (1, 4, 8), each cell strictly one at a time so cells never contend with each other. One CSV row per cell and concurrency, the last column naming what the row measured: ```text isl,osl,concurrency,first_call_ms,ttft_ms,prefill_tok_s,itl_p50_ms,itl_p99_ms,prefill_chunk_ms_p50,prefill_chunk_ms_max,decode_tok_s,speedup_vs_c1,batching_efficiency,admitted_tok,retires,peak_active_bytes,num_allocs,num_inflight_park_waits,status,reason,notes 128,128,1,685.7974,14.375232,8904.204,4.709812,11.978596,0,0,194.13823,1,1,112,3,3422640128,59180,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 128,128,4,937.06464,25.81384,4969.912,6.5161915,15.847513,0,0,552.3547,2.845162,0.7112905,448,12,3746307584,60982,6,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 128,128,8,745.5794,36.65184,3492.3213,4.9885244,10.634405,0,0,1434.3369,7.388225,0.92352813,896,24,4083747328,93121,12,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 2048,128,1,820.59033,67.9156,30155.074,5.575619,8.472811,0,0,159.87682,1,1,2032,3,4083747328,65325,20,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 2048,128,4,1258.1548,245.00293,8360.729,7.7572827,15.318566,0,0,414.64108,2.5935037,0.6483759,8128,12,4192593408,67393,41,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 2048,128,8,1308.3411,464.48145,4409.434,6.383371,11.002449,0,0,787.79083,4.927487,0.61593586,16256,24,4744088576,99717,63,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 128,1024,1,4980.8936,10.374589,12337.838,4.823417,8.758105,0,0,190.86696,1,1,112,3,4744088576,498596,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 128,1024,4,7377.311,26.746096,4836.729,6.6391444,13.775986,0,0,592.14624,3.102403,0.77560073,448,12,4744088576,508606,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) 128,1024,8,5360.8525,39.64144,3232.0154,5.407993,9.28407,0,0,1437.6843,7.5323896,0.9415487,896,24,4744088576,767141,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest) ``` The numbers are one machine's run (one workstation GPU, this model at its checkpoint dtype, bf16, cold prompts, the default); yours are the point. Stdout carries the CSV, ready for a spreadsheet or a script; diagnostics ride stderr. A failed cell reports `error` in `status` without invalidating the rows around it, and the `reason` column carries the model's own refusal or the cell's first failure in words (empty on an `ok` row). With `--output`, the report streams to disk as rows complete, so a run killed at its deadline keeps every finished cell. ## Shape the sweep to your workload - `--cells 512:256,4096:64`: your own `isl:osl` list, replacing the default three. `--isl`/`--osl` set a single cell directly. - `--concurrency 1,4,16`: the parallel-session ladder; `speedup_vs_c1` and `batching_efficiency` in the report tell you what the added concurrency actually bought. - `--shared-prefix 256`: how many prompt tokens every session shares (a system prompt, a RAG preamble); with it the `admitted_tok` and `retires` columns show the paged prefix cache engaging. Every LLM cell runs cold prompts by default (a fresh random prompt per session every iteration, so a paged prefix cache serves no repeat; `--unique-prompts` names that default and stays accepted). `--repeat-prompts` is the warm protocol: the warm-up and every iteration reuse one prompt set, so `prefill_tok_s` reads cache admits, and the row's `notes` say so. - `--warmup` and `--iters`: untimed full passes of each cell before it measures, then timed passes averaged. The default is two warm-up passes and one timed iteration; `--warmup 0` times a cold cell. A cell still cold after the default passes is a defect worth reporting, not a reason to raise the count. - The load knobs apply unchanged: `--kvcache paged|continuous` (each verb has its own default and `bench` runs paged), `--prefill-chunk`, `--step-token-budget`, a GGUF weight selector on the source, `--max-seq`. - `--memory-stages FILE` appends the runtime's memory-pool counters to a JSON file at each stage of the model's life (after the weights bind, after the one-forward warm-up, after the first generation), one row per device, with the checkpoint's size. It is the footprint diagnostic behind the figures on [Model requirements](../model-requirements.md). - `--output PREFIX` writes the report to files beside the terminal view; `--profile` adds the per-op profile on stderr (the summary opens with `Profiling Report Summary:`, its `top_ops:` block ranks op groups by self time, and the per-op table carries an `AvgSelf(us)` column, the per-op average that excludes nested spans), `--profile-pipeline` the per-node one. ## Reading the columns `ttft_ms` is the enqueue-to-first-token wall, the number an interactive user feels; `first_call_ms` is the cell's very first call, which pays the kernel loads a fresh process owes and is kept out of `ttft_ms` for that reason. `prefill_tok_s` is prompt ingest; `decode_tok_s` is the aggregate generation rate across sessions. The inter-token percentiles `itl_p50_ms`/`itl_p99_ms` sample decode-only gaps; on a chunked-prefill run the chunk boundary walls report separately (`prefill_chunk_ms_p50`/`_max`), so a decode tail is never polluted by an ingest wall. When the p99 sits far above the p50, the deployment story is a batching or cache-pressure story, and the concurrency ladder narrows down which. Three columns carry the memory story: `peak_active_bytes` is the runtime pool's high-water mark for the cell, `num_allocs` the allocation count behind it, and `num_inflight_park_waits` the number of times a request waited for a slot rather than proceeding, a direct read on admission pressure. The `-v` diagnostics carry the load and peak-memory figures the [Model requirements](../model-requirements.md) page's estimates are checked against. ## Audio and media axes A speech-to-text family's `bench` sweeps audio length instead of token shapes (`--audio-seconds 5,30,120`, with the same `--concurrency` ladder), reporting ingest as encoder frames over the end-to-end transcribe wall. Multimodal families add `mm_bench`, the media sweep over the serving path: `--images` per request, `--image-resolution`, `--video-frames`, with `effective_isl` reporting what the prefill actually ingested once media expansions counted, so a media-heavy row's ingest rate stays comparable to a text row's. Benchmark the quantization you intend to ship ([Run a specific GGUF quantization](run-a-gguf-quantization.md)); a `Q4_K_M` and a `Q8_0` of the same model can sit on different sides of a latency target. The serving flags that shaped a good bench row carry directly onto `serve` ([Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md)). --- # A conversational AI as one pipeline Compose speech-to-text, a chat model and text-to-speech into a single runtime::Pipeline: wav bytes in, the spoken answer out. Source: https://docs.clika.io/modelverse/how-to/conversational-ai-pipeline.md [Add your own node to a model pipeline](add-a-pipeline-node.md) put one custom node beside the zoo's text trio. This guide composes a complete application the same way: a conversational AI as one `ClikaRT::runtime::Pipeline`, a spoken question in as wav bytes, the spoken answer out. Eight steps, three of them the zoo's serving nodes, five of them yours: | step | node | in -> out | | --- | --- | --- | | listen | yours | wav bytes -> whisper features | | transcribe | yours | features -> question text | | template | yours | question -> the chat-templated prompt | | tokenize | zoo `TokenizerNode` | prompt -> input_ids | | decode | zoo `GenerativeDecoderModel` | input_ids -> tokens, continuous-batched | | detokenize | zoo `DetokenizerNode` | tokens -> answer text | | speak | yours | answer text -> waveform and its sample rate | | post | yours | waveform -> the peak-normalized waveform | The complete program is `examples/cpp/modelverse/03_conversational_pipeline.cpp` in the examples tree. ## The models Three registry loads. The TTS is a voice-cloning family and requires a voice reference; any short spoken wav works. `max_seq` sizes the decoder's KV pool, and a conversation turn needs nowhere near a 32768-token context window, so cap it: ```cpp constexpr const char* kSttRepo = "openai/whisper-large-v3-turbo"; constexpr const char* kLlmRepo = "Qwen/Qwen2.5-1.5B-Instruct"; constexpr const char* kTtsRepo = "ResembleAI/chatterbox-flash"; constexpr std::int64_t kLlmMaxSeq = 8192; ``` ```cpp const mv::ModelRegistry& registry = mv::ModelRegistry::builtin(); mv::LoadOptions load; load.device = device.value(); mv::LoadOptions llm_load = load; llm_load.max_seq = kLlmMaxSeq; const mv::SttModel stt = unwrap(registry.load_stt(kSttRepo, load)); const mv::GenerativeModel llm = unwrap(registry.load_generative(kLlmRepo, llm_load)); const mv::TtsModel tts = unwrap(registry.load_tts(kTtsRepo, load)); ``` **Placement.** `--device` puts all three models on one device (the CPU by default, `cuda` on request). A single request runs the steps in sequence, so the models share that device without contention. ## The custom nodes A node is the same three-member `runtime::Model` contract the [custom-node guide](add-a-pipeline-node.md) walks: `schema()` names the input and output tensors, `phases()` declares one `RunOnce` phase, `Phase_RunOnce` does the work. The speech nodes wrap the model handles they are given; the tensor specs they share are three small helpers over `spec::TensorSpec` (a byte string, a Float32 tensor of dynamic dims, an Int32 scalar): ```cpp spec::TensorSpec bytes_spec(const char* name) { return spec::TensorSpec{name, DataType::UInt8, {spec::TensorSpec::kDynamicDim}, /*optional=*/false}; } spec::TensorSpec f32_spec(const char* name, std::size_t rank) { return spec::TensorSpec{name, DataType::Float32, std::vector(rank, spec::TensorSpec::kDynamicDim), /*optional=*/false}; } spec::TensorSpec i32_scalar_spec(const char* name) { return spec::TensorSpec{name, DataType::Int32, {}, /*optional=*/false}; } ``` The listen node turns the request's encoded wav bytes into the whisper feature tensor through the model's own preprocessor configuration: ```cpp class ListenNode final : public crt::Model { public: explicit ListenNode(const mv::SttModel& stt) : stt_(&stt) { this->schema_.inputs.push_back(bytes_spec(kAudioIn)); this->schema_.outputs.push_back(f32_spec(kFeatures, 3)); this->phases_.push_back(crt::PhaseSpec{"listen", crt::PhaseKind::RunOnce}); } const crt::ModelSchema& schema() const override { return this->schema_; } ClikaRT::Span phases() const override { return this->phases_; } void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override { const std::vector wav = unwrap(ctx.inputs->get(kAudioIn)).item_as_vec(); ctx.outputs->set(kFeatures, this->stt_->features_from_bytes(wav)); } private: const mv::SttModel* stt_; crt::ModelSchema schema_; std::vector phases_; }; ``` The transcribe node is the same shape over `SttModel::transcribe_features`, and the speak node wraps `TtsModel::synthesize` with the required voice reference, emitting the waveform plus its sample rate as two outputs: ```cpp void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override { mv::SynthesizeOptions opts; opts.voice = this->voice_; const mv::SynthesizedAudio audio = this->tts_->synthesize( mvr::byte_tensor_to_string(unwrap(ctx.inputs->get(kAnswerText))), opts); const std::int32_t rate = audio.sample_rate; ctx.outputs->set(kSpeech, unwrap(Tensor::from_data( audio.samples.data(), {static_cast(audio.samples.size())}, DataType::Float32))); ctx.outputs->set(kSpeechRate, unwrap(Tensor::from_data(&rate, {}, DataType::Int32))); } ``` The post node peak-normalizes the waveform in tensor math, so it runs wherever the waveform lives; the sample rate rides from speak straight to the pipeline output. Request inputs are read-only on the node side, so the two ops that read `speech` allocate their results and only those are written in place: ```cpp void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override { const Tensor speech = unwrap(ctx.inputs->get(kSpeech)); Tensor peak = unwrap(ops::amax(unwrap(ops::abs(speech)))); CLIKA_CHECK(ops::clamp_(peak, kPeakFloor)); Tensor normalized = unwrap(ops::div(speech, peak)); CLIKA_CHECK(ops::mul_(normalized, kPeakTarget)); ctx.outputs->set(kAudioOut, std::move(normalized)); } ``` One build flag: a translation unit deriving from `runtime::Model` compiles with `-fno-rtti`, matching how the library builds. ## The chat template is the caller's job A pipeline consumes prompt bytes as-is; there is no hidden templating between nodes. The template node renders the transcript into the checkpoint's own chat framing, generation turn appended, so the decoder answers the question instead of continuing it: ```cpp void Phase_RunOnce(const crt::PhaseSpec&, crt::PhaseContext& ctx) override { ClikaRT::json::Json messages = ClikaRT::json::Json::array(); ClikaRT::json::Json turn = ClikaRT::json::Json::object(); turn["role"] = ClikaRT::json::Json("user"); turn["content"] = ClikaRT::json::Json( mvr::byte_tensor_to_string(unwrap(ctx.inputs->get(kTranscript)))); messages.push_back(std::move(turn)); ctx.outputs->set(mvr::kPrompt, mvr::string_to_byte_tensor(unwrap( this->tokenizer_->apply_chat_template(messages, /*add_generation_prompt=*/true)))); } ``` Skip this step and an instruction-tuned model greedy-decodes an untemplated instruction straight to its end-of-turn token; the reply is empty and nothing errors. The template node is where that goes right. ## Wiring The zoo trio constructs directly. A `PipelineStep` names the step, points at its node (non-owning; you keep the node alive), carries the step's execution knobs, and two rename maps: `input_map` from a pipeline name to a node input, `output_map` from a node output to a pipeline name. Edges otherwise wire by name, so one rename routes the detokenizer's generic `text` onto the conversational `answer_text` edge, where the speak node picks it up, and the decode step's scheduler admits `kMaxActive` sessions at a time: ```cpp crt::PipelineStep step(const char* name, crt::Model& model) { crt::PipelineStep st; st.name = name; st.model = &model; return st; } ``` ```cpp mvr::TokenizerNode tokenize(llm.tokenizer(), /*add_special_tokens=*/false); mvr::GenerativeDecoderModel decode(llm, gen, kMaxActive, mv::KVCacheMode::Continuous, /*prefill_chunk_tokens=*/0); mvr::DetokenizerNode detokenize(llm.tokenizer()); SpeakNode speak(tts, voice_path); PostNode post; std::vector steps; steps.push_back(step("listen", listen)); steps.push_back(step("transcribe", transcribe)); steps.push_back(step("template", prompt)); steps.push_back(step("tokenize", tokenize)); steps.push_back(step("decode", decode)); steps.back().exec_config.scheduler.max_active_sessions = kMaxActive; steps.push_back(step("detokenize", detokenize)); steps.back().output_map.emplace_back(mvr::kText, kAnswerText); steps.push_back(step("speak", speak)); steps.push_back(step("post", post)); crt::ModelSchema ext; ext.inputs.push_back(bytes_spec(kAudioIn)); ext.outputs.push_back(f32_spec(kAudioOut, 1)); ext.outputs.push_back(i32_scalar_spec(kSpeechRate)); ext.outputs.push_back(bytes_spec(kAnswerText)); crt::Pipeline pipeline = crt::Pipeline::create(std::move(steps), std::move(ext)); ``` The external schema is the caller's whole contract: only `audio` is settable from outside; `audio_out`, `speech_rate` and `answer_text` are readable; every edge between the nodes stays internal. ## One spoken question The request surface is the one every pipeline speaks. The response carries the answer's text, its normalized waveform and the waveform's sample rate; `io::save_audio` writes the wav file (`io::encode_audio` yields the same bytes in memory, for a serving payload): ```cpp crt::Request request; request.inputs.set(kAudioIn, bytes_to_tensor(question)); const crt::Response response = unwrap(pipeline.await(unwrap(pipeline.enqueue(std::move(request))))); const std::string answer = mvr::byte_tensor_to_string(unwrap(response.outputs.get(kAnswerText))); const Tensor reply = unwrap(response.outputs.get(kAudioOut)); const std::int32_t rate = unwrap(response.outputs.get(kSpeechRate)).item(); CLIKA_CHECK(ClikaRT::io::save_audio(reply, rate, reply_path)); ``` With a spoken "What is the capital of France?" as `question.wav` and any short spoken clip as `voice.wav` (`espeak-ng -v en-us -w question.wav "What is the capital of France?"` synthesizes one when no recording is at hand), `03_conversational_pipeline question.wav voice.wav reply.wav --device cuda` prints: ```text answer: The capital of France is Paris. spoken: 61440 samples at 24000 Hz -> reply.wav ``` The sample count varies from run to run; the TTS decodes with sampling. ## Streaming `enqueue` takes `RequestCallbacks` for event-driven consumption. The contract from `runtime/serving.h`: `on_chunk` runs per streamed chunk and the last clean-finish chunk carries `final = true`; `on_complete` runs once, at finalize, with the response; both are contained at the boundary, so a throw from one never unwinds into the engine; with callbacks set, no blocking `await` is needed. The request opts in with `streaming = true`: ```cpp crt::RequestCallbacks cbs; cbs.on_chunk = [&](const crt::ResponseChunk& c) { ++chunks; if (c.final) saw_final_marker = true; events.push_back(c.final ? 1 : 0); }; cbs.on_complete = [&](ClikaRT::Result) { ++completes; events.push_back(2); }; crt::Request r = audio_request(question_wav); r.streaming = true; crt::SessionId sid = 0; sid = pipe.enqueue(std::move(r), std::move(cbs)); ``` Callbacks run on the engine's pump thread; do not block in them. ## Concurrent sessions The decode step's `max_active_sessions` sizes its continuous batch; sessions beyond it queue rather than fail. With `max_active` 4, eight requests enqueued before any is awaited: ```cpp std::vector ids; for (std::size_t i = 0; i < args.burst_n; ++i) { ids.push_back(unwrap(pipe.enqueue(audio_request(question_wav)))); } std::size_t good = 0; for (const crt::SessionId sid : ids) { try { if (carries_answer(text_of(unwrap(pipe.await(sid)), convo::kAnswerText))) ++good; } catch (const ClikaRT::Error& e) { std::printf("[burst ] session %llu FAILED: %s\n", static_cast(sid), error_detail(e).c_str()); } } ``` Where a session silently returning a wrong answer under a burst would be costly, verify batched answers or serve one session at a time (`max_active` 1, the example's setting); Clika/ClikaRT#2058 is the tracker for that failure on the CUDA paged cache. ## Rebuilding, and the KV pool A live `GenerativeDecoderModel` holds its full KV pool on the device; two of them hold two pools, an out-of-memory at the second build on a 16 GB card. Scope the first pipeline and its nodes so they destruct before a rebuild (the example calls `pipeline.shutdown()` before it returns), and let `max_seq` size the pool for the conversation you serve. ## What the composition buys The composed pipeline builds its executors, its KV pool and its warmed kernels once and reuses them every turn: about 7 seconds per session, against about 22 seconds for the same three models driven in sequence with a fresh pipeline per turn. What a manual driver that kept its pipeline still would not get is the continuous batching. --- # Estimate depth, segment and detect The three image verbs: per-pixel depth maps, semantic segmentation masks, and open-vocabulary detection, from the command line and over HTTP. Source: https://docs.clika.io/modelverse/how-to/depth-segment-detect.md The vision families answer three per-image questions: how far away everything is (`depth`), what class each pixel belongs to (`segment`), and where the named things are (`detect`). The three verbs share their media handling (any image format the decoder accepts, one artifact per input) and their serving story, so this guide works all three. ## Estimate depth A depth family (Depth Anything and its relatives) maps each pixel to distance: ```bash clikart-cli depth-anything/Depth-Anything-V2-Small-hf depth photo.jpg ``` ```text photo.jpg: 64x64 min=1.877290 max=4.294917 ``` The payload is one line per image: the map's extent and its value range (relative maps: larger means closer). Artifacts opt in with `--output `, one per image, and each payload line then ends with ` -> `. The default `--output-format raw` writes the Float32 map as `_depth.npy`, the form a downstream consumer wants; `--output-format colormap` renders `_depth.png` instead, and `--colormap` picks its lookup table (`spectral` or `gray`). ## Segment an image A segmentation family labels every pixel with a class: ```bash clikart-cli facebook/detr-resnet-50-panoptic segment photo.jpg ``` The payload is the class-mask document: which classes appear, their pixel share, and the rendered overlay's path. `--mask-output mask.npy` also writes the raw class-id mask (UInt8, `[H, W]`) as a sidecar for programmatic consumers, and `--confidence` adds per-label confidence to the listing. ## Detect, with an open vocabulary or a trained class set The open-vocabulary detectors (OWL-ViT, Grounding DINO) find objects you NAME at run time through `--labels`; the closed-set DETR line (DETR, Deformable DETR, D-FINE, RF-DETR, RT-DETR) detects its checkpoint's trained classes and takes no label prompt. The open-vocabulary form: ```bash clikart-cli google/owlvit-base-patch32 detect photo.jpg \ --labels "a red bicycle, a street sign" --confidence 0.3 ``` ```text label score box (x1, y1, x2, y2) a red bicycle 0.812 (118, 204, 371, 468) a street sign 0.644 (402, 61, 455, 152) ``` One row per hit, in score order. The labels encode once and every image scores against them, so a many-image sweep pays the text encoding a single time; `--max-detections` caps the rows per image. ## Over HTTP Each verb has its serving twin in [the route table](serve-openai-compatible.md): `POST /v1/depth` (multipart image in, depth artifacts out), `POST /v1/segmentations` (image in, class-mask document out), and `POST /v1/detections` (image plus an open-vocabulary `prompt` in, box/score/label rows out): ```bash clikart-cli google/owlvit-base-patch32 serve --port 8000 ``` ```bash curl -s http://127.0.0.1:8000/v1/detections \ -F image=@photo.jpg -F prompt="a red bicycle, a street sign" ``` The operational story (binding, health, admission control) is the serve guide's, unchanged. Model sizing lives in [Model requirements](../model-requirements.md); the embedding-side image comparisons (which image is LIKE which, rather than what is IN one) are [Embed, compare and rerank](embed-and-rerank.md). --- # Embed, compare and rerank Dense vectors, similarity matrices and cross-encoder reranking from the command line and over HTTP: the retrieval building blocks. Source: https://docs.clika.io/modelverse/how-to/embed-and-rerank.md Retrieval systems stand on three operations: turn inputs into dense vectors, compare vectors, and rerank candidates against a query with a model that reads both together. Modelverse's embedding and reranking families provide all three as commands, and the same models serve them over HTTP. ## Compare texts Any text-embedding family provides `similar`: the positional texts embed in one batch, and the payload names the inputs by index and prints their cosine matrix: ```bash clikart-cli Qwen/Qwen3-Embedding-0.6B-GGUF:Q8_0 similar \ "How do I reset my password?" \ "Password reset instructions" \ "Quarterly revenue rose 4 percent" ``` ```text [0] How do I reset my password? [1] Password reset instructions [2] Quarterly revenue rose 4 percent scores (cosine; 1 = identical direction): [0] [1] [2] [0] 1.0000 0.7265 0.2591 [1] 0.7265 1.0000 0.3705 [2] 0.2591 0.3705 1.0000 ``` `--top-k N` switches the rendering to a per-input ranking (the nearest N others per row), the form you want when the list is long; `--json` emits the same result as a document for scripts. ## Compare images, or texts against images A dual-tower family (CLIP, SigLIP) embeds texts and images into one space, so `similar` grows two more forms: with `--image` attachments the matrix is texts against images, and with only images it is images against images: ```bash clikart-cli google/siglip-base-patch16-224 similar \ "a photo of a cat" "a photo of a dog" --image pet1.jpg --image pet2.jpg ``` For the raw vectors, `embed` takes images and prints a summary row per input (vector width, L2 norm, a quantized-vector signature, and each row's cosine against the first), with `--output PREFIX` writing `PREFIX.csv` (one row per image, the full vector as columns) and `PREFIX.json` (run configuration, summary and vectors) for whatever indexes them next: ```bash clikart-cli google/siglip-base-patch16-224 embed catalog/*.jpg --output catalog_vectors ``` ## Rerank documents against a query Embedding similarity is a coarse first pass; a reranker is a cross-encoder that reads the query and each document together and scores actual relevance. The reranking families provide `rerank`, with every document scored through one packed forward: ```bash clikart-cli Qwen/Qwen3-Reranker-0.6B rerank \ "how to reset a password" \ "Password reset instructions" \ "Changing your username" \ "Quarterly revenue rose 4 percent" ``` ```text 0.9231 [0] Password reset instructions 0.4106 [1] Changing your username 0.0312 [2] Quarterly revenue rose 4 percent ``` `--instruction` prepends the task instruction the checkpoint was trained with, when your use differs from the default retrieval phrasing. The everyday pipeline is both stages in order: `similar --top-k` over the corpus to shortlist, `rerank` over the shortlist to decide. ## Over HTTP The same models serve the retrieval routes ([the full route table](serve-openai-compatible.md)): ```bash clikart-cli Qwen/Qwen3-Embedding-0.6B-GGUF:Q8_0 serve --port 8000 ``` ```bash curl -s http://127.0.0.1:8000/v1/embeddings \ -H "Content-Type: application/json" \ -d '{"input": ["How do I reset my password?", "Password reset instructions"]}' ``` `POST /v1/embeddings` answers in the OpenAI shape, so `client.embeddings.create(...)` works against the base URL unchanged; `POST /v1/similarity` returns the cosine matrix directly, saving the round trip through raw vectors when the comparison is all you need. Embedding models are small ([Model requirements](../model-requirements.md) carries the figures), and the GGUF quantization selector on the source works here exactly as in [Run a specific GGUF quantization](run-a-gguf-quantization.md). --- # Add Modelverse to an existing CMake project find_package against the installed directory, staging the runtime beside your binary, and picking the dist on cross builds. Source: https://docs.clika.io/modelverse/how-to/existing-cmake-project.md Your project already builds; this guide adds the model zoo to it without restructuring anything. The installed directory ([Quick install](../getting-started/installation.md)) is the SDK: one `find_package`, one link target, C++17 or newer. ## The integration ```cmake find_package(Modelverse CONFIG REQUIRED PATHS "$ENV{MODELVERSE_INSTALL_DIR}/cmake") target_link_libraries(my_app PRIVATE Modelverse::modelverse) ``` That wires the public headers, your platform's `libclika_modelverse` with every model family registered, and the ClikaRT runtime installed in the same root. `PATHS` keeps the location out of the project files; setting `Modelverse_DIR` or extending `CMAKE_PREFIX_PATH` on the configure line works identically, which is the usual shape in CI: ```bash cmake -S . -B build -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake" ``` ## Run from your build directory The libraries live in the install root, and your freshly built binary does not know where that is. `modelverse_stage_runtime` copies the Modelverse library and the ClikaRT runtime next to the executable, so it runs from the build directory as-is, and the staged layout is exactly what you ship: ```cmake modelverse_stage_runtime(TARGET my_app) # beside the binary modelverse_stage_runtime(TARGET my_app DESTINATION lib) # or a subdirectory ``` ## When you already link ClikaRT A project already using the runtime keeps its arrangement: if `ClikaRT::ClikaRT` exists before `find_package(Modelverse ...)` runs, Modelverse links against it instead of resolving the runtime from the install root. The one constraint is a real pairing: each Modelverse build is built against the runtime of its own release, and a mismatched pair refuses at load with a readable error naming both versions rather than misbehaving later. ## Cross builds pick a dist The package selects the dist (platform/arch flavor) matching your target with the same detection ClikaRT ships, so a cross build that already configures ClikaRT correctly needs nothing extra. When the detection cannot see your intent, name the dist: ```bash cmake -S . -B build-android \ -DCMAKE_TOOLCHAIN_FILE="$NDK/build/cmake/android.toolchain.cmake" \ -DModelverse_DIR="$MODELVERSE_INSTALL_DIR/cmake" \ -DMODELVERSE_DIST=android-arm64 ``` An install built for a different platform fails at configure time, naming what it carries (`dist.json` records the platform); the fix is extracting the archive matching your target ([Get Modelverse](../getting-started/get-modelverse.md)) and pointing the configure at it. The program to write once the target links is [part 4 of the tutorial](../getting-started/first-model/04-use-it-from-code.mdx); for the runtime-only integration story (no model zoo), ClikaRT's own guide on [/clikart](/clikart) covers the same ground one layer down. --- # Register your own model family Teach the catalog an architecture it does not know, from the snapshot layout through identity matching to runnable commands. Source: https://docs.clika.io/modelverse/how-to/register-your-own-family.md The catalog covers the common architectures; your in-house model is not in it. Modelverse's answer is self-registration: a family is a translation unit that pushes a `ModelRegistration` into the shared registry at static-initialization time, and every built-in family registers exactly this way. There is no central list to edit. You compile one registration TU into your program against the public headers, and `list`, `info` and the loaders treat your family like any other. This page is the contract at full depth; for the guided walk from a hand-written checkpoint to a running family of your own, start with the [Adding your own model tutorial series](../getting-started/own-model/01-a-model-from-scratch.md); this page's worked example is a text-generation family, the detail-heavy case the tutorial deliberately avoids. ## What a snapshot must contain A model is a directory of files, and identity resolution reads only its configuration: - `config.json` with `model_type` and `architectures`; these two values are the identity keys a registration matches on. - The processor files the model needs (`tokenizer.json` and friends for text; preprocessor configs for image and audio). - The weights (`*.safetensors`, ONNX, or GGUF); untouched at identity time. A repo that ships no `config.json` can still be matched: a registration may carry a `probe_snapshot` function that recognizes the family from the file set alone and answers the `model_type` to route as. ## The registration, identity first One complete TU. With only this much, `list` shows the family (identity-only), `info` resolves your checkpoints to it, and every runnable command refuses precisely, naming what the family provides: ```cpp title="my_model_family.cpp" #include "clika_modelverse/models/registration.h" namespace { using namespace clika_modelverse; // The identity template: what a resolved checkpoint of this family IS. // The registry fills architecture, model_type and variant from config.json. class MyModel final : public Model { public: std::string_view family() const noexcept override { return "my-model"; } }; // The identity keys, matched against config.json. constexpr std::string_view kModelTypes[] = {"my_model"}; constexpr std::string_view kArchitectures[] = {"MyModelForCausalLM"}; // The Hugging Face pipeline tags the family's checkpoints carry, by the // constants of models/base/pipeline_tags.h (the registrar refuses a literal // outside that vocabulary). constexpr std::string_view kPipelineTags[] = {pipeline_tags::kTextGeneration}; // Static-init self-registration: constructing the registrar is the whole hookup. const ModelFamilyRegistrar kRegistrar{ [] { ModelRegistration registration{}; registration.metadata.family = "my-model"; registration.metadata.vendor = "my-org"; registration.metadata.pipeline_tags = kPipelineTags; registration.metadata.description = "the in-house my_model language models"; registration.metadata.input_combos = kTextOnlyCombos; registration.metadata.output_combos = kTextOnlyCombos; registration.hf_model_types = kModelTypes; registration.hf_architectures = kArchitectures; registration.factory = make_model(); return registration; }(), }; } // namespace ``` Compile that TU into any program that links `Modelverse::modelverse` (the command-dispatch functions in the library serve programs that embed them; the shipped `clikart-cli` binary knows only the built-in families). An incomplete declaration (no family name, no vendor, no pipeline tag or one outside the Hugging Face vocabulary, no modalities, no identity factory) is refused at process start, not papered over downstream. ```text my-model identity-only * Input Modalities: text * Output Modalities: text * Commands: none (identity-only) * Vendor: my-org ``` ## Making it runnable Two more fields turn identity into execution: - **`build_generative`** materializes the runnable model from a resolved snapshot: read the config, adapt the checkpoint's tensor names where they differ from what your forward expects (`ClikaRT::io::TensorsAdapter`, rename-at-lookup, no payload copies), assemble the forward from the `clika_modelverse/modules/` building blocks (attention, decoder stack, projections, rope cache) or your own `nn::Module`, and return a `GenerativeModel` wrapping it with the snapshot's tokenizer and generation defaults. The GGUF twin, `build_generative_gguf`, receives the already-loaded file so identity read and weight bind share one mapping. - **`cli`** declares the family's commands. The battery builders compose the standard text surface in one line, and with it your family's `prompt`, `serve` and `bench` behave exactly as the tutorial documents them: ```cpp registration.cli = CliSurface{batteries::cli::llm_cli, batteries::cli::llm_cli_verbs()}; registration.factory = make_model(); registration.build_generative = build_my_model; ``` A family wanting a different surface composes its own `FamilyApp` from the same public builders (`batteries/cli/prompt.h`, `serve.h`, `bench.h`), one verb per builder; the advertised verb list and the App are pinned to agree by the registration contract. Custom behavior around an EXISTING family needs none of this: [Add your own node to a model pipeline](add-a-pipeline-node.md) composes your logic with any catalog model. The complete field-by-field contract, including the speech, embedding and reranking factories, is `clika_modelverse/models/registration.h` in the installed headers. ## The nearest family to start from A checkpoint no family claims is refused by name: `clikart-cli info ` prints the `model_type` and `architectures[0]` it read and that no template serves them, and `clikart-cli list` prints every family with the keys it claims, so the family whose definition the checkpoint's `config.json` shares is the one to read first. A composite whose `config.json` carries a nested `text_config` naming a registered family composes that family's text encoder through the text-backbone seam (`build_text_backbone` on the registration), with no forward of its own for the tower. A decoder driven by the embeddings of an encoder you wrote has no such seam in this release: its forward is written from the modules, as the section above does for a text-generation family. --- # Run a specific GGUF quantization Pick one weight option of a repo that ships many, by tag, glob or exact file, and see the size table before anything downloads. Source: https://docs.clika.io/modelverse/how-to/run-a-gguf-quantization.md A GGUF repo usually publishes one model in several quantizations, from a 2-bit file that fits a phone to an 8-bit one that wants a workstation. Modelverse treats those as weight options of one source: it never guesses between them, the option table with sizes prints before any download, and a selector names the one you want. ## See the options Ask with `--dry`. A source with several weight options leads with the option table: ```bash clikart-cli info Qwen/Qwen3-4B-GGUF --dry ``` ```text Qwen/Qwen3-4B-GGUF family=qwen model_type= architecture=qwen3 variant: dense components: tokenizer=no image=no audio=no video=no license page: https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT model: Qwen/Qwen3-4B-GGUF provider: Qwen source: https://huggingface.co/Qwen/Qwen3-4B-GGUF license: apache-2.0 https://huggingface.co/Qwen/Qwen3-4B-GGUF/blob/main/LICENSE 'Qwen/Qwen3-4B-GGUF' ships 5 weight options; pick one: option files size selector Qwen3-4B-Q4_K_M 1 2.3 GiB :Q4_K_M Qwen3-4B-Q5_0 1 2.6 GiB :Q5_0 Qwen3-4B-Q5_K_M 1 2.7 GiB :Q5_K_M Qwen3-4B-Q6_K 1 3.1 GiB :Q6_K Qwen3-4B-Q8_0 1 4.0 GiB :Q8_0 pick by tag (":Q6_K" / --weights Q6_K; must match one option), by glob (":*Q8*"), by exact file (org/repo/.gguf, a /blob|/resolve URL, or --weights .gguf), or ":latest" for the hub default total: nothing is fetched until one option is selected repository: 14.7 GiB in 9 files; the rest is not fetched (other weight formats, files the fetch never takes) (dry; nothing downloaded) ``` Running the bare source refuses with the same table and exit code 2; the refusal is the answer to "what is there", not an obstacle. ## Select one A selector rides the source after a colon, and the same grammar works on every command the model provides: ```bash clikart-cli Qwen/Qwen3-4B-GGUF:Q6_K prompt "The capital of France is" ``` The refusal's own closing line is the whole grammar; the four forms, most specific last: - **A tag**, `Qwen/Qwen3-4B-GGUF:Q6_K`. An exact option name resolves outright; otherwise a case-insensitive contains-match runs over the option names and each option's first file name. The everyday form. - **A glob**, `Qwen/Qwen3-4B-GGUF:*Q4*`. Must match exactly one option; an ambiguous selector refuses with every `name (file)` candidate listed. - **`:latest`**, the repo's default option, an explicit opt-in rather than a silent fallback. - **An exact file**, `Qwen/Qwen3-4B-GGUF/Qwen3-4B-Q6_K.gguf`, a pasted Hugging Face Hub `blob`/`resolve` URL, or a unique trailing part of the file name. Unambiguous, the right form for scripts that must never drift when the repo adds options. `fetch` takes the same selection, either in the source or as `--weights Q6_K`, and prints the selected weight file's local path as its payload: ```bash WEIGHTS=$(clikart-cli fetch Qwen/Qwen3-4B-GGUF --weights Q6_K) ``` ## Pick the compute dtype On a quantized checkpoint, `--dtype` (`float16`, `bfloat16`, `float32`) selects the compute dtype: the precision every quantized weight decodes to and the activations and KV cache run at, behind the generative, embedding and reranker loads. The option pins the whole tree, and the token table of a quantized checkpoint stays quantized in memory, decoding each gathered row at the selected dtype, so a multi-gigabyte f32 table is never materialized. Absent, the checkpoint's own dtype rules. On a dense checkpoint the option is the weight cast applied at load: a llama checkpoint takes `float16`, `bfloat16` or `float32`, an object-detection checkpoint `float16` or `bfloat16`; a dense checkpoint of any other family refuses the option by name. ```bash clikart-cli Qwen/Qwen3-4B-GGUF:Q6_K prompt "The capital of France is" --dtype float16 ``` ## What counts as one option An option is a group, not always a file: a multi-part split (`-00001-of-00009.gguf`) is one option named by its stem (the table's `files` column counts its parts), and a subdirectory holding one quantization's parts is one option named by the subdirectory. Splits are listed with a `(split; not served yet)` note and are selectable by name, but loading one refuses precisely; the loader serves single-file GGUF. Sidecar files (mmproj, imatrix) are excluded from the options and the table says how many it left out. Which quantization fits which device is a sizing question; [Model requirements](../model-requirements.md) carries the figures. The loading path underneath (block-quantized weights decoded inside the kernels, served at their on-disk footprint) is ClikaRT's, documented in its GGUF guide on [/clikart](/clikart). --- # Run fully offline Fetch on a connected machine, move the model directory, and make any network touch an error instead of a surprise. Source: https://docs.clika.io/modelverse/how-to/run-fully-offline.md Modelverse needs a network exactly once per model, and not necessarily on the machine that runs it. A local directory is a first-class source everywhere a repo id is, and an offline switch turns any accidental network touch into an error. This is the deployment shape for air-gapped and restricted networks: the release archive installs by extraction alone, no package manager, and the models arrive the same way the archive did. ## Stage on a connected machine Fetch into a directory you control (not the shared cache), so the result is a self-contained folder: ```bash clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct --cache-dir staged ``` ```text staged/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775 ``` The payload path on stdout is the directory to ship. `fetch --dry` first shows what will be downloaded and how big it is; for a GGUF source, select the quantization at staging time (`--weights Q6_K`) so only the variant you deploy moves. A gated repo needs a token on this machine only (`HF_TOKEN`, or the token a Hub login stored); the token never travels with the files. ## Move it, run it Copy the directory however files reach the target (rsync over an approved channel, physical media; the `cp` below stands in for that transfer). The cache uses the Hugging Face hub layout, so capture the payload path instead of assuming it; fetching an already-cached source verifies and returns the same path immediately, which makes the capture free. On the offline machine, the copied directory is the source: ```bash SNAPSHOT=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct --cache-dir staged) cp -r "$SNAPSHOT" ./Qwen2.5-0.5B-Instruct clikart-cli ./Qwen2.5-0.5B-Instruct prompt "The capital of France is" ``` A local directory passes through resolution untouched, so `info`, `prompt`, `serve` and the library's snapshot call all accept it identically. No part of model resolution knows or cares that the machine has no route out. The license credential is the one thing that does, and the section below is that one thing. On an Android phone the same move works in two shapes. An app's cache root is its own cache directory (`ClikaRtAndroid.load` sets `XDG_CACHE_HOME` to it), and the model library's hub cache sits under it in the same layout as on the workstation, so a snapshot directory copied there serves the app with no network. The `clikart-cli` executable pushed to the phone over `adb` takes `--cache-dir ` for a cache you pushed beside it, or the copied snapshot directory itself as the source. ## Make offline a guarantee Trust but verify: add `--offline` and any operation that would touch the network fails with a readable error instead of hanging on a dead route: ```bash clikart-cli ./Qwen2.5-0.5B-Instruct serve --offline ``` `--offline` also goes before the command, where it covers the commands that name no model: `clikart-cli --offline info `, `--offline fetch`, `--offline list`. The environment forms hold the same guarantee process-wide, useful under systemd or in a container where flags are out of reach: `CLIKA_MODELVERSE_OFFLINE=1` (Modelverse's own switch) or `HF_HUB_OFFLINE=1` (honored for compatibility with other Hugging Face Hub tooling). With the switch set, a repo-id source still works when the snapshot is already in the cache; resolution is cache-only. A cached GGUF repository resolves its identity offline from its own weight file, so a GGUF source needs no hub metadata on the target either. ## The credential has to be the offline kind The runtime runs compute under a license, so the target machine needs a credential like any other ([Get Modelverse](../getting-started/get-modelverse.md#license-credential)). The kind matters here and nowhere else. An offline license is verified locally against Clika's signature and needs no route out, which is what makes it the one to ship to an air-gapped target. An online license needs a route to the platform, so a machine with no route cannot use one. Ask your platform for an offline license, and carry it to the target the way you carry the model directory. ```bash export CLIKA_RT_LICENSE=/srv/clika/clikart.license # the credential text, or a file holding it clikart-cli ./Qwen2.5-0.5B-Instruct prompt "The capital of France is" --offline ``` `bin/clikart-license-init ` stores it once under the user account instead, which suits a service account that starts without a shell profile. An offline license carries an end date, and [Runtime licenses](/platform/concepts/runtime-licenses) covers what happens on it and how to re-issue. ## What to verify on the target Three commands, no network, in order: `clikart-cli devices` proves the runtime loads and sees the hardware, `clikart-cli info ./` proves the model resolves, and a one-line `prompt` proves end to end, the credential included. If the documentation should travel too, the docs site you are reading has a self-contained offline bundle (the "Offline Docs" button in the navigation bar). Sizing the model to the offline hardware is the usual question in these deployments; [Model requirements](../model-requirements.md) and [Run a specific GGUF quantization](run-a-gguf-quantization.md) together answer it. --- # Script clikart-cli The contracts that make clikart-cli scriptable: exit codes, the stdout/stderr split, option templates and shell completion. Source: https://docs.clika.io/modelverse/how-to/script-the-cli.md The `clikart-cli` executable is built to be driven by scripts, and the contracts below are stable interfaces, not implementation details. A script that branches on them keeps working across releases. ## stdout is the payload, stderr is everything else Every command writes exactly its payload to stdout: generated text, a transcript, the catalog, a fetched path, CSV. Banners, progress bars, timings, warnings and errors all go to stderr. Capturing a result is therefore safe by construction: ```bash MODEL_DIR=$(clikart-cli fetch Qwen/Qwen2.5-0.5B-Instruct) ANSWER=$(clikart-cli "$MODEL_DIR" prompt "Two plus two is" --greedy --max-new-tokens 4) ``` `-q` mutes stderr down to warnings and errors, `-v` raises it to debug (including the effective-options report showing which value came from which source), and `--no-color` strips styling (also honored automatically for non-TTY output, `NO_COLOR`, and `TERM=dumb`). The runtime's own `[ClikaRT]` log lines follow the same rule: the level tag is colored only when stderr is a terminal, `NO_COLOR` is unset or empty, and `TERM` is not `dumb`, so a captured stderr holds no escape bytes. None of these change the payload. Machine-readable forms exist where the payload is tabular: `devices --json` and `devices --csv` print the hardware report as a document instead of the human view, `check --csv` prints the fit verdict table (one row per device, every term of the law a column) as CSV, `list --json` prints the registry as one document (per family: the vendor, the Hugging Face pipeline tags it serves, the modalities, the served and refused model types and GGUF architectures, the advertised verbs, whether it is runnable), and `info --json` / `info --csv` print the supported-models document, or its table as CSV, for one source (a one-entry document), several sources or `@FILE` (one source per line): per model the family, the resolved commit, the parameter count and where it was read from, the license, the gate, and one row per weight variant with the files a fetch lands and their sizes, which a catalog generated from the registry reads. The grammar is `clikart-cli info [--weights SEL] [--json | --csv] [--schema] [--dry]`: one source with no form flag prints the human description, `--schema` prints the document's JSON Schema and needs no source, and a source the hub refuses is a warning on stderr while the document carries the rest and the exit code is 1. ## The exit-code contract Four values, fixed: | Exit | Meaning | Example | | --- | --- | --- | | 0 | success | the payload is on stdout | | 1 | runtime failure | download interrupted, weights failed to load | | 2 | usage error | unknown flag, missing argument, unselected weight options | | 3 | the model cannot do this here | a verb the family does not provide, an identity-only family, an unsupported variant | The distinction between 2 and 3 is the useful one: 2 means fix the invocation, 3 means fix the model choice. A retry loop should retry 1 and never 2 or 3. ```bash SRC=openai/whisper-large-v3-turbo # the model this script expects FILE=meeting.wav # the input it transcribes if ! clikart-cli "$SRC" transcribe "$FILE" > out.txt; then case $? in 2) echo "bad invocation, check the flags" >&2; exit 2 ;; 3) echo "$SRC has no transcribe; pick a speech-to-text model" >&2; exit 3 ;; *) echo "runtime failure, retrying" >&2 ;; esac fi ``` ## Options as a file `generate-template` writes a JSON file with every root option at its default (null where there is no built-in default); `--template` loads one, and explicit flags always win over it. This is how a deployment pins its configuration in version control instead of a growing alias: ```bash clikart-cli generate-template defaults.json clikart-cli info "$SRC" --template defaults.json clikart-cli fetch "$SRC" --template defaults.json --cache-dir /data/models ``` A family's own verbs are their own contract; the template governs the root verbs. Secrets stay out of both: the Hugging Face Hub token comes from `HF_TOKEN` in the environment or from the token a Hub login stored, never a flag and never a template key, so neither shell history nor a committed template can leak it. ## Completion `clikart-cli` generates its own shell completion: ```bash source <(clikart-cli completion bash) # zsh and fish work the same way ``` Completion covers the root verbs and flags; model commands are discovered per source at run time, which completion cannot see, and that is the one place tab-completion ends and ` --help` takes over. --- # Serve a model over the OpenAI and Anthropic APIs Host any runnable model behind the built-in HTTP server, stream completions, and point existing OpenAI and Anthropic clients at it. Source: https://docs.clika.io/modelverse/how-to/serve-openai-compatible.md Any model whose family provides `serve` becomes an HTTP endpoint in one command, speaking the OpenAI request and response shapes and, for chat, the Anthropic Messages API beside them. Existing OpenAI and Anthropic SDK code connects by changing its base URL; nothing else about the client changes. ## Start the server ```bash clikart-cli Qwen/Qwen2.5-0.5B-Instruct serve --host 0.0.0.0 --port 8000 ``` ```text [2026-10-04 07:59:45.310] [modelverse] [info] serving on http://0.0.0.0:8000 (model=Qwen2.5-0.5B-Instruct); open it in a browser for the chat page [2026-10-04 07:59:45.310] [modelverse] [info] endpoints: POST /v1/chat/completions /v1/messages /v1/messages/count_tokens /v1/audio/transcriptions; GET /v1/models /health /props / (web UI) /dashboard [2026-10-04 07:59:45.310] [modelverse] [info] authentication: none; every route is open ``` The load flags from `prompt` apply unchanged (`--device`, `--max-seq`, a GGUF weight selector on the source); the operational flags are the server's own: - `--host` and `--port`: the bind address, `127.0.0.1:8000` by default. Loopback by default is deliberate; expose it with `--host 0.0.0.0` when you mean to. - `--no-web-ui`: disable the built-in chat page on `/`. - `--cors`: answer cross-origin requests, for a browser front end on another origin. - `--api-key KEY`: require this key on every route but `/health`, `/v1/health`, `GET /` and `GET /dashboard`, sent as `Authorization: Bearer ` (the OpenAI clients' form) or `x-api-key: ` (the Anthropic clients' form). A request without it, or with another key, answers 401 in the client's own error envelope with a `WWW-Authenticate: Bearer` challenge. Without the flag every route is open, so set it whenever the server listens beyond loopback; a key on the command line is readable in the process list. The two pages open without a key (a browser navigation carries no header); the chat page sends the key typed into its API key field on every call, and the dashboard reads the same stored key. **The key is a best-effort gate, not a security boundary:** it keeps a stray client on your network from using the model, and that is all it does. Access control for a server people you do not know can reach (identities, per-user keys and their rotation, rate limits, TLS, audit logs) belongs to an infrastructure layer in front of the serving process, a reverse proxy or an API gateway; the serving process itself stays a single-tenant engine behind it. - `--max-active` and `--max-queue`: admission control. Requests past the queue limit are shed and surface as HTTP 503 after retries with backoff, so an overloaded server degrades loudly instead of stalling silently. - `--max-upload-mib`: the largest request body accepted, 256 MiB by default. A larger upload answers 413 naming the size and the limit. Audio sizing: a 60-minute 16 kHz mono WAV is 111 MiB. - `--read-timeout`: seconds a connection may stay silent while its request is read, 30 by default. A stalled read closes the connection rather than holding a slot. - `--tool-parser`: how tool calls are read out of the model's text. By default the server picks, from the call formats the model's family declares, the one the checkpoint's chat template writes. The flag names another parser (`serve --help` lists them), or `none` turns extraction off. With no parser in use, a tool call stays in the reply as plain text. `/props` reports the parser in use as `tool_parser`, and `tool_parser_source` says where it came from: `family` or `option`. ## The routes What the server answers depends on the model's family; every serving model carries the common rows. | Route | Serves | Model families | | --- | --- | --- | | `POST /v1/chat/completions` | chat, streaming (SSE) and non-streaming | text and multimodal generators | | `POST /v1/messages` | chat in the Anthropic Messages API shape, streaming (SSE) and non-streaming | text and multimodal generators | | `POST /v1/messages/count_tokens` | the input token count the same `/v1/messages` request would ingest (text only) | text and multimodal generators | | `POST /v1/audio/transcriptions` | multipart audio file to transcript; `response_format` `json` (the default), `text`, `verbose_json` (the task, the language, the duration and the segments with their seconds and tokens), `srt` or `vtt` (one cue per segment); `language` the hint, `temperature` the decode's sampling (0, the default, is greedy; above 0 the decode samples under a fixed seed); `timestamp_granularities[]` takes `segment` (word timestamps are not produced); `prompt` is refused (no decode here conditions on a text prompt) | speech-to-text (Whisper) | | `POST /v1/audio/speech`; `POST`, `GET`, `DELETE /v1/voices` | text in, waveform out, in a voice registered first through the voice registry ([Speak text](speak-text.md)); `response_format` `wav` (the default) or `pcm` always, and `mp3`, `opus`, `aac` or `flac` when the host provides the `ffmpeg` tool (`CLIKA_RT_FFMPEG` names its path, else it is taken from the PATH; the tool is spawned, never linked); `GET /props` lists the served set as `speech_formats`, and the start-up log says which | text-to-speech (Chatterbox, Qwen3-TTS) | | `POST /v1/embeddings` | texts to dense vectors; `encoding_format` `float` (the default) or `base64` (each row as the base64 of its float32 little-endian bytes); `dimensions` shortens every row to its leading entries renormalized to unit length, refused past the model's own count | embedding models | | `POST /v1/similarity` | texts to a cosine matrix | embedding models | | `POST /v1/translations` | texts plus a language pair | translation models | | `POST /v1/images/generations`, `POST /v1/videos/generations` | prompt to a generated image or clip; `response_format` `b64_json` (the default) or `url`: the file is kept in memory under an unguessable id for an hour (at most 64 files, 256 MiB in all; the oldest leave first) and served at `GET /v1/files/` with no key, the id being the credential as a signed URL's is | diffusion models | | `POST /v1/depth` | multipart image to depth artifacts | depth estimators | | `POST /v1/segmentations` | multipart image to a class-mask document | segmentation models | | `POST /v1/detections` | multipart image to box/score/label rows (an open-vocabulary detector also takes a `prompt`) | detectors | | `GET /v1/models`, `GET /props` | what is loaded and how it is configured; the model listing answers the OpenAI shape (`object`, `created`, `owned_by`, `license`), or the Anthropic one (`type`, `display_name`, `created_at`, `has_more`) to a request carrying `anthropic-version` or `x-api-key` | all | | `GET /health`, `GET /v1/health` | `{"status":"ok"}`, the machine probe | all | | `GET /` | the built-in web chat page | all, unless `--no-web-ui` | | `GET /dashboard` | the live request dashboard | all | | any other path | 404 in the client's own envelope: the OpenAI one (`code` `unknown_url`, the path named) or, under `/v1/messages`, the Messages one (`not_found_error`) | all | ## Call it with the OpenAI API Non-streaming, from anything that can POST JSON: ```bash curl -s http://127.0.0.1:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"messages": [{"role": "user", "content": "One-line haiku about rain."}]}' ``` Streaming adds `"stream": true` and reads server-sent events, one `data:` line per delta, `data: [DONE]` last, the shape OpenAI clients already parse. Which means the SDKs work as-is: ```python from openai import OpenAI client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused") for chunk in client.chat.completions.create( model="qwen", messages=[{"role": "user", "content": "One-line haiku about rain."}], stream=True, ): print(chunk.choices[0].delta.content or "", end="", flush=True) ``` A multimodal model takes OpenAI-shaped image content parts in the same route; a Whisper model is called with a multipart file upload instead ([Transcribe audio](transcribe-audio.md) shows both of its forms). ## Call it with the Anthropic API `POST /v1/messages` takes the Messages API request body on the same server. `max_tokens` is required, as it is in the Anthropic API. The `anthropic-version` header is optional; `2023-06-01` is the version served, and any other value is refused. A request may name any model, and the reply names the one the server loaded. ```bash curl -s http://127.0.0.1:8000/v1/messages \ -H "Content-Type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{"model": "qwen", "max_tokens": 256, "messages": [{"role": "user", "content": "One-line haiku about rain."}]}' ``` The Anthropic SDK takes the server's address as its base URL, without `/v1`. Any API key works, because the server checks none: ```python from anthropic import Anthropic client = Anthropic(base_url="http://127.0.0.1:8000", api_key="unused") with client.messages.stream( model="qwen", max_tokens=256, messages=[{"role": "user", "content": "One-line haiku about rain."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) ``` A streamed reply follows the Messages API event order. It opens with `message_start` and a `ping`. Each content block then streams as `content_block_start`, its `content_block_delta` events and `content_block_stop`. The reply closes with `message_delta`, which carries the stop reason and the output usage, and then `message_stop`. There is no `[DONE]` line. A `ping` also goes out whenever the stream has been silent for 5 seconds. `max_tokens` is a cap. When the context window runs out first, the reply stops there with `stop_reason` `model_context_window_exceeded`. A prompt that fills the window on its own is refused with a message that starts `prompt is too long:`. The other stop reasons are `end_turn`, `max_tokens`, `stop_sequence` (with the matched `stop_sequence`) and `tool_use`. In `usage`, `input_tokens` counts the prompt tokens the prefix cache did not serve, `cache_read_input_tokens` counts the ones it did, and `output_tokens` counts the reply. Errors come back in the Anthropic error envelope, with `request_id` in the body, and every answer carries a `request-id` header. A refusal is a 400 `invalid_request_error` whose message starts with the path of the field it names. The other statuses: - 404 `not_found_error`: a path under `/v1/messages` that no route serves, such as the batches API. - 413 `request_too_large`: a body past `--max-upload-mib`. - 529 `overloaded_error`: the server is at capacity, for example when admission control sheds the request. The answer carries `retry-after` and `x-should-retry: true`. - 503 `api_error`: the device faulted. Restart the server; the answer carries `x-should-retry: false`. - 500 `api_error`: the generation failed. The `anthropic-version` header is read when present and taken as the served version when absent (a client that omits it is answered, not refused). ### What the Messages API route serves, accepts and refuses | Request | What the server does | | --- | --- | | text; base64 images (JPEG, PNG, GIF, WebP) to a multimodal model | served | | `system` as a string or as text blocks | served | | `stop_sequences` (up to 4, each up to 128 bytes), `temperature`, `top_p`, `top_k` | served | | custom tools, `tool_choice` `auto` or `none`, `tool_use` and `tool_result` blocks | served; calls come back as `tool_use` blocks when a parser reads the model's call format (the family's own by default, or the one `--tool-parser` names) | | `thinking` `enabled` (a `budget_tokens` of at least 1024, below `max_tokens`), `adaptive` or `disabled`; `output_config.effort` at a level the model declares | served by a model that reasons | | a final assistant turn (a response prefill: text only, with no trailing whitespace) | served; the reply carries only the new text | | `cache_control` | accepted, no effect; the server's prefix cache needs no marker | | `metadata`, `service_tier`, `output_config.task_budget` | accepted, no effect | | `context_management` | accepted, not applied; the history renders in full | | a tool's `strict` and `defer_loading`; the tool-search server tools | accepted; arguments are not schema-constrained, and every tool is offered up front | | an effort level the model does not declare | accepted; the model's default applies | | fields the server does not know, in a request that sends `anthropic-beta` | ignored | | `tool_choice` `any` or `tool` (a forced tool call) | refused | | other server tools, `container`, `output_config.format` (structured output) | refused | | `document`, `search_result` and other block types the server does not serve | refused | | an image `file` source or an image URL (the server fetches no remote payload) | refused | | `max_tokens` 0 | refused | | fields the server does not know, in a request without `anthropic-beta` | refused: `: Extra inputs are not permitted` | | a response prefill with `thinking` `enabled` or `adaptive`, or to a model that always reasons; `thinking` `disabled` on a model that always reasons and declares no effort levels | refused | | an image in a `count_tokens` request | refused; it counts text only | The server logs `context_management`, `strict`, `defer_loading`, the tool-search tools and an effort level the model does not declare once each, at Info, naming the field. ## Use it from Claude Code Claude Code speaks the Messages API, so it runs against this server once its environment variables point it there. Start the server with a context window that holds Claude Code's prompt, and put your own device in `--device`: ```bash clikart-cli Qwen/Qwen3-30B-A3B-Instruct-2507 serve --device vulkan:0 --max-seq 65536 ``` Then start Claude Code with these variables: ```bash unset ANTHROPIC_API_KEY export ANTHROPIC_BASE_URL=http://127.0.0.1:8000 export ANTHROPIC_AUTH_TOKEN=unused export ANTHROPIC_MODEL=Qwen3-30B-A3B-Instruct-2507 export ANTHROPIC_DEFAULT_HAIKU_MODEL=Qwen3-30B-A3B-Instruct-2507 export CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536 export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 export API_TIMEOUT_MS=600000 claude ``` - `ANTHROPIC_BASE_URL` is the server's address, without `/v1`. - `ANTHROPIC_AUTH_TOKEN` is the server's `--api-key` value, or any value when the server runs without one. Claude Code sends it as an `Authorization: Bearer` header, and it ranks above `ANTHROPIC_API_KEY` in Claude Code's [credential order](https://code.claude.com/docs/en/authentication). Unsetting `ANTHROPIC_API_KEY` keeps an Anthropic key out of a session that talks to another server. - `ANTHROPIC_MODEL` names the main model, and `ANTHROPIC_DEFAULT_HAIKU_MODEL` names the one Claude Code uses for background tasks. Behind a custom base URL, Claude Code passes any model name through unchecked, and this server answers every request with the model it loaded, so both name the served model, the id `/props` reports as `model_id`. - `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is the window Claude Code assumes for the model. For a model name it does not recognize, Claude Code can assume a window that differs from the one this server serves, so set it to the window `/props` reports as `n_ctx`. Claude Code then compacts the conversation before it reaches the server's limit ([model configuration](https://code.claude.com/docs/en/model-config)). - `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` turns off Claude Code's other traffic, such as auto-updates, telemetry and error reporting, for a network that reaches only this server. - `API_TIMEOUT_MS` is how long Claude Code waits for a reply, in milliseconds. 600000 (10 minutes) is its default; raise it when the first prefill of Claude Code's prompt takes longer than that on your device. To check the connection, run `/status` in Claude Code. Its `Anthropic base URL` line shows this server's address, and its `Auth token` line names `ANTHROPIC_AUTH_TOKEN` ([verify the connection](https://code.claude.com/docs/en/llm-gateway-connect#verify-the-connection)). Claude Code sends its whole system prompt and every tool definition with each request, about 16,000 tokens before the conversation starts, so the served window must hold them. The server caps the window at 4096 tokens by default, and with that cap Claude Code refuses the request itself with "Prompt is too long". `--max-seq` sets the window, and `/props` reports it as `n_ctx`. The tools reach the model through its chat template, and its calls come back to Claude Code as `tool_use` blocks when a parser reads the model's call format: the family's own by default, or the one `--tool-parser` names (above). `/props` shows which parser is in use. With `ANTHROPIC_BASE_URL` pointing at a host other than `api.anthropic.com`, Claude Code turns off features such as Remote Control and server-managed settings ([feature availability](https://code.claude.com/docs/en/feature-availability)). Anthropic's gateway documentation states that Anthropic "doesn't support routing Claude Code to non-Claude models through any gateway" ([LLM gateways](https://code.claude.com/docs/en/llm-gateway)). ## Operate it `GET /health` is the readiness probe for a load balancer. `GET /props` reports the loaded model and its effective options, the remote twin of `-v`'s effective-options report; its `device` field names the device the model actually landed on, which is what you check after `--device auto` or a fallback. Its `routes` list gives every API route the server answers, each with its method and its path, plus `api` (`openai` or `anthropic`) on a route that speaks one of those two APIs. The dashboard on `/dashboard` shows the same routes and the live requests; `/api/requests` is its JSON feed. Everything the server logs goes to stderr like the rest of the `clikart-cli` executable, so systemd or a container runtime captures it without configuration. For the same server inside your own process (your engines, your routes, the built-in UI swapped for your page), the `01_serve` program in [Additional examples](../examples.md) is the smallest complete consumer of the library's `ServeApi`. --- # Speak text Text-to-speech from the command line and over HTTP, with voice references and the synthesis knobs. Source: https://docs.clika.io/modelverse/how-to/speak-text.md A text-to-speech family turns text into a waveform. The catalog's runnable TTS families (Chatterbox, Qwen3-TTS) provide `speak`: the model synthesizes, the wav lands on disk, and the payload on stdout is the file's path, ready to pipe into the next command. The runnable Chatterbox checkpoint is `ResembleAI/chatterbox-flash` (the original `ResembleAI/chatterbox` snapshot ships a flow estimator the runtime does not serve, and the tool's refusal names the served one). CSM is cataloged for identity resolution today (`info` recognizes it); its runnable path is not shipped yet. ## From the command line ```bash clikart-cli ResembleAI/chatterbox-flash speak \ "The install works. Pick a model from the catalog." \ --voice reference.wav --output hello.wav ``` ```text hello.wav ``` Play it with anything; the file is 16-bit PCM at 24 kHz. The flags: - `--output PATH`: the wav path (`speech.wav` by default). The path is the stdout payload, so `aplay "$(clikart-cli speak "..." )"` works as a one-liner. - `--voice PATH`: REQUIRED for Chatterbox, a reference voice whose timbre the model matches (the family ships no default voice; `speak` without it refuses and says so). Where a family ships named voices, `--voices-dir` points at the set. - The synthesis knobs are the family's own (` speak --help` lists them); Chatterbox exposes `--exaggeration` (expressiveness), `--steps` (decode steps, quality against speed) and `--temperature`. - The load knobs apply unchanged: `--device`, `--cache-dir`, `--offline`. ## Over HTTP The same model serves `POST /v1/audio/speech` (text in, waveform out) behind a voice registry, rows in [the route table](serve-openai-compatible.md): `POST /v1/voices` registers a reference clip under a name, `GET /v1/voices` lists the registered voices, `DELETE /v1/voices/` frees a slot, and a speech request names one of them. Chatterbox ships no default voice, so a speech request before any registration is refused (`voice 'default' is not registered; POST /v1/voices first`). ```bash clikart-cli ResembleAI/chatterbox-flash serve --port 8000 ``` Register the reference clip once, as a multipart upload with `name` (the voice id) and `file` (the audio); the reply names the clip's decoded rate and duration: ```bash curl -s http://127.0.0.1:8000/v1/voices -F name=reference -F file=@reference.wav ``` ```text {"name":"reference","sample_rate":16000,"duration_seconds":11.0} ``` Then ask for speech in that voice; the reply body is the wav itself (`audio/wav`, 16-bit PCM at 24 kHz): ```bash curl -s http://127.0.0.1:8000/v1/audio/speech \ -H "Content-Type: application/json" \ -d '{"input": "The install works.", "voice": "reference"}' \ -o hello.wav ``` OpenAI SDK clients use their `audio.speech.create(...)` call against the base URL with `voice` set to a registered name; the server's built-in web page for a TTS model is a speech console rather than a chat window. The round trip with [Transcribe audio](transcribe-audio.md) is the natural smoke test: speak a sentence, transcribe the wav, and compare the text. Sizing lives in [Model requirements](../model-requirements.md). --- # Transcribe audio Speech-to-text with a Whisper model, from the command line and as an OpenAI-compatible HTTP endpoint. Source: https://docs.clika.io/modelverse/how-to/transcribe-audio.md A speech-to-text family provides `transcribe`, and the workflow is the same one the tutorial used for text: resolve the model, run its command, get the payload on stdout. This guide transcribes a file locally, then serves transcription over HTTP. ## From the command line ```bash clikart-cli openai/whisper-large-v3-turbo transcribe meeting.wav ``` ```text Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week. ``` The transcript is the entire stdout payload, ready to redirect into a file. Decoding progress and timing ride stderr. The audio loader decodes WAV, FLAC, MP3 and OGG (Vorbis) and resamples to the model's expected rate, so a file in any of those containers transcribes as is. A container outside that set (OGG Opus, M4A/AAC, WMA) is refused by name, `audio decode: : OGG (Opus) is not supported; this build decodes WAV, FLAC, MP3, OGG (Vorbis)`, with exit code 1; a damaged file, or a file that is not audio at all, is refused the same way, with the file named and the first bytes it actually found quoted back. The web UI is not bound by this set: it decodes a recording or an upload in the browser and sends 16 kHz WAV, so anything the browser plays transcribes. The load knobs from the tutorial apply unchanged: `--device cuda` places the encoder and decoder, `--cache-dir` controls where the snapshot lives, and a local directory works as the source for offline machines. Whisper checkpoints come in sizes from `whisper-tiny` (fits anywhere, fastest, roughest) to `whisper-large-v3` (the accuracy reference); [Model requirements](../model-requirements.md) has the figures. A clip longer than one 30 s window is windowed and transcribed whole. ## Over HTTP The same model serves the OpenAI transcription route: ```bash clikart-cli openai/whisper-large-v3-turbo serve --port 8000 ``` Call it as a multipart upload, the OpenAI shape: ```bash curl -s http://127.0.0.1:8000/v1/audio/transcriptions \ -F file=@meeting.wav \ -F model=whisper ``` ```text {"text":" Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week."} ``` OpenAI SDK clients use their `audio.transcriptions.create(...)` call against the server's base URL, unchanged. The server-side flags and operations story (binding, health, admission control) is the common one in [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md). For concurrent callers, `--concurrent-sessions N` admits N requests at once: Whisper keeps one encoder and mints decoder-only replicas over the same weights, encodes every window on the shared encoder with no lock held, and locks one replica for the decode alone, so N uploads run without a whole-request lock. The library form is `LoadOptions::concurrent_sessions`. For transcription inside your own process, the library's `SttModel` and the transcription engine mount on the same `ServeApi` the `01_serve` example demonstrates ([Additional examples](../examples.md)). --- # Translate text Neural machine translation from the command line and over HTTP, with the model's own language codes. Source: https://docs.clika.io/modelverse/how-to/translate-text.md A translation family provides `translate`, and the workflow is the transcription guide's sibling: resolve the model, run its command, get the payload on stdout. This guide translates sentences locally, then serves translation over HTTP. The model is `facebook/m2m100_418M`, a sequence-to-sequence translation model covering a hundred languages under two-letter codes; its model card is where its license and its language inventory are read, and the runtime prints the license at fetch, at load and in `info`. A chat model that translates (Hy-MT2 among them) answers through its `prompt` command instead, since its family provides no `translate`. ## From the command line ```bash clikart-cli facebook/m2m100_418M translate --to ko \ "The installation is complete." "Pick a model from the catalog." ``` ```text 설치가 완료되었습니다. 카탈로그에서 모델을 선택합니다. ``` Each source text prints as one line, in input order; the payload is only the translations, so redirecting into a file gives one translation per line. Decoding progress and timing ride stderr. The flags: - `--to LANG` (required): the target language code from the model's own inventory; M2M-100 uses two-letter codes (`en`, `ko`, `fr`, `de`, `ja`, `zh`, ...), and the server's `/props` lists them under `languages`. - `--from LANG`: the source language code; unset means the family's default source, which read the English above. - `--max-new-tokens N`: the decode budget in new tokens (`0`, the default, keeps the family default, clamped to the model's window). - The load knobs apply unchanged: `--device`, `--cache-dir`, `--offline`. ## When the budget binds first A translation that stops at its decode budget rather than at the model's end token still exits 0 and still prints its payload; what marks it is one warning on stderr naming the limit that bound it: ```bash clikart-cli facebook/m2m100_418M translate --to ko --max-new-tokens 8 \ "The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon." ``` ```text 회의는 10시에 시작되며 ``` ```text [2026-10-09 15:05:25.113] [m2m-100] [warning] translation stopped at the 8-token budget (--max-new-tokens); output may be incomplete ``` The output above is cut on purpose: an 8-token budget cannot hold the sentence, and the warning is the demonstration. A script that must catch truncation without parsing stderr uses the HTTP route below, where each row carries `finish_reason`. ## Over HTTP The same model serves `POST /v1/translations` ([the route table](serve-openai-compatible.md)): ```bash clikart-cli facebook/m2m100_418M serve --port 8000 ``` ```bash curl -s http://127.0.0.1:8000/v1/translations \ -H "Content-Type: application/json" \ -d '{"texts": ["Hello.", "The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon."], "source_lang": "en", "target_lang": "ko", "max_new_tokens": 8}' ``` ```text {"model":"m2m-100","source_lang":"en","target_lang":"ko","translations":[{"text":"안녕하세요","finish_reason":"stop"},{"text":"회의는 10시에 시작되며","finish_reason":"length"}],"usage":{"prompt_tokens":34,"completion_tokens":10,"total_tokens":44}} ``` Batch through the `texts` array; the response carries one translation per input, in order. Each row's `finish_reason` is the machine-readable form of the budget story above: `stop` means the model produced its end token, `length` means the decode budget cut the translation short (the small `max_new_tokens` above forces one of each). Omit `max_new_tokens` to keep the family default. `source_lang` is optional, as `--from` is. The server-side flags and operations story (binding, health, admission control) is the common one in [Serve a model over the OpenAI and Anthropic APIs](serve-openai-compatible.md). Sizing and quantization work as everywhere else ([Model requirements](../model-requirements.md)); for speech instead of text as the input, [Transcribe audio](transcribe-audio.md) chains into this guide. --- # Model requirements Memory and device figures per model variant, for sizing a deployment before anything downloads. Source: https://docs.clika.io/modelverse/model-requirements.md What a model needs to run is mostly decided before you fetch it: the variant and quantization fix the weight size, and the context length fixes the working memory on top. This page carries the figures for the models the documentation uses and the classes of hardware they fit; `clikart-cli info --dry` gives the on-disk number for any source, and the `bench` command measures the rest on your machine. `--memory-stages ` on any verb writes the runtime's memory-pool counters per device at each stage of the model's life, which is the measured form of the Run column. How to read the tables: **Disk** is the snapshot size (`fetch --dry` total). **Run** is the resident memory serving one session at the default context; longer contexts and concurrent sessions add KV cache on top, roughly linear in tokens. **Fits** names the smallest sensible device class: phone (4 GB), laptop or small board (8 GB), workstation GPU (8 GB VRAM and up). A model's license is not restated here: the model card on the Hugging Face Hub is its source of truth, and the runtime shows it to you where it matters, as `clikart-cli info ` prints the family's license page and the checkpoint's own license link, and the same lines print at fetch and at load. ## Text generation | Model | Variant | Disk | Run | Fits | | --- | --- | --- | --- | --- | | Qwen2.5 0.5B Instruct | fp16 (safetensors) | 0.9 GB | 1.2 GB | phone | | Qwen3 0.6B | fp16 (safetensors) | 1.4 GB | 1.9 GB | phone | | Qwen3 4B | GGUF Q6_K | 3.3 GB | 4.3 GB | laptop | | Qwen3 4B | GGUF Q8_0 | 4.3 GB | 5.4 GB | workstation GPU, 8 GB | The quantization ladder trades memory for fidelity in predictable steps: Q8_0 is near-fp16, Q6_K is the everyday default, Q4_K_M is the last stop before quality degrades noticeably in chat use. `--kv-quant` shrinks the per-token cache the same way when long contexts dominate the budget. ## Multimodal | Model | Variant | Disk | Run | Fits | | --- | --- | --- | --- | --- | | Gemma 4 E2B IT (vision) | fp16 (safetensors) | 9.5 GB | 12.4 GB | workstation GPU, 16 GB | | Gemma 4 E2B IT (vision) | GGUF Q4_0 (QAT) | 4.0 GB | 5.3 GB | laptop | A multimodal run holds the vision tower alongside the language model, and these figures include it: the GGUF row counts the projector file (`mmproj`) that carries the tower beside the weights. ## Speech-to-text | Model | Variant | Disk | Run | Fits | | --- | --- | --- | --- | --- | | Whisper base | fp16 (safetensors) | 0.3 GB | 0.6 GB | phone | | Whisper large-v3-turbo | fp16 (safetensors) | 1.6 GB | 2.2 GB | laptop | | Whisper large-v3 | fp16 (safetensors) | 3.1 GB | 4.0 GB | laptop | Transcription memory is stable with audio length (the model works in 30-second windows); throughput, not memory, is what a faster device buys. ## Embeddings | Model | Variant | Disk | Run | Fits | | --- | --- | --- | --- | --- | | Qwen3 Embedding 0.6B | GGUF Q8_0 | 0.7 GB | 1.1 GB | phone | Embedding models run one forward per input with no generation loop, so batch size is the only memory knob that matters in practice. ## After a model closes Closing a model returns its memory to the runtime's pool, not to the operating system: the pool keeps the bytes cached so the next allocation is cheap, and releases them under memory pressure or when asked. A program that closes one model and loads another therefore holds both until it asks in between: `ClikaRT::device::release_cached_memory(device)` from C++, `clika_runtime.empty_cache(device)` from Python, `Device.releaseCachedMemory()` from Kotlin. The Kotlin Modelverse binding makes the call at a model's close; a C++ or Python program makes it itself, after the close and before the next load, and `memory_stats(device)` shows the pool's cached bytes before and after. ## Reading a figure you do not see here For any other model: `clikart-cli check ` answers the question before anything is downloaded. It reads the model's documents and its file listing and prints one row per device: the weight bytes, the cache bytes per token of context, the cache dtype, the context length, the activation estimate, the total need against the device's memory, whether it fits, and the largest sequence length that fits (`--max-seq N` judges another length; `--assume-device phone=8G` judges a device that is not this machine; `--csv` for a script). `info --dry` prints the exact disk cost, and `devices` prints what this machine offers. When the estimate is close to the device's limit, measure with `bench` before committing the deployment; it reports peak resident memory along with throughput. On a board whose GPU and CPU share one memory (a Jetson, a phone), the device total `check` prices against is the whole machine's memory, which the operating system and every other process share, and the law takes no headroom off; so apply your own margin there, and size against what is free when the model starts rather than the total. `--memory-stages FILE` on any verb writes what a load holds on each device at each stage of the model's life, the measured form of the figure. A second model loaded beside the first takes its memory from the same pool. ```bash clikart-cli check Qwen/Qwen2.5-1.5B-Instruct --max-seq 4096 --assume-device phone=8G ``` For a catalog of models, `clikart-cli info --json` prints one document: each model (the base repository) with its variants, the base weights and every quantization among the sources, read from the hub's documents, listings and GGUF headers with nothing downloaded. Per variant it carries the quant tag as the file names it, the precision, the bits per weight, the files with their sizes and digests, what a runtime must decode to serve it (`requires`), the fit terms `check` prices from, and whether this build serves it. `info --schema` prints the JSON Schema the document validates against; a served model's `/props` carries the same block as `variant`, so a client matches the served weights to a catalog row. ```bash clikart-cli info unsloth/Qwen3-1.7B-GGUF --json clikart-cli info @sources.txt --json # one source per line clikart-cli info --schema ``` --- # Supported models The models the release serves, read off its supported-models document: one row per model with its weight variants and the backends the release's run record covers. Source: https://docs.clika.io/modelverse/supported-models.md The release ships its supported-models document beside the archives: `supported_models.json`, and `supported_models.csv`, the same document flat. It lists 109 checkpoints over 56 model families with 312 weight variants. This page lists the 80 of them whose weights carry a permissive license, the document's own class; the model card is where a checkpoint's license is read, and the runtime prints it at fetch, at load and in `info`. The whole document is a download of the platform with the release ([Get ClikaRT](/clikart/getting-started/get-clikart)). A model is named by its Hugging Face repository id, which `clikart-cli fetch ` and `AutoModel.from_pretrained(id)` take as they are. **Parameters** and **Context** are read off the checkpoint. **Weights** lists the weight variants the release serves for the model, base first, with their format. **Proven on** names the backends the release's own run record covers for the model; an empty cell is a model the release lists without a recorded run. ## Text generation and chat | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [arianraje/qwen3-4b-gdn-hybrid-stage1-align-dtfix](https://huggingface.co/arianraje/qwen3-4b-gdn-hybrid-stage1-align-dtfix) | `qwen3-next` | about 4.5 B | 40,960 | `BF16` (safetensors) | CPU, CUDA, Metal | | [DavidAU/Qwen3-MOE-4x0.6B-2.4B-Writing-Thunder](https://huggingface.co/DavidAU/Qwen3-MOE-4x0.6B-2.4B-Writing-Thunder) | `qwen` | about 1.5 B | 40,960 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [microsoft/phi-2](https://huggingface.co/microsoft/phi-2) | `phi` | about 2.8 B | 2,048 | `F16` (safetensors) | CPU, CUDA, Metal | | [microsoft/Phi-3-mini-4k-instruct](https://huggingface.co/microsoft/Phi-3-mini-4k-instruct) | `phi` | about 3.8 B | 4,096 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct) | `phi` | about 3.8 B | 131,072 | `IQ1_M` (gguf), `IQ1_S` (gguf), `IQ2_XS` (gguf), `IQ3_XS` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q3_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan | | [moonshotai/Kimi-Linear-48B-A3B-Instruct](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct) | `kimi-linear` | about 49.1 B | | `BF16` (safetensors) | CPU | | [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) | `gpt-oss` | about 20.9 B | 131,072 | `U8` (safetensors), `MXFP4` (gguf) | CPU, CUDA, Vulkan | | [OuteAI/Lite-Oute-1-300M](https://huggingface.co/OuteAI/Lite-Oute-1-300M) | `mistral` | about 300 M | 4,096 | `F32` (safetensors) | CPU, CUDA, Metal | | [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) | `qwen` | about 494 M | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) | `qwen` | about 752 M | 40,960 | `BF16` (gguf), `F8_E4M3` (safetensors), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q2_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ1_M` (gguf), `UD-IQ1_S` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_XXS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan | | [Qwen/Qwen3-0.6B-Base](https://huggingface.co/Qwen/Qwen3-0.6B-Base) | `qwen` | about 596 M | 32,768 | `Qwen3-Embedding-0.6B-f16` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan | | [SupraLabs/Supra2-100M-Base](https://huggingface.co/SupraLabs/Supra2-100M-Base) | `qwen` | about 101 M | 2,048 | `F32` (safetensors) | CPU, CUDA, Vulkan | | [zai-org/GLM-4-9B-0414](https://huggingface.co/zai-org/GLM-4-9B-0414) | `glm` | about 9.4 B | 32,768 | `BF16` (safetensors), `IQ2_M` (gguf), `IQ3_M` (gguf), `IQ3_XS` (gguf), `IQ3_XXS` (gguf), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q2_K` (gguf), `Q2_K_L` (gguf), `Q3_K_L` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q3_K_XL` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_L` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_L` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q6_K_L` (gguf), `Q8_0` (gguf), `THUDM_GLM-4-9B-0414-bf16` (gguf) | CPU, CUDA, Vulkan, Metal | | [zai-org/GLM-4.5-Air](https://huggingface.co/zai-org/GLM-4.5-Air) | `glm` | about 110.5 B | 131,072 | `BF16` (safetensors) | CPU | ## Vision-language chat | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) | `deepseek_ocr` | about 3.3 B | 32,768 | `BF16` (safetensors) | CPU, CUDA, Metal | | [deepseek-ai/DeepSeek-OCR](https://huggingface.co/deepseek-ai/DeepSeek-OCR) | `deepseek_ocr` | about 3.3 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal | | [deepseek-ai/DeepSeek-OCR-2](https://huggingface.co/deepseek-ai/DeepSeek-OCR-2) | `deepseek_ocr` | about 3.4 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal | | [deepseek-community/DeepSeek-OCR-2](https://huggingface.co/deepseek-community/DeepSeek-OCR-2) | `deepseek_ocr` | about 3.4 B | 8,192 | `BF16` (safetensors) | CPU, CUDA, Metal | | [florence-community/Florence-2-base](https://huggingface.co/florence-community/Florence-2-base) | `florence2` | about 232 M | 1,024 | `F16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [google/gemma-4-26B-A4B](https://huggingface.co/google/gemma-4-26B-A4B) | `gemma` | about 26.5 B | 262,144 | `BF16` (safetensors) | CPU, CUDA | | [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) | `gemma` | about 25.8 B | 262,144 | `BF16` (safetensors) | | | [google/gemma-4-26B-A4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized) | `gemma` | about 26.5 B | 262,144 | `BF16` (safetensors) | CUDA | | [google/gemma-4-31B](https://huggingface.co/google/gemma-4-31B) | `gemma` | about 32.7 B | 262,144 | `BF16` (safetensors) | CPU, CUDA | | [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) | `gemma` | about 31.3 B | 262,144 | `BF16` (safetensors) | CPU, CUDA, Vulkan | | [google/gemma-4-31B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-unquantized) | `gemma` | about 32.7 B | 262,144 | `BF16` (safetensors), `I32` (safetensors), `gemma-4-31B_q4_0-it` (gguf) | CPU, CUDA, Vulkan | | [meta-models/Muse-Glimmer-30B](https://huggingface.co/meta-models/Muse-Glimmer-30B) | `muse-glimmer` | about 29.8 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan | | [microsoft/Florence-2-base](https://huggingface.co/microsoft/Florence-2-base) | `florence2` | about 232 M | 1,024 | `F16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [Qwen/Qwen2-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen2-VL-2B-Instruct) | `qwen-vl` | about 2.2 B | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [Qwen/Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) | `qwen-vl` | about 2.1 B | 262,144 | `BF16` (safetensors), `F16` (gguf), `Q4_K_M` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal | | [Qwen/Qwen3-VL-30B-A3B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct) | `qwen-vl` | about 31.1 B | 262,144 | `F8_E4M3` (safetensors) | CPU, CUDA, Vulkan | | [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) | `qwen3-5` | about 873 M | 262,144 | `BF16` (safetensors), `IQ4_NL` (gguf), `IQ4_XS` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_0` (gguf), `Q4_1` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_XXS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan, Metal | | [Qwen/Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) | `qwen3-5-moe` | about 36 B | 262,144 | `BF16` (gguf), `I32` (safetensors), `MXFP4_MOE` (gguf), `Q3_K_M` (gguf), `Q3_K_S` (gguf), `Q4_K_M` (gguf), `Q4_K_S` (gguf), `Q5_K_M` (gguf), `Q5_K_S` (gguf), `Q6_K` (gguf), `Q8_0` (gguf), `UD-IQ2_M` (gguf), `UD-IQ2_XXS` (gguf), `UD-IQ3_S` (gguf), `UD-IQ3_XXS` (gguf), `UD-IQ4_NL` (gguf), `UD-IQ4_XS` (gguf), `UD-Q2_K_XL` (gguf), `UD-Q3_K_XL` (gguf), `UD-Q4_K_L` (gguf), `UD-Q4_K_XL` (gguf), `UD-Q5_K_XL` (gguf), `UD-Q6_K_S` (gguf), `UD-Q6_K_XL` (gguf), `UD-Q8_K_XL` (gguf) | CPU, CUDA, Vulkan | | [thinkingmachines/Inkling-Small](https://huggingface.co/thinkingmachines/Inkling-Small) | `inkling` | about 266 B | | `U8` (safetensors) | CPU | ## Multimodal (any to any) | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it) | `gemma` | about 12 B | 262,144 | `BF16` (safetensors) | CPU, CUDA, Vulkan | | [google/gemma-4-12B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized) | `gemma` | about 12 B | 262,144 | `gemma-4-12b-it-qat-q4_0` (gguf) | CPU, CUDA, Vulkan | | [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) | `gemma` | about 5.1 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [google/gemma-4-E2B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-unquantized) | `gemma` | about 5.1 B | 131,072 | `gemma-4-E2B_q4_0-it` (gguf) | CPU, CUDA, Vulkan | | [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) | `gemma` | about 8 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | | [google/gemma-4-E4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized) | `gemma` | about 7.9 B | 131,072 | `gemma-4-E4B_q4_0-it` (gguf) | CPU, CUDA, Vulkan | ## Speech to text | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [nvidia/canary-1b-v2](https://huggingface.co/nvidia/canary-1b-v2) | `canary` | about 979 M | | `F32` (safetensors) | | | [nvidia/parakeet-ctc-0.6b](https://huggingface.co/nvidia/parakeet-ctc-0.6b) | `parakeet` | about 609 M | | `F32` (safetensors) | | | [nvidia/parakeet-rnnt-0.6b](https://huggingface.co/nvidia/parakeet-rnnt-0.6b) | `parakeet` | about 617 M | | `F32` (safetensors) | | | [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | `parakeet` | about 627 M | | `F32` (safetensors) | | | [openai/whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo) | `whisper` | about 809 M | | `whisper-large-v3-turbo-q4_0` (gguf), `whisper-large-v3-turbo-q4_1` (gguf), `whisper-large-v3-turbo-q8_0` (gguf) | CPU, CUDA | | [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny) | `whisper` | about 38 M | 448 | `F32` (safetensors), `F16` (gguf), `Q4_K_M` (gguf), `Q5_K_M` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal | ## Text to speech | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [sesame/csm-1b](https://huggingface.co/sesame/csm-1b) | `csm` | about 1.6 B | 2,048 | `F32` (safetensors) | CPU, CUDA, Metal | ## Translation | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [tencent/Hy-MT2-1.8B](https://huggingface.co/tencent/Hy-MT2-1.8B) | `hunyuan` | about 2 B | 262,144 | `BF16` (safetensors), `Q4_K_M` (gguf), `Q6_K` (gguf), `Q8_0` (gguf) | CPU, CUDA, Vulkan, Metal | | [tencent/Hy-MT2-30B-A3B](https://huggingface.co/tencent/Hy-MT2-30B-A3B) | `hy-v3` | about 30.1 B | 262,144 | `F8_E4M3` (safetensors) | CPU, CUDA, Vulkan | ## Text embedding | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [TaylorAI/bge-micro-v2](https://huggingface.co/TaylorAI/bge-micro-v2) | `bert` | about 17 M | 512 | `F16` (safetensors) | Metal | ## Feature extraction | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) | `qwen-embedding` | about 596 M | 32,768 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Image embedding | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [facebook/dinov2-small](https://huggingface.co/facebook/dinov2-small) | `dinov2` | about 22 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Reranking | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [Qwen/Qwen3-Reranker-0.6B](https://huggingface.co/Qwen/Qwen3-Reranker-0.6B) | `qwen-reranker` | about 596 M | 40,960 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Object detection | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [facebook/detr-resnet-50](https://huggingface.co/facebook/detr-resnet-50) | `detr` | about 42 M | 1,024 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [microsoft/table-transformer-detection](https://huggingface.co/microsoft/table-transformer-detection) | `detr` | about 29 M | 1,024 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [PekingU/rtdetr_r18vd](https://huggingface.co/PekingU/rtdetr_r18vd) | `rt-detr` | about 20 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [PekingU/rtdetr_v2_r18vd](https://huggingface.co/PekingU/rtdetr_v2_r18vd) | `rt-detr` | about 20 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [Roboflow/rf-detr-nano](https://huggingface.co/Roboflow/rf-detr-nano) | `rf-detr` | about 30 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [ustc-community/dfine-nano-coco](https://huggingface.co/ustc-community/dfine-nano-coco) | `d-fine` | about 4 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Open-vocabulary object detection | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [google/owlv2-base-patch16](https://huggingface.co/google/owlv2-base-patch16) | `owlvit` | about 155 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [google/owlvit-base-patch32](https://huggingface.co/google/owlvit-base-patch32) | `owlvit` | about 153 M | 16 | `F32` (safetensors) | CPU, CUDA, Metal | | [IDEA-Research/grounding-dino-tiny](https://huggingface.co/IDEA-Research/grounding-dino-tiny) | `grounding-dino` | about 172 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [openmmlab-community/mm_grounding_dino_tiny_o365v1_goldg](https://huggingface.co/openmmlab-community/mm_grounding_dino_tiny_o365v1_goldg) | `grounding-dino` | about 173 M | 512 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Zero-shot image classification | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [google/siglip-base-patch16-224](https://huggingface.co/google/siglip-base-patch16-224) | `siglip` | about 203 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [google/siglip2-base-patch16-naflex](https://huggingface.co/google/siglip2-base-patch16-naflex) | `siglip` | about 375 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [wkcn/TinyCLIP-ViT-8M-16-Text-3M-YFCC15M](https://huggingface.co/wkcn/TinyCLIP-ViT-8M-16-Text-3M-YFCC15M) | `clip` | about 23 M | 77 | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Image segmentation | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [facebook/detr-resnet-50-panoptic](https://huggingface.co/facebook/detr-resnet-50-panoptic) | `detr-panoptic` | | 1,024 | `torch` (torch) | CPU, CUDA, Vulkan, Metal | | [Roboflow/rf-detr-seg-nano](https://huggingface.co/Roboflow/rf-detr-seg-nano) | `rf-detr-seg` | about 34 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Depth estimation | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [depth-anything/DA3-Small](https://huggingface.co/depth-anything/DA3-Small) | `da3` | about 34 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [depth-anything/Depth-Anything-V2-Small-hf](https://huggingface.co/depth-anything/Depth-Anything-V2-Small-hf) | `depth-anything` | about 25 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | | [Intel/dpt-large](https://huggingface.co/Intel/dpt-large) | `dpt` | about 342 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Image to text | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [zai-org/GLM-OCR](https://huggingface.co/zai-org/GLM-OCR) | `glm` | about 1.3 B | 131,072 | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Text to image | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [Qwen/Qwen-Image](https://huggingface.co/Qwen/Qwen-Image) | `qwenimage` | about 20.4 B | | `BF16` (safetensors) | CPU, CUDA, Vulkan | | [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | `z-image` | about 6.2 B | | `F32` (safetensors) | CPU, CUDA, Vulkan | ## Audio-language chat | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [Qwen/Qwen2-Audio-7B-Instruct](https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct) | `qwen2-audio` | about 8.4 B | | `BF16` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Image to 3D | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [depth-anything/DA3-BASE](https://huggingface.co/depth-anything/DA3-BASE) | `da3` | about 135 M | | `F32` (safetensors) | CPU, CUDA, Vulkan, Metal | ## Text to video | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [Wan-AI/Wan2.1-T2V-1.3B-Diffusers](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) | `wan` | about 1.4 B | | `F32` (safetensors) | CPU, CUDA, Vulkan | | [Wan-AI/Wan2.2-TI2V-5B-Diffusers](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers) | `wan` | about 5 B | | `F32` (safetensors) | CPU, CUDA, Vulkan | ## Other | Model | Family | Parameters | Context | Weights | Proven on | | --- | --- | --- | --- | --- | --- | | [facebook/m2m100_418M](https://huggingface.co/facebook/m2m100_418M) | `m2m-100` | | 1,024 | `torch` (torch) | CPU, CUDA, Vulkan, Metal | | [fromziro/ZeroS-v0.1-150M](https://huggingface.co/fromziro/ZeroS-v0.1-150M) | `qwen3-5` | about 152 M | 2,048 | `F32` (safetensors) | CPU, CUDA, Metal | | [rpatel622/mamba2-130m-hf-Q8_0-GGUF](https://huggingface.co/rpatel622/mamba2-130m-hf-Q8_0-GGUF) | `mamba2` | about 168 M | | `mamba2-130m-hf-q8_0` (gguf), `mamba2-130m-q8_0` (gguf) | CPU, CUDA, Vulkan | --- # Platform The CLIKA Platform runs benchmarks on real devices, keeps a fleet of them under management, serves models from them, and issues the licenses ClikaRT needs. Source: https://docs.clika.io/platform.md The CLIKA Platform is where a team measures models on the hardware they will actually ship on. You register your devices with the platform, point it at a model, and it runs the benchmark on each device and brings the numbers back: accuracy per test, tokens per second, time to first token, peak memory. The same platform keeps those devices under management afterwards, runs models on them as long-lived endpoints, and issues the licenses a ClikaRT runtime needs. One deployment carries three surfaces over the same data: a web application, a command-line tool (`clika-cli`), and an MCP server that lets Claude drive the platform for you in plain language. ## Why the Platform For the ML engineer. A benchmark answers a question about a model on a device, and both halves matter. Pick a model from the Hugging Face Hub or from your project's model list, pick the quality tests to run (MMLU, GSM8K, HumanEval+, IFEval, and the rest of the catalog), pick the devices, and the platform runs every model on every device and lines the results up side by side. You do not choose a runner, a container image, or a script: the platform maps each model and device pair to the job definition that is confirmed to work on that platform and architecture. For the team that owns the devices. A device runs one agent, which dials out to the platform over a single gRPC connection and needs no inbound port. From then on the device reports health every 15 seconds, accepts benchmark jobs and long-running services, exposes a terminal and a file browser, and can be reached through a remote desktop where the platform's providers can start one. Jetsons, x86 workstations, Windows machines, Raspberry Pi class boards and Android phones join the same fleet. For the operator running the deployment. The platform runs on your infrastructure, in your cloud or on your own hardware. An organization holds users, roles and projects; a project holds the devices, models, benchmarks and licenses of one piece of work. API keys carry a subset of their owner's permissions, so an integration gets exactly the access it needs. Every authorization decision is recorded. For whoever ships ClikaRT. A runtime credential issued by the platform is what a ClikaRT runtime presents to prove it may run. The platform issues those credentials per project, in an online form for runtimes that can reach the network and an offline bundle for the ones that cannot, and can revoke or rotate either one. ## What you get - **Devices.** A fleet with live health, hardware inventory, tags and per-device job history. Registration is one install command; the agent does the rest. - **Models.** A model is a reference to a Hugging Face repository plus the task it performs. Registering one is metadata only, so no weights move until a run needs them. The Models page also carries the Modelverse catalog, the models CLIKA has confirmed to run well on ClikaRT, each one click from registered. - **Benchmarks.** A benchmark run fans out over every model and device you picked, and returns one comparable result set: quality scores, throughput and latency, memory, and the per-sample inputs and outputs behind the score. - **Jobs and services.** A job is one benchmark execution on one device, with artifact push, setup, script, teardown and result collection. A service is a process the platform keeps running on a device and restarts when it dies. - **Model serving.** Run a model on your own devices as an OpenAI-compatible endpoint the platform proxies for you. - **Artifacts.** Versioned, checksummed files (datasets, weights, engine bundles, scripts) that the platform pushes to devices and deduplicates on arrival. - **Organizations, projects and access.** Members and roles at the organization, membership and resources per project, scoped API keys for integrations. - **Licensing.** The runtime credentials your ClikaRT projects issue, update and revoke, plus activation of the deployment's own platform license. - **CLI and Claude.** Every documented operation is a CLI command and an MCP tool, so the same workflow runs from a terminal, from CI, or from a conversation with Claude. ## Where to go next - [Getting started](getting-started/index.md): the five-minute mental model, then a tutorial that takes you from signing in to reading your first benchmark result. - [Concepts](concepts/index.md): one page each for organization, project, device, model, artifact, benchmark, job, service and deployment. - [How-to guides](how-to/index.md): job and service definition YAML, members and API keys, license activation, reading results, downloading the ClikaRT SDK. - [Using the CLI](cli/index.md): the platform from a terminal, with every `clika-cli` command, its arguments and its flags. - [Using with AI (MCP)](mcp/index.md): connecting Claude Desktop, Claude Code, Codex and other MCP clients to the platform. --- # Using the CLI clika-cli, the CLIKA Platform command-line client: install it, log in, and drive every platform operation from a terminal or a CI job. Source: https://docs.clika.io/platform/cli.md `clika-cli` is the CLIKA Platform from a terminal. It talks to your deployment's API and does the same things the web application does, so a benchmark you would start by clicking can be started from a shell script or a CI job, a device you would inspect on its page can be queried from your prompt, and an AI assistant can be given the platform as a set of tools. One binary covers the whole platform: devices, benchmarks, jobs, artifacts, models, projects and licenses are all subcommands, each generated from the deployment's own API description, so the CLI and the platform never disagree about what exists. New to it? [Get started with the CLI](get-started.md) takes you from nothing to your first commands in about ten minutes. The rest of this page is the reference: how to install it, how it authenticates, how the commands are shaped, and where each group of commands is documented. `clika-dm` is the sibling binary for a CLIKA Device Management deployment. It is built from the same source and behaves identically; only the set of resources differs. Everything on these pages describes `clika-cli`. ## What you can do with it - **Drive the platform from scripts.** List devices, launch a benchmark, wait for it, and fail the build if a leg did not complete. - **Keep definitions in git.** [`export`](apply-export.md) writes a job or service definition as YAML, [`apply`](apply-export.md) puts it back. The file you commit is the file the platform runs. - **Reach into a device.** Run a command on it, push a file to it, or open an SSH session, without leaving your terminal. - **Give an AI assistant access to the platform.** The [`mcp`](mcp.md) subcommand turns the CLI into a Model Context Protocol server that Claude Code, Cursor and similar tools can call. ## Install Every deployment serves the binaries built from its own commit, so the CLI you install always matches the orchestrator it talks to. The channel lives at `/files/installation/cli/`, where `` is the address of your deployment (for example `https://platform.clika.io`). **The download channel requires authentication.** That is deliberate, because the binaries are part of the deployment rather than a public download. You authenticate with one of three things: | Credential | Where it comes from | Best for | | --- | --- | --- | | Download token | The web app, under **Settings** then **Developer access**. Valid for 15 minutes. | A one-off install on your own machine. | | API key | The web app, same tab. Starts with `clika_` and does not expire until you revoke it. | CI runners and unattended installs. | | Session token | An existing browser session. | Rarely needed by hand. | ### macOS and Linux The **Developer access** settings tab prints a ready-to-paste one-liner with the token already filled in. It looks like this: ``` curl -fsSL "https://platform.clika.io/files/installation/cli/install.sh?token=" | CLIKA_BASE_URL="https://platform.clika.io" CLIKA_DOWNLOAD_TOKEN="" sh ``` The installer detects your operating system and CPU architecture, downloads the matching binary, checks it against the channel's `checksums.txt`, and installs it into `/usr/local/bin` when that is writable and `~/.local/bin` otherwise. Set `CLIKA_INSTALL_DIR` to choose a different directory. With an API key instead of a download token, set `CLIKA_API_KEY` in place of `CLIKA_DOWNLOAD_TOKEN`. ### Windows On Windows the same tab prints a PowerShell one-liner, which runs the installer's PowerShell twin, `install.ps1`: ``` $env:CLIKA_BASE_URL='https://platform.clika.io'; $env:CLIKA_DOWNLOAD_TOKEN=''; irm 'https://platform.clika.io/files/installation/cli/install.ps1?token=' | iex ``` It runs in Windows PowerShell 5.1 and PowerShell 7 and needs no administrator rights. It checks the download against `checksums.txt`, installs `clika-cli.exe` into `%LOCALAPPDATA%\Programs\clika\bin` (`CLIKA_INSTALL_DIR` chooses another directory), adds that directory to your user `PATH`, and puts it on the current session's `PATH` so the CLI runs straight away. On Windows on Arm it installs the native arm64 build. Running it again upgrades the CLI in place. Both installers also take `CLIKA_MCP`, a comma-separated list of AI clients (`claude-code`, `claude-desktop`, `codex`) to register the CLI's MCP server with once it is installed. [Get started with the CLI](get-started.md#register-your-ai-clients-in-the-same-step) shows the commands. ### Downloading one binary by hand Any platform can skip the installer and fetch a single asset. The asset name is `clika-cli--`, with `.exe` appended on Windows: ``` curl -fsSL -H "Authorization: Bearer clika_..." -o clika-cli "https://platform.clika.io/files/installation/cli/clika-cli-linux-amd64" && chmod +x clika-cli ``` The channel publishes these assets: | Operating system | Architecture | Asset | | --- | --- | --- | | Linux | x86_64 | `clika-cli-linux-amd64` | | Linux | ARM64 | `clika-cli-linux-arm64` | | macOS | Apple silicon | `clika-cli-darwin-arm64` | | Windows | x86_64 | `clika-cli-windows-amd64.exe` | | Windows | ARM64 | `clika-cli-windows-arm64.exe` | Alongside them the channel serves `latest.json` (the published version), `checksums.txt` and `signatures.txt`, which is what [`self-update`](self-update.md) reads. ## First login The CLI signs in with an API key: create one in the web app under **Settings** then **Developer access**, then run `login` once. It prompts for the key without echoing it, and saves the deployment address and the key to a profile file that every later command reads: ``` clika-cli --base-url https://platform.clika.io login ``` Email and password sign-in is for the web application only; the CLI never asks for a password. For CI, skip `login` and set `CLIKA_BASE_URL` and `CLIKA_API_KEY` in the environment instead. Then check that it worked: ``` clika-cli devices list ``` Full detail on credential types, profiles and multiple deployments is on the [authentication and profiles](authentication.md) page. ### Where credentials are stored `login` writes `$XDG_CONFIG_HOME/clika-cli/.json`, which on a default Linux or macOS setup is `~/.config/clika-cli/default.json`. The file is created with mode `0600` (readable only by you) and holds one deployment's base URL and credential. Use `--profile ` to keep several deployments side by side. ## Global flags These flags work on every command. Put them anywhere on the command line. | Flag | Type | Default | What it means | | --- | --- | --- | --- | | `--base-url` | string | from the profile | Address of the orchestrator to talk to, for example `https://platform.clika.io`. Reads `CLIKA_BASE_URL` when the flag is absent. | | `--api-key` | string | from the profile | A `clika_` API key, the CLI's only credential. Pass `--api-key=`, or a bare `--api-key` to be prompted without echo; when standard input is piped the value is read from it. Reads `CLIKA_API_KEY`. | | `--profile` | string | `default` | Which saved profile to load credentials from, or save them to. | | `-o`, `--output` | string | `table` | Output format: `table`, `json` or `yaml`. | | `--color` | string | `auto` | When to emit ANSI color: `auto` (only when writing to a terminal), `always`, or `never`. | | `--no-color` | bool | `false` | Turn color off. A non-empty `NO_COLOR` environment variable does the same. | | `--insecure-tls` | bool | `false` | Skip TLS certificate verification. Only for on-premise development stacks that use a throwaway certificate authority. | | `-y`, `--yes` | bool | `false` | Answer confirmation prompts with yes. Required for destructive deletes when standard input is not a terminal. | | `-h`, `--help` | bool | `false` | Print help for the command and exit. | | `-v`, `--version` | bool | `false` | Print the CLI version and exit. | Credentials resolve in this order, highest priority first: 1. The `--api-key` and `--base-url` flags. 2. The `CLIKA_API_KEY` and `CLIKA_BASE_URL` environment variables. 3. The saved profile. ## Output formats Every command that prints a response honors `-o`/`--output`. | Value | What you get | | --- | --- | | `table` (default) | A column table for any list response. Common resources (devices, jobs, artifacts, models, benchmark groups, job definitions, service definitions) have hand-picked columns; every other list derives its columns from the fields the rows actually carry, headed by the JSON field names. A single resource, or a shape the CLI does not recognise, prints as pretty JSON. Timestamps render as ages such as `3m ago`, and byte counts render with units. | | `json` | Pretty-printed JSON of the whole response, pagination envelope included. | | `yaml` | The same document as YAML. | Paginated endpoints wrap their array in `{items, total, page, page_size}`. Table mode unwraps that so you see only the rows; `json` and `yaml` keep it, so a script can still read the page state. An empty list in table mode writes a note such as `No devices found` to standard error and prints nothing to standard output, which keeps pipes clean. In `json` and `yaml` the empty document is printed as it came back. Color marks resource status only, and only when standard output is a terminal, so piped output never carries escape codes: ``` clika-cli devices list clika-cli devices list -o json | jq '.items[].name' clika-cli jobs list --color=always | less -R ``` Generated commands also accept `--raw`, which prints the response body byte for byte with no formatting at all. ## Names instead of UUIDs Most resources have both a UUID and a human name. Anywhere a command takes an id, you can type the name instead: ``` clika-cli devices get jetson-01 clika-cli job-definitions delete llm-latency ``` The CLI lists the collection once and matches on the `name` field. An argument that already looks like a UUID is used directly, with no extra request. Name lookup is enabled for devices, jobs, models, artifacts, job definitions, service definitions and benchmark groups, and names inside an [`apply`](apply-export.md) document resolve the same way. If two resources share a name, the CLI stops and lists the candidate ids rather than picking one, so a script never silently targets the wrong resource. ## Deletes ask first Deleting a device also deletes its job history, so `devices delete` and `devices batch-delete` prompt for confirmation: ``` $ clika-cli devices delete jetson-01 delete device "jetson-01" (1f0a...)? (this also deletes its job history) [y/N]: ``` Pass `-y`/`--yes` to skip the prompt. When standard input is not a terminal (a CI job, a pipeline), the CLI refuses to proceed without `--yes`, so a script has to state destructive intent explicitly. Other deletes accept `--yes` as a no-op, so you can pass it uniformly. Every delete reports what it removed, echoing both the name you typed and the id it resolved to: ``` $ clika-cli job-definitions delete smoke-test deleted job definition "smoke-test" (7a31...) ``` Under `-o json`, `-o yaml` or `--raw` that line moves to standard error and standard output carries the response body, so scripts keep parsing. ## Exit codes | Code | Meaning | | --- | --- | | `0` | The command succeeded. | | `1` | The command failed. The reason is printed to standard error, prefixed with `error:`. A benchmark watch whose legs did not all complete also exits `1`. | | the remote code | `devices exec` and `devices command` propagate the exit code of the command that ran on the device, the way `ssh` does. | | `130` | You interrupted a credential prompt with Ctrl-C. | ## The command tree Running `clika-cli --help` groups the tree into five sections. | Section | Commands | Documented in | | --- | --- | --- | | Getting started | `login`, `logout`, `self-update`, `version` | [Authentication and profiles](authentication.md), [self-update](self-update.md) | | Resources | 60 groups generated from the deployment's API description, one per resource family | The pages listed below | | Declarative | `apply`, `export` | [apply and export](apply-export.md) | | Advanced | `call`, `mcp`, `tools` | [MCP server and generic dispatch](mcp.md) | | Additional | `completion`, `help` | This page | The reference pages are grouped and ordered to follow the [concepts](../concepts/index.md), so a command sits where the thing it acts on is explained: | Group | Pages | | --- | --- | | Organization and project | [authentication and profiles](authentication.md), [organizations, projects and access](projects-and-orgs.md) | | Device | [devices](devices.md), [transfers](transfers.md), [events, metrics and alerts](monitoring.md) | | Benchmark | [benchmarks](benchmarks.md) | | Model deployment | [model deployment](model-deployment.md) | | Runtime licenses | [licensing](licensing.md) | | Artifact | [artifacts and models](artifacts.md) | | Job | [jobs](jobs.md), [job definitions](job-definitions.md) | | Service | [services and service definitions](services.md) | | Tools and automation | [apply and export](apply-export.md), [MCP and generic dispatch](mcp.md), [cloud instances](cloud.md), [keeping the CLI current](self-update.md), [other groups](other-resources.md), [cheat sheet](cheatsheet.md) | The CLI has no platform administration surface: an API key never carries the administrator step-up, so platform administration is done in the admin dashboard of the web app. Because the resource commands are generated, the mapping from the API to the command line is mechanical and worth knowing: | In the API | On the command line | | --- | --- | | A path parameter such as `{id}` | A positional argument, `` or `` | | A query parameter | A flag, `--` | | A request body | `--body ''` or `--body-file ` | | `GET` on a collection | `list` | | `GET` on one resource | `get ` | | `POST` to a collection | `create` | | `PUT` or `PATCH` on one resource | `update ` | | `DELETE` on one resource | `delete ` | | A verb in the path, such as `/restart` | A subcommand of that name | | A name that would collide | The method or the extra path parameter is appended, as in `services-create` or `services-svc_id` | Because the mapping is one to one, every resource subcommand is exactly one platform operation. The full command reference at the foot of each group page names that operation next to the command, with its method, its endpoint and its [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. For any command that takes a body, `--help` prints the body's field reference (name, type, whether it is required, description and permitted values) straight from the API description, plus a worked `--body` example when the body has required fields. ### Built-in prose help The binary carries its own short guides, which are the condensed form of what these pages cover: ``` clika-cli help credentials clika-cli help names clika-cli help output clika-cli help profiles clika-cli help apply clika-cli help self-update ``` ### Shell completion `completion` prints a completion script for `bash`, `zsh`, `fish` or `powershell`: ``` clika-cli completion zsh > "${fpath[1]}/_clika-cli" ``` ## How the CLI relates to the web app, the API and MCP All four are views onto one platform, and none of them can do something the others cannot see: - **The web app** is the interactive surface. It is where you create an organization, invite people, mint API keys and read charts. - **The platform API** is the contract. Every handler in it is annotated, and those annotations are compiled into a machine-readable description of the operations the platform offers. - **The CLI** is generated from that description. A new endpoint becomes a new subcommand with no hand-written mapping, which is why the resource tree is so large and so uniform. - **MCP**, the Model Context Protocol, exposes the same operations as tools an AI assistant can call. The CLI can serve them over standard input and output ([`clika-cli mcp`](mcp.md)), and the deployment serves them over HTTP at `/api/mcp`. Both are built from the same description, so a tool call and the matching CLI command do exactly the same thing. Authorization does not change between surfaces. A CLI command runs as the identity behind your token or API key and passes the same permission checks a browser request would. An API key can additionally be scoped to a subset of your own permissions, in which case it is refused anywhere that subset does not reach, whichever surface it is used from. ## Where to go next - [Authentication and profiles](authentication.md): the three credential types, and addressing several deployments from one machine. - [Cheat sheet](cheatsheet.md): the twenty commands you will use most. - [Devices](devices.md), [benchmarks](benchmarks.md), [jobs](jobs.md), [artifacts](artifacts.md): the day-to-day resource commands. - [Model deployment](model-deployment.md): serving a model from a device. - [Download the ClikaRT SDK](../how-to/download-the-clikart-sdk.md#download-with-the-platform-cli): `runtime-sdk list` and `runtime-sdk download`, the hand-written group that fetches a ClikaRT release archive and verifies its digest. - [apply and export](apply-export.md): the resource YAML round trip. - [MCP server and generic dispatch](mcp.md): give an AI assistant access to the platform. --- # apply and export clika-cli apply and export: the resource YAML round trip that keeps a definition in git identical to the definition the platform runs. Source: https://docs.clika.io/platform/cli/apply-export.md `apply` and `export` are two halves of one idea. `apply` reads a YAML file and creates or updates what it names. `export` fetches a resource and writes the YAML that `apply` accepts back for it. Together they let a definition live in your repository, be reviewed like code, and still be the exact thing the platform runs. This matters because the platform database is a real authoring surface too. Someone can create a job definition in the web app, and that definition is then only in the deployment. A definition edited in the dashboard and never committed disappears when the deployment is torn down; a definition edited in the repository and never applied never takes effect. The round trip is what keeps the two honest. ## `apply` ``` clika-cli apply -f [-f ...] [flags] ``` Reads one or more multi-document YAML files, routes each document by its `kind` field, and sends the resulting requests. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `-f`, `--file` | string, repeatable | none | The YAML file to apply. Repeat it to apply several. `-f -` reads standard input. | | `--dry-run` | bool | `false` | Resolve every reference and print the requests that would be sent, without sending them. | ### Semantics - **Definitions are created or updated by name.** `JobDefinition`, `ServiceDefinition` and `Model` are matched on their `name`, so applying a committed file twice does not create a duplicate. This is what makes `apply` safe to run from CI on every merge. - **Runs always create.** `Job`, `Service`, `Benchmark` and `BenchmarkGroup` are actions, not state, so each apply starts a new one. - **References may be names.** Write `devices: [jetson-01]` rather than a UUID; the same [name resolution](index.md#names-instead-of-uuids) the rest of the CLI uses applies inside the document. - **A `spec:` wrapper is accepted.** If your document nests the body under `spec:`, it is unwrapped. Keys beside `spec:` win over keys inside it. - **Documents apply in order.** Several documents in one file, separated by `---`, are applied top to bottom, so a document may reference something an earlier one created. ### Supported kinds | Kind | Endpoint | Semantics | Exportable | | --- | --- | --- | --- | | `JobDefinition` | `/api/job-definitions` | Create or update by name | Yes | | `ServiceDefinition` | `/api/service-definitions` | Create or update by name | Yes | | `Model` | `/api/models` | Create or update by name | Yes | | `Job` | `POST /api/jobs` | Always creates | Yes | | `Service` | `POST /api/devices/{id}/services` | Always creates, once per target device | No | | `Benchmark`, `BenchmarkGroup` | `POST /api/benchmark-groups` | Always creates | Yes | `Service` is apply-only, because there is no flat collection to fetch one service back from by id. Read a running service through `devices services-svc_id` instead. `clika-cli apply --help` prints the kinds this binary supports, which is the authoritative list for the deployment you are talking to. ### Reference fields Certain fields are resolved from a name to an id before the request goes out. | Field in your document | Looked up in | Sent as | | --- | --- | --- | | `job_definition` | `/api/job-definitions` | `job_definition_id` | | `device` | `/api/devices` | `device_id` | | `devices` (list) | `/api/devices` | `device_ids` | | `model` | `/api/models` | `model_id` | | `models` (list) | `/api/models` | `model_ids` | | `service_definition` | `/api/service-definitions` | The definition's fields, merged in as defaults | The canonical `*_id` fields accept a name too, so either spelling works. One thing that is deliberately **not** implemented: `device_selector`, the old group-based way of picking targets. `apply` reports it as unknown rather than quietly forwarding it, so list the devices explicitly. ### Examples Apply a definition and a run in one file: ```yaml kind: JobDefinition name: llm-latency output_path: /tmp/results.json result_type: structured script: command: - python3 - run.py timeout_sec: 3600 --- kind: Job job_definition: llm-latency devices: - jetson-01 - jetson-02 ``` ``` $ clika-cli apply -f llm-latency.yaml job definition "llm-latency" updated (7a31...) job "llm-latency-jetson-01" created (c4f2...) job "llm-latency-jetson-02" created (c4f3...) ``` Register a model and benchmark it in one file: ```yaml kind: Model huggingface_url: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct task: text-generation --- kind: Benchmark name: nightly-llm-sweep benchmark_type: llm_performance models: - Qwen2.5-0.5B-Instruct devices: - jetson-01 - orin-02 ``` Check what a file would do before it does it: ``` clika-cli apply -f defs.yaml -f run.yaml --dry-run ``` Generate a document and pipe it in: ``` cat run.yaml | clika-cli apply -f - ``` ## `export` ``` clika-cli export ``` Fetches one resource and writes it as canonical resource YAML on standard output. Server-managed fields are stripped: ids, timestamps, audit columns, run state and derived counters, and null fields are omitted. What comes out is what `apply` takes back in. Exportable kinds are `Benchmark`, `BenchmarkGroup`, `Job`, `JobDefinition`, `Model` and `ServiceDefinition`. ``` $ clika-cli export JobDefinition llm-latency kind: JobDefinition name: llm-latency description: Single-stream latency for text generation output_path: /tmp/results.json result_type: structured script: command: - python3 - run.py timeout_sec: 3600 required_resources: min_gpu_count: 1 ``` ``` clika-cli export ServiceDefinition vllm-server clika-cli export Job 1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 ``` ## The round trip Bringing a dashboard-authored definition into the repository: ``` clika-cli export JobDefinition llm-latency > job-defs/llm-latency/config.yaml git add job-defs/llm-latency/config.yaml && git commit -m "chore: commit llm-latency definition" ``` Pushing a repository edit back to a deployment: ``` clika-cli apply -f job-defs/llm-latency/config.yaml ``` Verifying that the two match, which is worth doing in CI: ``` clika-cli export JobDefinition llm-latency | diff -u job-defs/llm-latency/config.yaml - ``` Because `export` strips exactly the fields the platform owns, that `diff` is empty when the definition has not drifted, and shows precisely what changed when it has. ## Applying to several deployments A file is not tied to a deployment, so the same definition can be rolled out everywhere with a profile per target: ``` for p in cloud onprem; do clika-cli --profile "$p" apply -f job-defs/llm-latency/config.yaml; done ``` See [authentication and profiles](authentication.md) for setting those up. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli apply` #### `clika-cli apply` Applies one or more multi-document YAML files. Each document is routed by its `kind:` field; a `spec:` wrapper around the body is accepted and unwrapped. Definitions are created or updated by name, so re-applying a committed file is idempotent; runs (Job, Service, Benchmark) always create. References inside a document may be names rather than UUIDs. ``` clika-cli apply -f [-f ...] [flags] ``` Positional arguments: required ``, ``; optional `[-f ...]`. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--dry-run` | bool | `false` | print the resolved requests instead of sending them | | `-f, --file` | stringArray | none | YAML file to apply (repeatable; '-' reads stdin) | ### `clika-cli export` #### `clika-cli export` Fetches a resource and writes the YAML that `apply` accepts back for it, with server-managed fields (ids, timestamps, run state, counters) stripped. This is the download half of the job-definitions convention: what the platform emits should equal the file committed to the repo. ``` clika-cli export [flags] ``` Positional arguments: required ``, ``. ## Related - [job definitions](job-definitions.md): the field reference for what goes inside these documents. - [Jobs](jobs.md) and [benchmarks](benchmarks.md): what a `Job` or `Benchmark` document dispatches. - [How-to guides](../how-to/index.md): task-shaped walkthroughs for authoring a definition. - [CLI overview](index.md): global flags, output formats and exit codes. --- # Artifacts and models clika-cli artifacts and models: upload files the platform pushes to devices, register models from Hugging Face, and manage versions and retention. Source: https://docs.clika.io/platform/cli/artifacts.md An **artifact** is a file the platform stores for you: a model, a dataset, a script, a container image, a configuration file. Artifacts are what a [job definition](job-definitions.md) pushes to a device before it runs, and what a job leaves behind when it finishes. A **model** is a thinner thing: a registration that points at a model on the Hugging Face Hub, carrying the task metadata the platform needs to decide which benchmarks apply to it. Registering a model does not copy its weights into your artifact library. ## Uploading ### `artifacts upload` ``` clika-cli artifacts upload [flags] ``` Uploads one local file, or one local directory, to the artifact library. This command is hand-written rather than generated, because it streams, so a multi-gigabyte model never has to fit in memory. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--name` | string | the file name | Display name for the artifact. | | `--type` | string | none | What kind of thing this is: `model`, `dataset`, `script`, `docker_image`, `config` or `other`. | | `--version` | string | `v1` | Version label. Uploading the same name with a new label creates a new version in the same lineage. | | `--description` | string | none | Free text. | A **directory** is packed into a `tar.gz` first and registered with `content_format=directory`, so the platform pushes and unpacks it on a device as a single unit rather than as loose files. A **large file** is uploaded in chunks and resumes on its own. A dropped connection, a proxy timeout or a platform redeploy costs you one chunk, not the whole transfer, and re-running an interrupted upload of the same file continues where it stopped. Content the platform already holds is reused without sending it a second time. ``` $ clika-cli artifacts upload ./model.onnx --type model --version v2 uploading model.onnx (412 MB) ... created artifact "model.onnx" (9c14...) version v2 ``` ``` clika-cli artifacts upload ./dataset-dir --name imagenet-val --type dataset clika-cli artifacts upload ./run.py --type script --version v3 ``` ### `artifacts external` ``` clika-cli artifacts external --body '' ``` Registers something the platform pulls itself, rather than something you upload. The pull happens server side. | Body field | Type | Meaning | | --- | --- | --- | | `name` | string, required | Display name. | | `source_url` | string, required | Where to pull from. | | `type` | string, required | `docker_image` or `git_repo`. | | `credential_id` | string | A [registry credential](#registry-credentials) id. Required for a private registry. | | `version` | string | Version label, default `v1`. | | `description` | string | Free text. | | `tags` | object | Key and value metadata. | ``` clika-cli artifacts external --body '{"name":"detector-model","source_url":"registry.example.com/detector:latest","type":"docker_image"}' ``` ### `artifacts dataset` ``` clika-cli artifacts dataset --body '' ``` Registers a dataset artifact with the structure the platform needs to enumerate its entries, which is what makes `artifacts entries` and `artifacts entry` work below. ## Finding artifacts ``` clika-cli artifacts list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--type` | string | all | `model`, `dataset`, `script`, `docker_image`, `config` or `other`. | | `--search` | string | none | Full-text search on the artifact name. | | `--tag` | string, repeatable | none | Include only artifacts carrying this `key:value` tag. Repeating the flag narrows further. | | `--exclude_tag` | string, repeatable | none | Exclude artifacts carrying this `key:value` tag. | | `--latest_only` | string | `false` | Pass `true` to return only the newest version of each lineage. | | `--min_size_bytes` | string | none | Only artifacts at least this large. | | `--max_size_bytes` | string | none | Only artifacts at most this large. | | `--created_after` | string | none | Only artifacts created at or after this RFC 3339 timestamp. | | `--created_before` | string | none | Only artifacts created at or before it. | | `--sort` | string | `created_at` | Sort column: `created_at`, `name`, `size_bytes`, `updated_at` or `type`. | | `--dir` | string | `desc` | Sort direction, `asc` or `desc`. | | `--page` | string | `1` | Page number. | | `--page_size` | string | server default | Rows per page. | ``` $ clika-cli artifacts list --type model --latest_only true NAME TYPE VERSION SIZE CREATED qwen2.5-0.5b.onnx model v2 412 MB 2d ago detector.onnx model v1 12 MB 9d ago ``` ``` clika-cli artifacts list --tag project:vision --exclude_tag status:archived ``` `artifacts get ` prints one artifact in full. ## Downloading ``` clika-cli artifacts download ``` Streams the artifact's stored file to ``. When `` is an existing directory, the server-suggested filename is used inside it; anything else is taken as the file path to write. ``` clika-cli artifacts download imagenet-val ./data/ clika-cli artifacts download imagenet-val ./val.tar.gz ``` ## Versions, tags and expiry | Command | What it does | | --- | --- | | `artifacts versions ` | Every version in this artifact's lineage. | | `artifacts versions-create --body ''` | Register a new version. | | `artifacts update --body ''` | Change the metadata: name, description, type. | | `artifacts tags ` | The artifact's tags. | | `artifacts tags-update --body ''` | Set a tag. | | `artifacts tags-delete-tag ` | Remove one tag. | | `artifacts expiry --body ''` | Set or clear an expiry, after which the platform may reclaim the storage. | | `artifacts retention-policy` | The deployment's retention policy, which is what an unset expiry falls back to. | | `artifacts retention-policy-update --body ''` | Set that policy. | | `artifacts retention-policy-delete` | Clear it, restoring the default. | | `artifacts delete ` | Delete one version, or every version of the artifact with `--all_versions`. | | `artifacts batch-delete --body ''` | Delete several at once. | | `artifacts benchmarks ` | The benchmarks that used this artifact. | ## Dataset entries A dataset artifact is browsable without downloading the whole thing: ``` clika-cli artifacts entries imagenet-val clika-cli artifacts entry imagenet-val --body '' ``` `entries` lists what is inside; `entry` downloads one item out of it. ### Server-side dataset builds Preprocessing a dataset on the platform rather than on your laptop: ``` clika-cli dataset-builds create --body '' clika-cli dataset-builds get ``` `create` starts a build and returns its id; `get` reports its status. The output is a new dataset artifact. ## Chunked upload sessions `artifacts upload` manages resumable uploads for you, and these commands are the underlying session API, exposed for scripts that need to drive it directly: | Command | What it does | | --- | --- | | `artifacts uploads-create --body ''` | Open a chunked upload session. | | `artifacts uploads-chunks ` | Upload one chunk. | | `artifacts uploads-id ` | Session status, including which byte ranges are still missing. | | `artifacts uploads-delete-id ` | Abort the session. | Prefer `artifacts upload` unless you have a specific reason not to. ## Models ``` clika-cli models [command] ``` | Subcommand | Purpose | | --- | --- | | `list` | Registered models. | | `get ` | One model with its task and metadata. | | `preview --body ''` | Look up a Hugging Face model's metadata before registering it. | | `create --body ''` | Register a model. | | `update --body ''` | Change the model's task. | The create body takes `huggingface_url`, either a full URL or the `owner/repo` shorthand, and an optional `task`. Supply `task` only when the Hub has no pipeline tag of its own. It is recorded as a manual choice, and a real Hub tag is never overwritten by it. The task matters because it determines which benchmark types the model is compatible with. ``` clika-cli models create --body '{"huggingface_url":"Qwen/Qwen2.5-0.5B-Instruct"}' clika-cli models list --search qwen ``` Registering a model from a committed file works too: ```yaml kind: Model huggingface_url: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct task: text-generation ``` ## Registry credentials Private container registries and Git remotes need a credential before the platform can pull from them. | Command | What it does | | --- | --- | | `registry-credentials list` | Every credential, with its secret redacted. | | `registry-credentials get ` | One credential. | | `registry-credentials create --body ''` | Store one. | | `registry-credentials update --body ''` | Change one. | | `registry-credentials delete ` | Delete one. | Pass the resulting id as `credential_id` when you register an external artifact. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli artifacts` `artifacts` has 26 subcommands. #### `clika-cli artifacts batch-delete` Batch delete artifacts ``` clika-cli artifacts batch-delete [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--created_after` | string | none | Only artifacts created at or after this RFC 3339 timestamp | | `--created_before` | string | none | Only artifacts created at or before this RFC 3339 timestamp | | `--dry_run` | string | none | Preview only: report what would be deleted without deleting | | `--exclude_tag` | string | none | Exclude artifacts whose tags contain the given 'key:value' (repeatable, ANDed NOT) | | `--limit` | string | `100, max 500` | Maximum artifacts to act on in this call | | `--max_size_bytes` | string | none | Only artifacts at most this many bytes | | `--min_size_bytes` | string | none | Only artifacts at least this many bytes | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Full-text search on artifact name | | `--tag` | string | none | Include artifacts whose tags contain the given 'key:value' (repeatable, ANDed) | | `--type` | string | none | Filter by type (model\|dataset\|script\|docker_image\|config\|other) | Endpoint: `POST /api/artifacts/batch-delete`. MCP tool name: `post_artifacts_batch_delete`. #### `clika-cli artifacts benchmarks` List benchmarks for an artifact ``` clika-cli artifacts benchmarks [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}/benchmarks`. MCP tool name: `get_artifacts_id_benchmarks`. #### `clika-cli artifacts create` Upload artifact ``` clika-cli artifacts create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/artifacts`. MCP tool name: `post_artifacts`. #### `clika-cli artifacts dataset` Upload a dataset artifact ``` clika-cli artifacts dataset [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/artifacts/dataset`. MCP tool name: `post_artifacts_dataset`. #### `clika-cli artifacts delete` Delete artifact ``` clika-cli artifacts delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--all_versions` | string | `false` | Delete every version of the artifact, not only this one | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/artifacts/{id}`. MCP tool name: `delete_artifacts_id`. #### `clika-cli artifacts download` Streams the artifact's stored file to ``. When `` is an existing directory the server-suggested filename is used inside it. ``` clika-cli artifacts download [flags] ``` Positional arguments: required ``, ``. Endpoint: `GET /api/artifacts/{id}/download`. MCP tool name: `get_artifacts_id_download`. #### `clika-cli artifacts entries` List a dataset artifact's entries ``` clika-cli artifacts entries [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | `1` | 1-based page number | | `--per_page` | string | `100, max 1000` | Entries per page | | `--prefix` | string | none | Only entries whose path starts with this prefix | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}/entries`. MCP tool name: `get_artifacts_id_entries`. #### `clika-cli artifacts entry` Download a dataset artifact entry ``` clika-cli artifacts entry [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--path` | string | none | Entry path within the dataset (e.g. entries/clip-0.wav) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}/entry`. MCP tool name: `get_artifacts_id_entry`. #### `clika-cli artifacts expiry` Set or clear artifact expiry ``` clika-cli artifacts expiry [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/artifacts/{id}/expiry`. MCP tool name: `put_artifacts_id_expiry`. #### `clika-cli artifacts external` Create external artifact ``` clika-cli artifacts external [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/artifacts/external`. MCP tool name: `post_artifacts_external`. #### `clika-cli artifacts get` Get artifact ``` clika-cli artifacts get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}`. MCP tool name: `get_artifacts_id`. #### `clika-cli artifacts list` List artifacts ``` clika-cli artifacts list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--created_after` | string | none | Only artifacts created at or after this RFC 3339 timestamp | | `--created_before` | string | none | Only artifacts created at or before this RFC 3339 timestamp | | `--dir` | string | none | Sort direction: desc (default), asc | | `--exclude_tag` | string | none | Exclude artifacts whose tags contain the given 'key:value' (repeatable, ANDed NOT) | | `--latest_only` | string | none | Return only the latest version of each artifact lineage | | `--max_size_bytes` | string | none | Only artifacts at most this many bytes | | `--min_size_bytes` | string | none | Only artifacts at least this many bytes | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Full-text search on artifact name | | `--sort` | string | none | Sort column: created_at (default), name, size_bytes, updated_at, type | | `--tag` | string | none | Include artifacts whose tags contain the given 'key:value' (repeatable, ANDed) | | `--type` | string | none | Filter by type (model\|dataset\|script\|docker_image\|config\|other) | Endpoint: `GET /api/artifacts`. MCP tool name: `get_artifacts`. #### `clika-cli artifacts retention-policy` Get artifact retention policy ``` clika-cli artifacts retention-policy [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/retention-policy`. MCP tool name: `get_artifacts_retention_policy`. #### `clika-cli artifacts retention-policy-delete` Clear artifact retention policy ``` clika-cli artifacts retention-policy-delete [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/artifacts/retention-policy`. MCP tool name: `delete_artifacts_retention_policy`. #### `clika-cli artifacts retention-policy-update` Set artifact retention policy ``` clika-cli artifacts retention-policy-update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/artifacts/retention-policy`. MCP tool name: `put_artifacts_retention_policy`. #### `clika-cli artifacts tags` List artifact tags ``` clika-cli artifacts tags [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}/tags`. MCP tool name: `get_artifacts_id_tags`. #### `clika-cli artifacts tags-delete-tag` Delete artifact tag ``` clika-cli artifacts tags-delete-tag [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/artifacts/{id}/tags/{tag}`. MCP tool name: `delete_artifacts_id_tags_tag`. #### `clika-cli artifacts tags-update` Set artifact tag ``` clika-cli artifacts tags-update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/artifacts/{id}/tags`. MCP tool name: `put_artifacts_id_tags`. #### `clika-cli artifacts update` Update artifact metadata ``` clika-cli artifacts update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/artifacts/{id}`. MCP tool name: `put_artifacts_id`. #### `clika-cli artifacts upload` Uploads a local file to the artifact library. A directory is packed into a tar.gz first and registered with content_format=directory, so the platform pushes and unpacks it as a unit. ``` clika-cli artifacts upload [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--description` | string | none | optional description | | `--name` | string | none | artifact display name (defaults to the file name) | | `--type` | string | none | artifact type: model, dataset, script, docker_image, config, other | | `--version` | string | none | version label (server default: v1) | #### `clika-cli artifacts uploads-chunks` Upload one chunk of an artifact upload session ``` clika-cli artifacts uploads-chunks [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body-file` | string | none | request body (path to the raw file, sent as application/octet-stream) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/artifacts/uploads/{id}/chunks/{index}`. MCP tool name: `put_artifacts_uploads_id_chunks_index`. #### `clika-cli artifacts uploads-create` Create chunked artifact upload session ``` clika-cli artifacts uploads-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/artifacts/uploads`. MCP tool name: `post_artifacts_uploads`. #### `clika-cli artifacts uploads-delete-id` Abort artifact upload session ``` clika-cli artifacts uploads-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/artifacts/uploads/{id}`. MCP tool name: `delete_artifacts_uploads_id`. #### `clika-cli artifacts uploads-id` Get artifact upload session status ``` clika-cli artifacts uploads-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--seen` | string | none | The state the caller last observed; when the current state already differs, the request answers immediately instead of waiting. | | `--wait_s` | string | none | Long-poll: hold the request open up to this many seconds until the state changes (from seen when supplied, else from its value at arrival) or turns terminal, then answer with the current status either way, the response shape is identical to an immediate read, so a plain poll loop keeps working unchanged. Clamped server-side (default maximum 55). Omit or 0 to answer immediately. Chunk arrivals do not end the wait; only a state transition does. | Endpoint: `GET /api/artifacts/uploads/{id}`. MCP tool name: `get_artifacts_uploads_id`. #### `clika-cli artifacts versions` List artifact versions ``` clika-cli artifacts versions [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/artifacts/{id}/versions`. MCP tool name: `get_artifacts_id_versions`. #### `clika-cli artifacts versions-create` Create artifact version ``` clika-cli artifacts versions-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/artifacts/{id}/versions`. MCP tool name: `post_artifacts_id_versions`. ### `clika-cli models` `models` has 6 subcommands. #### `clika-cli models benchmark-groups` List benchmark groups that include a model ``` clika-cli models benchmark-groups [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | `20, max 50` | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/models/{id}/benchmark-groups`. MCP tool name: `get_models_id_benchmark_groups`. #### `clika-cli models benchmarks` List benchmarks run on a model ``` clika-cli models benchmarks [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | `25, max 200` | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--sort` | string | none | Row order: date (newest first, the default), latency (fastest first) or throughput (highest first) | Endpoint: `GET /api/models/{id}/benchmarks`. MCP tool name: `get_models_id_benchmarks`. #### `clika-cli models create` Register model ``` clika-cli models create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/models`. MCP tool name: `post_models`. #### `clika-cli models get` Get model ``` clika-cli models get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/models/{id}`. MCP tool name: `get_models_id`. #### `clika-cli models list` List models ``` clika-cli models list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--dir` | string | none | Sort direction: desc (default), asc | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Case-insensitive name substring filter | | `--sort` | string | none | Sort column: created_at (default), name, updated_at | Endpoint: `GET /api/models`. MCP tool name: `get_models`. #### `clika-cli models preview` Preview HuggingFace model metadata ``` clika-cli models preview [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/models/preview`. MCP tool name: `post_models_preview`. ### `clika-cli modelverse` `modelverse` has 3 subcommands. #### `clika-cli modelverse device-fit` Which catalogue entries fit the caller's devices ``` clika-cli modelverse device-fit [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/modelverse/device-fit`. MCP tool name: `get_modelverse_device_fit`. #### `clika-cli modelverse models` List the Modelverse catalogue ``` clika-cli modelverse models [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | `25, at most 500` | Entries per page | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/modelverse/models`. MCP tool name: `get_modelverse_models`. #### `clika-cli modelverse providers-logo` Get a Modelverse provider's logo ``` clika-cli modelverse providers-logo [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--v` | string | none | Cache version; ignored by the server | Endpoint: `GET /api/modelverse/providers/{id}/logo`. MCP tool name: `get_modelverse_providers_id_logo`. ### `clika-cli dataset-builds` `dataset-builds` has 2 subcommands. #### `clika-cli dataset-builds create` Start a server-side dataset preprocessing build ``` clika-cli dataset-builds create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/dataset-builds`. MCP tool name: `post_dataset_builds`. #### `clika-cli dataset-builds get` Get a dataset build's status ``` clika-cli dataset-builds get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/dataset-builds/{id}`. MCP tool name: `get_dataset_builds_id`. ### `clika-cli registry-credentials` `registry-credentials` has 5 subcommands. #### `clika-cli registry-credentials create` Create registry credential ``` clika-cli registry-credentials create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/registry-credentials`. MCP tool name: `post_registry_credentials`. #### `clika-cli registry-credentials delete` Delete registry credential ``` clika-cli registry-credentials delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/registry-credentials/{id}`. MCP tool name: `delete_registry_credentials_id`. #### `clika-cli registry-credentials get` Get registry credential ``` clika-cli registry-credentials get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/registry-credentials/{id}`. MCP tool name: `get_registry_credentials_id`. #### `clika-cli registry-credentials list` List registry credentials ``` clika-cli registry-credentials list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/registry-credentials`. MCP tool name: `get_registry_credentials`. #### `clika-cli registry-credentials update` Update registry credential ``` clika-cli registry-credentials update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/registry-credentials/{id}`. MCP tool name: `put_registry_credentials_id`. ## Related - [Transfers](transfers.md): following a push of an artifact to a device. - [Devices](devices.md): pushing an artifact to a device, and fetching files back. - [job definitions](job-definitions.md): declaring which artifacts a run needs. - [Benchmarks](benchmarks.md): models and datasets as benchmark inputs. - [Artifact concept](../concepts/artifact.mdx): what an artifact is, external artifacts included, in prose. --- # Authentication and profiles clika-cli login, logout and profiles: the API key the CLI signs in with, where it is stored, and how to address several deployments from one machine. Source: https://docs.clika.io/platform/cli/authentication.md Before the CLI can do anything it needs to know two things: which deployment to talk to, and who you are. This page covers both, the commands that set them (`login` and `logout`), and the API-key commands you use to mint a key from the terminal. ## The credential is an API key The CLI signs in with an API key and nothing else. Email and password sign-in is browser-only: the web application's sign-in carries the authenticator, email-verification and consent steps a terminal cannot complete, so the CLI never asks for a password and never accepts a pasted session token. | Credential | Lifetime | How you get it | Use it for | | --- | --- | --- | --- | | **API key** | Until you revoke it, or its expiry passes | The web app, under **Settings** then **Developer access**, or `auth api-key-create` from a terminal that is already signed in | Everything: your own terminal, CI jobs, automation, MCP hosts | Every credential is sent to the platform as an HTTP `Authorization: Bearer` header. An API key is prefixed `clika_`, is scoped when it is created (it can never do more than its owner may), and is revoked from the same settings page or with `auth api-key-delete-id`. A key never refreshes, because it does not need to. ## `login` Saves an API key and a base URL to a profile file, so later commands need neither flag. Run at a terminal with no flag, it prompts for the key without echoing it; scripts pass the key with `--api-key` or `CLIKA_API_KEY` instead. ``` clika-cli login [flags] ``` ### Flags | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--api-key` | string (global) | prompt | The `clika_` API key to store. Omitted at a terminal, `login` prompts for it without echo; omitted without a terminal and with no `CLIKA_API_KEY`, `login` stops, writes nothing and says where to create a key. A key given as a bare argument is refused, because it would land in shell history. | | `--base-url` | string (global) | none | The deployment to log in to. Required the first time. | | `--profile` | string (global) | `default` | Which profile file to write. | | `--insecure-tls` | bool (global) | `false` | Skip certificate verification, for an on-premise stack with a throwaway certificate authority. | ### Passing the key without leaking it At a terminal, run `login` with no flag and type the key at the prompt; nothing is echoed or recorded. Scripts pass it one of these ways: ``` clika-cli login --api-key clika_abc123... # visible in shell history clika-cli login --api-key=clika_abc123... # same CLIKA_API_KEY=clika_abc123... clika-cli login # from the environment ``` When standard input is piped, the value is read from it, which is the safe form for CI and Kubernetes: ``` echo "$CLIKA_API_KEY_FROM_SECRET_STORE" | clika-cli --base-url https://platform.clika.io login --api-key ``` ### Examples Log in at a terminal, prompted: ``` $ clika-cli --base-url https://platform.clika.io login API key: logged in, profile "default" ``` Log in to an on-premise development stack that uses a self-signed certificate, under its own profile: ``` clika-cli --base-url https://192.168.10.2 --insecure-tls login --profile onprem ``` ## Credential precedence A command resolves its credential and base URL from the first source that has them: 1. The `--api-key` and `--base-url` flags (a bare `--api-key` with piped standard input reads the key from it). 2. The `CLIKA_API_KEY` and `CLIKA_BASE_URL` environment variables. 3. The saved profile named by `--profile`, defaulting to `default`. `login` itself adds one more step after the environment: at a terminal it prompts; without a terminal it stops and names where to create a key. This is what makes a CI job easy: set two environment variables and never run `login` at all. ``` CLIKA_BASE_URL=https://platform.clika.io CLIKA_API_KEY=clika_... clika-cli devices list ``` ## Profiles A profile is one deployment's address plus one API key, saved in a file. It is what lets a single machine address a cloud deployment and an on-premise deployment without retyping anything. `login` writes `$XDG_CONFIG_HOME/clika-cli/.json`, which by default is `~/.config/clika-cli/.json`, with mode `0600` so only your account can read it. Every command accepts `--profile`, and it defaults to `default`. ``` clika-cli --base-url https://platform.clika.io login --profile cloud clika-cli --base-url https://192.168.10.2 --insecure-tls login --profile onprem clika-cli --profile onprem devices list clika-cli --profile cloud benchmarks watch nightly-llm-sweep ``` Because a profile carries the base URL as well as the credential, neither `--base-url` nor a key is needed again after the first login. ## `logout` Removes the credential saved in a profile. ``` clika-cli logout [flags] ``` It deletes the profile's credential file. The API key it held is not a session and stays valid until you revoke it, in the web app or with `auth api-key-delete-id `. A profile written by an earlier CLI version's email and password login still holds a session; that one is revoked on the server first, and the local file is cleared even when the server already considers it expired. ``` clika-cli logout clika-cli logout --profile onprem ``` ## Minting an API key for CI `auth api-key-create` mints a key without leaving the terminal. The response contains the key exactly once, so capture it immediately. The request body must set exactly one of `template` or `scopes`. | Body field | Type | Meaning | | --- | --- | --- | | `name` | string | The label you will see in the web app's key list. | | `template` | string | A ready-made scope: `viewer` (reads only), `worker` (reads plus create, update, cancel and invoke, which is the CI shape), or `admin` (every delegatable permission you yourself hold, deletes included). | | `scopes` | array of strings | An explicit permission allowlist instead of a template, for example `["artifacts:read","jobs:read","jobs:write"]`. Every entry must be a delegatable permission and one you hold yourself. | | `expires_in_days` | integer | How long the key lives. Omit the field for the platform default of 90 days, pass `0` for a key that never expires, or any number of days up to 3650. | A key can never grant more than you have. Requests made with it resolve to the intersection of your own permissions and the key's scope, so narrowing your account later narrows the key with it. Device shell access, tunnels and remote desktop are in no template at all and have to be requested explicitly in `scopes`. ``` $ clika-cli auth api-key-create --body '{"name":"nightly-ci","template":"worker","expires_in_days":365}' { "id": "8f2c...", "name": "nightly-ci", "key": "clika_...", "expires_at": "2027-09-03T00:00:00Z" } ``` `auth api-key-scope-catalog` describes every scope a key may carry, which is the list to read before writing a `scopes` array by hand. List and revoke keys with the neighbouring commands: ``` clika-cli auth api-key clika-cli auth api-key-delete-id 8f2c1d34-5678-90ab-cdef-1234567890ab ``` ## Related commands | Command | What it does | | --- | --- | | `auth me` | Prints the account the current credential belongs to. The quickest way to answer "who am I logged in as". | | `auth me-orgs` | Lists the organizations your account belongs to. | | `auth me-projects` | Lists the projects you can see. | | `auth sessions` | Lists your active sessions, browser sessions included. | | `auth sessions-delete-id ` | Revokes one session by id. | | `auth switch-org`, `auth switch-project` | Move the current session to another organization or project. | | `auth cli-download-token` | Mints the short-lived token that the CLI install channel accepts. Useful for scripting an install. | | `auth change-password` | Changes your own password; the CLI itself never uses the password. | ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli login` #### `clika-cli login` Saves an API key to a profile file for later use by other commands. ``` clika-cli login [flags] ``` ### `clika-cli logout` #### `clika-cli logout` Removes the profile's credentials file. The API key it held is not a session and stays valid until revoked, in the web app, or with `clika-cli auth api-key-delete-id `. ``` clika-cli logout [flags] ``` ### `clika-cli auth` `auth` has 39 subcommands. #### `clika-cli auth accept-invite` Accept invitation ``` clika-cli auth accept-invite [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/accept-invite`. MCP tool name: `post_auth_accept_invite`. #### `clika-cli auth api-key` List API keys ``` clika-cli auth api-key [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/api-key`. MCP tool name: `get_auth_api_key`. #### `clika-cli auth api-key-create` Create API key ``` clika-cli auth api-key-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/api-key`. MCP tool name: `post_auth_api_key`. #### `clika-cli auth api-key-delete-id` Revoke API key ``` clika-cli auth api-key-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/auth/api-key/{id}`. MCP tool name: `delete_auth_api_key_id`. #### `clika-cli auth api-key-rotate` Rotate API key ``` clika-cli auth api-key-rotate [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/api-key/{id}/rotate`. MCP tool name: `post_auth_api_key_id_rotate`. #### `clika-cli auth api-key-scope-catalog` Describe the API-key scope catalogue ``` clika-cli auth api-key-scope-catalog [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/api-key/scope-catalog`. MCP tool name: `get_auth_api_key_scope_catalog`. #### `clika-cli auth capabilities` The caller's effective capabilities ``` clika-cli auth capabilities [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/capabilities`. MCP tool name: `get_auth_capabilities`. #### `clika-cli auth captcha` Captcha requirement for the credential forms ``` clika-cli auth captcha [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/captcha`. MCP tool name: `get_auth_captcha`. #### `clika-cli auth change-password` Change password ``` clika-cli auth change-password [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/change-password`. MCP tool name: `post_auth_change_password`. #### `clika-cli auth check-email` Check whether an email is already registered ``` clika-cli auth check-email [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | A signup code; when it admits the address, signup_enabled is true | | `--email` | string | none | Email address to check | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/check-email`. MCP tool name: `get_auth_check_email`. #### `clika-cli auth cli-download-token` Mint a CLI download token ``` clika-cli auth cli-download-token [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/cli-download-token`. MCP tool name: `post_auth_cli_download_token`. #### `clika-cli auth forgot-password` Request password reset ``` clika-cli auth forgot-password [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/forgot-password`. MCP tool name: `post_auth_forgot_password`. #### `clika-cli auth invite-info` Get invitation info ``` clika-cli auth invite-info [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--token` | string | none | Invitation token from the invite email | Endpoint: `GET /api/auth/invite-info`. MCP tool name: `get_auth_invite_info`. #### `clika-cli auth login` Login ``` clika-cli auth login [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/login`. MCP tool name: `post_auth_login`. #### `clika-cli auth logout` Logout ``` clika-cli auth logout [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/logout`. MCP tool name: `post_auth_logout`. #### `clika-cli auth me` Get current user ``` clika-cli auth me [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/me`. MCP tool name: `get_auth_me`. #### `clika-cli auth me-orgs` List my organizations ``` clika-cli auth me-orgs [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/me/orgs`. MCP tool name: `get_auth_me_orgs`. #### `clika-cli auth me-projects` List my projects ``` clika-cli auth me-projects [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/me/projects`. MCP tool name: `get_auth_me_projects`. #### `clika-cli auth oauth2-callback` OAuth2/OIDC callback (platform-configured provider) ``` clika-cli auth oauth2-callback [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | OAuth2 authorization code | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--redirect_uri` | string | none | Redirect URI to present at code exchange; must equal the one used at init. Defaults to this endpoint's own URL | | `--state` | string | none | OAuth2 state parameter, as returned by the identity provider | Endpoint: `GET /api/auth/oauth2/{provider}/callback`. MCP tool name: `get_auth_oauth2_provider_callback`. #### `clika-cli auth refresh` Refresh access token ``` clika-cli auth refresh [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/refresh`. MCP tool name: `post_auth_refresh`. #### `clika-cli auth resend-verification` Resend verification email ``` clika-cli auth resend-verification [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/resend-verification`. MCP tool name: `post_auth_resend_verification`. #### `clika-cli auth reset-password` Reset password with token ``` clika-cli auth reset-password [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/reset-password`. MCP tool name: `post_auth_reset_password`. #### `clika-cli auth saml-acs` SAML Assertion Consumer Service ``` clika-cli auth saml-acs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/saml/{provider}/acs`. MCP tool name: `post_auth_saml_provider_acs`. #### `clika-cli auth saml-by-id-acs` SAML Assertion Consumer Service (organization provider) ``` clika-cli auth saml-by-id-acs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/saml/by-id/{provider_id}/acs`. MCP tool name: `post_auth_saml_by_id_provider_id_acs`. #### `clika-cli auth saml-by-id-login` Initiate SAML login (organization provider) ``` clika-cli auth saml-by-id-login [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--redirect_url` | string | none | URL to redirect to after successful login | Endpoint: `GET /api/auth/saml/by-id/{provider_id}/login`. MCP tool name: `get_auth_saml_by_id_provider_id_login`. #### `clika-cli auth saml-by-id-metadata` SAML SP metadata (organization provider) ``` clika-cli auth saml-by-id-metadata [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/saml/by-id/{provider_id}/metadata`. MCP tool name: `get_auth_saml_by_id_provider_id_metadata`. #### `clika-cli auth saml-login` Initiate SAML login ``` clika-cli auth saml-login [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--redirect_url` | string | none | URL to redirect to after successful login | Endpoint: `GET /api/auth/saml/{provider}/login`. MCP tool name: `get_auth_saml_provider_login`. #### `clika-cli auth saml-metadata` SAML SP metadata ``` clika-cli auth saml-metadata [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/saml/{provider}/metadata`. MCP tool name: `get_auth_saml_provider_metadata`. #### `clika-cli auth sessions` List active sessions ``` clika-cli auth sessions [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/sessions`. MCP tool name: `get_auth_sessions`. #### `clika-cli auth sessions-delete-id` Revoke session ``` clika-cli auth sessions-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/auth/sessions/{id}`. MCP tool name: `delete_auth_sessions_id`. #### `clika-cli auth signup` Sign up a new user ``` clika-cli auth signup [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/signup`. MCP tool name: `post_auth_signup`. #### `clika-cli auth signup-codes-preview` Preview a signup code ``` clika-cli auth signup-codes-preview [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | The signup code, with or without its dashes | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/signup-codes/preview`. MCP tool name: `get_auth_signup_codes_preview`. #### `clika-cli auth sso-authorize` Initiate SSO login (OAuth2/OIDC) ``` clika-cli auth sso-authorize [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--redirect_url` | string | none | URL to redirect to after successful login | Endpoint: `GET /api/auth/sso/{provider}/authorize`. MCP tool name: `get_auth_sso_provider_authorize`. #### `clika-cli auth sso-callback` OAuth2/OIDC callback ``` clika-cli auth sso-callback [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | OAuth2 authorization code | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--state` | string | none | OAuth2 state parameter | Endpoint: `GET /api/auth/sso/{provider}/callback`. MCP tool name: `get_auth_sso_provider_callback`. #### `clika-cli auth sso-check-email` Check email domain for SSO ``` clika-cli auth sso-check-email [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--email` | string | none | Email address to check | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/sso/check-email`. MCP tool name: `get_auth_sso_check_email`. #### `clika-cli auth sso-providers` List public SSO providers ``` clika-cli auth sso-providers [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/auth/sso/providers`. MCP tool name: `get_auth_sso_providers`. #### `clika-cli auth switch-org` Switch organization ``` clika-cli auth switch-org [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/switch-org`. MCP tool name: `post_auth_switch_org`. #### `clika-cli auth switch-project` Switch project ``` clika-cli auth switch-project [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/switch-project`. MCP tool name: `post_auth_switch_project`. #### `clika-cli auth verify-email` Verify email ``` clika-cli auth verify-email [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/auth/verify-email`. MCP tool name: `post_auth_verify_email`. ## Related - [CLI overview](index.md): install, global flags, output formats and exit codes. - [self-update](self-update.md): keeping the binary in step with the deployment. - [MCP server](mcp.md): the same credentials, used by an AI assistant. --- # Benchmarks clika-cli benchmarks and benchmark-groups: launch a benchmark across devices, watch it to completion, and read the measured results. Source: https://docs.clika.io/platform/cli/benchmarks.md A **benchmark** measures how a model performs on real hardware: throughput, latency, memory, accuracy. You point the platform at one or more models and one or more devices, and it creates a **benchmark group**: a parent record with one **leg** per model and device pair. Each leg is a [job](jobs.md) that runs on its target and reports back. Two command groups cover this: - `benchmarks` is a small hand-written workflow: launch, then `watch` or `results`. It is what you use in a script. - `benchmark-groups` is the generated resource group, with the full create, read, update and delete surface plus sharing. The remaining `benchmark-*` groups are read-only catalogs describing what the platform can measure. ## The workflow ``` clika-cli apply -f benchmark.yaml clika-cli benchmarks watch nightly-llm-sweep clika-cli benchmarks results nightly-llm-sweep -o json ``` Launching through [`apply`](apply-export.md) keeps the definition of the run in a file you can commit. You can equally create the group directly with `benchmark-groups create --body ''`; the two reach the same endpoint. A text-generation model maps to more than one type, so `benchmark_type` is required for an LLM run; the refusal is `400 BENCHMARK_TYPE_REQUIRED`, and it does not list the candidates. `benchmark-types list` prints them (the speed test for an LLM is `llm_performance`), and `benchmark-compatibility` says which can run on a given device. Benchmark jobs do not appear in `jobs list`; read them through `benchmark-groups` and `benchmarks results`. ### `benchmarks watch` ``` clika-cli benchmarks watch [flags] ``` Polls the group's legs, both child jobs on registered devices and hosted-device legs, prints a line to standard error whenever a leg changes state, and then renders the final table on standard output in whatever `-o` format you chose. **It exits non-zero when any leg ends in a state other than `completed`**, which is what makes it usable as a CI gate. Expect silence between state changes: a quick LLM run on a small machine spends about five minutes in `pushing_artifacts` and ten in `running` with no line between. `jobs execution-log ` prints the device's progress in the meantime, and stopping `watch` does not stop the run. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--interval` | duration | `5s` | How often to poll. Accepts Go duration syntax, for example `30s` or `2m`. | ``` $ clika-cli benchmarks watch nightly-llm-sweep jetson-01 / Qwen2.5-0.5B-Instruct: queued -> running orin-02 / Qwen2.5-0.5B-Instruct: queued -> running jetson-01 / Qwen2.5-0.5B-Instruct: running -> completed orin-02 / Qwen2.5-0.5B-Instruct: running -> completed DEVICE MODEL STATUS DURATION jetson-01 Qwen2.5-0.5B-Instruct completed 4m12s orin-02 Qwen2.5-0.5B-Instruct completed 3m48s ``` Gate a pipeline on it: ``` clika-cli apply -f benchmark.yaml && clika-cli benchmarks watch nightly-llm-sweep ``` Poll less often on a long run: ``` clika-cli benchmarks watch nightly-llm-sweep --interval 30s ``` ### `benchmarks results` ``` clika-cli benchmarks results ``` Fetches the same legs together with their measured metrics, at any time, whether or not the run has finished. Unlike `watch`, it does not block and does not set a failure exit code. ``` $ clika-cli benchmarks results nightly-llm-sweep DEVICE MODEL STATUS TOKENS/S TTFT PEAK MEM jetson-01 Qwen2.5-0.5B-Instruct completed 142.6 210ms 1.8 GB orin-02 Qwen2.5-0.5B-Instruct completed 118.3 265ms 1.8 GB ``` ``` clika-cli benchmarks results nightly-llm-sweep -o json | jq '.[] | {device, tokens_per_second}' ``` ## Creating a group directly ``` clika-cli benchmark-groups create --body '' ``` The body says what to run, on what, and where. Two constraints are enforced by the platform: at least one of `device_ids` or `hosted_device_arns` must be set, and exactly one of `model_ids` or `huggingface_url`. | Body field | Type | Meaning | | --- | --- | --- | | `name` | string | Display name. Generated for you when omitted, but naming it is what lets you say `benchmarks watch nightly-llm-sweep` later. | | `model_ids` | array of strings | One or more registered model ids. Mutually exclusive with `huggingface_url`. | | `huggingface_url` | string | A convenience that registers and benchmarks a single Hugging Face model in one step. Mutually exclusive with `model_ids`. | | `device_ids` | array of strings | Registered devices to run on. Each must be online and benchmark capable. | | `hosted_device_arns` | array of strings | Rented phones from `hosted-devices list` instead of, or alongside, your own devices. 1 to 5 identifiers. One hosted run is created per selected model and linked to the group; its results appear marked `execution: "hosted"`. | | `hosted_backend` | string | Which hosted cloud the identifiers belong to, `device_farm` or `test_lab`. Only needed when the deployment has more than one configured. | | `hosted_android_version` | string | The Android version a hosted run targets. Required by a backend that schedules a device model plus an OS version; ignored by one whose identifier already encodes it. | | `benchmark_type` | string | Which benchmark to run. Required only when the platform cannot infer it, which happens when the model's task maps to more than one compatible type. | | `job_definition_id` | string | A power-user override that pins the [job definition](job-definitions.md) used. Normally omitted: the platform maps the benchmark type and the device to the right definition through its compatibility matrix. Applies to registered devices only. | | `force` | boolean | Skip the device resource compatibility check. Use it when you know a device can take the model and the check disagrees. | | `transmit_detailed_outputs` | boolean | Whether each leg sends back the per-sample input, model output and expected answer that power the detailed-outputs view. Defaults to true; set it to false to keep results lightweight. | ``` clika-cli benchmark-groups create --body '{ "name": "nightly-llm-sweep", "huggingface_url": "https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct", "device_ids": ["1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8"] }' ``` The same run as a YAML file you can commit, applied with [`apply`](apply-export.md), which additionally lets you name devices and models instead of pasting UUIDs: ```yaml kind: Benchmark name: nightly-llm-sweep benchmark_type: llm_performance models: - Qwen2.5-0.5B-Instruct devices: - jetson-01 - orin-02 ``` ## Reading and managing a group | Command | What it does | | --- | --- | | `benchmark-groups list` | Every group, newest first. Takes `--page` and `--page_size`. | | `benchmark-groups get ` | One group with its status summary. | | `benchmark-groups jobs ` | The group's legs as jobs, which is what `benchmarks watch` polls. | | `benchmark-groups results ` | The group's per-leg results, the same data `benchmarks results` renders. | | `benchmark-groups update --body '{"name":"..."}'` | Rename the group. | | `benchmark-groups delete ` | Delete the group. | ### Sharing results A benchmark group can be published behind an opaque link so someone without an account can read it. | Command | What it does | | --- | --- | | `benchmark-groups share-mark ` | Create the share link. Only one live link exists per group. | | `benchmark-groups share ` | Show the current link. | | `benchmark-groups share-card ` | Upload the preview card image shown when the link is unfurled. | | `benchmark-groups share-mark-share_id ` | Revoke the share link. | | `public benchmark-shares ` | The read side of a share link, which is what an anonymous reader hits. | | `public benchmark-shares-card.png ` | The card image behind that link. | ## The benchmark catalogs These read-only lists describe what this deployment can measure. They are worth checking before you build a request by hand, because they tell you the exact slugs the API expects. | Command | What it lists | | --- | --- | | `benchmark-types list` | The benchmark types available, by slug. This is where a `benchmark_type` value comes from. | | `benchmark-categories list` | How those types are grouped for display. | | `benchmark-metrics list` | Every metric the platform can report, with its unit and meaning. | | `benchmark-io-schemas list` | The input and output schemas benchmark results conform to. | | `benchmark-task-categories list` | The map from a model task to its benchmark category. | | `benchmark-compatibility list --model_id ` | Which benchmark types are compatible with one model, and on what hardware. Run this when a group creation is refused as incompatible. | | `model-recommendations list` | Models the platform suggests for a given target. | | `huggingface-tasks list` | The Hugging Face task names the platform understands, which is what maps a model to its compatible benchmark types. | ## Comparing runs ``` clika-cli comparisons list clika-cli comparisons create --body '' clika-cli comparisons get clika-cli comparisons delete ``` A comparison is a saved side-by-side of several benchmark results, the thing you keep when you want to show that a change helped. ## Rescoring Some benchmarks are scored on the server after the run. If the scoring logic changed, or a scorer failed, you can score an existing job again without re-running it on the device: ``` clika-cli jobs rescore ``` ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli benchmarks` `benchmarks` has 2 subcommands. #### `clika-cli benchmarks results` Fetches the group's legs, child jobs and cloud hosted device legs, with their measured metrics. ``` clika-cli benchmarks results [flags] ``` Positional arguments: required ``. #### `clika-cli benchmarks watch` Polls a benchmark group's legs, child jobs on registered devices and cloud hosted device legs, and prints a line whenever one changes state, then renders the final table. Exits non-zero when any leg ends in a state other than completed, so it gates CI. ``` clika-cli benchmarks watch [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--interval` | duration | `5s` | poll interval | ### `clika-cli benchmark-groups` `benchmark-groups` has 13 subcommands. #### `clika-cli benchmark-groups cancel` Cancel benchmark group ``` clika-cli benchmark-groups cancel [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/benchmark-groups/{id}/cancel`. MCP tool name: `post_benchmark_groups_id_cancel`. #### `clika-cli benchmark-groups create` Create benchmark group ``` clika-cli benchmark-groups create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/benchmark-groups`. MCP tool name: `post_benchmark_groups`. #### `clika-cli benchmark-groups delete` Delete benchmark group ``` clika-cli benchmark-groups delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--delete_outputs` | string | none | Also delete the artifacts the child jobs' pipelines collected (each job's output file and execution log). Default false: the outputs stay in the artifact library, tagged with their job id, for the record. | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/benchmark-groups/{id}`. MCP tool name: `delete_benchmark_groups_id`. #### `clika-cli benchmark-groups get` Get benchmark group ``` clika-cli benchmark-groups get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups/{id}`. MCP tool name: `get_benchmark_groups_id`. #### `clika-cli benchmark-groups jobs` List benchmark group jobs ``` clika-cli benchmark-groups jobs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups/{id}/jobs`. MCP tool name: `get_benchmark_groups_id_jobs`. #### `clika-cli benchmark-groups list` List benchmark groups ``` clika-cli benchmark-groups list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups`. MCP tool name: `get_benchmark_groups`. #### `clika-cli benchmark-groups outputs` List benchmark group output artifacts ``` clika-cli benchmark-groups outputs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups/{id}/outputs`. MCP tool name: `get_benchmark_groups_id_outputs`. #### `clika-cli benchmark-groups results` List benchmark group results ``` clika-cli benchmark-groups results [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups/{id}/results`. MCP tool name: `get_benchmark_groups_id_results`. #### `clika-cli benchmark-groups share` Get benchmark share link ``` clika-cli benchmark-groups share [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-groups/{id}/share`. MCP tool name: `get_benchmark_groups_id_share`. #### `clika-cli benchmark-groups share-card` Upload benchmark share card ``` clika-cli benchmark-groups share-card [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/benchmark-groups/{id}/share/card`. MCP tool name: `put_benchmark_groups_id_share_card`. #### `clika-cli benchmark-groups share-mark` Create benchmark share link ``` clika-cli benchmark-groups share-mark [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/benchmark-groups/{id}/share`. MCP tool name: `post_benchmark_groups_id_share`. #### `clika-cli benchmark-groups share-mark-share_id` Revoke benchmark share link ``` clika-cli benchmark-groups share-mark-share_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/benchmark-groups/{id}/share/{share_id}`. MCP tool name: `delete_benchmark_groups_id_share_share_id`. #### `clika-cli benchmark-groups update` Rename benchmark group ``` clika-cli benchmark-groups update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PATCH /api/benchmark-groups/{id}`. MCP tool name: `patch_benchmark_groups_id`. ### `clika-cli benchmark-categories` `benchmark-categories` has 1 subcommands. #### `clika-cli benchmark-categories list` List benchmark categories ``` clika-cli benchmark-categories list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-categories`. MCP tool name: `get_benchmark_categories`. ### `clika-cli benchmark-compatibility` `benchmark-compatibility` has 2 subcommands. #### `clika-cli benchmark-compatibility batch` Get benchmark compatibility for several models over chosen devices ``` clika-cli benchmark-compatibility batch [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/benchmark-compatibility/batch`. MCP tool name: `post_benchmark_compatibility_batch`. #### `clika-cli benchmark-compatibility list` Get benchmark compatibility for a model ``` clika-cli benchmark-compatibility list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--model_id` | string | none | Model ID (UUID) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-compatibility`. MCP tool name: `get_benchmark_compatibility`. ### `clika-cli benchmark-io-schemas` `benchmark-io-schemas` has 1 subcommands. #### `clika-cli benchmark-io-schemas list` List benchmark I/O schemas ``` clika-cli benchmark-io-schemas list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-io-schemas`. MCP tool name: `get_benchmark_io_schemas`. ### `clika-cli benchmark-metrics` `benchmark-metrics` has 1 subcommands. #### `clika-cli benchmark-metrics list` List benchmark metrics ``` clika-cli benchmark-metrics list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-metrics`. MCP tool name: `get_benchmark_metrics`. ### `clika-cli benchmark-task-categories` `benchmark-task-categories` has 1 subcommands. #### `clika-cli benchmark-task-categories list` List benchmark task-category map ``` clika-cli benchmark-task-categories list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-task-categories`. MCP tool name: `get_benchmark_task_categories`. ### `clika-cli benchmark-types` `benchmark-types` has 1 subcommands. #### `clika-cli benchmark-types list` List benchmark types ``` clika-cli benchmark-types list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/benchmark-types`. MCP tool name: `get_benchmark_types`. ### `clika-cli comparisons` `comparisons` has 4 subcommands. #### `clika-cli comparisons create` Create job comparison ``` clika-cli comparisons create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/comparisons`. MCP tool name: `post_comparisons`. #### `clika-cli comparisons delete` Delete comparison ``` clika-cli comparisons delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/comparisons/{id}`. MCP tool name: `delete_comparisons_id`. #### `clika-cli comparisons get` Get comparison ``` clika-cli comparisons get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/comparisons/{id}`. MCP tool name: `get_comparisons_id`. #### `clika-cli comparisons list` List comparisons ``` clika-cli comparisons list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/comparisons`. MCP tool name: `get_comparisons`. ### `clika-cli huggingface-tasks` `huggingface-tasks` has 1 subcommands. #### `clika-cli huggingface-tasks list` List HuggingFace pipeline tasks ``` clika-cli huggingface-tasks list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/huggingface-tasks`. MCP tool name: `get_huggingface_tasks`. ### `clika-cli model-recommendations` `model-recommendations` has 1 subcommands. #### `clika-cli model-recommendations list` Model recommendations for a benchmark type ``` clika-cli model-recommendations list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--benchmark_type` | string | none | Benchmark type slug | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/model-recommendations`. MCP tool name: `get_model_recommendations`. ## Related - [Jobs](jobs.md): a benchmark leg is a job, and the job commands read its logs, outputs and per-sample results. - [Devices](devices.md): choosing and preparing the hardware to benchmark on. - [apply and export](apply-export.md): keeping a benchmark definition in git. - [Artifacts](artifacts.md): the models and datasets a benchmark consumes. - [Benchmark concept](../concepts/benchmark.mdx): what the platform measures, and the datasets it ships. --- # Cheat sheet The twenty clika-cli commands worth memorising, in one table. Source: https://docs.clika.io/platform/cli/cheatsheet.md Twenty commands cover most of what anyone does with `clika-cli`. Every one of them is documented in full on the page linked in the last column. | # | Command | What it does | | --- | --- | --- | | 1 | `clika-cli --base-url https://platform.clika.io login --api-key` | Log in once and save the deployment and credential to a profile. The bare `--api-key` prompts, so the secret never reaches your shell history. [More](authentication.md#login) | | 2 | `clika-cli whoami` | Show the platform, the key prefix, the principal and the capability count of the current credential (`auth me` takes a user session, which the CLI does not hold). [More](authentication.md#related-commands) | | 3 | `clika-cli --profile onprem devices list` | Run any command against a second deployment, using its saved profile. [More](authentication.md#profiles) | | 4 | `clika-cli devices list --status online` | See which devices are up right now. [More](devices.md#devices-list) | | 5 | `clika-cli devices get jetson-01` | Inspect one device in full, by name rather than UUID. [More](devices.md#devices-get) | | 6 | `clika-cli devices exec jetson-01 'nvidia-smi'` | Run a command on a device and get its output and exit code back. [More](devices.md#devices-exec) | | 7 | `clika-cli devices push jetson-01 ./run.sh /opt/clika/run.sh --permissions 0755` | Copy a local file onto a device. [More](devices.md#devices-push) | | 8 | `clika-cli devices files-fetch jetson-01 /var/log/agent.log ./agent.log` | Pull a file back off a device, checksum verified. [More](devices.md#devices-files-fetch) | | 9 | `clika-cli devices connect jetson-01` | Open an SSH session to a device without looking up its address. [More](devices.md#devices-connect) | | 10 | `clika-cli apply -f benchmark.yaml` | Create or update whatever a YAML file names, idempotently for definitions. [More](apply-export.md#apply) | | 11 | `clika-cli benchmarks watch nightly-llm-sweep` | Follow a benchmark to completion and exit non-zero if any leg failed. The CI gate. [More](benchmarks.md#benchmarks-watch) | | 12 | `clika-cli benchmarks results nightly-llm-sweep -o json` | Read a benchmark group's measured metrics, in a form a script can parse. [More](benchmarks.md#benchmarks-results) | | 13 | `clika-cli jobs list --status failed` | Find the jobs that went wrong. [More](jobs.md#jobs-list) | | 14 | `clika-cli jobs execution-log ` | Read what actually happened on the device during a job. [More](jobs.md#getting-results-back) | | 15 | `clika-cli artifacts upload ./model.onnx --type model --version v2` | Put a file in the artifact library, resumably, however large it is. [More](artifacts.md#artifacts-upload) | | 16 | `clika-cli artifacts download imagenet-val ./data/` | Stream an artifact back out to a local path. [More](artifacts.md#downloading) | | 17 | `clika-cli export JobDefinition llm-latency > config.yaml` | Bring a definition authored in the dashboard into your repository. [More](apply-export.md#export) | | 18 | `clika-cli events tail --types job_completed` | Follow the platform's durable event feed live, resumable from a cursor. [More](monitoring.md#events-tail) | | 19 | `clika-cli mcp --toolset device-ops` | Serve the platform to an AI assistant as a curated set of tools. [More](mcp.md#mcp) | | 20 | `clika-cli self-update` | Pull the CLI build this deployment publishes, checksum and signature verified. [More](self-update.md#self-update) | ## Flags worth knowing | Flag | Effect | | --- | --- | | `-o json` | JSON instead of a table, pagination envelope included, for piping into `jq`. | | `--profile ` | Address a different deployment. | | `-y` | Answer confirmation prompts with yes. Required for a destructive delete in a script. | | `--insecure-tls` | Skip certificate verification, for an on-premise stack with a throwaway certificate authority. | | `--dry-run` | On `apply`, print the requests instead of sending them. | | `--raw` | On a generated command, print the response body byte for byte. | ## When you do not know the command ``` clika-cli --help clika-cli devices --help clika-cli tools --tag devices clika-cli help credentials ``` The first two walk the tree, `tools` lists every operation by tool name, and `help ` prints the built-in prose guides. See the [complete command index](other-resources.md#complete-command-index) for every top-level command and the page that documents it. --- # Cloud instances clika-cli cloud-instances, cloud-credentials, cloud-images and vpn: renting a machine in a cloud account and enrolling it as a device. Source: https://docs.clika.io/platform/cli/cloud.md Not every device has to be hardware you own. The platform can provision a machine in **your** cloud account, install the device agent on it, and register it as an ordinary [device](devices.md). From that point on it behaves like any other target: it takes jobs, runs benchmarks, and appears in `devices list`. Three groups cover this, in the order you use them: a credential for your cloud account, an image catalogue, and the instances themselves. `vpn` is the network side, for when an instance has to reach a private network. ## Before the first instance A cloud instance needs a stored cloud credential, and the web application is where one is added: **Settings**, then **Credentials**, choosing the provider (Azure, Google Cloud or AWS). For Azure the form takes a Tenant ID, a Client ID and Client Secret (a service principal), the Subscription ID and a Default Location, and its **Required permissions** panel names the roles the principal needs on the subscription. The platform creates and uses a resource group named `clika-cloud-`; there is no field for it. **Test** on the stored credential answers "credential validated with provider" in a few seconds. Then, on the web, **Devices**, the **Cloud helpers** tab (which shows no control until a credential exists), **Add Cloud Device**: pick the GPU class, press **Find options**, choose a type and region, and **Create**. Pick a type with more memory than the quick benchmark peaks at (a 135M-parameter model peaked above 1 GB), set the region to the credential's location yourself, and expect the first instance to create a VPN gateway that the platform keeps, and bills, until you delete it. The hourly rate shows in the CLI's JSON rather than on the configure step. ## Cloud credentials The platform needs permission to create machines in your account. That permission is stored as a cloud credential. | Command | What it does | | --- | --- | | `cloud-credentials list` | Stored credentials, with their secrets redacted. | | `cloud-credentials get ` | One credential. | | `cloud-credentials create --body ''` | Store one. | | `cloud-credentials update --body ''` | Change one. | | `cloud-credentials test ` | Check that it still works, before you find out during a provision. | | `cloud-credentials delete ` | Remove one. Organization owners and admins may remove any of the organization's credentials; a credential stored for your own user is yours to remove. | Run `cloud-credentials test` after any rotation on the cloud side. It is the difference between a clear "this credential is stale" and a provision that fails halfway. ## Images ``` clika-cli cloud-images refresh ``` The platform caches the machine images available in each region. `refresh` re-reads them from the provider, which you need after a new image is published and before it can be selected. ## Instances ``` clika-cli cloud-instances [command] ``` | Subcommand | Purpose | | --- | --- | | `catalog` | The instance types, regions and images you can provision, given your stored credentials. Read this before writing a create body. | | `list` | Instances the platform has provisioned. | | `get ` | One instance with its current state. | | `create --body ''` | Provision an instance. | | `start ` | Start a stopped instance. | | `stop ` | Stop it without destroying it, which usually stops the compute charge while keeping the disk. | | `terminate ` | Destroy it. | | `delete ` | Remove the platform's record of it. | | `retry ` | Retry a provision that failed partway. | | `refresh-post` | Re-read every instance's state from the provider. | | `refresh-post-id ` | Re-read one instance's state. | A typical first run: ``` clika-cli cloud-credentials test 5b2e0a1c-3d4e-5f60-7182-93a4b5c6d7e8 clika-cli cloud-instances catalog clika-cli cloud-instances create --body '' clika-cli cloud-instances list ``` Provisioning is asynchronous. `create` returns as soon as the request is accepted, and the instance moves through its states afterwards. Watch it with `cloud-instances get `, or with the [event feed](monitoring.md#events), which reports the device coming online once the agent registers. `stop`, `terminate` and `delete` are three different things, and the distinction costs money if you get it wrong. `stop` keeps the instance and its disk. `terminate` destroys the instance at the provider. `delete` removes the platform's record, which is what you want only after the instance is already gone. An instance in status `error` still accepts `stop`, which stops the VM's billing while you decide; to be rid of it, `terminate` it, then `delete` the record once it is gone. The web page of a failed instance offers only **Retry** and **Delete**, so stop and terminate are CLI commands. ## VPN ``` clika-cli vpn list clika-cli vpn setup --body '' clika-cli vpn verify clika-cli vpn delete ``` A VPN gateway lets provisioned instances, or devices on a private network, reach the platform and be reached by it without exposing anything publicly. `vpn verify` checks the tunnel end to end, which is worth running before you conclude a device is broken. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli cloud-credentials` `cloud-credentials` has 6 subcommands. #### `clika-cli cloud-credentials create` Create cloud credential ``` clika-cli cloud-credentials create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-credentials`. MCP tool name: `post_cloud_credentials`. #### `clika-cli cloud-credentials delete` Delete cloud credential ``` clika-cli cloud-credentials delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/cloud-credentials/{id}`. MCP tool name: `delete_cloud_credentials_id`. #### `clika-cli cloud-credentials get` Get cloud credential ``` clika-cli cloud-credentials get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/cloud-credentials/{id}`. MCP tool name: `get_cloud_credentials_id`. #### `clika-cli cloud-credentials list` List cloud credentials ``` clika-cli cloud-credentials list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/cloud-credentials`. MCP tool name: `get_cloud_credentials`. #### `clika-cli cloud-credentials test` Test cloud credential ``` clika-cli cloud-credentials test [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-credentials/{id}/test`. MCP tool name: `post_cloud_credentials_id_test`. #### `clika-cli cloud-credentials update` Update cloud credential ``` clika-cli cloud-credentials update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/cloud-credentials/{id}`. MCP tool name: `put_cloud_credentials_id`. ### `clika-cli cloud-images` `cloud-images` has 1 subcommands. #### `clika-cli cloud-images refresh` Refresh cloud image cache ``` clika-cli cloud-images refresh [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-images/refresh`. MCP tool name: `post_cloud_images_refresh`. ### `clika-cli cloud-instances` `cloud-instances` has 14 subcommands. #### `clika-cli cloud-instances catalog` Get cloud instance catalog ``` clika-cli cloud-instances catalog [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--credential_id` | string | none | Cloud credential ID to use for API calls | | `--provider` | string | none | Cloud provider: aws, gcp or azure | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--region` | string | none | Region to list instance types and images for | | `--zone` | string | none | Zone to list instance and accelerator types for | Endpoint: `GET /api/cloud-instances/catalog`. MCP tool name: `get_cloud_instances_catalog`. #### `clika-cli cloud-instances catalog-refresh` Refresh the cloud catalog cache ``` clika-cli cloud-instances catalog-refresh [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/catalog/refresh`. MCP tool name: `post_cloud_instances_catalog_refresh`. #### `clika-cli cloud-instances create` Provision a cloud instance ``` clika-cli cloud-instances create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances`. MCP tool name: `post_cloud_instances`. #### `clika-cli cloud-instances delete` Delete cloud instance ``` clika-cli cloud-instances delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/cloud-instances/{id}`. MCP tool name: `delete_cloud_instances_id`. #### `clika-cli cloud-instances facets` List the accelerators and instance shapes the cloud catalog cache can offer, for a cascading form ``` clika-cli cloud-instances facets [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--backend` | string | none | Backend the engine must run on: cuda, cuda:12.4, none, rocm, decides the refusal reason and which CUDA versions are listed | | `--cuda` | string | none | Minimum CUDA version wanted (the older spelling of backend=cuda:``) | | `--gpu` | string | none | The accelerator, as the accelerator list spells it (t4, a100 80gb, rtx pro 6000) or none / cpu for no GPU; absent answers the accelerator list | | `--gpu_count` | string | none | Keep the shapes that carry at least this many GPUs | | `--min_memory_gb` | string | none | Minimum memory in GB | | `--min_vcpu` | string | none | Minimum vCPUs | | `--provider` | string | none | Keep one provider's shapes only: aws, gcp or azure | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--region` | string | none | A region id keeps that region alone; any other text is a case-insensitive part of a region id or name | Endpoint: `GET /api/cloud-instances/facets`. MCP tool name: `get_cloud_instances_facets`. #### `clika-cli cloud-instances get` Get cloud instance ``` clika-cli cloud-instances get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/cloud-instances/{id}`. MCP tool name: `get_cloud_instances_id`. #### `clika-cli cloud-instances list` List cloud instances ``` clika-cli cloud-instances list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--status` | string | none | Filter by status (pending, provisioning, installing_agent, registered, running, terminating, terminated, error) | Endpoint: `GET /api/cloud-instances`. MCP tool name: `get_cloud_instances`. #### `clika-cli cloud-instances options` Find cloud instance options from a plain description ``` clika-cli cloud-instances options [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--backend` | string | none | Backend the engine must run on: cuda, cuda:12.4, cuda:13, none (CPU engine, nothing installed for the GPU), rocm (refused with a reason until the engine has it). Default: any, CUDA on an NVIDIA GPU | | `--cuda` | string | none | Minimum CUDA version wanted, e.g. 12.4 or 13 (the older spelling of backend=cuda:``; may be combined with backend=cuda, must not disagree with it) | | `--gpu` | string | none | GPU wanted, in plain words: L4, nvidia t4, A100 80GB, RTX PRO 6000, or none / cpu for a device without a GPU. A name the platform does not recognise answers 200 with an empty list and a warning saying so, never 400 | | `--gpu_count` | string | `1 when a GPU is named` | How many GPUs | | `--instance_type` | string | none | Keep one exact instance type only (case-insensitive), how a form resolves the shape a person picked into the option it launches | | `--min_memory_gb` | string | none | Minimum memory in GB | | `--min_vcpu` | string | none | Minimum vCPUs | | `--provider` | string | none | Keep one provider's machines only: aws, gcp or azure | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--region` | string | none | Region preference: a region id keeps that region alone; any other text is a case-insensitive part of a region id or name (europe, korea) | Endpoint: `GET /api/cloud-instances/options`. MCP tool name: `get_cloud_instances_options`. #### `clika-cli cloud-instances refresh-post` Refresh all cloud instance statuses ``` clika-cli cloud-instances refresh-post [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/refresh`. MCP tool name: `post_cloud_instances_refresh`. #### `clika-cli cloud-instances refresh-post-id` Refresh single cloud instance status ``` clika-cli cloud-instances refresh-post-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/{id}/refresh`. MCP tool name: `post_cloud_instances_id_refresh`. #### `clika-cli cloud-instances retry` Retry failed cloud instance ``` clika-cli cloud-instances retry [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/{id}/retry`. MCP tool name: `post_cloud_instances_id_retry`. #### `clika-cli cloud-instances start` Start cloud instance ``` clika-cli cloud-instances start [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/{id}/start`. MCP tool name: `post_cloud_instances_id_start`. #### `clika-cli cloud-instances stop` Stop cloud instance ``` clika-cli cloud-instances stop [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/{id}/stop`. MCP tool name: `post_cloud_instances_id_stop`. #### `clika-cli cloud-instances terminate` Terminate cloud instance ``` clika-cli cloud-instances terminate [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/cloud-instances/{id}/terminate`. MCP tool name: `post_cloud_instances_id_terminate`. ### `clika-cli vpn` `vpn` has 4 subcommands. #### `clika-cli vpn delete` Delete VPN gateway ``` clika-cli vpn delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/vpn/{id}`. MCP tool name: `delete_vpn_id`. #### `clika-cli vpn list` List VPN configs ``` clika-cli vpn list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/vpn`. MCP tool name: `get_vpn`. #### `clika-cli vpn setup` Setup VPN gateway ``` clika-cli vpn setup [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/vpn/setup`. MCP tool name: `post_vpn_setup`. #### `clika-cli vpn verify` Verify VPN connectivity ``` clika-cli vpn verify [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/vpn/{id}/verify`. MCP tool name: `post_vpn_id_verify`. ## Related - [Devices](devices.md): what a provisioned instance becomes once its agent registers. - [Events, metrics and alerts](monitoring.md): watching a provision complete. - [Organizations, projects and access](projects-and-orgs.md): who may store a cloud credential. --- # Devices clika-cli devices: find the machines registered with your platform, run commands on them, move files, and manage their desired state. Source: https://docs.clika.io/platform/cli/devices.md A **device** is a machine registered with your platform: a Jetson on a bench, a phone in a drawer, a workstation, a cloud instance. Each one runs the CLIKA device agent, which keeps a connection open to the orchestrator so the platform can send it work and read its state back. The `devices` group is how you inspect and drive them from a terminal. `clika-cli devices --help` lists 59 subcommands. Most are generated one-for-one from the API, and the table at the end of this page names every one. The sections below cover the ones you will actually type. ## Registering a device from the CLI The web dialog's registration exists as two commands here: mint an enrollment token (the body needs a `name`), then print the install script for the device's operating system and run it on the device. `install list --raw` prints the whole script, so on the device itself pipe it to `sh`. ``` clika-cli enrollment-tokens create --body '{"name":"my-laptop","max_uses":1}' ``` ``` clika-cli install list --os linux --token --raw | sh ``` The device is listed within a minute of the installer finishing. If the platform refuses it, the installer prints `==> ERROR: The platform refused this device` with the reason and exits 1: a used token (`TOKEN_EXHAUSTED`), a device name already registered on another identity (set `DEVICE_NAME=` in front of the command and run it again), or the plan's device limit. `enrollment-codes create` mints a short activation code instead of a token, the kind the web dialog's command carries; see [Register a device by code](../how-to/register-a-device-by-code.md). ## Finding a device ### `devices list` ``` clika-cli devices list [flags] ``` Lists the devices your project can see, newest registration first. Every filter is a flag, and filters combine with AND unless noted. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--status` | string | all | `online` or `offline`. A device is online when the orchestrator has heard from its agent recently. | | `--platform` | string | all | `linux`, `macos`, `windows` or `android`. | | `--search` | string | none | Free text matched against the device name, its group name, or any of its tags. | | `--term` | string, repeatable | none | A committed free-text filter. Each term must match on its own, so terms narrow each other. | | `--tag` | string, repeatable | none | Filter by tag. Tags that share a key (the text before the first `=`) are OR-ed together; different keys are AND-ed. A tag with no `=` narrows on its own. | | `--search_or` | bool | `false` | Widen instead of narrow: OR the `--search` text against the tag and term filters rather than AND-ing it. Ignored when there are no tag or term filters. | | `--group` | string | none | Filter by group name. | | `--accelerator` | string | none | Filter by accelerator type. | | `--gpu_model` | string | none | Filter by GPU model. | | `--min_gpu_vram` | string | none | Minimum GPU video memory, in bytes. | | `--min_memory` | string | none | Minimum total system memory, in bytes. | | `--attestation_type` | string | all | The key storage the device proved it has: `tpm_attested`, `keystore_attested`, `file_backed` or `unverified`. | | `--dirty` | string | all | `true` or `false`. A dirty device has local changes the platform did not make, and holds queued jobs until you clear the flag. | | `--summary` | string | `false` | Pass `true` to return only `{id, name, status}` per device. Much cheaper on a large fleet, because the health snapshot, network interfaces, capabilities, resources and tags are all skipped. | | `--sort` | string | `registered_at` | Sort column: `registered_at`, `name`, `status`, `platform` or `last_heartbeat`. | | `--dir` | string | `desc` | Sort direction, `asc` or `desc`. | | `--page` | string | `1` | Page number. | | `--page_size` | string | server default | Rows per page. | | `--raw` | bool | `false` | Print the response body byte for byte. | ``` $ clika-cli devices list --status online NAME STATUS PLATFORM LAST HEARTBEAT TAGS jetson-01 online linux 12s ago site=lab,role=bench orin-02 online linux 8s ago site=lab mac-mini-1 online macos 31s ago site=office ``` ``` clika-cli devices list --tag site=lab --tag role=bench -o json | jq -r '.items[].name' ``` ### `devices get` ``` clika-cli devices get ``` Prints one device in full: its identity, platform, hardware inventory, capabilities, tags, network interfaces and the latest health snapshot. The argument accepts the device name, so you rarely need its UUID. ``` clika-cli devices get jetson-01 clika-cli devices get jetson-01 -o yaml ``` ## Working on a device These four commands are hand-written rather than generated, because each of them does something a plain API call cannot: propagate an exit code, stream a file, or hand your terminal over to `ssh`. ### `devices exec` ``` clika-cli devices exec [flags] ``` Runs one shell command on the device through its agent. Remote standard output goes to your standard output and remote standard error to your standard error, and the CLI exits with the remote exit code, so it composes in a script exactly like `ssh` does. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--working-dir` | string | agent default | Directory to run the command in, on the device. | | `--shell` | string | platform default | Shell binary to run the command with. The agent picks a sensible one per platform when empty. | | `--timeout` | int | server default | Execution timeout in seconds. The server caps it at 300. | | `--detach` | bool | `false` | Start the command and print its id instead of waiting for it. | ``` $ clika-cli devices exec jetson-01 'nvidia-smi --query-gpu=name,memory.used --format=csv' name, memory.used [MiB] Orin, 1834 MiB ``` ``` clika-cli devices exec jetson-01 'ls -la' --working-dir /opt/clika clika-cli devices exec jetson-01 'sleep 5; echo done' --timeout 30 ``` For anything longer than a couple of minutes, detach and pick the result up later. The command keeps running on the device whether or not the CLI is still attached, so a dropped connection no longer loses the transcript: ``` $ clika-cli devices exec jetson-01 'make -j all' --detach 5f0c1a2b-3c4d-5e6f-7081-92a3b4c5d6e7 ``` ### `devices command` ``` clika-cli devices command [flags] ``` Reads one command back out of the device's history: its status (`pending`, `running`, `exited` or `lost`), its exit code, and the stored transcript. Once the command has exited, this exits with the remote exit code too. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--wait` | string | not set | Block until the command reaches a terminal state. Bare `--wait` waits indefinitely; `--wait=5m` or `--wait=300` bounds it. | | `--tail` | int | `0` | Print only the last N lines of each stream. | | `--stdout-offset` | int | `0` | Byte offset to resume reading standard output from. The offsets reached are printed to standard error, so a poller can pass them back. | | `--stderr-offset` | int | `0` | The same, for standard error. | ``` clika-cli devices command jetson-01 5f0c1a2b-3c4d-5e6f-7081-92a3b4c5d6e7 --wait clika-cli devices command jetson-01 5f0c1a2b-3c4d-5e6f-7081-92a3b4c5d6e7 --tail 50 ``` `devices commands ` lists the history rather than one entry. ### `devices push` ``` clika-cli devices push [flags] ``` Uploads one local file to an absolute path on the device. A `` ending in `/` keeps the local file name. The upload streams, so a multi-gigabyte file never has to fit in memory. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--permissions` | string | agent default | Octal file mode to set on the device, for example `0755`. | | `--mkdir-parents` | bool | `false` | Create missing parent directories on the device. | ``` clika-cli devices push jetson-01 ./run.sh /opt/clika/run.sh --permissions 0755 clika-cli devices push jetson-01 ./config.yaml /etc/clika/ --mkdir-parents ``` Pushing a file the platform already stores is a different command. `devices files-push-from-library` sends an existing [artifact](artifacts.md) to the device without it passing through your machine at all. ### `devices files-fetch` ``` clika-cli devices files-fetch [dest] [flags] ``` Downloads a file or a directory from the device. With no `[dest]` the file lands in the current directory under its remote name. The download is verified end to end. A single file is checked against the SHA-256 the device computed while reading its own disk, echoed on success so you can keep it; a `--recursive` directory is checked against the checksum of the archive the server assembled. An older agent that declares no checksum completes the fetch unverified rather than failing it. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--recursive` | bool | `false` | Fetch a directory, packed as a `tar.gz`. | | `--tail-bytes` | int | `0` | Print only the last N bytes to standard output, instead of writing a file. Useful for peeking at a log. | ``` clika-cli devices files-fetch jetson-01 /var/log/agent.log ./agent.log clika-cli devices files-fetch jetson-01 /data/results/run-42 --recursive clika-cli devices files-fetch jetson-01 /var/log/agent.log --tail-bytes 4096 ``` `devices files --path /some/dir` lists a directory without downloading anything. ### `devices connect` ``` clika-cli devices connect [flags] ``` Fetches the device's SSH details from the orchestrator and then replaces the CLI process with the `ssh` command it suggests, so your terminal behaves exactly as if you had typed that `ssh` yourself. On Windows, which has no process-replacing exec, the `ssh` client is launched as a child process instead. This needs an `ssh` client on your `PATH` and direct network reachability to the device. The orchestrator reports whether it believes you are on the same network, but it cannot route for you. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--dry-run` | bool | `false` | Print the `ssh` command instead of running it. | ``` $ clika-cli devices connect jetson-01 --dry-run ssh -p 22 clika@192.168.10.24 ``` ## Desired state The platform keeps a **desired state** per device: the configuration, the jobs and the services it should be running. The agent reconciles the device towards it and reports what actually happened. That is why a change here is a request, not an instruction that takes effect the moment the command returns. | Command | What it does | | --- | --- | | `devices desired-state ` | The whole desired state in one document. | | `devices desired-state-config ` | The desired agent configuration on its own. | | `devices desired-state-config-update --body ''` | Set the desired agent configuration. | | `devices desired-state-jobs ` | The jobs the device should be running. | | `devices desired-state-jobs-update-job_name --body ''` | Add or change one desired job. | | `devices desired-state-jobs-delete-job_name ` | Remove one desired job. | | `devices desired-state-services ` | The services the device should be running. | | `devices desired-state-services-update-svc_id --body ''` | Set one desired service state. | | `devices desired-state-services-delete-svc_id ` | Remove one desired service state. | | `devices reconcile ` | Ask the agent to reconcile now instead of at its next cycle. | | `devices reconciliation-log ` | What the agent did on its recent reconciliation passes, and why. | When a device has been changed outside the platform, the orchestrator marks it **dirty** and holds its queued jobs rather than running them against an unknown state. Inspect it, then release the queue: ``` clika-cli devices list --dirty true clika-cli devices clear-dirty jetson-01 ``` ## Health, power and hardware | Command | What it does | | --- | --- | | `devices health ` | The latest health snapshot: reachability, load, temperature, disk and memory as the agent last reported them. | | `devices health-history ` | The same, over time. | | `devices health-probe ` | Ask for a fresh probe now rather than waiting for the next report. | | `devices resources ` | Refresh the stored hardware inventory (CPU, memory, accelerators). | | `devices power --body '{"policy":"prevent_sleep"}'` | Set the power policy. `prevent_sleep` keeps a laptop or phone awake for long runs; an empty policy restores the default. | | `devices job-queue ` | The jobs queued for this device. | | `devices job-queue-delete ` | Clear the queue. | | `devices benchmarks ` | The benchmarks that have run on this device. | ## Services on a device A **service** is a long-running process the platform keeps alive on a device, as opposed to a job, which runs and finishes. See [job definitions](job-definitions.md) for how one is described. | Command | What it does | | --- | --- | | `devices services ` | List the services on the device. | | `devices services-create --body ''` | Start a service. | | `devices services-svc_id ` | Get one service. | | `devices services-restart ` | Restart it. | | `devices services-stop ` | Stop it. | | `devices services-delete-svc_id ` | Remove it. | | `devices services-update --body ''` | Update the artifacts it runs. | | `devices services-artifacts ` | Whether its artifacts have arrived on the device. | | `devices services-events ` | Its lifecycle events. | ## Agent updates and certificates Each device agent can be updated from the platform, and each one holds a client certificate that identifies it. | Command | What it does | | --- | --- | | `devices update-create ` | Trigger an agent update on one device. | | `devices batch-update` | Trigger an agent update on many devices at once. | | `devices update-history ` | Past agent updates and their outcome. | | `devices update-policy ` | The device's update policy. | | `devices update-policy-update --body ''` | Set that policy. | | `devices cert-status ` | The device certificate's validity and expiry. | | `devices revoke-cert ` | Revoke it. The device cannot reconnect until it re-enrolls. | | `devices ssh-info ` | The SSH connection details `devices connect` uses. | | `devices ssh-keys --body ''` | Replace the authorized keys the agent installs. | | `devices engine-package ` | Provision the inference engine package onto the device. | ## Tags and tag scripts Tags are `key=value` labels you attach to devices and then filter on. A **tag script** is a script the platform runs automatically on every device carrying a given tag, which is how a fleet stays consistent without anyone logging in. | Command | What it does | | --- | --- | | `devices tags` | Every tag in use across the fleet. | | `devices tags-starred` | The tags pinned to the top of the web app's filters. | | `devices tags-starred-update --body ''` | Replace the starred set. | | `devices run-tag-scripts ` | Run this device's tag scripts now. | | `devices tag-script-runs ` | The device's tag-script run history. | | `tag-scripts list` | All tag scripts. | | `tag-scripts get ` | One tag script. | | `tag-scripts create --body ''` | Create one. | | `tag-scripts update --body ''` | Change one. | | `tag-scripts toggle ` | Enable or disable one without deleting it. | | `tag-scripts delete ` | Delete one. | ## Enrolling a new device A device joins the platform by presenting an **enrollment token**: a single credential that authorizes it to register and receive its own certificate. You mint the token, the device uses it once, and the certificate takes over from there. | Command | What it does | | --- | --- | | `enrollment-tokens create --body ''` | Mint a token. | | `enrollment-tokens list` | Every token, used and unused. | | `enrollment-tokens get ` | One token. | | `enrollment-tokens qr ` | The QR code form, which is how a phone enrolls. | | `enrollment-tokens revoke ` | Revoke it. The row is kept with a `revoked_at` timestamp rather than deleted, so the audit trail survives. | ## Hosted devices Not every target has to be a machine you own. **Hosted devices** are phones and boards rented from a cloud device farm for the length of a run. | Command | What it does | | --- | --- | | `hosted-devices backends` | Which hosted-device backends this deployment has enabled. | | `hosted-devices list` | The hosted fleet available to you. | | `hosted-devices runs` | Hosted-device runs, past and present. | | `hosted-devices runs-create --body ''` | Start a benchmark run on hosted devices. | | `hosted-devices runs-id ` | One run. | | `hosted-devices runs-stop ` | Stop a run early. | | `hosted-devices runs-delete-id ` | Delete a run record. | | `hosted-devices runnable-types` | The benchmark types hosted devices can run. | | `hosted-devices legs-telemetry` | A hosted leg's self-reported telemetry. | A benchmark group can mix both kinds of target, and `benchmarks watch` follows registered-device legs and hosted-device legs alike. See [benchmarks](benchmarks.md). ## Deleting devices ``` clika-cli devices delete clika-cli devices batch-delete --body '' ``` Deleting a device also deletes its job history, so both commands prompt for confirmation. Pass `-y` to skip the prompt; in a script the CLI refuses to proceed without it. ## Subcommand index Every subcommand, as a quick index. The full detail for each one, with its flags and the API operation it dispatches, is in the [full command reference](#full-command-reference) below. | Subcommand | Purpose | | --- | --- | | `batch-delete` | Delete several devices at once. | | `batch-update` | Update the agent on several devices at once. | | `benchmarks` | Benchmarks that ran on this device. | | `cert-status` | Device certificate validity. | | `clear-dirty` | Clear the dirty flag and release queued jobs. | | `command` | Read one recorded command: status, exit code, transcript. | | `commands` | List the device's command history. | | `commands-cid` | Get one command by id, generated form of `command`. | | `connect` | Open an SSH session to the device. | | `delete` | Delete the device and its job history. | | `desired-state` | Full desired state. | | `desired-state-config` | Desired agent configuration. | | `desired-state-config-update` | Set desired agent configuration. | | `desired-state-jobs` | Desired jobs. | | `desired-state-jobs-delete-job_name` | Remove one desired job. | | `desired-state-jobs-update-job_name` | Set one desired job. | | `desired-state-services` | Desired services. | | `desired-state-services-delete-svc_id` | Remove one desired service. | | `desired-state-services-update-svc_id` | Set one desired service. | | `engine-package` | Provision the engine package. | | `exec` | Run a shell command on the device. | | `files` | List a directory on the device. | | `files-fetch` | Download a file or directory from the device. | | `files-push` | Generated file upload, form fields as flags. | | `files-push-from-library` | Push a stored artifact to the device. | | `get` | Get one device. | | `health` | Latest health snapshot. | | `health-history` | Health over time. | | `health-probe` | Probe health now. | | `job-queue` | Jobs queued for the device. | | `job-queue-delete` | Clear that queue. | | `list` | List devices. | | `power` | Set the power policy. | | `push` | Upload a local file to the device. | | `reconcile` | Force a reconciliation pass. | | `reconciliation-log` | Reconciliation history. | | `resources` | Refresh the hardware inventory. | | `revoke-cert` | Revoke the device certificate. | | `run-tag-scripts` | Run the device's tag scripts. | | `services` | Services on the device. | | `services-artifacts` | Artifact status for one service. | | `services-create` | Start a service. | | `services-delete-svc_id` | Delete a service. | | `services-events` | Service lifecycle events. | | `services-restart` | Restart a service. | | `services-stop` | Stop a service. | | `services-svc_id` | Get one service. | | `services-update` | Update a service's artifacts. | | `ssh-info` | SSH connection details. | | `ssh-keys` | Replace authorized keys. | | `tag-script-runs` | Tag-script runs for this device. | | `tags` | All device tags. | | `tags-starred` | Starred tags. | | `tags-starred-update` | Replace starred tags. | | `update-create` | Trigger an agent update. | | `update-history` | Agent update history. | | `update-policy` | Get the update policy. | | `update-policy-update` | Set the update policy. | | `update-put` | Update device metadata such as its name. | ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli devices` `devices` has 67 subcommands. #### `clika-cli devices batch-delete` Batch delete devices ``` clika-cli devices batch-delete [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices batch-update` Batch update agent on multiple devices ``` clika-cli devices batch-update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices benchmark-groups` List benchmark groups that ran on a device ``` clika-cli devices benchmark-groups [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | `25, max 200` | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices benchmarks` List benchmarks run on a device ``` clika-cli devices benchmarks [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices cert-status` Get device cert status ``` clika-cli devices cert-status [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices clear-dirty` Clear device dirty state ``` clika-cli devices clear-dirty [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices command` Reads one command from the device's history: status (pending|running|exited|lost), exit code, and the stored transcript. Remote stdout goes to stdout and stderr to stderr, and once the command has exited this exits with the remote exit code. ``` clika-cli devices command [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--stderr-offset` | int | none | byte offset to read stderr from (from a previous read) | | `--stdout-offset` | int | none | byte offset to read stdout from (from a previous read) | | `--tail` | int | none | print only the last N lines of each stream | #### `clika-cli devices commands` List device command history ``` clika-cli devices commands [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page (default: 20, max: 500) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices commands-cid` Get one device command ``` clika-cli devices commands-cid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--seen` | string | none | Last observed status; the wait also ends as soon as the current status differs | | `--stderr_offset` | string | `0` | Byte offset to read stderr from | | `--stdout_offset` | string | `0` | Byte offset to read stdout from | | `--tail_lines` | string | none | Return only the last N lines of each stream (mutually exclusive with offsets) | | `--wait_s` | string | none | Long-poll up to this many seconds for the command to reach a terminal state (clamped to the deployment ceiling) | #### `clika-cli devices connect` Fetches the device's SSH info from the orchestrator and executes the suggested ssh command, replacing this process. Requires an ssh client on PATH and network reachability to the device, the orchestrator reports whether it believes you are on the same network. ``` clika-cli devices connect [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--dry-run` | bool | `false` | print the ssh command instead of running it | #### `clika-cli devices delete` Delete device ``` clika-cli devices delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state` Get device desired state ``` clika-cli devices desired-state [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-config` Get desired config for device ``` clika-cli devices desired-state-config [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-config-update` Set desired config for device ``` clika-cli devices desired-state-config-update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-jobs` Get desired jobs for device ``` clika-cli devices desired-state-jobs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-jobs-delete-job_name` Remove desired job for device ``` clika-cli devices desired-state-jobs-delete-job_name [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-jobs-update-job_name` Set desired job for device ``` clika-cli devices desired-state-jobs-update-job_name [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-services` List desired services for device ``` clika-cli devices desired-state-services [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-services-delete-svc_id` Remove desired service state ``` clika-cli devices desired-state-services-delete-svc_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices desired-state-services-update-svc_id` Set desired service state ``` clika-cli devices desired-state-services-update-svc_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices engine-package` Provision the engine package onto a device ``` clika-cli devices engine-package [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices exec` Runs a command on the device via its agent and prints the result. Remote stdout goes to stdout and stderr to stderr, and this command exits with the remote exit code, so it composes in scripts. ``` clika-cli devices exec [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--detach` | bool | `false` | start the command and print its id instead of waiting | | `--shell` | string | none | shell binary (agent picks a platform default when empty) | | `--timeout` | int | none | execution timeout in seconds (server max 300) | | `--working-dir` | string | none | working directory on the device | #### `clika-cli devices files` List directory on device ``` clika-cli devices files [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--path` | string | none | Absolute path to list on device | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices files-fetch` Downloads a file from the device to `` (default: the current directory, keeping the remote file name). The download is verified end-to-end from the response trailer: a single file against the SHA-256 the device computed while reading its disk (echoed on success for your own records), a --recursive directory against the server-assembled tar.gz checksum. Old agents declare no checksum, the fetch then completes unverified rather than failing. With --tail-bytes N only the last N bytes are printed to stdout. ``` clika-cli devices files-fetch [dest] [flags] ``` Positional arguments: required ``, ``; optional `[dest]`. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--recursive` | bool | `false` | fetch a directory as a tar.gz archive | | `--tail-bytes` | int | none | print only the last N bytes (text) to stdout | #### `clika-cli devices files-push` Push file to device ``` clika-cli devices files-push [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices files-push-from-library` Push artifact from library to device ``` clika-cli devices files-push-from-library [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--async` | string | none | When true, return 202 with a transfer_id immediately and run the transfer in the background; poll GET /api/transfers/{transfer_id}. When absent, the call blocks until completion (the pre-existing behaviour). | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices get` Get device ``` clika-cli devices get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices health` Get device health ``` clika-cli devices health [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices health-history` Get device health history ``` clika-cli devices health-history [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--duration` | string | none | Relative window: 1h, 6h, 24h, 7d, 30d | | `--from` | string | none | Range start (RFC3339, used when duration absent) | | `--interval` | string | none | Downsampling bucket size in seconds. Floored at the range divided by 2000, so one read answers at most 2000 points; a smaller value is widened, not refused. | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--to` | string | none | Range end (RFC3339, used when duration absent) | #### `clika-cli devices health-probe` Probe device health on demand ``` clika-cli devices health-probe [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--wait_s` | string | `5, clamped to 1-10` | Seconds to wait for the agent's reply | #### `clika-cli devices job-queue` Get device job queue ``` clika-cli devices job-queue [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices job-queue-delete` Clear device job queue ``` clika-cli devices job-queue-delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices list` List devices ``` clika-cli devices list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--accelerator` | string | none | Filter by accelerator type | | `--attestation_type` | string | none | Filter by the device's verified key-storage medium: tpm_attested, keystore_attested, file_backed, unverified | | `--dir` | string | none | Sort direction: desc (default), asc | | `--dirty` | string | none | Filter by dirty state (true/false) | | `--gpu_model` | string | none | Filter by GPU model | | `--group` | string | none | Filter by group name | | `--min_gpu_vram` | string | none | Minimum GPU VRAM bytes | | `--min_memory` | string | none | Minimum total memory bytes | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--platform` | string | none | Filter by platform: linux, macos, windows, android | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Search by name, group name or tag | | `--search_or` | search | none | When true, search is OR-ed against the combined tag+term filters instead of AND-ed, so in-flight text widens the committed filter set. Ignored when there are no tag or term filters. | | `--sort` | string | none | Sort column: registered_at (default), name, status, platform, last_heartbeat | | `--status` | string | none | Filter by status: online, offline, provisioning, error (the last two are cloud devices whose VM is still launching or failed to launch) | | `--summary` | string | none | Pass 'true' to return only {id, name, status} per device instead of the full row. Filtering, sorting and pagination are unchanged; the health snapshot, network interfaces, capabilities, resources and tags are omitted, and the health lookup they require is skipped. Any other value returns the full row. | | `--tag` | string | none | Filter by tag (repeatable). Tags sharing a key (the text before the first '=') are OR-ed; different keys are AND-ed. Tags without '=' each narrow the result on their own. | | `--term` | string | none | Committed free-text filter (repeatable). Each term must match the device's name, group name or a tag, so terms AND with each other and with the tag filters. | #### `clika-cli devices power` Set device power policy ``` clika-cli devices power [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices push` Uploads a local file to an absolute path on the device. A `` ending in '/' keeps the local file name. ``` clika-cli devices push [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--mkdir-parents` | bool | `false` | create missing parent directories on the device | | `--permissions` | string | none | octal file mode on the device (e.g. 0755) | #### `clika-cli devices reconcile` Force reconciliation ``` clika-cli devices reconcile [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices reconciliation-log` Get reconciliation log ``` clika-cli devices reconciliation-log [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | none | Max records to return (default: 100) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices resources` Update device resource inventory ``` clika-cli devices resources [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices revoke-cert` Revoke device certificate ``` clika-cli devices revoke-cert [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices run-tag-scripts` Run tag scripts on device ``` clika-cli devices run-tag-scripts [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services` List services on device ``` clika-cli devices services [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-artifacts` Get service artifact status ``` clika-cli devices services-artifacts [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-create` Start service on device ``` clika-cli devices services-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-delete-svc_id` Delete service on device ``` clika-cli devices services-delete-svc_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-events` List service events ``` clika-cli devices services-events [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | none | Max events to return (default: 100, max: 1000) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-restart` Restart service on device ``` clika-cli devices services-restart [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-stop` Stop service on device ``` clika-cli devices services-stop [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-svc_id` Get service on device ``` clika-cli devices services-svc_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices services-update` Update service artifacts on device ``` clika-cli devices services-update [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices session` An exec session keeps one shell alive on the device: cd and exports persist between commands, background processes stay anchored to the live shell (and visible to the next command), and commands may run unbounded. The session survives this CLI exiting, reconnect and keep using the same session id. #### `clika-cli devices sessions` List exec sessions ``` clika-cli devices sessions [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--all` | string | none | Include closed sessions | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices sessions-create` Open an exec session ``` clika-cli devices sessions-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices sessions-delete-sid` Close an exec session ``` clika-cli devices sessions-delete-sid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices sessions-exec` Run a command in an exec session ``` clika-cli devices sessions-exec [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices sessions-sid` Get one exec session ``` clika-cli devices sessions-sid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices sessions-signal` Signal an exec session's running command ``` clika-cli devices sessions-signal [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices ssh-info` Get SSH connection info ``` clika-cli devices ssh-info [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices ssh-keys` Update SSH authorized keys ``` clika-cli devices ssh-keys [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices tag-script-runs` List tag script runs for device ``` clika-cli devices tag-script-runs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | none | Max records to return (default: 100, max: 1000) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices tags` List device tags ``` clika-cli devices tags [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices tags-starred` List starred device tags ``` clika-cli devices tags-starred [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices tags-starred-update` Replace starred device tags ``` clika-cli devices tags-starred-update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices update` Update device ``` clika-cli devices update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices update-agent` Trigger agent update on device ``` clika-cli devices update-agent [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices update-history` Get device agent update history ``` clika-cli devices update-history [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | none | Max records to return (default: 50) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices update-policy` Get device update policy ``` clika-cli devices update-policy [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli devices update-policy-update` Set device update policy ``` clika-cli devices update-policy-update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli enrollment-tokens` `enrollment-tokens` has 5 subcommands. #### `clika-cli enrollment-tokens create` Create enrollment token ``` clika-cli enrollment-tokens create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/enrollment-tokens`. MCP tool name: `post_enrollment_tokens`. #### `clika-cli enrollment-tokens get` Get enrollment token ``` clika-cli enrollment-tokens get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/enrollment-tokens/{id}`. MCP tool name: `get_enrollment_tokens_id`. #### `clika-cli enrollment-tokens list` List enrollment tokens ``` clika-cli enrollment-tokens list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/enrollment-tokens`. MCP tool name: `get_enrollment_tokens`. #### `clika-cli enrollment-tokens qr` Get enrollment QR code ``` clika-cli enrollment-tokens qr [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/enrollment-tokens/{id}/qr`. MCP tool name: `get_enrollment_tokens_id_qr`. #### `clika-cli enrollment-tokens revoke` Revoke enrollment token ``` clika-cli enrollment-tokens revoke [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/enrollment-tokens/{id}`. MCP tool name: `delete_enrollment_tokens_id`. ### `clika-cli enrollment-codes` `enrollment-codes` has 2 subcommands. #### `clika-cli enrollment-codes create` Create enrollment code ``` clika-cli enrollment-codes create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/enrollment-codes`. MCP tool name: `post_enrollment_codes`. #### `clika-cli enrollment-codes delete` Revoke enrollment code ``` clika-cli enrollment-codes delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/enrollment-codes/{id}`. MCP tool name: `delete_enrollment_codes_id`. ### `clika-cli hosted-devices` `hosted-devices` has 9 subcommands. #### `clika-cli hosted-devices backends` List the enabled hosted-device backends ``` clika-cli hosted-devices backends [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/hosted-devices/backends`. MCP tool name: `get_hosted_devices_backends`. #### `clika-cli hosted-devices legs-telemetry` Get a hosted leg's self-reported telemetry ``` clika-cli hosted-devices legs-telemetry [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/hosted-devices/legs/{id}/telemetry`. MCP tool name: `get_hosted_devices_legs_id_telemetry`. #### `clika-cli hosted-devices list` List the hosted device fleet ``` clika-cli hosted-devices list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--backend` | string | none | Which hosted-device backend to list (device_farm \| test_lab); omit when only one is enabled, required when more than one is | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--refresh` | string | none | Set to 1 to bypass the server-side cache and re-fetch the fleet | Endpoint: `GET /api/hosted-devices`. MCP tool name: `get_hosted_devices`. #### `clika-cli hosted-devices runnable-types` List the benchmark types hosted devices can run ``` clika-cli hosted-devices runnable-types [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/hosted-devices/runnable-types`. MCP tool name: `get_hosted_devices_runnable_types`. #### `clika-cli hosted-devices runs` List hosted-device runs ``` clika-cli hosted-devices runs [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--benchmark_group_id` | string | none | Filter to the runs created by one benchmark group (UUID) | | `--page` | string | none | Page number (default: 1) | | `--per_page` | string | none | Items per page (default: 20) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--status` | string | none | Filter by status (pending\|uploading\|scheduling\|queued\|running\|collecting\|completed\|failed\|stopped) | Endpoint: `GET /api/hosted-devices/runs`. MCP tool name: `get_hosted_devices_runs`. #### `clika-cli hosted-devices runs-create` Start a benchmark run on hosted devices ``` clika-cli hosted-devices runs-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/hosted-devices/runs`. MCP tool name: `post_hosted_devices_runs`. #### `clika-cli hosted-devices runs-delete-id` Delete a hosted-device run ``` clika-cli hosted-devices runs-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/hosted-devices/runs/{id}`. MCP tool name: `delete_hosted_devices_runs_id`. #### `clika-cli hosted-devices runs-id` Get a hosted-device run ``` clika-cli hosted-devices runs-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/hosted-devices/runs/{id}`. MCP tool name: `get_hosted_devices_runs_id`. #### `clika-cli hosted-devices runs-stop` Stop a hosted-device run ``` clika-cli hosted-devices runs-stop [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/hosted-devices/runs/{id}/stop`. MCP tool name: `post_hosted_devices_runs_id_stop`. ### `clika-cli tag-scripts` `tag-scripts` has 6 subcommands. #### `clika-cli tag-scripts create` Create tag script ``` clika-cli tag-scripts create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/tag-scripts`. MCP tool name: `post_tag_scripts`. #### `clika-cli tag-scripts delete` Delete tag script ``` clika-cli tag-scripts delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/tag-scripts/{id}`. MCP tool name: `delete_tag_scripts_id`. #### `clika-cli tag-scripts get` Get tag script ``` clika-cli tag-scripts get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/tag-scripts/{id}`. MCP tool name: `get_tag_scripts_id`. #### `clika-cli tag-scripts list` List tag scripts ``` clika-cli tag-scripts list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/tag-scripts`. MCP tool name: `get_tag_scripts`. #### `clika-cli tag-scripts toggle` Toggle tag script enabled state ``` clika-cli tag-scripts toggle [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/tag-scripts/{id}/toggle`. MCP tool name: `post_tag_scripts_id_toggle`. #### `clika-cli tag-scripts update` Update tag script ``` clika-cli tag-scripts update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/tag-scripts/{id}`. MCP tool name: `put_tag_scripts_id`. ## Related - [Benchmarks](benchmarks.md): running work across a set of devices and reading the results. - [Jobs](jobs.md): one unit of work on one device. - [Artifacts](artifacts.md): the files you push to devices and the files they produce. - [job definitions](job-definitions.md): what the platform runs on a device. - [Device concept](../concepts/device.mdx): what a device is and what you can do with one, in prose. - [CLI overview](index.md): global flags, output formats and exit codes. --- # Get started with the CLI Ten minutes from nothing to your first clika-cli commands: install the CLI from your deployment, log in, look around, and read a benchmark run from the terminal. Source: https://docs.clika.io/platform/cli/get-started.md `clika-cli` is the platform from a terminal. Everything you can do in the web application, from listing devices to starting a benchmark and reading its results, has a command, and the commands are generated from your deployment's own API, so they always match what the platform can do. This page takes you from nothing to your first useful commands. The [reference](index.md) then covers every command. ## What you need - An account on a CLIKA Platform deployment, and its address, the one you type into the browser. - A terminal on Linux, macOS or Windows. The CLI takes about 15 MB of disk and about 70 MB of RAM ([system requirements](/clikart/system-requirements#platform-software)). ## 1. Install it The CLI is downloaded from your own deployment, so the version you install always matches the platform you talk to. 1. Sign in to the web application and open **Settings**, then **Developer access**. 2. On the **Install the CLI** card, pick the tab for your computer and copy the install command shown there. The command needs no token: the CLI download is open to anyone who can reach the platform, and the installer checks the file against the published checksum. (`clika-cli auth cli-download-token` still mints an attribution token for scripted installs; it is optional.) 3. Paste it into your terminal and run it. On Linux and macOS the command runs the shell installer, which installs `clika-cli` into `/usr/local/bin` when that is writable and `~/.local/bin` otherwise: ``` curl -fsSL "https://platform.clika.io/files/installation/cli/install.sh" | CLIKA_BASE_URL="https://platform.clika.io" sh ``` On Windows it is a PowerShell command. It runs in Windows PowerShell 5.1 and PowerShell 7, needs no administrator rights, installs `clika-cli.exe` into `%LOCALAPPDATA%\Programs\clika\bin`, and adds that directory to your user `PATH`: ``` $env:CLIKA_BASE_URL='https://platform.clika.io'; irm 'https://platform.clika.io/files/installation/cli/install.ps1' | iex ``` Both installers detect your CPU architecture, check the download against the published checksum, and upgrade an installed CLI in place when you run them again. Check that it runs: ``` clika-cli version ``` ## 2. Log in Create an API key on the same **Developer access** tab, then save it in a profile together with your deployment's address: ``` clika-cli login --base-url https://platform.clika.io --api-key ``` The command prompts for the key without echoing it and writes the profile to `~/.config/clika-cli/default.json`, readable by you alone. Every later command reads it, so you type the address and the key once. If you work with more than one deployment, give each its own profile with `--profile ` and pick it the same way on every command. On an on-premise deployment with a self-signed certificate, add `--insecure-tls` to the login and to later commands, or set `CLIKA_INSECURE_TLS=1` in the environment. The profile does not record it. ### Register your AI clients in the same step The installers can also connect Claude Code, Claude Desktop and Codex to the platform once the CLI is in place. Set `CLIKA_MCP` to the clients you use, separated by commas, and the installer runs [`clika-cli mcp install`](mcp.md#mcp-install-uninstall-status-and-headers) for each: ``` curl -fsSL "https://platform.clika.io/files/installation/cli/install.sh" | CLIKA_BASE_URL="https://platform.clika.io" CLIKA_MCP=claude-code,claude-desktop,codex sh ``` ``` $env:CLIKA_MCP='claude-code,claude-desktop,codex'; $env:CLIKA_BASE_URL='https://platform.clika.io'; irm 'https://platform.clika.io/files/installation/cli/install.ps1' | iex ``` Each registration needs a logged-in profile. At a terminal, `mcp install` offers the login and prompts for the API key without echoing it, so step 2 happens on the way. For an unattended install, set `CLIKA_API_KEY` as well: the installer first runs `login`, which saves that key to the default profile, then registers the clients. If a registration fails, the CLI stays installed, the installer names the clients that failed, and `clika-cli mcp install ` retries one. [Using with AI (MCP)](../mcp/index.md) covers each client. ## 3. Look around ``` clika-cli devices list clika-cli models list clika-cli benchmark-groups list ``` Each command prints a table of what your current project holds, the same rows the web application shows. Add `-o json` to any command to get the raw answer for a script, and `--help` to any command or group to see its flags. ## 4. Read a benchmark from the terminal Start a benchmark in the web application, or pick one that already finished, then follow and read it from the shell: ``` clika-cli benchmarks watch "my-first-run" clika-cli benchmarks results "my-first-run" clika-cli benchmarks results "my-first-run" -o json ``` `watch` prints the run's progress until every leg has finished and exits non-zero if one did not, which is what makes a benchmark a step in a CI pipeline. `results` prints the headline metrics, and with `-o json` the full result for further processing. ## Where next - [Benchmarks](benchmarks.md): start a run from the terminal, choose its models, tests and devices, and cancel or share it. - [Devices](devices.md): register, inspect and control devices, run a command on one, push a file to it. - [apply and export](apply-export.md): keep job and service definitions as YAML files in git. - [Using with AI (MCP)](../mcp/index.md): give Claude Desktop, Claude Code, Codex or another assistant the platform as tools. - [Cheat sheet](cheatsheet.md): the commands people reach for most, on one page. --- # Job definitions clika-cli job-definitions: the reusable description of a piece of work the platform runs on a device. Source: https://docs.clika.io/platform/cli/job-definitions.md A **job definition** describes a piece of work the platform knows how to run: which command, with which files pushed to the device first, what resources the device must have, and where the result file lands. Dispatching it creates a [job](jobs.md). Definitions are the reusable half of the platform. You author one once, then run it against any device that meets its requirements. Because they are also the thing most worth reviewing before it runs on a fleet, they round-trip to YAML through [`apply` and `export`](apply-export.md), so the file in your repository and the definition in the platform can be kept identical. The long-running counterpart is a [service definition](services.md). The difference in one line: a job finishes, a service does not. ## The commands ``` clika-cli job-definitions [command] ``` | Subcommand | Purpose | | --- | --- | | `list` | Every definition. Takes `--search ` for a partial name match. | | `get ` | One definition in full. | | `create --body ''` | Create one. | | `update --body ''` | Change one. | | `delete ` | Delete one. | ## The body Three fields are required: `name`, `output_path` and `script`. | Field | Type | Meaning | | --- | --- | --- | | `name` | string, required | Display name, and the handle you use everywhere else in place of the UUID. | | `output_path` | string, required | Path on the device where the run writes its result file. The platform collects this file when the job finishes. | | `script` | object, required | What to run. `script.command` is the argument vector, `script.working_dir` the directory to run it in, `script.env` an object of environment variables, and `script.timeout_sec` a wall-clock limit. | | `description` | string | Free text for whoever reads the definition next. | | `artifacts` | array of objects | Files to push to the device before the run: models, datasets, runner bundles. See [artifacts](artifacts.md). | | `benchmark_type_id` | string | The benchmark type this definition implements. Setting it links the definition into the catalogue and turns on server-side scoring. | | `required_resources` | object | The minimum device this definition can run on: `min_cpu_cores`, `min_disk_bytes`, `min_gpu_count`, further minimums, an `accelerators` list, a `custom` object, and `engine`. | | `cleanup_policy` | object | What to remove from the device afterwards: `remove_artifacts` (unspecified means true), `remove_output`, and `custom_paths`. | Two notes on `required_resources` that catch people out. The numeric minimums are a **headroom warning**. A dispatch that fails them can be forced through with `force`, on the assumption that you know something the inventory does not. `required_resources.engine` is **not** forceable. It names a licensed inference engine the definition cannot run without, today only `modelverse`, which is delivered to devices as a separate entitlement-gated package. An absent engine is a missing component rather than a tight fit, so the dispatch is refused. ## Examples ``` $ clika-cli job-definitions list --search latency NAME BENCHMARK TYPE UPDATED llm-latency performance 3d ago vision-latency performance 3d ago ``` ``` clika-cli job-definitions create --body '{ "name": "GPU Benchmark v2", "output_path": "/tmp/benchmark_results.json", "script": {"command": ["python3", "run.py"], "timeout_sec": 1800} }' ``` Editing an existing definition is usually easier through the YAML round trip than by writing a JSON patch: ``` clika-cli export JobDefinition llm-latency > job-defs/llm-latency/config.yaml $EDITOR job-defs/llm-latency/config.yaml clika-cli apply -f job-defs/llm-latency/config.yaml ``` ## Keeping definitions in git The platform database and your repository are two halves of one workflow. A definition can be authored in either place, but it should end up in both. A definition edited in the dashboard and never committed disappears when the deployment is torn down, and a definition edited in the repository and never applied never takes effect. The round trip is [`export`](apply-export.md) to bring the platform's copy down, and [`apply`](apply-export.md) to push your copy up. Because `export` strips server-managed fields (ids, timestamps, run state, counters), what the platform emits equals the file you committed. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli job-definitions` `job-definitions` has 5 subcommands. #### `clika-cli job-definitions create` Create job definition ``` clika-cli job-definitions create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/job-definitions`. MCP tool name: `post_job_definitions`. #### `clika-cli job-definitions delete` Delete job definition ``` clika-cli job-definitions delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/job-definitions/{id}`. MCP tool name: `delete_job_definitions_id`. #### `clika-cli job-definitions get` Get job definition ``` clika-cli job-definitions get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/job-definitions/{id}`. MCP tool name: `get_job_definitions_id`. #### `clika-cli job-definitions list` List job definitions ``` clika-cli job-definitions list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Filter by name (partial match) | Endpoint: `GET /api/job-definitions`. MCP tool name: `get_job_definitions`. #### `clika-cli job-definitions update` Update job definition ``` clika-cli job-definitions update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/job-definitions/{id}`. MCP tool name: `put_job_definitions_id`. ## Related - [Jobs](jobs.md): dispatching a definition and reading what it produced. - [Service definitions](services.md): the long-running counterpart. - [apply and export](apply-export.md): the resource YAML format and its round trip. - [Artifacts and models](artifacts.md): the files a definition pushes to a device. - [Job concept](../concepts/job.mdx): what a job is, in prose. --- # Jobs clika-cli jobs: dispatch work to a device, follow it, and pull back its logs, outputs and per-sample results. Source: https://docs.clika.io/platform/cli/jobs.md A **job** is one unit of work the platform runs on one device: a benchmark leg, a script, a data collection pass. It has a lifecycle (queued, running, finished), it produces an execution log and an output file, and it may produce per-sample results the platform scores afterwards. Most jobs are created for you. Launching a [benchmark group](benchmarks.md) creates one job per model and device pair; a [job definition](job-definitions.md) pushed as desired state creates jobs on a schedule. You create one by hand when you want an ad-hoc run. ``` clika-cli jobs [command] ``` ## `jobs list` ``` clika-cli jobs list [flags] ``` Lists jobs, newest first. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--status` | string | all | One of `pending`, `starting`, `running`, `completed`, `failed`, `cancelled`, `pushing_artifacts`, `running_setup`, `running_teardown`, `collecting_output`, `cleaning_artifacts` or `interrupted`. The intermediate values are useful when a job seems stuck: they say which phase it is in. | | `--device_id` | string | all | Only jobs on one device, by UUID. | | `--artifact_id` | string | all | Only jobs involving one artifact, by UUID. | | `--search` | string | none | Full-text search on the job name. | | `--sort` | string | `created_at` | Sort column: `created_at`, `status`, `name` or `device`. | | `--dir` | string | `desc` | Sort direction, `asc` or `desc`. | | `--page` | string | `1` | Page number. | | `--page_size` | string | server default | Rows per page. | | `--raw` | bool | `false` | Print the response body byte for byte. | ``` $ clika-cli jobs list --status running NAME STATUS DEVICE CREATED nightly-llm-sweep-jetson-01 running jetson-01 2m ago nightly-llm-sweep-orin-02 running orin-02 2m ago ``` ``` clika-cli jobs list --status failed -o json | jq -r '.items[] | "\(.name)\t\(.error)"' ``` ## `jobs get` ``` clika-cli jobs get ``` One job in full: its status and phase, the device it ran on, timings, the definition it came from, and any error. Accepts the job name as well as its UUID. ## `jobs create` ``` clika-cli jobs create --body '' [flags] ``` Dispatches a job. The body must set at least one of `job_definition_id`, `script` or `artifact_id`, which are the three ways of saying what to run. | Body field | Type | Meaning | | --- | --- | --- | | `job_definition_id` | string | The preferred form: run an existing [job definition](job-definitions.md). Accepts the definition's name. | | `script` | object | An inline script configuration, when you do not want a stored definition. Requires `output_path`. | | `artifact_id` | string | The legacy proxy-based form. Prefer a script or a definition. | | `device_id` | string | The single device to run on. | | `device_ids` | array of strings | Several devices, dispatching one job each. | | `model_id` | string | The model artifact to run against. | | `model_ids` | array of strings | Several models. | | `huggingface_url` | string | Source the model from a Hugging Face repository instead of an uploaded artifact. Valid only alongside `job_definition_id`. | | `name` | string | Job name. Generated when omitted. | | `output_path` | string | Where on the device the job writes its output. Required for an inline script job. | | `result_type` | string | `raw` (default) or `structured`. Structured results are parsed and scored; raw ones are stored as they are. | | `interpreter` | string | Which result interpreter to score with: `llm`, `vision` or `audio`. | | `idle_timeout_sec` | integer | Abandon the job after this many seconds with no progress. | | `retry_config` | object | `max_retries`, `retry_delay_ms` and `retry_on`, a list of result statuses that trigger a retry. | | `artifacts` | array of objects | Artifacts to push to the device, overriding the definition's own list. | | `force` | boolean | Skip the device resource compatibility check. Also available as the `--force` flag. | ``` clika-cli jobs create --body '{ "job_definition_id": "llm-latency", "device_ids": ["jetson-01", "orin-02"], "model_id": "Qwen2.5-0.5B-Instruct" }' ``` Running a definition from a file instead, so the request is reviewable in git, is what [`apply`](apply-export.md) is for: ```yaml kind: Job job_definition: llm-latency devices: - jetson-01 - jetson-02 ``` ## Following a job There is no `jobs watch`. For a benchmark group use [`benchmarks watch`](benchmarks.md), which follows every leg at once. For a single job, poll `jobs get`, or read its log as it grows: ``` clika-cli jobs execution-log nightly-llm-sweep-jetson-01 ``` `jobs execution-log` downloads what the agent captured while the job ran, which is the first thing to read when a job failed. ## Getting results back A job produces up to three kinds of output, and each has its own command. | Command | What it returns | | --- | --- | | `jobs output ` | The job's output file, as the job wrote it. | | `jobs execution-log ` | The execution log captured on the device. | | `jobs results ` | The per-sample proxy results, paginated. Takes `--status` (`success`, `timeout`, `queue_full`, `server_error`, `proxy_error`), `--page` and `--page_size`, which defaults to 50. | | `jobs results-body ` | One proxy result's full response body, which the list view truncates. | | `jobs results-summary ` | The aggregate: counts, rates and the headline numbers. | | `jobs results-throughput ` | Throughput over the life of the run, for plotting. | | `jobs benchmark-results ` | The measured benchmark metrics for a benchmark job. | | `jobs benchmark-results-io ` | The per-sample input, model output and expected answer, when the run was created with detailed outputs enabled. | | `jobs benchmark-results-samples ` | The stored benchmark result samples. | | `jobs benchmark-results-summary ` | The aggregate of those benchmark results. | | `jobs results-latency-histogram ` | Latency distribution across the run, for plotting. | ``` $ clika-cli jobs results-summary nightly-llm-sweep-jetson-01 { "total": 500, "success": 498, "timeout": 2, "p50_latency_ms": 210, "p95_latency_ms": 470 } ``` ``` clika-cli jobs results nightly-llm-sweep-jetson-01 --status timeout -o json ``` ## Rescoring ``` clika-cli jobs rescore ``` Runs the server-side scoring pass again over an existing benchmark job's stored results. Use it when the scoring logic changed, or when a scorer failed while the run itself was fine. The device is not involved, so this is cheap compared with re-running the job. ## Stopping and removing ``` clika-cli jobs cancel clika-cli jobs delete ``` `cancel` stops a job that is queued or running. `delete` removes a finished job and its stored results. ## Subcommand index Every subcommand, as a quick index. The full detail for each one, with its flags and the API operation it dispatches, is in the [full command reference](#full-command-reference) below. | Subcommand | Purpose | | --- | --- | | `list` | List jobs with filters. | | `get` | Get one job. | | `create` | Dispatch a job. | | `cancel` | Cancel a queued or running job. | | `delete` | Delete a job and its results. | | `output` | Download the job's output file. | | `execution-log` | Download the execution log. | | `results` | List per-sample proxy results. | | `results-body` | Download one result's response body. | | `results-summary` | Aggregate result summary. | | `results-throughput` | Throughput over time. | | `benchmark-results` | Measured benchmark metrics. | | `benchmark-results-io` | Per-sample input and output samples. | | `benchmark-results-samples` | Stored benchmark result samples. | | `benchmark-results-summary` | Aggregate of the benchmark results. | | `results-latency-histogram` | Latency histogram over the run. | | `rescore` | Re-run server-side scoring. | ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli jobs` `jobs` has 19 subcommands. #### `clika-cli jobs benchmark-results` Get benchmark results ``` clika-cli jobs benchmark-results [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/benchmark-results`. MCP tool name: `get_jobs_id_benchmark_results`. #### `clika-cli jobs benchmark-results-io` Get benchmark result I/O samples ``` clika-cli jobs benchmark-results-io [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--correct` | string | none | Filter by correctness, accuracy-style schemas only | | `--filter_field` | string | none | Field filter (repeatable, aligned with filter_op / filter_value): a field the schema declares, correct, latency_ms, outcome, input.``, output.`` or expected_output.`` | | `--filter_op` | string | none | Field filter operator (repeatable, aligned): gte / lte for a range filter, eq for a boolean or categorical one | | `--filter_value` | string | none | Field filter value (repeatable, aligned): a number for gte / lte, true / false for a boolean eq, a string for a categorical eq | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page (default: 50, max: 500) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--status` | string | none | Filter by sample status (e.g. success, failure) | Endpoint: `GET /api/jobs/{id}/benchmark-results/io`. MCP tool name: `get_jobs_id_benchmark_results_io`. #### `clika-cli jobs benchmark-results-samples` Get benchmark result samples ``` clika-cli jobs benchmark-results-samples [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page (default: 50) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--status` | string | none | Filter by sample status | Endpoint: `GET /api/jobs/{id}/benchmark-results/samples`. MCP tool name: `get_jobs_id_benchmark_results_samples`. #### `clika-cli jobs benchmark-results-summary` Get benchmark results summary ``` clika-cli jobs benchmark-results-summary [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/benchmark-results/summary`. MCP tool name: `get_jobs_id_benchmark_results_summary`. #### `clika-cli jobs cancel` Cancel job ``` clika-cli jobs cancel [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/jobs/{id}/cancel`. MCP tool name: `post_jobs_id_cancel`. #### `clika-cli jobs create` Create job ``` clika-cli jobs create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--force` | string | none | Skip resource compatibility check (default: false) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/jobs`. MCP tool name: `post_jobs`. #### `clika-cli jobs delete` Delete job ``` clika-cli jobs delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--delete_outputs` | string | none | Also delete the artifacts this job's pipeline collected (its output file and execution log). Default false: the outputs stay in the artifact library, tagged with the job id, for the record. | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/jobs/{id}`. MCP tool name: `delete_jobs_id`. #### `clika-cli jobs execution-log` Download job execution log ``` clika-cli jobs execution-log [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/execution-log`. MCP tool name: `get_jobs_id_execution_log`. #### `clika-cli jobs get` Get job ``` clika-cli jobs get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}`. MCP tool name: `get_jobs_id`. #### `clika-cli jobs list` List jobs ``` clika-cli jobs list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--artifact_id` | string | none | Filter by artifact ID (UUID) | | `--device_id` | string | none | Filter by device ID (UUID) | | `--dir` | string | none | Sort direction: desc (default), asc | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Full-text search on job name | | `--sort` | string | none | Sort column: created_at (default), status, name, device | | `--status` | string | none | Filter by status (pending\|starting\|running\|completed\|failed\|cancelled\|pushing_artifacts\|running_setup\|running_teardown\|collecting_output\|cleaning_artifacts\|interrupted) | Endpoint: `GET /api/jobs`. MCP tool name: `get_jobs`. #### `clika-cli jobs output` Download job output ``` clika-cli jobs output [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/output`. MCP tool name: `get_jobs_id_output`. #### `clika-cli jobs outputs` List job output artifacts ``` clika-cli jobs outputs [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/outputs`. MCP tool name: `get_jobs_id_outputs`. #### `clika-cli jobs rescore` Re-run server-side scoring for a benchmark job ``` clika-cli jobs rescore [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/jobs/{id}/rescore`. MCP tool name: `post_jobs_id_rescore`. #### `clika-cli jobs results` List job proxy results ``` clika-cli jobs results [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page (default: 50) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--status` | string | none | Filter by result status: success, timeout, queue_full, server_error, proxy_error | Endpoint: `GET /api/jobs/{id}/results`. MCP tool name: `get_jobs_id_results`. #### `clika-cli jobs results-body` Download one proxy result's response body ``` clika-cli jobs results-body [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/results/{result_id}/body`. MCP tool name: `get_jobs_id_results_result_id_body`. #### `clika-cli jobs results-latency-histogram` Job latency histogram ``` clika-cli jobs results-latency-histogram [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--buckets` | string | none | Number of histogram buckets (default: 20, max: 200) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/results/latency-histogram`. MCP tool name: `get_jobs_id_results_latency_histogram`. #### `clika-cli jobs results-summary` Job results summary ``` clika-cli jobs results-summary [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/{id}/results/summary`. MCP tool name: `get_jobs_id_results_summary`. #### `clika-cli jobs results-throughput` Job throughput over time ``` clika-cli jobs results-throughput [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--window_sec` | string | none | Aggregation window in seconds (default: 10, max: 86400) | Endpoint: `GET /api/jobs/{id}/results/throughput`. MCP tool name: `get_jobs_id_results_throughput`. #### `clika-cli jobs status-counts` Count jobs by status ``` clika-cli jobs status-counts [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/jobs/status-counts`. MCP tool name: `get_jobs_status_counts`. ## Related - [Benchmarks](benchmarks.md): the usual way jobs get created, and how to watch a whole group of them. - [job definitions](job-definitions.md): the reusable description a job runs from. - [Devices](devices.md): where jobs run, and the queue and dirty-state rules that decide when. - [Artifacts](artifacts.md): the inputs jobs consume and the outputs they leave behind. - [Job concept](../concepts/job.mdx): what a job is, in prose. --- # Licensing clika-cli license, projects licenses and entitlement-profiles: activating a deployment, and issuing the credentials a runtime uses. Source: https://docs.clika.io/platform/cli/licensing.md Two different things are called a license on this platform, and keeping them apart saves a lot of confusion. | | **Platform license** | **Runtime credential** | | --- | --- | --- | | What it licenses | The deployment itself | One project's use of the runtime | | Who issues it | CLIKA, through the license portal | Your own platform, to your own project | | Where it is applied | Once, to the whole deployment | Handed to each ClikaRT runtime | | Commands | `license *` | `projects licenses*` | The CLI has no command that issues a runtime credential: `clika-cli licenses` (plural) prints the third-party software notices, not licenses, and issuing, rotating, revealing and revoking a runtime credential happen in the web application, signed in as a user (the reason is under [Runtime credentials](#runtime-credentials)). A fresh deployment starts unactivated, and its `/api/*` surface is gated until a platform license is applied. A runtime credential is a separate artifact issued afterwards, and what it may grant is capped by what the platform license grants. ## The platform license ``` clika-cli license [command] ``` | Subcommand | Purpose | | --- | --- | | `status` | The short answer: is this deployment licensed, and until when. | | `list` | The full license record, entitlements included. | | `state` | The enforcing state, which is what the platform actually acts on. | | `activate --body '{"license_key":"..."}'` | Apply a license key. | | `update --body ''` | Replace the license, for a renewal or a plan change. | ``` $ clika-cli license status STATE PLAN EXPIRES active enterprise 2027-03-01 ``` ``` clika-cli license activate --body '{"license_key":""}' ``` `status` is the one to reach for first when a deployment answers `403` on everything, because an unactivated or expired platform license looks like a permission problem until you check. The difference between `list` and `state` is worth knowing. `list` shows the record as stored; `state` shows what the platform enforces, which is derived from the signed license rather than from editable columns. When the two disagree, `state` is the truth. ## Runtime credentials A runtime credential is issued to a **project**, which is why it lives under `projects`. ``` clika-cli projects licenses ``` | Subcommand | Purpose | | --- | --- | | `licenses ` | List a project's credentials. | | `licenses-lid ` | Get one, with its key redacted. | | `licenses-create --body ''` | Issue one. | | `licenses-update-lid --body ''` | Change its name, note or expiry. | | `licenses-reveal ` | Show the key material once, for handing to a runtime. | | `licenses-rotate ` | Issue fresh key material for the same credential. | | `licenses-revoke ` | Revoke it. | ### Two kinds `kind` is the only required field on issue, and it decides how the runtime will verify itself. | `kind` | How it works | Use it when | | --- | --- | --- | | `simple_key` | The runtime calls the platform to check its entitlement. | The runtime can reach the platform. | | `cert_bundle` | The entitlement is signed into an offline artifact the runtime verifies locally. | The runtime is air-gapped, or must keep working through a network outage. | The web application calls `simple_key` **Online** and `cert_bundle` **Offline**; both are handed to the runtime as a `CLIKA1-...` license key. | Body field | Type | Meaning | | --- | --- | --- | | `kind` | string, required | `simple_key` or `cert_bundle`. | | `name` | string | Operator-facing display name. | | `note` | string | Free-form note for whoever reads the credential list next. | | `expires_at` | string | Absolute expiry, RFC 3339, and it must be in the future. Omitted, a certificate bundle takes the 90-day default and a simple key gets no expiry. A certificate bundle is also clamped to the platform license's own window. | | `entitlements` | object | What the credential grants. Capped at issue time against the organization's and the platform license's ceilings, so you cannot issue yourself more than you hold. Omitted, a certificate bundle inherits an empty bag capped to the ceiling and a simple key inherits the project's active certificate entitlements. | | `machine_binding` | object | A signed hardware-lock policy, for `cert_bundle` only. The platform validates its shape and signs it onto the bundle unchanged; the runtime is what evaluates seats and fingerprints. Passing it with `simple_key` is refused rather than silently ignored, because a simple key has nowhere to carry it. | ``` $ clika-cli projects licenses-create 3c9a1b2d-4e5f-6071-8293-a4b5c6d7e8f9 --body '{"kind":"cert_bundle","name":"edge-fleet-2026"}' issued project license "edge-fleet-2026" (b71e...), kind cert_bundle, expires 2026-12-02 ``` ``` clika-cli projects licenses-reveal 3c9a1b2d-4e5f-6071-8293-a4b5c6d7e8f9 b71e0f3a-1234-5678-9abc-def012345678 ``` Treat the output of `licenses-reveal` as a secret. It is the material a runtime authenticates with, and this documentation deliberately shows no real key. `licenses-create`, `licenses-rotate`, `licenses-revoke` and `licenses-reveal` take `project_licenses:write`, which no API key can carry, and the CLI signs in only with an API key, so under a key they answer `403 API_KEY_SCOPE_INSUFFICIENT`. Issue, rotate, reveal and revoke in the web application; the read commands work from the CLI. Rotation and revocation are separate on purpose. `licenses-rotate` replaces the key material while the credential and its entitlements stay in place, and `licenses-revoke` ends the credential entirely. ## Entitlement profiles ``` clika-cli entitlement-profiles [command] ``` An entitlement profile is a named, reusable entitlement bag, so an organization does not have to hand-write the same set on every issue. | Subcommand | Purpose | | --- | --- | | `list` | Every profile. | | `get ` | One profile. | | `create --body ''` | Create one. | | `update --body ''` | Change one. | | `delete ` | Delete one. | ## Activation codes ``` clika-cli activation-codes list clika-cli activation-codes create --body '' clika-cli activation-codes delete ``` An activation code is what a device presents at first contact so it can enrol without an operator typing credentials into it. See also [enrollment tokens](devices.md#enrolling-a-new-device), which are the same idea for the device agent. ## Certificates ``` clika-cli pki ca-cert clika-cli crl list ``` `pki ca-cert` downloads the certificate authority certificate a client needs to verify the platform's issued certificates. `crl list` downloads the certificate revocation list, which names the certificates that have been revoked since. ## The runtime-facing API A few small groups are not really a user-facing surface. They are what a ClikaRT runtime or a device calls for itself, exposed as CLI subcommands because every documented operation is, and occasionally useful for debugging an integration. | Subcommand | Who normally calls it | | --- | --- | | `register create` | A runtime registering itself. | | `renew create` | A runtime renewing its certificate over mutual TLS. | | `revoke create` | Revoking a device license. | | `entitlements list` | A runtime refreshing its entitlement token over mutual TLS. | | `enroll create` | A device enrolling with an activation code. | | `deployments heartbeat ` | A runtime's periodic heartbeat. | | `crl list` | The revocation list. | The model serving surface has its own page: see [model deployment](model-deployment.md). ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli license` `license` has 5 subcommands. #### `clika-cli license activate` Activate platform license ``` clika-cli license activate [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/license/activate`. MCP tool name: `post_license_activate`. #### `clika-cli license list` Get full license details ``` clika-cli license list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/license`. MCP tool name: `get_license`. #### `clika-cli license state` Get the enforcing platform license state ``` clika-cli license state [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/license/state`. MCP tool name: `get_license_state`. #### `clika-cli license status` Get platform license status ``` clika-cli license status [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/license/status`. MCP tool name: `get_license_status`. #### `clika-cli license update` Update platform license ``` clika-cli license update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/license`. MCP tool name: `put_license`. ### `clika-cli licenses` #### `clika-cli licenses` Prints the copyright and license texts of every open-source component compiled into this binary (its THIRD_PARTY_NOTICES.txt). The same file is served next to the binaries on the download channel, at ``/files/installation/cli/THIRD_PARTY_NOTICES.txt. ``` clika-cli licenses [flags] ``` ### `clika-cli entitlement-profiles` `entitlement-profiles` has 5 subcommands. #### `clika-cli entitlement-profiles create` Create entitlement profile ``` clika-cli entitlement-profiles create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/entitlement-profiles`. MCP tool name: `post_entitlement_profiles`. #### `clika-cli entitlement-profiles delete` Delete entitlement profile ``` clika-cli entitlement-profiles delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/entitlement-profiles/{id}`. MCP tool name: `delete_entitlement_profiles_id`. #### `clika-cli entitlement-profiles get` Get entitlement profile ``` clika-cli entitlement-profiles get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/entitlement-profiles/{id}`. MCP tool name: `get_entitlement_profiles_id`. #### `clika-cli entitlement-profiles list` List entitlement profiles ``` clika-cli entitlement-profiles list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/entitlement-profiles`. MCP tool name: `get_entitlement_profiles`. #### `clika-cli entitlement-profiles update` Update entitlement profile ``` clika-cli entitlement-profiles update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/entitlement-profiles/{id}`. MCP tool name: `put_entitlement_profiles_id`. ### `clika-cli activation-codes` `activation-codes` has 3 subcommands. #### `clika-cli activation-codes create` Create activation code ``` clika-cli activation-codes create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/activation-codes`. MCP tool name: `post_activation_codes`. #### `clika-cli activation-codes delete` Delete activation code ``` clika-cli activation-codes delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/activation-codes/{id}`. MCP tool name: `delete_activation_codes_id`. #### `clika-cli activation-codes list` List activation codes ``` clika-cli activation-codes list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/activation-codes`. MCP tool name: `get_activation_codes`. ### `clika-cli pki` `pki` has 2 subcommands. #### `clika-cli pki ca-cert` Download CA certificate ``` clika-cli pki ca-cert [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/pki/ca-cert`. MCP tool name: `get_pki_ca_cert`. #### `clika-cli pki ca-chain` Download the device-enrollment CA chain ``` clika-cli pki ca-chain [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/pki/ca-chain`. MCP tool name: `get_pki_ca_chain`. ### `clika-cli register` `register` has 1 subcommands. #### `clika-cli register create` Register a runtime ``` clika-cli register create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli renew` `renew` has 1 subcommands. #### `clika-cli renew create` Renew device certificate via mTLS ``` clika-cli renew create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli revoke` `revoke` has 1 subcommands. #### `clika-cli revoke create` Revoke device license ``` clika-cli revoke create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli entitlements` `entitlements` has 1 subcommands. #### `clika-cli entitlements list` Refresh entitlement token via mTLS ``` clika-cli entitlements list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli enroll` `enroll` has 1 subcommands. #### `clika-cli enroll create` Enroll device via activation code ``` clika-cli enroll create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli deployments` `deployments` has 1 subcommands. #### `clika-cli deployments heartbeat` Runtime heartbeat ``` clika-cli deployments heartbeat [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ### `clika-cli crl` `crl` has 1 subcommands. #### `clika-cli crl list` Download Certificate Revocation List ``` clika-cli crl list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | ## Related - [Organizations, projects and access](projects-and-orgs.md): the projects credentials are issued to. - [Devices](devices.md): enrollment tokens and device certificates. - [Model deployment](model-deployment.md): the deployed model that runs under a runtime license. - [job definitions](job-definitions.md): `required_resources.engine`, the entitlement-gated inference engine a definition can require. - [Runtime licenses concept](../concepts/runtime-licenses.mdx): what a license means on this platform, in prose. - [License the runtime](/clikart/how-to/license-the-runtime): where a ClikaRT or Modelverse runtime reads the credential that `licenses-reveal` shows. --- # MCP server and generic dispatch clika-cli mcp, tools and call: expose the platform to an AI assistant as Model Context Protocol tools, or dispatch any API operation by name. Source: https://docs.clika.io/platform/cli/mcp.md The Model Context Protocol, MCP, is a standard way for an AI assistant to call external tools. `clika-cli mcp` turns the CLI into an MCP server, so an assistant such as Claude Code or Cursor can read your devices, launch a benchmark and fetch its results by calling tools rather than by shelling out. `tools` and `call` are the same machinery exposed to you directly. `tools` lists every operation by its tool name, and `call` invokes one. They are useful for scripting, for reaching an endpoint that has no convenient subcommand, and for reproducing by hand exactly what an assistant did. Connecting Claude Desktop, Claude Code, Codex or another client end to end is covered in [Using with AI (MCP)](../mcp/index.md). This page documents the commands and their flags. ## `mcp` ``` clika-cli mcp [flags] ``` Runs an MCP server on standard input and output. You do not run this at a terminal yourself. You configure an AI harness to launch it, and the harness speaks the protocol over the pipe. Every documented platform operation is a tool with a JSON schema, **except the platform administration API** (`/api/admin/**`), which no toolset serves. By default the server lists the `api-key` toolset, the operations an API key can be granted. `--toolset full` lists every operation (see [Toolsets](#toolsets)). | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--toolset` | string | `api-key` | Which named toolset to serve. | | `--remote` | bool | `false` | Proxy to the orchestrator's own HTTP MCP endpoint instead of dispatching from the embedded API description. | | `--skip-update-check` | bool | `false` | Do not check the deployment's release channel for a newer CLI at start. | The global flags all apply, and the ones that matter here are `--profile`, `--token`, `--api-key`, `--base-url` and `--insecure-tls`, because the server authenticates to the platform exactly as any other command does. ``` clika-cli mcp clika-cli mcp --toolset full clika-cli mcp --toolset device-ops clika-cli mcp --profile onprem --insecure-tls clika-cli mcp --remote --profile onprem --insecure-tls ``` ### Toolsets A toolset is a named selection of the platform's tools for one kind of work. | Toolset | What it covers | | --- | --- | | `api-key` (default) | The operations an API key can be granted. Sign-in flows, the capabilities only a browser session can hold (users, roles, SSO, license and billing changes) and multipart-only uploads are in `full` only. | | `full` | Every documented operation except the platform administration API. | | `device-ops` | Day-to-day device work: list and inspect devices, read health, run commands, push and fetch files, browse artifacts. | | `benchmark-ops` | Benchmark runs end to end: the type catalog, the model pool, the compatibility check, picking registered or hosted targets, launching a group, polling it, reading results. Also the ClikaRT archive listing and download link, so an assistant can fetch the runtime the run used ([Download the ClikaRT SDK](../how-to/download-the-clikart-sdk.md)). | Ask a deployment what it serves: ``` $ curl -H "Authorization: Bearer clika_..." https://platform.clika.io/api/mcp/toolsets ``` A toolset is only offered where the product's API covers all of it, and an unknown or unavailable name fails with the servable names listed. In local mode the server also reads the credential's own capabilities from the platform at start and lists only the tools that credential can call, and says on standard error how many of the toolset's tools it serves. **A toolset is not a permission boundary.** It changes only which tools are *listed*. Every call still passes the platform's full authentication, licensing and permission checks, and a client that already knows a tool name can call it through the full endpoint. To restrict what a credential may do, scope the API key instead. See [authentication](authentication.md#minting-an-api-key-for-ci). ### Local mode versus `--remote` By default the CLI answers from its own embedded copy of the API description, and dispatches each call as an HTTP request to the platform. That works everywhere, including against a self-signed on-premise deployment, because the CLI terminates TLS itself and `--insecure-tls` is scoped to its own connections. With `--remote`, the CLI instead proxies to the orchestrator's Streamable HTTP MCP endpoint at `/api/mcp`, fetching the tool list from the server when it connects. Two consequences: - **The tool surface follows the server.** A platform that gains an operation exposes it immediately, with no CLI update. - **`--insecure-tls` is scoped to one connection.** This is the reason `--remote` exists. An MCP host that cannot be told to trust a self-signed certificate can point at the local CLI instead, and the CLI makes the one connection that needs the exception. The blunt alternative, disabling certificate verification for the host's entire process, is worth avoiding. `--toolset` composes with `--remote`, selecting the server's `/api/mcp/` endpoint. ### Connecting without the CLI at all Each orchestrator serves the same tools directly over Streamable HTTP at `POST /api/mcp`, built from its own copy of the API description. A tool call is dispatched in-process back through the API router, so it passes the same authentication, licensing and permission checks a direct API request would. The endpoint is stateless, and a request without a valid credential is rejected before the MCP handler runs. Point any HTTP MCP client at it with a session token or a `clika_` API key in the `Authorization` header. Where the deployment uses a self-signed certificate, the client, not the CLI, is the one that has to trust the certificate authority, which is exactly the case `mcp --remote` exists to rescue. ## `mcp install`, `uninstall`, `status` and `headers` ``` clika-cli mcp install [flags] clika-cli mcp uninstall [flags] clika-cli mcp status ``` `mcp install` registers the platform's MCP server with an AI client by writing an entry into the client's own config file, so nobody edits JSON or TOML by hand. `` is `claude-code`, `claude-desktop` or `codex`. `mcp uninstall` removes what `install` added, and `mcp status` reports, for each client, whether the server is registered, under which name and profile, which config file was read and why that file. Neither `claude` nor `codex` is run. The CLI edits the file directly. | Flag | Commands | Type | Default | Meaning | | --- | --- | --- | --- | --- | | `--name` | `install` | string | `clika-platform` | The name of a new entry. An entry this profile installed earlier keeps its name. | | `--config-path` | `install`, `uninstall` | string | none | The client's config file, instead of the detected one. | | `--dry-run` | `install`, `uninstall` | bool | `false` | Print the change and write nothing. | The global flags that matter here are `--profile`, whose key the entry uses, `--base-url`, which must match the profile's deployment, and `--insecure-tls`, which changes the entry's shape (below). A profile does not record `--insecure-tls`, so pass it to `install` whenever the deployment needs it. ``` clika-cli mcp install claude-code clika-cli mcp install claude-desktop --profile onprem --insecure-tls clika-cli mcp install claude-code --name clika-staging --profile staging clika-cli mcp install codex clika-cli mcp install claude-code --dry-run clika-cli mcp status clika-cli mcp uninstall claude-code ``` ### The login check Before writing anything, `install` checks the profile's API key with one request to the deployment (`GET /api/auth/capabilities`). It stops when the profile has no saved key, when it belongs to a different deployment than `--base-url`, or when the deployment answers 401. At a terminal it then offers to run the login and continues once the new key is accepted. Without a terminal it prints the exact `login` command to run. A certificate error ends the check with a hint to pass `--insecure-tls`. Nothing is written until the key works. ### What each client gets | Client | Config file | Entry | | --- | --- | --- | | `claude-code` | The user-scope `mcpServers` of `~/.claude.json`, or `$CLAUDE_CONFIG_DIR/.claude.json` when that is set. | HTTP to `/api/mcp`, with its `Authorization` header from `headersHelper`. With `--insecure-tls`, a stdio entry running `mcp --remote`. | | `claude-desktop` | `claude_desktop_config.json`, in `~/Library/Application Support/Claude/` on macOS and `$XDG_CONFIG_HOME/Claude/` or `~/.config/Claude/` on Linux. On Windows, `%LOCALAPPDATA%\Packages\Claude_*\LocalCache\Roaming\Claude\` for the Microsoft Store install, found by asking Windows which `Claude_*` package is installed for your user, and `%APPDATA%\Claude\` otherwise. | A stdio entry running `mcp --remote`. | | `codex` | `[mcp_servers.]` in `~/.codex/config.toml`, or `$CODEX_HOME/config.toml` when that is set (`%USERPROFILE%\.codex\` on Windows). The Codex CLI, its IDE extension and ChatGPT Desktop all read this file. | HTTP to `/api/mcp`, with its `Authorization` header from `http_headers_helper`. With `--insecure-tls`, a stdio entry running `mcp --remote`. | The command in an entry is this CLI's own absolute path with symlinks resolved, and `install` warns when that path is a temporary location that will not last. On Windows, when more than one Claude package is installed, `install` cannot tell which one you run and requires `--config-path`. ### No key in the client's config An entry names the profile, never the key. A stdio entry runs `clika-cli --profile mcp --remote`, which reads the key from the profile when the client starts it. An HTTP entry's header helper runs `mcp headers`: ``` clika-cli --profile mcp headers ``` `mcp headers` prints `{"Authorization":"Bearer "}` from the profile and makes no request to the platform. It is the command `headersHelper` (Claude Code) and `http_headers_helper` (Codex) run, and it is not listed in `clika-cli mcp --help`. Because every entry reads the key from the profile, rotating or revoking a key needs a new `login` and no change to any client. ### Only marked entries are touched Each entry `install` writes carries a marker, `clika-cli:`: in the `CLIKA_MCP_MANAGED_BY` environment variable of a stdio entry, or in the static `X-Clika-Managed-By` header of an HTTP entry (`headers` in Claude Code, `http_headers` in Codex). The server ignores both. `install`, `uninstall` and `status` act on marked entries only: - Running `install` again updates this profile's entry in place, under whatever name it has now. - A new entry is created under `--name`. A name already held by an entry without this profile's marker is refused, and nothing is changed. - `uninstall` removes every entry carrying this profile's marker. An entry written by hand, or by another profile, is left alone. The rest of the file is kept as it was: only the entry's bytes change, and every other key keeps its order and formatting. The write is atomic and keeps the file's permissions, and when something changed the previous content is saved next to it as `.bak`. ### Restarting the client A client reads its config when it starts. When the client's app is running (a Claude Code session, Claude Desktop, a Codex session or ChatGPT Desktop), `install` and `uninstall` print one line saying so and that a restart loads or removes the server. Restart it. Nothing is printed when it is not running. ## `tools` ``` clika-cli tools [flags] ``` Lists every documented operation as ` `. The tool name is what `call` takes, and what an MCP `tools/call` takes, so this is how you find the name for an operation you can see in the API. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--tag` | string | none | Filter by the operation's API tag, for example `devices` or `jobs`. | ``` $ clika-cli tools --tag devices get_devices GET /api/devices post_devices POST /api/devices get_devices_id GET /api/devices/{id} put_devices_id PUT /api/devices/{id} delete_devices_id DELETE /api/devices/{id} post_devices_id_exec POST /api/devices/{id}/exec ``` ``` clika-cli tools | wc -l clika-cli tools --tag jobs ``` ## `call` ``` clika-cli call [key=value ...] [flags] ``` Invokes one operation by tool name, with flat `key=value` arguments. Path and query parameters come from the API description; the request body is `body=` or `body=@path/to/file.json`. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | Print the response body byte for byte, with no pretty-printing. | ``` clika-cli call get_devices clika-cli call get_devices_id id=1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 clika-cli call post_job_definitions body=@./job-def.json clika-cli call post_models body='{"huggingface_url":"Qwen/Qwen2.5-0.5B-Instruct"}' ``` `call` uses the same dispatch path as the MCP server, so `call ` and an assistant's `tools/call ` produce identical results. That makes it the right way to reproduce something an assistant did, or to check what a tool will return before handing it to one. For everyday use prefer the resource subcommands. `clika-cli devices list` is `call get_devices` with tables, filters as flags and name resolution added. ## What an assistant sees A `tools/list` returns every tool with its input schema. A `tools/call` dispatches the matching HTTP request and returns the response body as text, pretty-printed when it is JSON. A non-2xx response comes back as an in-band error result rather than a protocol failure, so the model can read the message and correct itself. Operations that require a `multipart/form-data` upload are flagged in their tool descriptions. An assistant cannot call those, since a tool call is sent as JSON. A file still reaches the platform through the chunked upload session (`POST /api/artifacts/uploads`, then one base64-encoded chunk per call) or an external source the platform fetches itself (`POST /api/artifacts/external`), and from a terminal with `artifacts upload` or `devices push`. See [artifacts](artifacts.md) and [devices](devices.md). ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli mcp` `mcp` has 3 subcommands. #### `clika-cli mcp install` Adds an entry for the platform's MCP server to the client's own config file (claude-code, claude-desktop, codex). The entry names the profile, never the key: the client gets the key from this CLI when it connects. Run `login` first, or answer yes when install offers to. ``` clika-cli mcp install [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--config-path` | string | none | the client's config file, instead of the detected one | | `--dry-run` | bool | `false` | print the change and write nothing | | `--name` | string | `"clika-platform"` | entry name for a new entry (an entry this profile installed keeps its name) | #### `clika-cli mcp status` For each client (claude-code, claude-desktop, codex), says whether `mcp install` registered the server there, under which entry name and profile, which config file was read, and why that file. ``` clika-cli mcp status [flags] ``` #### `clika-cli mcp uninstall` Removes every entry in the client's config that `mcp install` added for this profile (claude-code, claude-desktop, codex). An entry written by hand, or by another profile, is left alone. ``` clika-cli mcp uninstall [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--config-path` | string | none | the client's config file, instead of the detected one | | `--dry-run` | bool | `false` | print the change and write nothing | ### `clika-cli tools` #### `clika-cli tools` Lists every documented operation outside the platform-admin API as ` `. The tool name is what `call` and an MCP `tools/call` take. ``` clika-cli tools [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--tag` | string | none | filter by OpenAPI tag (e.g. devices, jobs) | ### `clika-cli call` #### `clika-cli call` Invokes the named tool (see the `tools` subcommand). Arguments are key=value pairs. Path/query parameters come from the spec; the request body is provided as body=`` or body=@path/to/file.json. ``` clika-cli call [key=value ...] [flags] ``` Positional arguments: required ``; optional `[key=value ...]`. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | ## Related - [Using with AI (MCP)](../mcp/index.md): connecting Claude Desktop, Claude Code, Codex and other clients, step by step. - [Authentication and profiles](authentication.md): the credential the server presents, and scoping an API key. - [CLI overview](index.md): how the CLI, the API and MCP relate to one another. --- # Model deployment clika-cli servings: deploy a model behind an endpoint on a device, and call it through the platform. Source: https://docs.clika.io/platform/cli/model-deployment.md A **model deployment**, called a **serving** in the API and the CLI, is a model running behind an endpoint on one of your devices. The platform starts it, keeps it up, and proxies calls to it, so a client never has to reach the device directly. It is a [service](services.md) whose command the platform supplies rather than you: the model runs through the CLIKA runtime on the device, and the endpoints a client calls are derived from the model's type rather than configured by hand. The commands sit under the `servings` group. ## Managing a serving | Command | Purpose | | --- | --- | | `servings list` | List servings. | | `servings get ` | Get one, with its state and its endpoint. | | `servings create --body ''` | Create one. | | `servings update --body ''` | Change one. | | `servings start ` | Start it. | | `servings stop ` | Stop it without removing it. | | `servings restart ` | Restart it. | | `servings delete ` | Delete it. | | `servings key ` | Reveal the deployment key the proxy wants. | | `servings fit` | Which devices a model can be deployed to. | | `projects servings ` | List one project's servings. | ``` clika-cli servings list clika-cli servings get 4b7e0c1a-2d3e-4f50-6172-8394a5b6c7d8 ``` Starting and stopping are separate from creating and deleting for the same reason they are on a service: stopping frees the device's accelerator while keeping the deployment and its configuration, so you can bring the same endpoint back without redefining it. `servings create` creates the serving `pending` and starts nothing; run `servings start ` next. A serving that has sat in `pending` for minutes with no device state has not been started. A start reads `starting` for about two to three minutes when a benchmark has already put the engine on the device, and for five to nine minutes on a device that has never run it; on a device busy with a job the start waits behind that job. ## Calling a serving through the platform ``` clika-cli servings devices-proxy ``` The proxy is how a client reaches a deployed model without a route to the device. You address the platform, name the serving, the device and the upstream path, and the platform forwards the call to the model and the response back. `` is the path on the model's own endpoint. Which paths exist depends on the model's type, because the runtime exposes the endpoints that type implies rather than a fixed set. Four variants cover the HTTP methods, since a proxied call has to carry the caller's own verb: | Command | Method it proxies | | --- | --- | | `servings devices-proxy` | `GET` | | `servings devices-proxy-create` | `POST` | | `servings devices-proxy-update` | `PUT` | | `servings devices-proxy-update-patch` | `PATCH` | | `servings devices-proxy-delete` | `DELETE` | The proxy wants the deployment's own key on the `Authorization` header, which `servings key ` prints; the proxy subcommands send no request body. For anything beyond a check that the endpoint answers, call the proxy endpoint from your own client rather than through the CLI. The CLI's value here is discovery: `servings get` tells you the serving is up and what its endpoint is. ## What it needs A model deployment consumes a licensed inference engine on the device, so the same entitlement gate that applies to a [job definition](job-definitions.md) applies here. A device without the engine package cannot host one, and the refusal is not forceable. Provision the engine with `devices engine-package `, and see [runtime licenses](licensing.md) for the entitlement that permits it. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli servings` `servings` has 17 subcommands. #### `clika-cli servings api-reference` Download a model serving's offline API reference ``` clika-cli servings api-reference [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--format` | string | none | json (OpenAPI 3, the default), yaml (OpenAPI 3) or markdown | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings create` Create model serving ``` clika-cli servings create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings delete` Delete model serving ``` clika-cli servings delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--purge_model` | string | none | Also remove the deployment's model from each device's engine model cache where no other deployment on that device uses it. Default false: every device keeps its cached model. | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings devices-proxy` Call a model serving on one device through the platform ``` clika-cli servings devices-proxy [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings devices-proxy-create` Call a model serving on one device through the platform ``` clika-cli servings devices-proxy-create [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings devices-proxy-delete` Call a model serving on one device through the platform ``` clika-cli servings devices-proxy-delete [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings devices-proxy-update` Call a model serving on one device through the platform ``` clika-cli servings devices-proxy-update [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings devices-proxy-update-patch` Call a model serving on one device through the platform ``` clika-cli servings devices-proxy-update-patch [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings fit` Which devices a model can be deployed to ``` clika-cli servings fit [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--model_id` | string | none | Model ID (UUID) | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings get` Get model serving ``` clika-cli servings get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings key` Reveal a model serving's deployment key ``` clika-cli servings key [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings key-rotate` Rotate a model serving's deployment key ``` clika-cli servings key-rotate [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings list` List model servings ``` clika-cli servings list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings restart` Restart model serving ``` clika-cli servings restart [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings start` Start model serving ``` clika-cli servings start [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings stop` Stop model serving ``` clika-cli servings stop [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli servings update` Update model serving ``` clika-cli servings update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | ## Related - [Services](services.md): the free-form counterpart, where you supply the command. - [Runtime licenses](licensing.md): the entitlement a deployed model runs under. - [Devices](devices.md): the device that hosts the deployment. - [Artifacts and models](artifacts.md): registering the model to deploy. - [Model deployment concept](../concepts/model-deployment.mdx): what a deployment is, in prose. --- # Events, metrics and alerts clika-cli events, metrics, alerts, notifications, audit and webhooks: watching what the platform is doing and being told when it changes. Source: https://docs.clika.io/platform/cli/monitoring.md Four surfaces tell you what the platform is doing, and they answer different questions. | Group | Question it answers | | --- | --- | | `events` | What has happened recently, as a live feed? | | `metrics` | What are the numbers, over time? | | `alerts` | What condition should tell me, and did it? | | `audit` | Who did what, and when? | `notifications` and `webhooks` are the delivery side. Notifications go to people in the web app, and webhooks go to your own systems. ## Events The event feed is durable and ordered. Every job, transfer and command completion, every device going online or offline, every alert firing lands in it with a position, and you resume from a position rather than from a timestamp. That is what makes it safe for a script to disconnect and come back without missing anything. ### `events tail` ``` clika-cli events tail [flags] ``` Follows the feed, printing one JSON line per event. It prints everything retained after `--cursor` first, then blocks on the server's long poll and prints each new event the moment it lands. There is no polling interval to tune, because it is not polling. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--cursor` | int | `0` | Start after this feed position. `0` means the start of the retention window. | | `--types` | string | all | Comma-separated event names to include, for example `transfer_completed,job_completed`. | | `--once` | bool | `false` | Print what the feed already holds after `--cursor` and exit, instead of following. | The `id` on the last line you printed is the cursor to resume from. ``` $ clika-cli events tail --types job_completed {"id":1041,"type":"job_completed","job":"nightly-llm-sweep-jetson-01","status":"completed","at":"2026-09-03T09:12:44Z"} {"id":1042,"type":"job_completed","job":"nightly-llm-sweep-orin-02","status":"completed","at":"2026-09-03T09:13:02Z"} ``` ``` clika-cli events tail --cursor 1041 --once ``` ### `events list` ``` clika-cli events list [flags] ``` Reads one page and exits. Use it when you want a page rather than a stream. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--cursor` | int | `0` | Return events after this position. | | `--limit` | int | server default 100 | Events per page, up to 500. | | `--types` | string | all | Comma-separated event names. | | `--wait_s` | int | `0` | When nothing is past `--cursor` yet, hold the read open this many seconds waiting for the next event. | | `--stream` | bool | `false` | Stream the raw live server-sent-events push instead of reading the feed. There is no replay, no filtering, and a buffering ingress can hold it back indefinitely, which the hosted platform's did. Prefer `events tail`, which long-polls the durable feed and delivers through any proxy. | Resume from the response's `next_cursor`, or from a row's `id`. ``` clika-cli events list --types job_completed --limit 20 clika-cli events list --cursor 1041 --wait_s 30 ``` ## Metrics ``` clika-cli metrics data [flags] clika-cli metrics log --body '' ``` `metrics data` queries the stored data points for one metric by name. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--start` | string | none | Range start, RFC 3339. | | `--end` | string | none | Range end, RFC 3339. | | `--limit` | string | `1000` | Maximum points to return. | | `--offset` | string | `0` | Offset, for paging through a long range. | ``` clika-cli metrics data tokens_per_second --start 2026-09-01T00:00:00Z --end 2026-09-03T00:00:00Z ``` `metrics log` writes data points in, which is how an external system feeds numbers to the platform that no job produced. ## Alerts ``` clika-cli alerts rules clika-cli alerts rules-id clika-cli alerts rules-create --body '' clika-cli alerts rules-update-id --body '' clika-cli alerts rules-delete-id clika-cli alerts history ``` An alert rule is a condition over the platform's state or metrics. `alerts history` is what fired and when, which is the first thing to read when someone asks whether an alert actually went off. ## Notifications Notifications are the in-app inbox for a person. | Command | What it does | | --- | --- | | `notifications list` | Your notifications. | | `notifications mark-read ` | Mark one read. | | `notifications mark-all-read` | Mark them all read. | | `notifications delete-delete-id ` | Delete one. | | `notifications delete-delete` | Clear them all. | ## Webhooks Webhooks are the outbound half. The platform posts to a URL you own when something happens. | Command | What it does | | --- | --- | | `webhooks list` | Configured webhooks. | | `webhooks get ` | One webhook. | | `webhooks create --body ''` | Add one. | | `webhooks update --body ''` | Change one. | | `webhooks test ` | Send a test delivery, so you can confirm the endpoint before relying on it. | | `webhooks delete ` | Remove one. | Prefer a webhook when another system should react. Prefer `events tail` when a script of yours should react and you would rather not run a listening endpoint. ## Audit ``` clika-cli audit list clika-cli audit get ``` `audit list` is the activity log across the organization. `audit get` narrows it to one resource, for example every change ever made to one device, which is the query you want during an investigation. Licensing actions (issue, reveal, rotate, revoke) are recorded in the same trail. See [licensing](licensing.md). ## Deployment health Three unauthenticated-style probes, exposed as commands like everything else: | Command | What it reports | | --- | --- | | `health list` | The overall health check, dependencies included. | | `livez list` | Liveness: the process is up. | | `ready list` | Readiness: the process is up and able to serve. | They are most useful when you are debugging a deployment rather than using one. A `ready` that fails while `livez` passes usually means a dependency, such as the database, is not available yet. Each body carries `version`, the running build's version (the value the start-up log names, `dev` on a local build), so a script can tell which release a deployment runs without reading its image tag. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli events` `events` has 2 subcommands. #### `clika-cli events list` Reads one page of the durable org event feed (GET /api/events?cursor=) and exits: job/transfer/command completions, device online/offline edges, alert triggers and more. Resume from a response's next_cursor (or a row's ID) with --cursor; --wait_s holds an empty read open until the next event lands. Use 'events tail' to follow the feed, or --stream for the raw live SSE push (no replay, unfiltered). ``` clika-cli events list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--cursor` | int | none | return events after this feed position (0 = from the start of the retention window) | | `--limit` | int | none | maximum events per page (server default 100, maximum 500) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--stream` | bool | `false` | stream the live SSE push instead of reading the feed (no replay, no filters; a proxy that buffers responses still blocks it, so prefer 'events tail') | | `--types` | string | none | comma-separated event names to include (e.g. transfer_completed,job_completed) | | `--wait_s` | int | none | when nothing is past --cursor yet, hold the read open up to this many seconds for the next event | Endpoint: `GET /api/events`. MCP tool name: `get_events`. #### `clika-cli events tail` Streams the durable org event feed (GET /api/events?cursor=) as one JSON line per event: job/transfer/command completions, device online/offline edges, alert triggers and more. Prints everything retained after --cursor (default 0 = the whole retention window), then blocks on the server's wait_s long-poll and prints each new event the moment it lands, no tick-polling. The last line's "id" is the cursor to resume from. ``` clika-cli events tail [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--cursor` | int | none | start after this feed position (0 = from the start of the retention window) | | `--once` | bool | `false` | print what the feed holds after --cursor and exit instead of following | | `--types` | string | none | comma-separated event names to include (e.g. transfer_completed,job_completed) | ### `clika-cli metrics` `metrics` has 2 subcommands. #### `clika-cli metrics data` Query metric data points ``` clika-cli metrics data [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--end` | string | none | Range end (RFC3339) | | `--limit` | string | none | Max points to return (default: 1000) | | `--offset` | string | none | Offset for pagination | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--start` | string | none | Range start (RFC3339) | Endpoint: `GET /api/metrics/{name}/data`. MCP tool name: `get_metrics_name_data`. #### `clika-cli metrics log` Log metric data points ``` clika-cli metrics log [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/metrics/log`. MCP tool name: `post_metrics_log`. ### `clika-cli alerts` `alerts` has 6 subcommands. #### `clika-cli alerts history` List alert history ``` clika-cli alerts history [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page (default: 50, max: 500) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/alerts/history`. MCP tool name: `get_alerts_history`. #### `clika-cli alerts rules` List alert rules ``` clika-cli alerts rules [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/alerts/rules`. MCP tool name: `get_alerts_rules`. #### `clika-cli alerts rules-create` Create alert rule ``` clika-cli alerts rules-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/alerts/rules`. MCP tool name: `post_alerts_rules`. #### `clika-cli alerts rules-delete-id` Delete alert rule ``` clika-cli alerts rules-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/alerts/rules/{id}`. MCP tool name: `delete_alerts_rules_id`. #### `clika-cli alerts rules-id` Get alert rule ``` clika-cli alerts rules-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/alerts/rules/{id}`. MCP tool name: `get_alerts_rules_id`. #### `clika-cli alerts rules-update-id` Update alert rule ``` clika-cli alerts rules-update-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/alerts/rules/{id}`. MCP tool name: `put_alerts_rules_id`. ### `clika-cli notifications` `notifications` has 5 subcommands. #### `clika-cli notifications delete-delete` Clear all notifications ``` clika-cli notifications delete-delete [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/notifications`. MCP tool name: `delete_notifications`. #### `clika-cli notifications delete-delete-id` Delete notification ``` clika-cli notifications delete-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/notifications/{id}`. MCP tool name: `delete_notifications_id`. #### `clika-cli notifications list` List notifications ``` clika-cli notifications list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | `50, max 200` | Page size | | `--offset` | string | `0` | Page offset | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--unread` | string | none | When true, return only unread notifications | Endpoint: `GET /api/notifications`. MCP tool name: `get_notifications`. #### `clika-cli notifications mark-all-read` Mark all notifications as read ``` clika-cli notifications mark-all-read [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/notifications/mark-all-read`. MCP tool name: `post_notifications_mark_all_read`. #### `clika-cli notifications mark-read` Mark notification as read ``` clika-cli notifications mark-read [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/notifications/{id}/mark-read`. MCP tool name: `post_notifications_id_mark_read`. ### `clika-cli audit` `audit` has 2 subcommands. #### `clika-cli audit get` Resource audit history ``` clika-cli audit get [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | `100, max 500` | Page size | | `--offset` | string | `0` | Page offset | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/audit/{target_type}/{target_id}`. MCP tool name: `get_audit_target_type_target_id`. #### `clika-cli audit list` List audit / activity events ``` clika-cli audit list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--action` | string | none | Filter by action name | | `--actor_id` | string | none | Filter by actor id | | `--category` | string | none | mutation\|authz\|access\|utilization | | `--limit` | string | `100, max 500` | Page size | | `--offset` | string | `0` | Page offset | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--since` | string | none | Inclusive lower bound (RFC3339) | | `--target_id` | string | none | Filter by target id | | `--target_type` | string | none | Filter by target type | | `--until` | string | none | Inclusive upper bound (RFC3339) | Endpoint: `GET /api/audit`. MCP tool name: `get_audit`. ### `clika-cli webhooks` `webhooks` has 6 subcommands. #### `clika-cli webhooks create` Create webhook ``` clika-cli webhooks create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/webhooks`. MCP tool name: `post_webhooks`. #### `clika-cli webhooks delete` Delete webhook ``` clika-cli webhooks delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/webhooks/{id}`. MCP tool name: `delete_webhooks_id`. #### `clika-cli webhooks get` Get webhook ``` clika-cli webhooks get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/webhooks/{id}`. MCP tool name: `get_webhooks_id`. #### `clika-cli webhooks list` List webhooks ``` clika-cli webhooks list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/webhooks`. MCP tool name: `get_webhooks`. #### `clika-cli webhooks test` Test webhook ``` clika-cli webhooks test [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/webhooks/{id}/test`. MCP tool name: `post_webhooks_id_test`. #### `clika-cli webhooks update` Update webhook ``` clika-cli webhooks update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/webhooks/{id}`. MCP tool name: `put_webhooks_id`. ### `clika-cli health` `health` has 1 subcommands. #### `clika-cli health list` Health check ``` clika-cli health list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/health`. MCP tool name: `get_health`. ### `clika-cli livez` `livez` has 1 subcommands. #### `clika-cli livez list` Liveness check ``` clika-cli livez list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/livez`. MCP tool name: `get_livez`. ### `clika-cli ready` `ready` has 1 subcommands. #### `clika-cli ready list` Readiness check ``` clika-cli ready list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/ready`. MCP tool name: `get_ready`. ## Related - [Jobs](jobs.md) and [benchmarks](benchmarks.md): the work these events and metrics describe. - [Devices](devices.md): device health, which is per-device rather than deployment-wide. - [Organizations, projects and access](projects-and-orgs.md): who is allowed to read the audit log. --- # Other resource groups The remaining clika-cli resource groups, and a complete index of every top-level command with the page that documents it. Source: https://docs.clika.io/platform/cli/other-resources.md The pages before this one cover the groups you reach for daily. This one covers what is left, and ends with an index of every top-level command so nothing in the tree is undocumented. ## Agent and engine packages The device agent and the inference engine are two separately versioned things the platform distributes to devices. | Command | What it does | | --- | --- | | `agent-versions list` | Agent versions this deployment publishes. | | `agent-versions latest` | The newest stable one, which is what a device updates to by default. | | `agent-versions create --body ''` | Register a new agent version. | | `agent-versions delete ` | Remove one. | | `engine-packages list` | Inference engine package versions. | | `engine-packages latest` | The newest enabled one. | | `engine-packages create --body ''` | Register a version. | | `engine-packages delete ` | Remove one. | The engine package is entitlement-gated and delivered to a device separately from the agent, which is why a [job definition](job-definitions.md) can require it through `required_resources.engine` and that requirement cannot be forced past. Provisioning the engine onto one device is `devices engine-package `; see [devices](devices.md#agent-updates-and-certificates). ## `agent self` ``` clika-cli agent self ``` Deregisters the calling device. This is an endpoint a device agent calls for itself, not something you normally run, because from your own terminal there is no calling device. Deregister a device you administer with `devices delete` instead. ## `internal mqtt-authn` ``` clika-cli internal mqtt-authn --body '' ``` Answers the message broker's authentication callback, so a device is admitted on the enrolment session token it already holds. The broker calls this, not you. It is documented here only because every documented operation becomes a subcommand. ## Single sign-on, deployment-wide | Command | What it does | | --- | --- | | `sso providers` | Identity providers configured on the deployment. | | `sso providers-id ` | One provider. | | `sso providers-create --body ''` | Add one. | | `sso providers-update-id --body ''` | Change one. | | `sso providers-delete-id ` | Remove one. | The per-organization equivalents live under `orgs sso-providers*`, and the public read used by a sign-in page is `auth sso-providers`. See [organizations, projects and access](projects-and-orgs.md#single-sign-on). ## Billing catalogue ``` clika-cli billing catalog ``` The plans and prices this deployment offers. The subscription commands that act on them are under `orgs`; see [billing and plans](projects-and-orgs.md#billing-and-plans). `billing stripe-webhook` is the payment provider's inbound webhook, called by the provider rather than by you. ## Public share reads ``` clika-cli public benchmark-shares clika-cli public benchmark-shares-card.png ``` Reads a benchmark result through a share token, which is the same path an anonymous visitor takes. Create the link with `benchmark-groups share-mark`; see [sharing results](benchmarks.md#sharing-results). ## Support windows against your organization ``` clika-cli support-grants list ``` The read-only view of any support window currently open against your organization: who, why, and until when. Opening and revoking a window is a platform administrator action, performed from the admin dashboard in the web app. ## Your own profile | Command | What it does | | --- | --- | | `profile update --body ''` | Update your own profile. | | `profile avatar-create` | Upload a profile picture. | | `profile avatar-delete` | Remove it. | | `profile avatar-user_id ` | Read another user's avatar. | ## Complete command index Every top-level command, in the order `clika-cli --help` prints it, with where it is documented. | Command | Documented in | | --- | --- | | `login` | [Authentication and profiles](authentication.md) | | `logout` | [Authentication and profiles](authentication.md) | | `self-update` | [Keeping the CLI current](self-update.md) | | `version` | [Keeping the CLI current](self-update.md) | | `activation-codes` | [Licensing](licensing.md#activation-codes) | | `agent` | This page | | `agent-versions` | This page | | `alerts` | [Events, metrics and alerts](monitoring.md#alerts) | | `artifacts` | [Artifacts and models](artifacts.md) | | `audit` | [Events, metrics and alerts](monitoring.md#audit) | | `auth` | [Authentication and profiles](authentication.md) | | `benchmark-categories` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmark-compatibility` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmark-groups` | [Benchmarks](benchmarks.md) | | `benchmark-io-schemas` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmark-metrics` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmark-task-categories` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmark-types` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `benchmarks` | [Benchmarks](benchmarks.md) | | `billing` | This page | | `cloud-credentials` | [Cloud instances](cloud.md#cloud-credentials) | | `cloud-images` | [Cloud instances](cloud.md#images) | | `cloud-instances` | [Cloud instances](cloud.md#instances) | | `comparisons` | [Benchmarks](benchmarks.md#comparing-runs) | | `dataset-builds` | [Artifacts and models](artifacts.md#server-side-dataset-builds) | | `devices` | [Devices](devices.md) | | `engine-packages` | This page | | `enrollment-tokens` | [Devices](devices.md#enrolling-a-new-device) | | `entitlement-profiles` | [Licensing](licensing.md#entitlement-profiles) | | `events` | [Events, metrics and alerts](monitoring.md#events) | | `health` | [Events, metrics and alerts](monitoring.md#deployment-health) | | `hosted-devices` | [Devices](devices.md#hosted-devices) | | `huggingface-tasks` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `install` | [Keeping the CLI current](self-update.md#the-device-agent-is-a-separate-channel) | | `internal` | This page | | `job-definitions` | [Job definitions](job-definitions.md) | | `jobs` | [Jobs](jobs.md) | | `license` | [Licensing](licensing.md#the-platform-license) | | `livez` | [Events, metrics and alerts](monitoring.md#deployment-health) | | `metrics` | [Events, metrics and alerts](monitoring.md#metrics) | | `model-recommendations` | [Benchmarks](benchmarks.md#the-benchmark-catalogs) | | `models` | [Artifacts and models](artifacts.md#models) | | `notifications` | [Events, metrics and alerts](monitoring.md#notifications) | | `orgs` | [Organizations, projects and access](projects-and-orgs.md#organizations) | | `permissions` | [Organizations, projects and access](projects-and-orgs.md#roles-and-permissions) | | `pki` | [Licensing](licensing.md#certificates) | | `profile` | This page | | `projects` | [Organizations, projects and access](projects-and-orgs.md#projects) | | `public` | This page | | `ready` | [Events, metrics and alerts](monitoring.md#deployment-health) | | `registry-credentials` | [Artifacts and models](artifacts.md#registry-credentials) | | `roles` | [Organizations, projects and access](projects-and-orgs.md#roles-and-permissions) | | `runtime-sdk` | [Download the ClikaRT SDK](../how-to/download-the-clikart-sdk.md#download-with-the-platform-cli) | | `service-definitions` | [Services and service definitions](services.md#service-definitions) | | `services` | [Services and service definitions](services.md#running-services) | | `sso` | This page | | `support-grants` | This page | | `tag-scripts` | [Devices](devices.md#tags-and-tag-scripts) | | `transfers` | [Transfers](transfers.md) | | `users` | [Organizations, projects and access](projects-and-orgs.md#users) | | `v1` | [Licensing](licensing.md#the-runtime-facing-api) for the runtime endpoints, [model deployment](model-deployment.md) for the servings | | `vpn` | [Cloud instances](cloud.md#vpn) | | `webhooks` | [Events, metrics and alerts](monitoring.md#webhooks) | | `apply` | [apply and export](apply-export.md) | | `export` | [apply and export](apply-export.md) | | `call` | [MCP server and generic dispatch](mcp.md#call) | | `mcp` | [MCP server and generic dispatch](mcp.md) | | `tools` | [MCP server and generic dispatch](mcp.md#tools) | | `completion` | [CLI overview](index.md#shell-completion) | | `help` | [CLI overview](index.md#built-in-prose-help) | ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli agent` `agent` has 3 subcommands. #### `clika-cli agent jobs-run-credential` Fetch the calling device's run credential for a job ``` clika-cli agent jobs-run-credential [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/agent/jobs/{id}/run-credential`. MCP tool name: `get_agent_jobs_id_run_credential`. #### `clika-cli agent self` Deregister the calling device ``` clika-cli agent self [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/agent/self`. MCP tool name: `delete_agent_self`. #### `clika-cli agent services-run-credential` Fetch the calling device's run credential for a model serving instance ``` clika-cli agent services-run-credential [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/agent/services/{service_id}/run-credential`. MCP tool name: `get_agent_services_service_id_run_credential`. ### `clika-cli agent-versions` `agent-versions` has 3 subcommands. #### `clika-cli agent-versions latest` Get latest stable agent version ``` clika-cli agent-versions latest [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/agent-versions/latest`. MCP tool name: `get_agent_versions_latest`. #### `clika-cli agent-versions list` List agent versions ``` clika-cli agent-versions list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--fields` | string | none | Comma-separated optional fields to include on each row: checksums, download_urls, signatures | | `--limit` | string | `50, maximum 200; a larger value is clamped` | Rows to return | | `--offset` | string | `0` | Rows to skip from the newest | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/agent-versions`. MCP tool name: `get_agent_versions`. #### `clika-cli agent-versions rollbacks` List devices rolled back from an agent version ``` clika-cli agent-versions rollbacks [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | none | Page number (default: 1) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/agent-versions/{version}/rollbacks`. MCP tool name: `get_agent_versions_version_rollbacks`. ### `clika-cli billing` `billing` has 2 subcommands. #### `clika-cli billing catalog` Billing catalog ``` clika-cli billing catalog [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/billing/catalog`. MCP tool name: `get_billing_catalog`. #### `clika-cli billing stripe-webhook` Stripe webhook intake ``` clika-cli billing stripe-webhook [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/billing/stripe/webhook`. MCP tool name: `post_billing_stripe_webhook`. ### `clika-cli internal` `internal` has 1 subcommands. #### `clika-cli internal mqtt-authn` Authenticate an MQTT client on behalf of the broker ``` clika-cli internal mqtt-authn [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/internal/mqtt/authn`. MCP tool name: `post_internal_mqtt_authn`. ### `clika-cli legal` `legal` has 4 subcommands. #### `clika-cli legal documents` Current legal documents ``` clika-cli legal documents [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/legal/documents`. MCP tool name: `get_legal_documents`. #### `clika-cli legal documents-kind` Current legal document ``` clika-cli legal documents-kind [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/legal/documents/{kind}`. MCP tool name: `get_legal_documents_kind`. #### `clika-cli legal documents-versions` Legal document versions ``` clika-cli legal documents-versions [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/legal/documents/{kind}/versions`. MCP tool name: `get_legal_documents_kind_versions`. #### `clika-cli legal documents-versions-version` Legal document version ``` clika-cli legal documents-versions-version [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/legal/documents/{kind}/versions/{version}`. MCP tool name: `get_legal_documents_kind_versions_version`. ### `clika-cli profile` `profile` has 6 subcommands. #### `clika-cli profile avatar-create` Upload profile avatar ``` clika-cli profile avatar-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/profile/avatar`. MCP tool name: `post_profile_avatar`. #### `clika-cli profile avatar-delete` Delete profile avatar ``` clika-cli profile avatar-delete [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/profile/avatar`. MCP tool name: `delete_profile_avatar`. #### `clika-cli profile avatar-user_id` Get profile avatar ``` clika-cli profile avatar-user_id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--exp` | string | none | Unix-seconds expiry (matched by the HMAC) | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--size` | string | none | Variant size; 256 (default) or 64 | | `--token` | string | none | HMAC signature token | Endpoint: `GET /api/profile/avatar/{user_id}`. MCP tool name: `get_profile_avatar_user_id`. #### `clika-cli profile update` Update own profile ``` clika-cli profile update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/profile`. MCP tool name: `put_profile`. #### `clika-cli profile walkthroughs` List completed walkthroughs ``` clika-cli profile walkthroughs [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/profile/walkthroughs`. MCP tool name: `get_profile_walkthroughs`. #### `clika-cli profile walkthroughs-update-key` Mark a walkthrough completed ``` clika-cli profile walkthroughs-update-key [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/profile/walkthroughs/{key}`. MCP tool name: `put_profile_walkthroughs_key`. ### `clika-cli public` `public` has 2 subcommands. #### `clika-cli public benchmark-shares` Get shared benchmark result (public) ``` clika-cli public benchmark-shares [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/public/benchmark-shares/{token}`. MCP tool name: `get_public_benchmark_shares_token`. #### `clika-cli public benchmark-shares-card.png` Get shared benchmark card image (public) ``` clika-cli public benchmark-shares-card.png [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/public/benchmark-shares/{token}/card.png`. MCP tool name: `get_public_benchmark_shares_token_card.png`. ### `clika-cli runtime` `runtime` has 3 subcommands. #### `clika-cli runtime sdk` List the ClikaRT SDK downloads offered, by selection ``` clika-cli runtime sdk [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/runtime/sdk`. MCP tool name: `get_runtime_sdk`. #### `clika-cli runtime sdk-download-token` Mint a download link for a ClikaRT SDK selection ``` clika-cli runtime sdk-download-token [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/runtime/sdk/{dist}/download-token`. MCP tool name: `post_runtime_sdk_dist_download_token`. #### `clika-cli runtime sdk-pip-index-token` Mint the pip index link for the ClikaRT Python wheels ``` clika-cli runtime sdk-pip-index-token [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/runtime/sdk/pip-index-token`. MCP tool name: `post_runtime_sdk_pip_index_token`. ### `clika-cli runtime-sdk` `runtime-sdk` has 3 subcommands. #### `clika-cli runtime-sdk download` Downloads the selection named by `` (an id from `runtime-sdk list`, such as linux-amd64-cpp-cu12, or a platform id such as linux-amd64 for the vendor's whole archive) or by the facet flags (--platform with --build, and where they apply --cuda and --python) to ``, which defaults to the current directory. When `` is an existing directory the file keeps its name inside it; any other `` is the file path to write. ``` clika-cli runtime-sdk download [] [dest] [flags] ``` Positional arguments: required ``; optional `[]`, `[dest]`. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--build` | string | none | what is built: python, cpp, cli, android, rust, go, kotlin (Linux only), or everything for the vendor's whole archive | | `--cuda` | string | none | the CUDA image on Linux: 12, 13, both (an archive build) or none | | `--platform` | string | none | the platform the program runs on: linux-amd64, linux-arm64, darwin-arm64, windows-amd64, windows-arm64 or android-arm64 | | `--python` | string | none | the interpreter for a Python wheel download: cp310, cp311, cp312, cp313 or cp314 | #### `clika-cli runtime-sdk list` Lists one row per selection the deployment can deliver now: the id `download` takes, the platform, the build, the CUDA image, the interpreter, the backends inside and the size. A selection whose files are not staged on the deployment is left out; --all lists every selection with its state. The facet flags narrow the rows. An older platform that lists only the whole archives shows them as before. ``` clika-cli runtime-sdk list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--all` | bool | `false` | list every selection, the ones not staged on this deployment included, with its state | | `--build` | string | none | what is built: python, cpp, cli, android, rust, go, kotlin (Linux only), or everything for the vendor's whole archive | | `--cuda` | string | none | the CUDA image on Linux: 12, 13, both (an archive build) or none | | `--platform` | string | none | the platform the program runs on: linux-amd64, linux-arm64, darwin-arm64, windows-amd64, windows-arm64 or android-arm64 | | `--python` | string | none | the interpreter for a Python wheel download: cp310, cp311, cp312, cp313 or cp314 | #### `clika-cli runtime-sdk pip-command` Mints a short-lived link to the platform's own package index for the ClikaRT Python packages and prints the one pip command that installs clika_runtime and clika_modelverse at the engine version this deployment runs. pip picks the wheel for the interpreter it runs under; --cuda picks the CUDA image of the runtime wheel on Linux (12 or 13; none is the CPU and Vulkan build). The link stops working ten minutes after it was minted, so run the command soon. ``` clika-cli runtime-sdk pip-command [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--cuda` | string | none | the CUDA image of the runtime wheel: 12, 13 or none | ### `clika-cli sso` `sso` has 5 subcommands. #### `clika-cli sso providers` List SSO providers (admin) ``` clika-cli sso providers [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/sso/providers`. MCP tool name: `get_sso_providers`. #### `clika-cli sso providers-create` Create SSO provider ``` clika-cli sso providers-create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/sso/providers`. MCP tool name: `post_sso_providers`. #### `clika-cli sso providers-delete-id` Delete SSO provider ``` clika-cli sso providers-delete-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/sso/providers/{id}`. MCP tool name: `delete_sso_providers_id`. #### `clika-cli sso providers-id` Get SSO provider ``` clika-cli sso providers-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/sso/providers/{id}`. MCP tool name: `get_sso_providers_id`. #### `clika-cli sso providers-update-id` Update SSO provider ``` clika-cli sso providers-update-id [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/sso/providers/{id}`. MCP tool name: `put_sso_providers_id`. ### `clika-cli support-grants` `support-grants` has 3 subcommands. #### `clika-cli support-grants approve` Approve a support access request ``` clika-cli support-grants approve [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/support-grants/{id}/approve`. MCP tool name: `post_support_grants_id_approve`. #### `clika-cli support-grants deny` Deny a support access request ``` clika-cli support-grants deny [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/support-grants/{id}/deny`. MCP tool name: `post_support_grants_id_deny`. #### `clika-cli support-grants list` Support windows against your organization ``` clika-cli support-grants list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--active` | string | none | Only live windows | | `--pending` | string | none | Only requests awaiting your organization's decision | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/support-grants`. MCP tool name: `get_support_grants`. ## Related - [CLI overview](index.md): global flags, output formats and exit codes. - [Cheat sheet](cheatsheet.md): the twenty commands worth memorising. --- # Organizations, projects and access clika-cli orgs, projects, users, roles and permissions: the tenancy and access-control commands. Source: https://docs.clika.io/platform/cli/projects-and-orgs.md Everything on the platform belongs to an **organization**, which is the tenant boundary: its own members, its own devices, its own billing. Inside an organization, a **project** groups resources and scopes who can see them. Every session and every API key is bound to one organization and one project at a time, which is why almost every other command implicitly means "in the project I am currently in". Access is granted through **roles**. A role is a named bundle of **permissions**, and permissions are the strings the platform checks on each request, for example `devices:read` or `jobs:write`. Most of this is comfortable in the web app. The commands here exist so you can script onboarding, audit access, and read the current state without clicking. ## Where am I ``` clika-cli auth me clika-cli auth me-orgs clika-cli auth me-projects ``` `auth me` prints the account behind your current credential; the other two list what it can reach. Move the session: ``` clika-cli auth switch-org --body '{"org_id":"..."}' clika-cli auth switch-project --body '{"project_id":"..."}' ``` ## Projects ``` clika-cli projects [command] ``` | Subcommand | Purpose | | --- | --- | | `list` | Projects you can see. | | `get ` | One project. | | `create --body ''` | Create a project. Needs an organization `owner` or `admin`; a project role does not confer it. | | `update --body ''` | Rename or reconfigure one. | | `delete ` | Delete one. Resources belonging to it are reassigned or removed according to the platform's rules, so read `get` first. | | `members ` | Who is on the project. | | `members-create --body ''` | Add a member. | | `members-update-user_id --body ''` | Change a member's project role. | | `members-delete-user_id ` | Remove a member. | An organization always has a default project, and the last owner of a project cannot be removed, so you cannot lock yourself out by mistake. Project licenses are covered on the [licensing page](licensing.md); they are `projects licenses*` because a runtime credential is issued to a project. ## Organizations ``` clika-cli orgs [command] ``` The organization surface is large. It divides into five parts. ### Membership | Subcommand | Purpose | | --- | --- | | `list` | Organizations you can see. | | `get ` | One organization. | | `create --body ''` | Create one. | | `update --body ''` | Change its settings. | | `delete ` | Delete it. | | `members ` | Its members and their roles. | | `members-create --body ''` | Invite a member. | | `members-update-user_id --body ''` | Change a member's role. | | `members-delete-user_id ` | Remove a member. | | `members-reinvite ` | Send the invitation again, which supersedes the previous one. | | `transfer-ownership --body ''` | Hand ownership to another member. | ### Teams Teams group members inside an organization so access can be granted to a group rather than person by person. | Subcommand | Purpose | | --- | --- | | `teams ` | List teams. | | `teams-team_id ` | Get one team. | | `teams-create --body ''` | Create a team. | | `teams-update-team_id --body ''` | Rename or reconfigure it. | | `teams-delete-team_id ` | Delete it. | | `teams-members ` | List its members. | | `teams-members-create --body ''` | Add a member. | | `teams-members-delete-user_id ` | Remove a member. | ### Single sign-on | Subcommand | Purpose | | --- | --- | | `sso-providers ` | The identity providers configured for this organization. | | `sso-providers-id ` | One provider. | | `sso-providers-create --body ''` | Add one. | | `sso-providers-update-id --body ''` | Change one. | | `sso-providers-delete-id ` | Remove one. | | `sso-domains ` | Email domains verified for this organization. | | `sso-domains-verifications --body ''` | Start verifying a domain. | | `sso-domains-verifications-confirm --body ''` | Confirm the verification once the DNS record is in place. | | `sso-domains-delete-id ` | Remove a verified domain. | A verified domain is what lets someone with an address at that domain be routed to your identity provider automatically. People who sign in through SSO have no password, so they authenticate the CLI with an [API key](authentication.md#minting-an-api-key-for-ci). The `sso` group is the deployment-wide counterpart: `sso providers`, `sso providers-create` and `sso providers-id`. ### Billing and plans | Subcommand | Purpose | | --- | --- | | `subscription ` | The current subscription. | | `subscription-checkout --body ''` | Start a checkout to subscribe. | | `subscription-change-plan --body ''` | Move to another plan. | | `subscription-cancel ` | Cancel at the end of the period. | | `subscription-resume ` | Undo a pending cancellation. | | `credits-ledger ` | Credit balance and history. | | `credits-checkout --body ''` | Buy more credits. | | `billing-portal ` | A link to the hosted billing portal. | | `plan --body ''` | Assign the organization's entitlement plan, where an organization administrator is allowed to self-select. | | `billing catalog` | The plans and prices this deployment offers. | ### Users `users` is the account surface. On most deployments it is restricted to administrators, because it reaches accounts across the organization. | Subcommand | Purpose | | --- | --- | | `list` | Accounts. | | `get ` | One account. | | `create --body ''` | Create an account. | | `update --body ''` | Change one. | | `delete ` | Delete one. | | `roles ` | The roles assigned to an account. | | `roles-create --body ''` | Assign a role. | | `roles-delete-role_id ` | Revoke a role. | | `export ` | Export an account's personal data, the data-subject access request. | | `erase ` | Erase an account, the right-to-erasure request under GDPR Article 17. Distinct from `delete`, and irreversible. | ## Roles and permissions ``` clika-cli permissions list clika-cli roles list ``` `permissions list` is the catalogue: every permission string the platform checks, with what it governs. Read it before writing a role or scoping an API key, because these strings are the exact values both take. | Subcommand | Purpose | | --- | --- | | `roles list` | Every role. | | `roles get ` | One role and its permissions. | The role catalogue is fixed: roles are read, assigned and unassigned, never created or edited. Two rules are worth holding on to, because they explain most refusals: - **A credential can never exceed its owner.** An API key resolves to the intersection of your permissions and the key's scope, so narrowing an account narrows every key it minted. - **The surface does not matter.** The web app, the CLI and MCP all pass the same checks. If a command is refused, the same action would be refused in a browser. ## Cross-tenant support access When CLIKA support needs to look at your organization's data, it is granted through a time-boxed, reason-carrying grant rather than a shared account: ``` clika-cli support-grants list ``` The grant is scoped to one organization, attenuates the capabilities it confers, and expires on its own. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli projects` `projects` has 21 subcommands. #### `clika-cli projects create` Create project ``` clika-cli projects create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects`. MCP tool name: `post_projects`. #### `clika-cli projects delete` Delete project ``` clika-cli projects delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/projects/{id}`. MCP tool name: `delete_projects_id`. #### `clika-cli projects get` Get project ``` clika-cli projects get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects/{id}`. MCP tool name: `get_projects_id`. #### `clika-cli projects licenses` List project licenses ``` clika-cli projects licenses [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects/{id}/licenses`. MCP tool name: `get_projects_id_licenses`. #### `clika-cli projects licenses-create` Issue project license ``` clika-cli projects licenses-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/licenses`. MCP tool name: `post_projects_id_licenses`. #### `clika-cli projects licenses-delete-lid` Delete project license ``` clika-cli projects licenses-delete-lid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/projects/{id}/licenses/{lid}`. MCP tool name: `delete_projects_id_licenses_lid`. #### `clika-cli projects licenses-hardware` List a project license's hardware seats ``` clika-cli projects licenses-hardware [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--page` | string | `1` | Page number | | `--page_size` | string | `25, max 200` | Page size | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects/{id}/licenses/{lid}/hardware`. MCP tool name: `get_projects_id_licenses_lid_hardware`. #### `clika-cli projects licenses-hardware-delete-hardware_id` Release a project license's hardware seat ``` clika-cli projects licenses-hardware-delete-hardware_id [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/projects/{id}/licenses/{lid}/hardware/{hardware_id}`. MCP tool name: `delete_projects_id_licenses_lid_hardware_hardware_id`. #### `clika-cli projects licenses-hardware-release` Release a project license's hardware seat by its id in the body ``` clika-cli projects licenses-hardware-release [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/licenses/{lid}/hardware/release`. MCP tool name: `post_projects_id_licenses_lid_hardware_release`. #### `clika-cli projects licenses-lid` Get project license ``` clika-cli projects licenses-lid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects/{id}/licenses/{lid}`. MCP tool name: `get_projects_id_licenses_lid`. #### `clika-cli projects licenses-reveal` Reveal project license key ``` clika-cli projects licenses-reveal [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/licenses/{lid}/reveal`. MCP tool name: `post_projects_id_licenses_lid_reveal`. #### `clika-cli projects licenses-revoke` Revoke project license ``` clika-cli projects licenses-revoke [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/licenses/{lid}/revoke`. MCP tool name: `post_projects_id_licenses_lid_revoke`. #### `clika-cli projects licenses-rotate` Rotate project license ``` clika-cli projects licenses-rotate [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/licenses/{lid}/rotate`. MCP tool name: `post_projects_id_licenses_lid_rotate`. #### `clika-cli projects licenses-update-lid` Update project license ``` clika-cli projects licenses-update-lid [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PATCH /api/projects/{id}/licenses/{lid}`. MCP tool name: `patch_projects_id_licenses_lid`. #### `clika-cli projects list` List projects ``` clika-cli projects list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects`. MCP tool name: `get_projects`. #### `clika-cli projects members` List project members ``` clika-cli projects members [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/projects/{id}/members`. MCP tool name: `get_projects_id_members`. #### `clika-cli projects members-create` Add project member ``` clika-cli projects members-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/projects/{id}/members`. MCP tool name: `post_projects_id_members`. #### `clika-cli projects members-delete-user_id` Remove project member ``` clika-cli projects members-delete-user_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/projects/{id}/members/{user_id}`. MCP tool name: `delete_projects_id_members_user_id`. #### `clika-cli projects members-update-user_id` Update project member role ``` clika-cli projects members-update-user_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PATCH /api/projects/{id}/members/{user_id}`. MCP tool name: `patch_projects_id_members_user_id`. #### `clika-cli projects servings` List a project's model servings ``` clika-cli projects servings [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli projects update` Update project ``` clika-cli projects update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PATCH /api/projects/{id}`. MCP tool name: `patch_projects_id`. ### `clika-cli orgs` `orgs` has 38 subcommands. #### `clika-cli orgs billing-portal` Open billing portal ``` clika-cli orgs billing-portal [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/billing-portal`. MCP tool name: `post_orgs_org_id_billing_portal`. #### `clika-cli orgs credits-checkout` Start credit top-up checkout ``` clika-cli orgs credits-checkout [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/credits/checkout`. MCP tool name: `post_orgs_org_id_credits_checkout`. #### `clika-cli orgs credits-ledger` Org credit ledger ``` clika-cli orgs credits-ledger [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--limit` | string | `50, max 200` | Page size | | `--offset` | string | `0` | Page offset | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/credits/ledger`. MCP tool name: `get_orgs_org_id_credits_ledger`. #### `clika-cli orgs delete` Delete organization ``` clika-cli orgs delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--acknowledge_live_resources` | string | none | The details.fingerprint of the LIVE_CLOUD_RESOURCES refusal being acknowledged: the delete proceeds only while the live cloud resources are still exactly those, which are then left running | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}`. MCP tool name: `delete_orgs_org_id`. #### `clika-cli orgs get` Get organization ``` clika-cli orgs get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}`. MCP tool name: `get_orgs_org_id`. #### `clika-cli orgs list` List organizations ``` clika-cli orgs list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs`. MCP tool name: `get_orgs`. #### `clika-cli orgs members` List organization members ``` clika-cli orgs members [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/members`. MCP tool name: `get_orgs_org_id_members`. #### `clika-cli orgs members-create` Add organization member ``` clika-cli orgs members-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/members`. MCP tool name: `post_orgs_org_id_members`. #### `clika-cli orgs members-delete-user_id` Remove organization member ``` clika-cli orgs members-delete-user_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}/members/{user_id}`. MCP tool name: `delete_orgs_org_id_members_user_id`. #### `clika-cli orgs members-password-reset-link` Issue a password-reset link for an organization member ``` clika-cli orgs members-password-reset-link [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/members/{user_id}/password-reset-link`. MCP tool name: `post_orgs_org_id_members_user_id_password_reset_link`. #### `clika-cli orgs members-reinvite` Re-invite organization member ``` clika-cli orgs members-reinvite [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/members/{user_id}/reinvite`. MCP tool name: `post_orgs_org_id_members_user_id_reinvite`. #### `clika-cli orgs members-update-user_id` Update organization member role ``` clika-cli orgs members-update-user_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/orgs/{org_id}/members/{user_id}`. MCP tool name: `put_orgs_org_id_members_user_id`. #### `clika-cli orgs plan` Assign org entitlement plan (org admin self-select) ``` clika-cli orgs plan [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/orgs/{org_id}/plan`. MCP tool name: `put_orgs_org_id_plan`. #### `clika-cli orgs sso-domains` List verified SSO domains for an org ``` clika-cli orgs sso-domains [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/sso-domains`. MCP tool name: `get_orgs_org_id_sso_domains`. #### `clika-cli orgs sso-domains-delete-id` Delete a verified SSO domain ``` clika-cli orgs sso-domains-delete-id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}/sso-domains/{id}`. MCP tool name: `delete_orgs_org_id_sso_domains_id`. #### `clika-cli orgs sso-domains-verifications` Initiate SSO domain verification ``` clika-cli orgs sso-domains-verifications [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/sso-domains/verifications`. MCP tool name: `post_orgs_org_id_sso_domains_verifications`. #### `clika-cli orgs sso-domains-verifications-confirm` Confirm SSO domain verification ``` clika-cli orgs sso-domains-verifications-confirm [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/sso-domains/verifications/confirm`. MCP tool name: `post_orgs_org_id_sso_domains_verifications_confirm`. #### `clika-cli orgs sso-providers` List org SSO providers ``` clika-cli orgs sso-providers [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/sso/providers`. MCP tool name: `get_orgs_org_id_sso_providers`. #### `clika-cli orgs sso-providers-create` Create org SSO provider ``` clika-cli orgs sso-providers-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/sso/providers`. MCP tool name: `post_orgs_org_id_sso_providers`. #### `clika-cli orgs sso-providers-delete-id` Delete org SSO provider ``` clika-cli orgs sso-providers-delete-id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}/sso/providers/{id}`. MCP tool name: `delete_orgs_org_id_sso_providers_id`. #### `clika-cli orgs sso-providers-id` Get org SSO provider ``` clika-cli orgs sso-providers-id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/sso/providers/{id}`. MCP tool name: `get_orgs_org_id_sso_providers_id`. #### `clika-cli orgs sso-providers-update-id` Update org SSO provider ``` clika-cli orgs sso-providers-update-id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/orgs/{org_id}/sso/providers/{id}`. MCP tool name: `put_orgs_org_id_sso_providers_id`. #### `clika-cli orgs subscription` Org subscription ``` clika-cli orgs subscription [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/subscription`. MCP tool name: `get_orgs_org_id_subscription`. #### `clika-cli orgs subscription-cancel` Cancel subscription at period end ``` clika-cli orgs subscription-cancel [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/subscription/cancel`. MCP tool name: `post_orgs_org_id_subscription_cancel`. #### `clika-cli orgs subscription-change-plan` Change subscription plan ``` clika-cli orgs subscription-change-plan [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/subscription/change-plan`. MCP tool name: `post_orgs_org_id_subscription_change_plan`. #### `clika-cli orgs subscription-checkout` Start subscription checkout ``` clika-cli orgs subscription-checkout [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/subscription/checkout`. MCP tool name: `post_orgs_org_id_subscription_checkout`. #### `clika-cli orgs subscription-resume` Resume subscription ``` clika-cli orgs subscription-resume [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/subscription/resume`. MCP tool name: `post_orgs_org_id_subscription_resume`. #### `clika-cli orgs subscription-usage` Org subscription usage ``` clika-cli orgs subscription-usage [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/subscription/usage`. MCP tool name: `get_orgs_org_id_subscription_usage`. #### `clika-cli orgs teams` List teams ``` clika-cli orgs teams [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/teams`. MCP tool name: `get_orgs_org_id_teams`. #### `clika-cli orgs teams-create` Create team ``` clika-cli orgs teams-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/teams`. MCP tool name: `post_orgs_org_id_teams`. #### `clika-cli orgs teams-delete-team_id` Delete team ``` clika-cli orgs teams-delete-team_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}/teams/{team_id}`. MCP tool name: `delete_orgs_org_id_teams_team_id`. #### `clika-cli orgs teams-members` List team members ``` clika-cli orgs teams-members [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/teams/{team_id}/members`. MCP tool name: `get_orgs_org_id_teams_team_id_members`. #### `clika-cli orgs teams-members-create` Add team member ``` clika-cli orgs teams-members-create [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/teams/{team_id}/members`. MCP tool name: `post_orgs_org_id_teams_team_id_members`. #### `clika-cli orgs teams-members-delete-user_id` Remove team member ``` clika-cli orgs teams-members-delete-user_id [flags] ``` Positional arguments: required ``, ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/orgs/{org_id}/teams/{team_id}/members/{user_id}`. MCP tool name: `delete_orgs_org_id_teams_team_id_members_user_id`. #### `clika-cli orgs teams-team_id` Get team ``` clika-cli orgs teams-team_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/orgs/{org_id}/teams/{team_id}`. MCP tool name: `get_orgs_org_id_teams_team_id`. #### `clika-cli orgs teams-update-team_id` Update team ``` clika-cli orgs teams-update-team_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/orgs/{org_id}/teams/{team_id}`. MCP tool name: `put_orgs_org_id_teams_team_id`. #### `clika-cli orgs transfer-ownership` Transfer organization ownership ``` clika-cli orgs transfer-ownership [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/orgs/{org_id}/transfer-ownership`. MCP tool name: `post_orgs_org_id_transfer_ownership`. #### `clika-cli orgs update` Update organization ``` clika-cli orgs update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/orgs/{org_id}`. MCP tool name: `put_orgs_org_id`. ### `clika-cli users` `users` has 10 subcommands. #### `clika-cli users create` Create user ``` clika-cli users create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/users`. MCP tool name: `post_users`. #### `clika-cli users delete` Delete user ``` clika-cli users delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/users/{id}`. MCP tool name: `delete_users_id`. #### `clika-cli users erase` Erase user (GDPR Art 17) ``` clika-cli users erase [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/users/{id}/erase`. MCP tool name: `post_users_id_erase`. #### `clika-cli users export` Export my personal data ``` clika-cli users export [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/users/{id}/export`. MCP tool name: `get_users_id_export`. #### `clika-cli users get` Get user ``` clika-cli users get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/users/{id}`. MCP tool name: `get_users_id`. #### `clika-cli users list` List users ``` clika-cli users list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/users`. MCP tool name: `get_users`. #### `clika-cli users roles` List user roles ``` clika-cli users roles [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/users/{id}/roles`. MCP tool name: `get_users_id_roles`. #### `clika-cli users roles-create` Assign role to user ``` clika-cli users roles-create [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/users/{id}/roles`. MCP tool name: `post_users_id_roles`. #### `clika-cli users roles-delete-role_id` Revoke role from user ``` clika-cli users roles-delete-role_id [flags] ``` Positional arguments: required ``, ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/users/{id}/roles/{role_id}`. MCP tool name: `delete_users_id_roles_role_id`. #### `clika-cli users update` Update user ``` clika-cli users update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/users/{id}`. MCP tool name: `put_users_id`. ### `clika-cli roles` `roles` has 2 subcommands. #### `clika-cli roles get` Get role ``` clika-cli roles get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/roles/{id}`. MCP tool name: `get_roles_id`. #### `clika-cli roles list` List roles ``` clika-cli roles list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/roles`. MCP tool name: `get_roles`. ### `clika-cli permissions` `permissions` has 1 subcommands. #### `clika-cli permissions list` List permissions ``` clika-cli permissions list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/permissions`. MCP tool name: `get_permissions`. ## Related - [Authentication and profiles](authentication.md): the credentials these permissions apply to, and how to scope an API key. - [Licensing](licensing.md): project licenses, which are issued per project. - [Organization and project concept](../concepts/organization-and-project.mdx): what an organization and a project are, in prose. - [CLI overview](index.md): global flags and output formats. --- # Keeping the CLI current clika-cli self-update and version: pull a newer CLI from the deployment's own release channel, with checksum and signature verification. Source: https://docs.clika.io/platform/cli/self-update.md Each deployment serves the CLI binaries built from its own commit, so the right version of `clika-cli` is whichever one that deployment publishes. `self-update` fetches it, and `version` tells you what you are running now. ## `version` ``` clika-cli version ``` Prints the release version stamped into the binary when it was built. It talks to nothing, so it works with no credentials and no network. ``` $ clika-cli version 1.4.2 ``` A binary you built yourself with a plain `go build` reports `dev`. That is not a cosmetic difference. A non-semantic version cannot be compared against the channel's published version, so `self-update` refuses to act on it. ## `self-update` ``` clika-cli self-update [flags] ``` Reads the latest published version from the deployment's release channel at `/files/installation/cli`, and if it is newer than the running binary, downloads the matching asset for your operating system and architecture, verifies it, and atomically replaces the binary in place. Re-run `clika-cli` afterwards; the process that performed the update is still the old one. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--check` | bool | `false` | Report the latest published version and stop. Nothing is downloaded or replaced. | ``` $ clika-cli self-update --check running 1.4.2, latest 1.5.0 ``` ``` $ clika-cli self-update downloading clika-cli-linux-amd64 1.5.0 ... checksum ok signature ok updated to 1.5.0, re-run clika-cli ``` ### Coming from `clika-rt` Until September 2026 this program was named `clika-rt`. The rename is a clean cut: an installed `clika-rt` does not update itself into `clika-cli`, the old asset names are no longer served, and a profile saved under `~/.config/clika-rt/` is not read. Install `clika-cli` from the deployment's **Settings → CLI** tab (or the asset table on the [install page](./index.md)), run `clika-cli login` once, and remove the old binary. ### It needs a credential The release channel is authenticated, exactly like the rest of the deployment. `self-update` sends the same credential every other command uses (flag, then environment, then profile) on **every** fetch it makes: the version file, the checksum list, the signature list and the binary itself. A CLI that has never logged in cannot update itself, and says so with a `401`. ``` clika-cli login --base-url https://platform.clika.io --api-key clika-cli self-update ``` On an on-premise deployment with a self-signed certificate, pass `--insecure-tls`. This is one place the CLI cannot decide for you whether a certificate authority is trustworthy: ``` clika-cli --profile onprem --insecure-tls self-update ``` ### What is verified Two independent checks, both of which must pass: - **The checksum**, taken from the channel's `checksums.txt`. - **A detached signature** over the binary, verified against a public key embedded in the CLI at build time. The signature is the one that matters, because whoever serves the binary also serves the checksum. A checksum alone proves the download was not corrupted, not that it came from CLIKA. A release build refuses an update that is unsigned, tampered with, or signed by a different key. A binary you built locally embeds no trust key. It logs that signature enforcement is off and falls back to the checksum alone, which is fine for development and is not what you should be shipping to anyone. Transport is HTTPS only, and a redirect that would downgrade the connection is refused. ## Installing in the first place `self-update` replaces an installed CLI; it cannot bootstrap one. For the first install, and for the per-platform asset names, see [the install section of the CLI overview](index.md#install). ## The device agent is a separate channel Do not confuse updating the CLI with updating the **device agent**, the program that runs on each of your devices. They use the same download and verification machinery but are different artifacts on different schedules. | Command | What it does | | --- | --- | | `install list --os ` | Prints the agent install script for a platform, which is what you run on a new device to enroll it. | | `install availability --os ` | Reports whether an agent build is available for that platform. | | `install enrollment-payload` | The Android enrolment QR payload, which is what a phone scans. | | `devices update-create ` | Triggers an agent update on one device. | | `devices batch-update` | Triggers an agent update on many devices. | | `agent-versions list` | The agent versions this deployment publishes. | | `agent-versions latest` | The newest one. | See [devices](devices.md#agent-updates-and-certificates) for the update policy and history commands. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli self-update` #### `clika-cli self-update` Fetches the latest clika-cli version from the orchestrator's release channel (``/files/installation/cli), verifies its checksum and signature, and atomically replaces this binary. Re-run clika-cli after a successful update. ``` clika-cli self-update [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--check` | bool | `false` | only report the latest version, do not update | ### `clika-cli version` #### `clika-cli version` Prints the release version stamped into this binary at build time. A plain `go build` leaves it `dev`, which disables self-update. ``` clika-cli version [flags] ``` ### `clika-cli install` `install` has 3 subcommands. #### `clika-cli install availability` Agent install availability ``` clika-cli install availability [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--os` | string | none | Platform to probe (android, linux, windows). Defaults via User-Agent. | | `--raw` | bool | `false` | print raw response without pretty-printing | #### `clika-cli install enrollment-payload` Android enrollment QR payload ``` clika-cli install enrollment-payload [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | Activation code the device will enroll with (case and the dash do not matter); the payload then carries the addresses and TLS posture without a token | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--token` | string | none | Enrollment token (JWT) to embed in the payload; one of token or code is required | #### `clika-cli install list` Get agent install script ``` clika-cli install list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--code` | string | none | Activation code the installed agent enrolls with, in place of a token (as shown in the Add Device dialog; case and the dash do not matter) | | `--os` | string | none | Override platform detection: linux, windows, android | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--token` | string | none | Pre-fill enrollment token in the generated script | ## Related - [CLI overview](index.md): the install channel, its assets, and the credentials it accepts. - [Authentication and profiles](authentication.md): getting a credential that the channel accepts. - [Devices](devices.md): updating the agents on your fleet. --- # Services and service definitions clika-cli service-definitions and services: the long-running processes the platform keeps alive on a device. Source: https://docs.clika.io/platform/cli/services.md A **service** is a long-running process the platform keeps alive on a device: an HTTP server, a collector, anything that is supposed to stay up. A **service definition** describes one, the way a [job definition](job-definitions.md) describes a piece of work that finishes. The difference in one line: a job finishes, a service does not. This page covers the free-form service feature, where you supply the command. A service that serves one of your models through the CLIKA runtime is a [model deployment](model-deployment.md), which the platform defines for you. ## Service definitions ``` clika-cli service-definitions [command] ``` | Subcommand | Purpose | | --- | --- | | `list` | Every service definition. | | `get ` | One definition. | | `create --body ''` | Create one. | | `update --body ''` | Change one. | | `delete ` | Delete one. | A service definition carries the fields a long-running process needs that a job does not. | Field | Type | Meaning | | --- | --- | --- | | `name` | string | Display name and handle. | | `description` | string | Free text. | | `command` | array of strings | The argument vector to run. Unlike a job definition, it sits at the top level rather than under `script`. | | `env` | object | Environment variables. | | `port` | integer | The port the service listens on, which is what the platform advertises and health-checks. | | `health_check` | object | How to tell whether the process is actually serving, as opposed to merely running. | | `auto_restart` | boolean | Restart the process when it exits. | | `artifacts` | array of objects | Files pushed to the device before the service starts. | | `cleanup_steps` | array of objects | What to run when the service is stopped or removed. | | `required_resources` | object | The same minimums, accelerator list and `engine` gate as a [job definition](job-definitions.md#the-body). | | `project_id` | string | Scopes the definition to one project. | ## Running services `services` and `devices services-*` are the runtime side. They act on the processes actually running, as opposed to the definitions describing them. ``` clika-cli services list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--device_id` | string | all | Only services on one device. | | `--state` | string | all | Filter by service state. | | `--search` | string | none | Match a service id, a device name, or a service type. | | `--page` | string | `1` | Page number. | | `--page_size` | string | server default | Rows per page. | That is an organization-wide view. To act on one device's services, use the [device commands](devices.md#services-on-a-device): ``` clika-cli devices services jetson-01 clika-cli devices services-create jetson-01 --body '{"service_definition_id":"vllm-server"}' clika-cli devices services-restart jetson-01 clika-cli devices services-stop jetson-01 ``` Starting a service from a committed file works the same way as a job: ```yaml kind: Service service_definition: vllm-server devices: - jetson-01 ``` `Service` is apply-only. There is no flat collection to fetch one service by id across the fleet, so `export` does not support the kind. Read a running service back through `devices services-svc_id` instead. ## Desired state keeps a service up A service is not started once and forgotten. The platform holds it in the device's **desired state**, and the agent reconciles towards it, which is what brings the process back after a device reboots. The desired-state commands are on the [devices page](devices.md#desired-state). ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli service-definitions` `service-definitions` has 5 subcommands. #### `clika-cli service-definitions create` Create service definition ``` clika-cli service-definitions create [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `POST /api/service-definitions`. MCP tool name: `post_service_definitions`. #### `clika-cli service-definitions delete` Delete service definition ``` clika-cli service-definitions delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/service-definitions/{id}`. MCP tool name: `delete_service_definitions_id`. #### `clika-cli service-definitions get` Get service definition ``` clika-cli service-definitions get [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/service-definitions/{id}`. MCP tool name: `get_service_definitions_id`. #### `clika-cli service-definitions list` List service definitions ``` clika-cli service-definitions list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `GET /api/service-definitions`. MCP tool name: `get_service_definitions`. #### `clika-cli service-definitions update` Update service definition ``` clika-cli service-definitions update [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--body` | string | none | request body (inline JSON) | | `--body-file` | string | none | request body (path to a JSON file) | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `PUT /api/service-definitions/{id}`. MCP tool name: `put_service_definitions_id`. ### `clika-cli services` `services` has 1 subcommands. #### `clika-cli services list` List running services (org-wide) ``` clika-cli services list [flags] ``` | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--device_id` | string | none | Filter by device ID (UUID) | | `--page` | string | none | Page number (1-based) | | `--page_size` | string | none | Items per page | | `--raw` | bool | `false` | print raw response without pretty-printing | | `--search` | string | none | Match service id, device name, or type | | `--state` | string | none | Filter by service state | Endpoint: `GET /api/services`. MCP tool name: `get_services`. ## Related - [Model deployment](model-deployment.md): the pre-defined service that serves a model through the CLIKA runtime. - [Job definitions](job-definitions.md): the counterpart for work that finishes. - [Devices](devices.md): starting, stopping and inspecting a service on one device. - [apply and export](apply-export.md): the resource YAML round trip. - [Service concept](../concepts/service.mdx): what a service is, in prose. --- # Transfers clika-cli transfers: follow a file transfer between the platform and a device to completion, and cancel one that is no longer wanted. Source: https://docs.clika.io/platform/cli/transfers.md Moving a file to or from a device is not instant, and it does not happen inside the command that started it. When the platform pushes an artifact to a device, or pulls a file back, it creates a **transfer**: a record with a state, a byte count, a rate and, if it went wrong, an error. The `transfers` group is how you follow one. You get a transfer id from whichever command started the move, for example `devices files-push-from-library`, which sends a stored [artifact](artifacts.md) to a device without it passing through your machine. There are two subcommands. ## `transfers get` ``` clika-cli transfers get [flags] ``` Reads a transfer's status: its state, how many bytes have moved, the current rate, and the error if it failed. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--wait` | string | not set | Block until the transfer reaches a terminal state, printing each state change as it happens. A bare `--wait` waits indefinitely; a value bounds it, in seconds (`--wait=600`) or as a duration (`--wait=10m`). | Three behaviours are worth knowing before you put this in a script: - **It gates scripts.** With `--wait`, the command exits non-zero when the transfer ends `failed` or `canceled`, so `&&` does the right thing. - **A bounded wait that expires still reports.** The latest status is printed and the exit code is non-zero, so a timeout is not silently indistinguishable from success. - **An `unknown` state is waited on, not failed.** `unknown` means the platform temporarily lost track of the transfer, which is not the same as the transfer having gone wrong. ``` $ clika-cli transfers get 1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 STATE BYTES RATE UPDATED running 318 MB/412 MB 22.4 MB/s 2s ago ``` Wait for it and only then start the run that needs the file: ``` clika-cli transfers get 1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 --wait=10m \ && clika-cli jobs create --body '{"job_definition_id":"llm-latency","device_id":"jetson-01"}' ``` ## `transfers delete` ``` clika-cli transfers delete ``` Cancels a transfer that is still in flight. The device stops receiving, and the transfer ends in the `canceled` state, which a concurrent `--wait` will see and exit non-zero on. ``` $ clika-cli transfers delete 1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 canceled transfer 1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8 ``` ## Which command creates a transfer | Command | What moves | | --- | --- | | `devices files-push-from-library --body ''` | A stored artifact goes from the platform to the device. | | `devices services-update --body ''` | The artifacts a running service needs are refreshed on the device. | | A job dispatch whose definition declares `artifacts` | Each declared artifact is pushed before the job starts. | `devices push` and `devices files-fetch` are different. They stream through your own machine and finish inside the command, so they produce no transfer record to follow. ## Full command reference Every command below is generated from the deployment's own API description, so one subcommand is exactly one platform operation. Each entry names the method, the endpoint and the [MCP](mcp.md) tool name, so the same operation is identifiable whichever surface you drive it from. Path parameters are positional arguments, query parameters are flags, and a request body is `--body` or `--body-file`. The hand-written commands, the ones that stream, propagate an exit code, or hand your terminal to `ssh`, carry no operation line. The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before. ### `clika-cli transfers` `transfers` has 2 subcommands. #### `clika-cli transfers delete` Cancel a transfer ``` clika-cli transfers delete [flags] ``` Positional arguments: required ``. | Flag | Type | Default | Meaning | | --- | --- | --- | --- | | `--raw` | bool | `false` | print raw response without pretty-printing | Endpoint: `DELETE /api/transfers/{id}`. MCP tool name: `delete_transfers_id`. #### `clika-cli transfers get` Reads a device file transfer's status (state, bytes, rate, error). ``` clika-cli transfers get [flags] ``` Positional arguments: required ``. Endpoint: `GET /api/transfers/{id}`. MCP tool name: `get_transfers_id`. ## Related - [Artifacts and models](artifacts.md): what is being moved, and how it got into the library. - [Devices](devices.md): the file commands on a device, including the ones that do not create a transfer. - [job definitions](job-definitions.md): declaring the artifacts a run needs, which is what triggers most transfers. --- # Concepts The objects the platform is built from, what each one is for, and how they connect. Source: https://docs.clika.io/platform/concepts.md The platform is built from a small set of objects. This section has a page per object, each answering the same four questions: what it is, why it exists, what it relates to, and where you see it in the web application, the CLI and through Claude. Read them in this order the first time. Each page assumes the ones above it. ## Start here 1. [Organization and project](organization-and-project.mdx): who your account belongs to, and the box that holds one piece of work. [Roles](roles.md) lists what each organization and project role may do. 2. [Device](device.mdx): a machine running the CLIKA agent, what it reports, and everything you can do with one. ## Runtime What you do with a model once the platform can reach your hardware. 3. [Benchmark](benchmark.mdx): a run over models, tests and devices, and every benchmark in the catalogue. 4. [Model deployment](model-deployment.mdx): one of your models running on your devices behind an endpoint the platform proxies. 5. [Runtime licenses](runtime-licenses.mdx): the credential a ClikaRT runtime presents, its entitlements, and how you issue and revoke one. ## Device management The machinery underneath, and what you reach for when the built-in paths are not enough. 6. [Artifact](artifact.mdx): the versioned files the platform stores and puts on devices, including external Docker and Git sources. 7. [Job](job.mdx): the single execution of one benchmark on one device, and its job-definition template. 8. [Service](service.mdx): a process you tell the platform to keep running on a device. ## How they fit together ```text Organization one company or team; holds users and roles └── Project one piece of work; holds its own resources ├── Model a Hugging Face repository plus its task ├── Benchmark run models x tests x devices, one result set │ └── Job one model on one device, one execution │ └── Result quality scores, performance metrics, per-item data ├── Model deployment a model served on devices, behind the platform proxy ├── Service a process kept running on a device ├── Artifact a versioned file the platform puts on devices └── Runtime license what a ClikaRT runtime presents to run Devices belong to the organization, not to a project, so every project of an organization sees the same fleet. ``` The direction of causality is worth holding on to. A **benchmark run** is what you create; the platform fans it out into one **job** per model and device pair, each job runs a **job definition** the platform picked for that device, and each job produces one **result**. You never create a job by hand for an ordinary benchmark, and you never pick the job definition. ## The vocabulary that trips people up | Term | What it means here | | --- | --- | | Benchmark | One run you created, over one or more models, tests and devices. The API calls it a benchmark group. | | Quality test | One named test inside a run (MMLU, GSM8K, IFEval). The API calls it a benchmark type. | | Job | One execution of one model on one device inside a run. | | Job definition | The template a job executes: artifacts, setup, script, teardown, output path. | | Deployment (in the web application) | One of your models served on your devices. See [Model deployment](model-deployment.mdx). | | Serving | The API's name for the same object. | | Runtime license | The credential a ClikaRT runtime presents. Issued by your platform, per project. See [Runtime licenses](runtime-licenses.mdx). | | Platform license | A different artifact: what CLIKA issues to activate one deployment of the platform itself. See [Activate a deployment license](../how-to/activate-a-deployment-license.md). | ## Where to go next - [Getting started](../getting-started/index.md) if you have not run anything yet. - [How-to guides](../how-to/index.md) once the vocabulary is familiar. --- # Artifact The versioned files the platform stores and moves onto devices: uploads, directories, external Docker and Git sources, tags, deduplication and retention. Source: https://docs.clika.io/platform/concepts/artifact.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; An artifact is a file or directory the platform stores and can put on a device: a dataset, a set of model weights, a script, a config file, a compiled engine bundle, an archive. Anything a job or a service needs on disk arrives as an artifact. ## Why it exists A benchmark is only reproducible if the inputs are. An artifact gives a file an identity (a name, a version, a checksum) so a run can say exactly what it used, and so the platform can decide whether the file already needs transferring at all. It is also the boundary for cleanup: the platform knows what it put on a device, which is what lets it take it away again. ## Fields | Field | What it is for | | --- | --- | | `name`, `description` | How you find it again. | | `version` | Your own version string (`v1.0`, `2026-03-05`, `q4-quantized`). Optional and worth setting. | | `type` | `dataset`, `model_weight`, `script`, `config`, `docker_context`, `binary`, `archive`, `docker_image`, `git_repo`, `other`. A label for humans, not a behavior switch. | | `size_bytes`, `checksum_sha256` | Computed on upload. The checksum is why a repeat push transfers nothing. | | `content_format` | `file` (default) or `directory`. A directory artifact is stored as a tar.gz and extracted into the destination path on the device. | | `tags` | Key and value pairs you filter on. | | `expires_at` | When the artifact is swept, if a retention window applies. | ## Versions and tags Uploading a new version keeps the artifact in the same lineage and moves the `latest` tag to it. You can add your own tags (`stable`, `v2`, `golden`) and point them at any version in the lineage, the way Docker tags work. A job or service definition refers to an artifact by lineage plus an optional tag, and the platform resolves the tag to a concrete version at dispatch. That is the useful part: a definition that says `tag: latest` picks up your new upload without being edited, and a definition that says `tag: golden` keeps running the version you blessed. A tag that does not exist fails the dispatch rather than falling back to something else. ## Transfer and deduplication Before pushing an artifact the agent checks whether the destination path already holds a file with the same SHA-256. If it matches, the transfer is skipped entirely. This is worth knowing because it shapes how definitions are written. A benchmark that stages a multi-gigabyte engine bundle is slow the first time on a device and fast afterwards, which is why definitions that push large engines usually leave them in place instead of cleaning them up after every run. ## External artifacts An artifact does not have to be a file the platform stores. It can also be a reference the device fetches itself, which is how a container image or a repository of code joins the same library as an uploaded file. Two kinds exist. **A Docker image** (`docker_image`) carries a standard image reference: `nginx:1.25`, `ghcr.io/org/app:v2.1`, `my-registry.example.com:5000/team/model:latest`. At dispatch the agent runs `docker pull`, optionally for a named platform on a multi-architecture image, and reports the digest and size it actually got. **A Git repository** (`git_repo`) carries an HTTPS clone URL. The agent runs a shallow clone into the destination path, at a branch, tag or commit when the reference names one, and reports the commit it landed on. What you can do with them: - **Keep a benchmark harness in Git rather than in an upload.** Point a job definition at the repository and every run clones the current state of the branch you named, so fixing the harness is a push rather than a re-upload. - **Run a container as a managed service.** A Docker image artifact plus a service definition of type `compose` or `process` is how a device runs something the platform did not build. - **Version them like any other artifact.** A new version of an external artifact points at a different tag or ref, so the lineage records which image or commit each run used. - **Authenticate to a private registry or repository.** The artifact stores a reference to a registry credential rather than the secret. The platform resolves the credential at dispatch and hands the agent what it needs for that one pull, so rotating the credential does not mean editing artifacts. Three properties follow from the platform holding no file: - The artifact's size is zero until the agent reports what it pulled, and its checksum is the digest or commit the agent saw rather than one the platform computed. - There is nothing to download from the platform, so the download endpoint answers `404` for an external artifact. - The device has to be able to do the pull. An agent reports whether it has Docker and Git available when it registers, and a dispatch that needs a capability the agent lacks is refused at dispatch time rather than failing halfway through a run. ## Retention and cleanup Artifacts accumulate, and a benchmark campaign's intermediates can reach tens of gigabytes, so the library has an explicit lifecycle. - **Filters that express a cleanup.** The artifact list filters on type, tag, search text, creation date and size, so "older than 30 days and bigger than 1 GB" is a query rather than a manual review. - **Batch delete over the same filters**, bounded per call and with a dry run that shows you what would go. An artifact still referenced by a job or service is reported as in use and skipped, so a batch never has to be all-or-nothing. - **A retention window.** An organization can set a default retention in days, which stamps an expiry on every upload; an individual artifact can carry its own expiry or none at all. Platform-generated artifacts (job outputs, dataset builds) are not stamped. - **A sweeper** removes expired artifacts after a grace window, skipping anything still referenced or mid-transfer, and every deletion is recorded in the audit log. ## Job output becomes an artifact When a benchmark job finishes, the platform fetches the file at the job's declared output path and stores it as an artifact linked back to the job. If the job declared structured results, that same file is also parsed into the result you read in the UI. The raw file stays downloadable either way, which is what you reach for when you want the numbers in their original form. ## Where you see it
**Artifacts** in the sidebar is the library: upload, register an external source, browse versions and tags, set an expiry. A device's **Files** tab can also push straight from the library, and save a file fetched from a device back into it.
```bash clika-cli artifacts list clika-cli artifacts upload ./dataset.tar.gz --name prompts --type dataset clika-cli artifacts download prompts ./prompts.tar.gz ``` `upload` takes a file or a directory; a directory is packed and marked so the agent extracts it on arrival.
Ask Claude: > What artifacts do we have over 1 GB that nothing has used this month? That reaches the `get_artifacts` and `get_artifacts_id` tools, which the `device-ops` toolset carries alongside the device operations.
## Related pages - [Job](job.mdx): where artifacts are pushed and output is collected. - [Service](service.mdx): the other consumer of artifacts. - [Write a job definition](../how-to/write-a-job-definition.md): how a definition refers to an artifact and its tag. --- # Benchmark What a benchmark run is, how the platform turns it into work on your devices, and every benchmark in the catalogue: what it is, what it tests, and how it is scored. Source: https://docs.clika.io/platform/concepts/benchmark.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A benchmark is one run you created: a set of models, a set of quality tests, and a set of devices. The platform executes every combination and returns one result set you can read side by side. The API calls a run a **benchmark group**, and the web application shows a run as a single row under **Benchmarks**. ## Why it exists Measuring a model on a device is a small job. Comparing two models on three devices is bookkeeping, and doing it by hand is where mistakes come from: a different script here, a different prompt set there, and the numbers stop being comparable. A benchmark run makes the comparison the unit of work. Every leg runs the same test with the same harness, and every result carries the model, the device and the test it belongs to. ## What you choose, and what the platform chooses You choose three things: the **models**, the **quality tests**, and the **devices**. That is the whole input. The platform chooses how to run each combination. A benchmark type plus a device's platform and architecture resolve to exactly one job definition, through a compatibility table the deployment owns. If a device shape has no entry for that test, the dispatch refuses that device with `DEVICE_NOT_RUNNABLE` rather than falling back to something unproven. There is no "pick your runner" step, and a benchmark that is confirmed on Linux arm64 but not yet on Android says so instead of producing a number you cannot trust. Two more checks happen before anything downloads. A model whose weights will not fit the device's free disk is refused up front, with the figure the device reported. And a device keeps the models it already downloaded between runs, within a size and age bound, so a repeat run on the same device starts in seconds instead of fetching again. ## Two shapes of benchmark A **performance** benchmark measures how fast a model runs: tokens per second, time to first token, latency, peak memory. A **quality** benchmark measures whether its answers are right, and each one scores that differently. A single run can carry both. Quality scoring on this platform is deterministic wherever it can be. Multiple-choice answers are matched exactly against the gold letter, free-text answers by normalized exact match, code by executing it against the benchmark's own tests, translations and summaries by their published overlap metrics. No benchmark uses a model to grade another model. ## The catalogue What a deployment offers depends on which benchmarks it has installed and which datasets have been staged on it, so treat the list below as the shape of the catalogue rather than a promise about your own platform. `GET /api/benchmark-types` answers that question for a specific deployment, and the picker only offers the tests whose task matches the models you selected. The catalogue is grouped the way the platform groups models: by the **task family** a benchmark applies to. A benchmark is offered for a model when the model's task belongs to its family, which is why an automatic-speech-recognition model is never asked to answer multiple-choice questions. One property cuts across the families and is worth knowing before you read any quality score. Several benchmarks run a public Hugging Face model through `transformers` on the device rather than through the ClikaRT engine, which makes them a measurement of the model and not of the engine or its quantization. Each entry says so where it applies. ### LLM Text-generation models: the family most people mean by "a model". These benchmarks send prompts and read completions, and they split into one performance test and a set of quality tests that score the text that comes back. How the scores read: the quality numbers here are the ones with the widest published baselines, which makes them tempting to compare against a leaderboard. Several of these datasets are also present in most models' pre-training data, so the honest use is a **relative** one, comparing the same model across engines, quantizations or devices, rather than an absolute claim about a model's intelligence. #### LLM Performance Text-generation throughput and latency under a representative inference workload. Reports tokens per second, time to first token and inter-token latency, with peak memory alongside. This is the benchmark to run when the question is how fast a model is on a device rather than how good its answers are. #### Accuracy (MMLU) Multiple-choice questions across academic and professional subjects, in the style of MMLU. Scored by exact match of the predicted choice letter against the gold answer, with throughput and latency reported alongside. The per-question record shows the subject, the choices in the order the model saw them, and the letter it picked, so a score can be traced to the questions behind it. #### GPQA Diamond Graduate-level science questions written to be hard for a non-expert with a search engine (198 items in the Diamond subset). Multiple choice, scored by exact match of the choice letter. #### ARC Challenge Grade-school science questions from the AI2 Reasoning Challenge, challenge set. Multiple choice, scored by exact match. Widely present in pre-training corpora, so read it as a relative yardstick rather than an absolute quality claim. #### HellaSwag Commonsense sentence completion: pick the ending that plausibly continues a described situation. Multiple choice on the validation split, scored by exact match, with the same caveat about pre-training exposure as ARC. #### TruthfulQA mc1 Resistance to common misconceptions, using the single-correct-answer (mc1) scoring. Scored by exact match of the choice letter. The published mc2 figure is a different metric and is not comparable to this one. #### GSM8K Grade-school mathematical word problems answered with chain-of-thought reasoning. Scored on numeric equality of the final number, so `1,203`, `$1203` and `1203.0` all count as the same answer. An answer containing no number at all is counted unparseable rather than wrong, which keeps a model that never answers distinguishable from one that answers badly. #### AA-LCR Long-context reasoning over multi-document source material, answered as free text and scored by normalized exact match. This is the benchmark for the question "does the model still reason when the prompt is long". #### Coding (HumanEval+) Functional correctness on HumanEval+ from EvalPlus: the model writes a function, and the platform runs the benchmark's tests against it. Scored by execution, not by resemblance to a reference solution. #### Coding (MBPP+) The same harness over MBPP+ (399 tasks), a different distribution of programming problems. It is a separate benchmark rather than more items in the same one, because a pass rate mixed across two task distributions is not a comparable number. #### Coding (SciCode) Scientific-computing code generation, scored by executing the generated program against the benchmark's tests. The tasks assume domain context, so it separates models that can write code from models that can write the right code for a scientific problem. #### Spreadsheet (SpreadsheetBench 2) Business spreadsheet tasks: the model produces a script, the platform executes it, and the resulting workbook is diffed against the golden one over the cell ranges the task grades. Closer to real office automation than a coding puzzle. #### Instruction Following (IFEval) Prompts carrying machine-checkable constraints (word counts, forbidden characters, required formats). The answer is scored by running those checkers, with no reference text and no judge model involved. Reported as the strict per-prompt rate, where a prompt counts only when every one of its instructions is satisfied, with the per-instruction rate alongside. #### Summarization (CNN/DailyMail, ROUGE) News-article summarization scored by ROUGE against the reference highlights, with ROUGE-L F1 as the headline number. ROUGE measures n-gram overlap rather than meaning, so a correct summary worded differently scores low. Read it as a regression yardstick on one model, not as a summarization-quality claim. ### Multimodal Vision-language models: the ones that take an image (or sampled video frames) alongside text and answer in text. The task is `image-text-to-text` or document question answering, and the model under test is still a chat model, which is why these read like the LLM quality tests with a picture attached. How the scores read: every benchmark in this family runs a public Hugging Face model through `transformers` on the device rather than through the ClikaRT engine, so a score measures the model, not the engine or its quantization. Free-text answers are scored by normalized exact match, which is strict about wording, so a lower number here does not always mean a worse answer. #### MMMU multiple-choice College-level, multi-discipline questions over images (MMMU validation split, multiple-choice rows). Scored by exact match of the choice letter. #### MMMU-Pro The harder MMMU-Pro variant, standard four-option split. A separate benchmark from MMMU, and the two numbers are not comparable. #### CharXiv reasoning Chart-understanding questions over scientific figures, answered as free text and scored by normalized exact match. The question is whether the model can read a plot rather than recognize an object. #### MathVision Visual mathematics over figure images (testmini, open-ended rows), answered as free text and scored by normalized exact match. #### ZeroBench Deliberately hard multimodal reasoning, 100 items, one greedy sample per item. Free text, normalized exact match. The published figure for this dataset is a pass rate over five samples, so a single-sample score here reads lower by construction. #### BabyVision free-response Fine-grained visual discrimination, free-response rows only, scored by normalized exact match. #### DocVQA Document visual question answering: read the answer out of a scanned page. Scored by ANLS, a normalized edit-distance similarity against each acceptable answer, zeroed below a threshold so an OCR near-miss keeps partial credit while a wrong answer earns none. #### OmniDocBench page OCR Whole-page document transcription to markdown over annotated PDF pages. The device produces one transcription per page and a scorer compares it against ground truth covering text, formulas and tables in reading order, reporting a mean page-level normalized edit distance where lower is better. It needs server-side scoring enabled on the deployment. #### VQAv2 Open-ended visual question answering scored by the official consensus rule: a prediction agreeing with three or more of the ten human annotators earns full credit, one or two earns partial. #### WorldVQA Real-world visual question answering over staged images, free text, normalized exact match. ### Vision Classifiers, detectors and segmenters: models whose output is a label, a box or a mask rather than text. There is no prompt and no completion here, which is why these benchmarks report the metrics computer vision uses rather than an accuracy over answers. How the scores read: these also run a public Hugging Face model through `transformers` on the device rather than through the ClikaRT engine. Preprocessing follows each checkpoint's own image processor rather than a fixed recipe, so a figure is a reliable yardstick for the same model across devices or quantizations and a poor one for ranking two publishers' models against each other. Several of the datasets are licensed for non-commercial research only, and a deployment has to stage them before the benchmark can run. #### ImageNet-1k Image classification over the ILSVRC-2012 validation split, reporting top-1 and top-5 accuracy. The dataset is account-gated and restricted to non-commercial research and education. #### COCO detection Object detection over the COCO 2017 validation split (80 classes), scored by mean average precision at IoU .5 to .95, with mAP@50 and mAP@75 alongside. #### ADE20K segmentation Semantic segmentation over ADE20K scene parsing, scored by mean intersection-over-union with pixel accuracy alongside. Released for non-commercial research and educational use. #### Cityscapes segmentation Semantic segmentation over urban street scenes, same metrics as ADE20K. Licensed for non-commercial use, and its canonical distribution requires registration. #### Image Classification (top-1) Top-1 accuracy against a labeled image set through the ClikaRT engine. Not runnable today: no engine classifier verb exists yet, so no device shape matches it and a dispatch refuses rather than substituting something else. ### Audio Speech and audio models: transcription, voice generation and audio-event classification. Two of these run through the ClikaRT engine, which makes them the audio family's answer to both questions at once, how fast and how good. How the scores read: the transcription metrics are error rates, so **lower is better**, which inverts the reading of every other quality number in the catalogue. Speed is reported as a real-time factor, the seconds of audio processed per second of wall clock, so a figure above one means faster than real time. #### Speech-to-Text (WER) Speech-to-text accuracy on read speech, scored by word error rate and character error rate. Real-time factor and audio throughput are reported alongside, and the per-clip record carries the audio, the reference and a word-level diff against the prediction. #### STT Performance Speech-to-text throughput and latency alone: the real-time factor and audio-seconds per second under a representative workload, with no accuracy scoring. #### Text-to-Speech (voice generation) Voice generation from a fixed sentence set through the ClikaRT engine. Each generated clip is captured with its prompt, and the run reports real-time factor, latency and throughput plus a round-trip word error rate obtained by transcribing the generated audio back to text. #### AudioSet tagging Multi-label audio-event classification over AudioSet's sound-event classes, scored by mean average precision. The model under test is an audio classifier, and the score is computed over the whole run rather than per item, so the number of classes actually present in the subset is reported beside it: a high score over few classes is an easy ranking rather than a good model. Runs through `transformers` on the device. ### Image generation Text-to-image models. The output is a picture, so there is no answer to mark right or wrong and the score measures whether the image matches the prompt. #### Image Generation (diffusion) Text-to-image generation from a fixed prompt set through the ClikaRT engine. Each image is captured with its prompt, and the run reports a CLIP-cosine alignment score between prompt and image alongside images per second and latency. ### Embedding Embedding models: the output is a vector, and quality is a question about ranking rather than about generated text. These benchmarks embed a corpus and a query set and score how well the ranking puts the right documents first. How the scores read: nDCG@10 is reported as a 0 to 1 ratio, so a published two-digit figure is this value times 100. Both datasets carry usage restrictions, and both run a public Hugging Face model through `transformers` on the device rather than through the ClikaRT engine. #### BEIR NFCorpus Zero-shot medical information retrieval over a document corpus and query set, scored by nDCG@10 with Recall@10 and Recall@100 alongside. #### BEIR SciFact Zero-shot scientific-claim retrieval, same metrics and reporting as NFCorpus. The upstream dataset is non-commercial. ## Reading a run A completed run gives you the headline metrics per model and device, the per-item record behind each score, and the run's own metadata. [Read the results](../getting-started/first-benchmark/06-read-the-results.md) walks the result view, and [Read benchmark results](../how-to/read-benchmark-results.md) covers pulling the same numbers from the CLI or an assistant. **Share Results** on a run mints a link that shows it to someone without an account. There is one live link per run, and revoking it invalidates that link immediately. ## Where you see it
**Benchmarks** in the sidebar lists your runs. **New Benchmark** opens the two-step flow: devices and models side by side first, then the quality tests and the dataset mode over the pairs that fit. A run's own page carries the summary, the per-metric comparison and the per-item output viewer.
```bash clika-cli benchmark-types list clika-cli benchmark-groups list clika-cli benchmarks watch "nightly-llm-sweep" clika-cli benchmarks results "nightly-llm-sweep" ``` `watch` exits non-zero when any leg ends in anything but completed, which is what makes it a CI gate.
Ask Claude: > Run an LLM performance benchmark of Qwen2.5-0.5B-Instruct on my fastest online Linux device, wait for it, and tell me the tokens per second. That reaches `get_benchmark_types` and `get_benchmark_compatibility` to work out what can run, `post_benchmark_groups` to launch it, `get_benchmark_groups_id` to poll, and `get_benchmark_groups_id_results` to read it. The `benchmark-ops` toolset carries exactly this loop.
## Related pages - [Job](job.mdx): what one leg of a run does on the device. - [Artifact](artifact.mdx): the datasets and engine bundles a run stages. - [Run a benchmark](../getting-started/first-benchmark/05-run-a-benchmark.md) in the tutorial. --- # Device A machine running the CLIKA agent: how it registers, what it reports, the states it moves through, and what blocks it from taking work. Source: https://docs.clika.io/platform/concepts/device.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A device is a machine that runs the CLIKA agent and has registered itself with the platform. A Jetson on a bench, an x86 workstation, a Windows laptop, a single-board computer, an Android phone: same agent, same fleet, same benchmark surface. ## Why it exists Numbers measured on a development machine do not tell you what a model does on the hardware you ship. The device object is how the platform reaches real hardware, so a benchmark result carries the device it was measured on and two results can be compared honestly. ## How a device joins You install the agent with one command from the **Register Device** dialog. That command carries a short activation code, so the device proves it may join without anyone typing credentials on it. The agent then dials out to the platform and holds one gRPC connection open for everything that follows: commands, file transfer, terminal, job dispatch. Nothing connects inward to the device, so no inbound port has to be opened. By default the code enrolls one device and expires 30 minutes after the dialog made it. **Reusable, many devices** switches to a batch command that any number of devices can register with for ten minutes, which is how a rack or a shelf of phones joins in one pass. The CLI and the API mint enrollment tokens as well, the other credential a device can enroll with. At registration the agent reports what the machine is: platform (`linux`, `darwin`, `windows`, `android`) and architecture (`amd64`, `arm64`, `arm`), agent version, network interfaces, and a hardware inventory (CPU cores, memory, GPU model and VRAM, disk, and detected accelerators such as `cuda`, `tensorrt`, `coreml` or `nnapi`). The platform assigns a device id on first registration and confirms the same id on every later connection. ## What it reports Every 15 seconds the agent sends a heartbeat: CPU, memory, disk, temperature, uptime, battery where there is one, network type, and the state of the processes and services the platform asked it to watch. The platform stores the heartbeat as health history, updates the live view, evaluates alert rules against it, and pushes it to anyone watching the device page. A device with no heartbeat for 60 seconds is marked offline. Services already running on it keep running: going offline is a statement about the connection, not about the machine. ## States | State | Meaning | | --- | --- | | `ONLINE` | Connected, heartbeats arriving, available for work. | | `BUSY` | Running a benchmark job. | | `SUSPENDED` | Taken out of service by an operator. | | `ERROR` | The agent reported an error state. | | `OFFLINE` | No heartbeat for 60 seconds. | ### The dirty flag Every benchmark job ends with teardown steps, and those steps always run, including after a failure. If teardown itself fails, the platform flags the device **dirty** and records which job did it and why. A dirty device takes no new benchmark jobs, because leftover state from a failed run is how one bad job turns into a series of misleading results. Dirty clears in one of two ways: the agent retries the teardown when it reconnects and reports success, or an operator who has cleaned the machine clears it from the device page. ## Resources and capabilities Two different things get reported, and jobs use them differently. - **Resources** are typed hardware facts (`cpu_cores`, `memory_bytes`, `gpu_vram_bytes`, `accelerators`). A job definition can declare `required_resources`, and the platform refuses to dispatch to a device that does not meet them, naming the requirement it failed. - **Capabilities** are platform metadata (terminal mode, SSH status, feature strings). They describe what the agent can do on this machine rather than how much hardware it has. You can also declare your own resources in the agent's configuration (a rack label, an engine name), and they merge with what auto-detection found. ## What you can do with one A registered device is not only a benchmark target. Everything below runs through the platform over the agent's own connection, so none of it needs an inbound route to the machine. - **Run benchmarks.** The device takes benchmark jobs, one at a time, and reports each state transition as it goes. Its **Jobs** tab is the history of everything the platform has run on it, with the outcome of each. - **Run managed services.** A [service](service.mdx) is a process the platform keeps running on the device and restarts when it exits. The **Services** tab lists what the device runs, with state, uptime and restart count, and a live log view per service. - **Serve a model.** A [model deployment](model-deployment.mdx) is the built-in service: one of your registered models, running on the device behind an endpoint the platform proxies for you. - **Browse and transfer files.** The **Files** tab is a file manager for the device. Navigate its filesystem, upload from your machine or from the artifact library, download a file, and save what you downloaded back into the library as a new artifact. Uploads are chunked and resumable, so a multi-gigabyte push survives a dropped connection. - **Run a command.** The **Commands** tab runs a one-off command and records it with its exit code, its output and when it ran. A recorded command stays readable afterwards, which makes it usable as evidence and not only as a convenience. - **Open a terminal.** An interactive shell in the browser over the same connection. On Linux, macOS and rooted Android it is a real PTY; on Windows and unrooted Android it is more limited. - **Open a remote desktop.** Where a provider exists for that platform, the device's screen is viewable in the browser. When no session can be served, the device reports a reason and the page names the remedy (install a VNC server, enable Screen Sharing, grant the Android consent prompt) instead of hiding the feature. - **Read its health.** CPU, memory, disk, temperature, uptime, battery and network arrive every 15 seconds, with history per metric and alert rules evaluated against them. - **Tag it.** Free-form key and value tags travel with the device and filter the fleet list, which is how a large fleet stays navigable (`rack=lab-a`, `owner=vision-team`). - **Update its agent.** The platform tracks the agent version per device and can trigger a self-update, one device or a batch at a time. - **Suspend or retire it.** A suspended device stays registered and takes no work. Deleting it removes it from the fleet. ## Where you see it
**Devices** in the sidebar lists the fleet. A device's own page carries the Overview, Jobs, Services, Files and Commands tabs, with Terminal and Remote Desktop in its header.
```bash clika-cli devices list clika-cli devices get "Linux x86 Workstation" clika-cli devices exec "Linux x86 Workstation" "nvidia-smi" clika-cli devices push "Linux x86 Workstation" ./model.gguf /tmp/model.gguf ``` Every command takes a device name or its id. See [Using the CLI](../cli/index.md).
Ask Claude: > Which of my devices are online, and how much GPU memory does each have? That reaches the `get_devices` and `get_devices_id` tools. `get_devices_id_health` returns the live health and `post_devices_id_exec` runs a command; the `device-ops` toolset exists for exactly this work.
## Related pages - [Job](job.mdx): what actually runs on a device. - [Service](service.mdx): what keeps running on it. - [Register or pick a device](../getting-started/first-benchmark/03-pick-a-device.md) in the tutorial. --- # Job One execution of one benchmark on one device: the definition it runs, the states it moves through, and what it guarantees on the way out. Source: https://docs.clika.io/platform/concepts/job.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A job is one execution of one benchmark on one device. A benchmark run over two models and three devices creates six jobs, each independent: its own status, its own result, its own log. ## The job definition A job definition is the template a job executes. It is self-contained, so anyone reading it can tell exactly what will happen on the device. | Part | What it declares | | --- | --- | | `artifacts` | Files to push before anything runs, each with a destination path on the device. | | `setup_steps` | Commands to run after the files land (create a virtual environment, install dependencies, start a helper service). | | `script` | The benchmark itself: command, working directory, environment, timeout. | | `output_path` | Where the script must write its result file. | | `result_type` | `raw` (keep the file) or `structured` (parse it into a result the platform can chart and compare). | | `teardown_steps` | Commands that always run at the end, successful or not. | | `cleanup_policy` | What of the pushed files, output and scratch paths is removed afterwards. | | `required_resources` | Minimum CPU, memory, VRAM, disk, GPU count and accelerators the device must report. | For an ordinary benchmark you never write one of these and never pick one: the platform resolves the definition from the benchmark type and the device's platform and architecture. You write a definition when you are adding a new kind of benchmark to a deployment. [Write a job definition](../how-to/write-a-job-definition.md) covers the file in full. ## What the platform does, and what your script does The platform handles logistics: get the device into the right state, hand the script its inputs and an output path, run it, and collect what it produced. What happens inside the script is yours. The platform does not interpret your inference code, your timing, or how you talk to a local server. The one thing it requires is a result file at `output_path`. ## States ```text QUEUED -> PUSHING_ARTIFACTS -> RUNNING_SETUP -> RUNNING -> RUNNING_TEARDOWN -> DONE |-> FAILED |-> PARTIAL |-> INTERRUPTED ``` | State | What is happening | | --- | --- | | `QUEUED` | Waiting for the device to finish its current job. | | `PUSHING_ARTIFACTS` | Transferring files, skipping any whose checksum already matches on the device. | | `RUNNING_SETUP` | Executing the setup steps in order. | | `RUNNING` | Executing the script, streaming its output. | | `RUNNING_TEARDOWN` | Executing the teardown steps. | | `DONE` | Script succeeded, teardown succeeded, output collected. | | `FAILED` | A step failed, or the declared output file was never written. | | `PARTIAL` | The script succeeded and produced results, but teardown failed. The results are usable; the device is flagged dirty. | | `INTERRUPTED` | The device disconnected mid-run. Teardown is retried when it reconnects, and a late result can still complete the job. | A job dispatched to a device that is already busy is requeued with backoff rather than failed, and gives up after ten minutes of trying. The per-device queue is visible on the device page. ## Two guarantees worth knowing **Teardown always runs.** Whether the script succeeded, failed, timed out or was cancelled, the teardown steps execute. That is what stops a failed run from leaving a service running, a GPU held, or a temporary directory filling the disk. If teardown itself fails, the device is flagged [dirty](device.mdx#the-dirty-flag) and takes no further benchmark jobs until it is cleared. **A cancel stops the whole process tree.** Every step runs as its own process group, so cancelling a job (or hitting the script timeout) stops what the step started and everything it spawned. A job the platform reports as cancelled is not still holding the device's memory. ## Collecting the output When the script exits, the platform fetches the file at `output_path`, stores it as an artifact, and links it to the job. Four outcomes are deliberately distinguished. - **Nothing at the declared path**: the job **fails**, naming the path that was expected. Declaring an output path and not writing it is a broken contract with the platform, even when the script exited zero. - **A file that is not a result envelope**: the job **completes** with a diagnostic, and the raw file stays attached. The format was the script's own choice, and its exit code was its own verdict. - **A valid envelope**: the job completes and the result is parsed, charted and comparable. - **The fetch failed for another reason** (the device went away mid-transfer): the job completes and the failure is logged, because nothing was established about what is on the device. ## Cleanup After teardown, cleanup runs in a fixed order: your teardown steps first, then the pushed artifacts (unless the definition keeps them), then any custom paths, then the output path (only if the definition asks for it). The default removes pushed artifacts and keeps the output. Definitions that push a large engine bundle usually keep it, because the checksum skip makes the next run fast only if the file is still there. ## Where you see it
**Jobs** in the sidebar lists every job in the project. A device's **Jobs** tab lists what ran on that device, and a benchmark run's **Details** tab lists the jobs that run fanned out into.
```bash clika-cli jobs list clika-cli jobs get clika-cli apply -f job.yaml # kind: Job, for a definition-based dispatch ```
Ask Claude: > Did any job fail on my devices today, and what did it say? That reaches `get_jobs` and `get_jobs_id`, whose payload carries the failing step and the tail of what the script printed.
## Related pages - [Benchmark](benchmark.mdx): what creates jobs. - [Device](device.mdx): where they run. - [Write a job definition](../how-to/write-a-job-definition.md). --- # Model deployment Running one of your models on your own devices as an endpoint: the platform proxy, the endpoints a served model exposes, the deployment page, and how ClikaRT does the work. Source: https://docs.clika.io/platform/concepts/model-deployment.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A model deployment runs one of your registered models on one or more of your devices and puts an HTTP endpoint in front of it. The model loads once and stays loaded, so requests are answered without paying the load cost each time. The web application calls this surface **Deployments**, and the **Deploy** button on a benchmark result leads to it. In the API the object is a **serving**. ## Why it exists A benchmark tells you how a model performs on a device. A deployment is how you then use it, with an application, a notebook or a colleague sending requests and the model answering from the hardware you chose. It is also the shortest path from "this model won the benchmark" to "this model is answering", because the platform builds the whole service for you: it stages the runner, starts the process, watches its health, restarts it when it dies, and gives you an address to call. ## What you choose A deployment takes a model, the devices to run it on, and how much of each device to use. | Choice | What it does | | --- | --- | | Model | Any model registered in the project. | | Devices | One or more. Each device runs its own copy of the model and answers on its own endpoint. | | Compute | `auto`, `cpu` or a GPU. New deployments run `auto`, which lets the engine place the model, and the deployment records where it landed; a deployment created before that default keeps its pinned value. | | Serve settings | The context window, the number of parallel requests and the engine's other serve parameters. The platform derives them from the device's memory and the model, shows each with its basis in the **Configure** dialog of the deployment's page, and lets you override any of them there; **Save** keeps the change for the next start, **Save and restart** applies it now, and a saved-but-not-applied change shows as restart pending under the page's header. | A deployment is admitted only where it fits: a start on a GPU whose memory cannot hold the model's weights, its cache and the engine's working reserve is refused with the arithmetic that says what holds the card; a device whose free disk cannot take the weights is refused before anything downloads; and a device whose platform cannot host a deployment is refused at creation, with the reason shown in the picker. Deploying from the web application means create and start in one action. Through the API the two are separate: a serving is created `pending`, and a start call brings it up. ## The proxy The model server on a device binds the device's loopback interface, so nothing outside the device can reach it directly. Every call goes through the platform instead, at a per-device address: ```text /api/v1/servings/{serving_id}/devices/{device_id}/proxy/{path} ``` Everything after `proxy/` is passed through to the model server verbatim, the query string travels with it, and responses stream rather than being buffered, so token-by-token output arrives as it is produced. Three things are worth knowing about calling it. - **You authenticate to the platform, not to the device.** A session or an API key on the `Authorization` header is what the proxy checks, and the credential is consumed there and never forwarded to the device. The capability required is `model_serving:invoke`, which is deliberately separate from being able to see a deployment: reading a list is not the same as spending a device's GPU. - **The address exists only while the instance is running.** A device's proxy address appears when its instance reports running and is absent otherwise, because an address that cannot answer is worse than no address. - **It is an operator and integration surface, not a production front door.** The platform applies a request-size limit, a per-deployment rate limit and a per-device concurrency cap, and refuses with `429` and a `Retry-After` rather than queueing without bound. A refusal names its reason: `SERVING_NOT_RUNNING` when the instance is not up, `DEVICE_OFFLINE` when the device is gone, `SERVING_CONCURRENCY_LIMIT` when the device is already busy with as many requests as it will take. ## The endpoints a deployment exposes The endpoint set depends on what the model does. Every deployment answers the three that describe the server itself: ```text GET /v1/models what this server is serving GET /v1/health liveness, which the platform polls to decide the instance is up GET /props the server's own properties ``` On top of those, the model's task decides what is meaningful. | The model does | Endpoint | | --- | --- | | Text generation, including vision-language chat | `POST /v1/chat/completions` | | Speech to text | `POST /v1/audio/transcriptions` | | Embeddings or sentence similarity | `POST /v1/embeddings` | | Text to speech | `POST /v1/audio/speech`, plus `POST /v1/voices` to register a voice | The shapes are OpenAI-compatible, so an existing client library usually works by pointing its base URL at the proxy address. Two details save a confused half hour. Send the model id that `GET /v1/models` returns as the `model` value, because the server derives that id from the checkpoint it loaded rather than from the name you gave the deployment. And for a vision model, send images as content parts carrying a `data:` URI: the server fetches no remote URLs. The server itself mounts every route and refuses the ones its model cannot serve with an explanatory `400`, so a call to the wrong endpoint tells you what happened. Where the platform does not know the model's task, it lists only the three universal endpoints and points you at `GET /v1/models` and `GET /props` to see what the server is actually serving. ## How ClikaRT does the work The process on the device is the CLIKA inference engine, `clikart-cli`, running its serve mode with ClikaRT underneath it. The platform stages a small wrapper script as an artifact, starts it as a managed service, and the wrapper finds the engine and hands it the model. Where the engine comes from, in the order the wrapper looks: an explicit path in the device's environment, then the engine bundle root that benchmark runs already use, then the newest engine bundle staged under the device's staging directory, then whatever is on the path. A device already set up for benchmarking therefore serves with no additional configuration. The model reaches the device the same way a benchmark's model does. For a Hugging Face model the engine fetches the checkpoint itself, honoring the [access token](../getting-started/first-benchmark/04-add-a-model.md#the-hugging-face-access-token) when the repository is gated; for an uploaded model the platform delivers the artifact. A checkpoint already staged on the device is used as it is. ### The engine's own license Each instance of a deployment runs the engine under a runtime credential the platform mints for it when the instance starts and revokes when it stops; when the device reconnects after losing the platform, the agent fetches a fresh one and restarts the engine under it. You never issue this credential yourself, and it does not count against the credentials you mint for your own applications. The agent launches the engine directly, so a device needs no Python or shell for a deployment. ## Lifecycle A deployment's status is a roll-up of its devices: `running` when all of them are, `partial` when some are, `failed` when none came up, `stopped` when all were stopped, and `pending` before the first start. Each device reports its own state (`starting`, `running`, `failed`, `stopped`) with its process id, port, uptime and restart count, and a failed instance carries its exit code and the tail of what it printed, so a red badge always comes with an explanation. Auto-restart is on. The device restarts a crashed instance with an exponential backoff that starts at a few seconds and caps at five minutes, gives up after ten consecutive failures and settles into `failed`, and resets its failure count once an instance has stayed up for a minute. That combination keeps a flapping deployment from restarting forever while letting a long-lived one recover from an occasional crash. Stopping a deployment removes it from the device's desired state first, so the reconciler does not bring it back, then stops the process group. The configuration is kept, so starting it again re-applies what it had. Deleting it stops every instance and cleans up what the platform staged. ## The deployment page A deployment's page replaces the Deployments list in place and reads top to bottom in the order an operator needs it. **Header.** A back link to the Deployments list (or to the model the page was opened from), then the deployment's name, renamed in place from the pencil beside it, with its status and when the serve last reported. To the right, **Configure** opens the serve settings dialog; **Stop** or **Start** and **Restart** act on every device; **Open inference UI** opens the engine's own chat page and dashboard in a new tab, through the platform proxy. The last one needs the `model_serving:invoke` capability and a running instance; with several running devices it asks which one. The kebab at the top right holds **Delete**, with the choice of also removing the model's files from the devices where no other deployment uses them. A saved serve setting that is not applied yet shows as a restart-pending line under the header. **What runs where.** A card for the model and one for each device, each a shortcut to its own page. A device card carries the instance's uptime and the device's live readings (CPU, memory, disk and temperature) as on the device page; a device that is not reporting says so with when it was last seen. Where the engine placed the model is on the card's hover. A failed or restarted instance adds a line under the cards with its state, its restart count, the failure in one sentence and the engine's log behind a disclosure. A model whose checkpoint carries no chat template is noted here too, because its chat endpoint continues text rather than answering. **Server logs.** The inference server's output, one viewer per device, following the tail; **Pause Auto-scroll** holds the view while you read. **API.** The address to call the deployment at, one per device that is serving, through the platform (works from anywhere that reaches the platform, with an API key), and under a rule the device's own address on its network. Where the server binds the device's loopback interface, which is the default, that address answers only from the device itself, and the card says so. **View API reference** opens the endpoint list for the model's task: each row is a method, a path and a description, and opens to the full request URL, a request and response example, the streaming shape where the endpoint streams, and a `curl` you can paste after exporting your API key. ## Where you see it
**Deployments** in the sidebar lists the project's deployments with their status and devices. **Deploy** on a benchmark result creates one from the model that run measured. A deployment's own page is described in [The deployment page](#the-deployment-page).
```bash clika-cli servings list # the project's deployments clika-cli servings get # one deployment, with per-device status clika-cli servings start # start it; stop and restart alongside clika-cli servings devices-proxy /v1/models ``` The last one calls a running deployment through the platform proxy, which is the quickest way to check what a server is actually serving. See the [Using the CLI](../cli/index.md).
Ask Claude: > Which models are deployed right now, and on which devices? That reaches `get_v1_servings` and `get_v1_servings_id`. Starting or stopping a deployment stays a deliberate action in the web application or the CLI.
## Related pages - [Service](service.mdx): the free-form version of the same machinery, for a process the platform does not build for you. - [Device](device.mdx): what a deployment runs on. - [Benchmark](benchmark.mdx): how you choose which model to deploy. --- # Organization and project The two containers every resource sits in: the organization that owns users and devices, and the project that scopes one piece of work. Source: https://docs.clika.io/platform/concepts/organization-and-project.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; Every request you make to the platform is answered in the context of one organization and one project. Knowing which is which explains most of what you can and cannot see. ## Organization An organization is the tenant: one company, or one team that bills and administers itself. It owns the user accounts, the roles those users hold, the device fleet, and every project underneath it. Your account belongs to at least one organization, and an account with no organization is refused before any other check runs. An account is created in one of two ways: an invitation from an organization's owner or admin, which adds you to that organization, or a signup code, which creates a new organization with you as its owner ([Sign up with a signup code](../getting-started/sign-up-with-a-signup-code.md)). Public self-service signup is closed during the alpha, so the sign-in page routes an unknown email address to support rather than to a sign-up form. Within an organization a member holds one of three roles. | Role | What it carries | | --- | --- | | `owner` | The organization's ultimate administrator. Creates projects and may remove any of the organization's cloud credentials. | | `admin` | Administers members and organization settings. Creates projects and may remove any of the organization's cloud credentials. | | `member` | Ordinary access to the organization's work, which includes reading the device fleet. In the projects they belong to, members register models, start and cancel their own benchmarks and deploy; read-only access is the project `viewer` role. | Creating a project is an organization owner's or admin's action. A project role, however senior, does not confer it. These sit alongside the platform's permission system, which grants capabilities such as `devices:write` or `benchmarks:read` through roles bound to your account. What you can do is the intersection: the capability must be granted, and the resource must be inside an organization and project you belong to. ### Deleting an organization Deleting an organization ends every member's access to it at once: sessions, API keys and device connections stop working, and its projects, devices and artifacts disappear from the web application. The data itself is purged by a sweep after a grace period of seven days. An organization that still has cloud machines running is not deleted straight away: the refusal lists every machine with its cloud account and region, and the deletion proceeds only once you acknowledge that list, so a machine is never left running unbilled and unseen. ## Project A project is the box one piece of work lives in: its models, benchmark runs, jobs, services, artifacts and runtime credentials. Every organization starts with a default project, and the project switcher in the top bar decides which one you are looking at. Nothing you create lands outside a project. Project membership is separate from organization membership, and it is enforced: a member of the organization who is not a member of the project cannot read the project's resources. A project member holds one of four roles. | Role | What it carries | | --- | --- | | `owner` | Full control, including deleting the project. The last owner cannot be removed or demoted. | | `admin` | Manages members and resources. | | `developer` | Creates and runs the project's work. | | `viewer` | Reads it. | Project roles are about the project's own resources. None of them carries anything about devices, and none of them lets you create another project. There is no organization-level viewer role: for a person who should only read, give them `member` at the organization and `viewer` in each project. [Roles](roles.md) lists what each role may do. ### Devices are the exception Devices belong to the organization, not to a project. Every project of an organization sees the same fleet, and a device registered once is available to all of them. This is deliberate: a physical device is a piece of shared hardware, and forcing it into one project would mean re-registering it for the next. Reading the fleet comes with organization membership itself. Registering a device, which starts with an enrollment token, and changing or removing one need a device capability granted at the organization level. No project role supplies it, so a project admin who is only a member of the organization cannot enroll a device. Everything else is project-scoped. A resource in another project answers `404`, exactly as a resource in another organization does, so you cannot learn what exists elsewhere by probing ids. ## Where you see it
The organization switcher sits top left, the project switcher top right. **Settings**, then **Organization**, holds the organization's members and its identity-provider settings. A project's own members are on its dashboard, under **Members**.
```bash clika-cli login --profile # prompts for your API key, saves it with the deployment URL clika-cli projects list clika-cli orgs list ``` A profile carries one deployment and one credential, so switching deployments is a `--profile` flag rather than a re-login.
Ask Claude: > Which projects can I see, and who are the members of the one I am in? That reaches the `get_projects` and `get_projects_id_members` tools. See [Using with AI (MCP)](../mcp/index.md).
## Related pages - [Roles](roles.md), what each organization and project role may do. - [Device](device.mdx), the org-scoped exception. - [Runtime licenses](runtime-licenses.mdx), for the runtime credentials a project issues. - [Manage members and API keys](../how-to/manage-members-and-api-keys.md), for the day-to-day operations. --- # Roles The organization roles and the project roles as the platform enforces them: what each may do, where each is set, and the points still being clarified. Source: https://docs.clika.io/platform/concepts/roles.md Access has two levels. An **organization role** says what you may do with the organization itself: its members, its devices, its projects. A **project role** says what you may do inside one project: its models, benchmarks, deployments and licenses. Everyone holds exactly one organization role and one role per project they belong to. ## Organization roles Set under **Settings**, then **Organization**. The invite dialog offers Admin and Member; ownership is transferred with **Make owner** in a member's row menu. | Role | May do | | --- | --- | | **Owner** | Everything below, plus transfer ownership and delete the organization. The web application offers the deletion under **Settings**, then **Organization** (**Delete organization**, owner only; type the organization's name to confirm, and every member loses access at once); the CLI does not, because the deletion takes a signed-in session, which the CLI does not hold. An organization always has one owner. | | **Admin** | Invite members (as Member), remove members, create projects, register, rename, tag and delete devices, open a terminal on a device, cancel anyone's benchmark, read the whole organization's activity. An admin cannot promote a Member to Admin; ask the owner. | | **Member** | Work in the projects they belong to with their project role. A member sees the device list and a device's page but cannot register a device or open its terminal (`devices:exec`), and reads only their own rows of the activity view. | There is no organization-level viewer. For a person who should only read, give them **Member** at the organization and **Viewer** in each project. ## Project roles Set on the project's dashboard under **Members**, then **Manage Members**. A new member of the organization joins its default project as **Developer**; change it there. | Role | May do | | --- | --- | | **Owner** | Everything below, plus delete the project. The last owner cannot be removed or demoted. | | **Admin** | Add and remove project members, and issue, rotate and revoke the project's runtime licenses. | | **Developer** | Register models, start benchmarks, cancel their own, create deployments, read results. | | **Viewer** | Read models, benchmarks, results and deployments. Deploy, Stop, Reveal key and Issue are hidden or disabled. | A project admin's rights hold inside that project only: an organization Member who is Admin of one project cannot invite organization members, and cannot manage another project they are only Developer in. ## What the words in the app's Roles table mean The Roles table under **Settings**, then **Organization**, lists the platform's own role catalog, including `user`, `platform_admin` and `support_admin`. `platform_admin` and `support_admin` are the deployment's staff roles, never assigned to a customer account; `user` is the base every account holds. They are not choices for your team; the two tables above are. ## API keys and roles A key carries a subset of its owner's capabilities and never more, so a key created by a Member reaches what that Member reaches. Changing who is a member, their role, or the organization's ownership takes a signed-in session whatever the key's scope. [Manage members and API keys](../how-to/manage-members-and-api-keys.md) has the rules. ## Being clarified - Whether an organization Member may reveal a deployment's key. A Member could on the hosted platform in early October 2026, with no confirmation. - Whether a project Viewer should see **Add Model** at all; the dialog it opened was a stale one. - The Member's rights on a device's **Commands** tab, which no role check has exercised. ## Related pages - [Organization and project](organization-and-project.mdx): the two containers the roles apply to. - [Manage members and API keys](../how-to/manage-members-and-api-keys.md): the day-to-day operations. --- # Runtime licenses The credential a ClikaRT runtime presents to prove it may run: the two kinds, every entitlement it can carry, entitlement profiles, and the lifecycle. Source: https://docs.clika.io/platform/concepts/runtime-licenses.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A runtime license is the credential a ClikaRT runtime presents to prove it is allowed to run. Your platform issues it, per project, and you hand it to the code you ship. Creating, updating and revoking these credentials is one of the reasons the platform exists. ## Why you need one, and how you use it [ClikaRT](/clikart) and [Modelverse](/modelverse) do not run unlicensed. When the runtime starts inside your application, it looks for a credential and checks that the credential is valid and that it grants what the application is about to do: which inference backends it may load, which operating systems it may run on, whether it may run in a container or a virtual machine, and whether it may compress models or export their state. Without a valid credential the runtime does not serve the application, and with a credential that lacks a grant the corresponding call is refused with a typed error. The credential is not tied to a person, and by default not to a machine either; a **hardware seat budget** (below) can cap how many machines hold one. It is tied to a **project** on your platform: the project is the thing you license, and every runtime that ships under that project presents the project's credential. That is why issuing one is done from inside a project, and why revoking a credential stops every runtime that presents it. Using one is a three-step loop, and the diagram below is that loop: 1. **Issue it on the platform.** Open the project, go to **Licenses**, and issue a credential. You choose the kind (see [the two kinds](#the-two-kinds)) and the entitlements it grants, or start from an entitlement profile. An Offline bundle is shown once, at creation, so copy it then; an Online key can be read back later. 2. **Hand it to your code.** The artifact you distribute is the `CLIKA1-...` text the platform shows as the license's license key, for both kinds. Put it where your ClikaRT or Modelverse installation reads it, exactly as you would any other secret your application needs at start-up: [License the runtime](/clikart/how-to/license-the-runtime) describes the places for each language binding and packaging, the `clikart-cli` command and an Android app included; the credential itself is the same for both products, because both are the same runtime. 3. **Let the runtime prove itself.** An Online key is presented to the platform when the runtime starts and again at intervals, so the platform can confirm it, record when it was last seen, and refuse it the moment you revoke it: a runtime starting afresh is refused at once, and a runtime already running is cut off at its next check-in, which the runtime makes every five minutes. An Offline bundle is verified locally against Clika's signature, which is what lets a runtime work with no route to the platform at all. ![The runtime-license loop: a project on the Clika Platform issues a credential; you place it with your application; the ClikaRT or Modelverse runtime inside the application presents it. An Online key is confirmed by the platform at start and at intervals; an Offline bundle is verified locally against Clika's signature.](/img/platform/concepts/runtime-license-flow.svg) Rotate when a credential may have leaked or when its entitlements have to change; revoke when a project is retired. Both are covered under [Lifecycle](#lifecycle). ## The two kinds The choice is about how the runtime authenticates, not about how strong the credential is. | Kind | What it is | What it needs | What it gives up | | --- | --- | --- | --- | | **Online** | A `CLIKA1-...` license key the runtime presents to the platform on each run | A route from the runtime to the platform | Nothing important | | **Offline** | A signed `CLIKA1-...` license bundle the runtime verifies locally | Nothing at run time | Real-time revocation | Both kinds arrive as the same `CLIKA1-...` text, and that text is the whole artifact you hand the runtime; a bare `clika_rk_...` API key is not a credential and is refused. Most people arrive expecting the offline bundle to be the safer artifact. It is not. A runtime holding one on a machine with no route to the platform keeps verifying it successfully until the bundle expires, because nothing can tell it otherwise; a runtime that can reach the platform picks the revocation up at its next start (within seconds in our tests), but nothing forces a runtime to reach it. **If revocation has to actually stop a running runtime, issue Online credentials.** Rotation and revocation differ by kind for the same reason. Rotating an Online key issues a new key and the platform stops accepting the old one at once, or at the end of a handover grace you choose when rotating (up to seven days), so running SDKs can switch; revoking one stops it on its next call. Rotating or revoking an Offline bundle marks it here and puts its certificate serial on the deployment's revocation list (`GET /api/v1/crl`), but a runtime already holding the bundle keeps verifying it until that list reaches it or the bundle expires: redeploy the replacement, and deliver the list where revocation has to bite before expiry. Two rules follow from an offline bundle being signed material: - A project holds as many active Offline bundles as it has runtimes or devices; one per node is the node-locked shape, and each bundle has its own leaf certificate and serial, rotated and revoked on its own. The bound is the organization's plan: its Offline-license limit caps the active bundles across the organization's projects, and issuing past it is refused with `OFFLINE_LICENSE_LIMIT_REACHED` naming the limit. Rotating a bundle replaces it, so a rotation never counts against the limit. - An offline bundle's entitlements cannot be edited in place, because they sit inside a signature. Rotating re-signs them, which is the only honest way to change them. The credential's name is the one field you can edit, precisely because it is not signed. Online credentials have no such bound: a project can hold as many as it has runtimes, and revoking one is enforced by the platform at once and reaches a running runtime at its next check-in, within five minutes. ## Entitlements An entitlement is a permission carried inside the credential. A runtime credential carries a flat set of grants and no quotas: every entitlement is on or off, and off is the default for anything not granted. All of them are always present in the credential, so "not granted" is explicit rather than an absence somebody has to interpret. ### Backends Which inference backend the runtime may load. A backend the credential does not grant will not load, whatever the hardware offers. | Entitlement | Grants | | --- | --- | | **CPU** | CPU inference. Always granted: ClikaRT cannot execute without it, so the platform pins it on even when your organization's ceiling would have removed it. | | **NVIDIA CUDA** | GPU acceleration on NVIDIA hardware. | | **Vulkan** | GPU acceleration through Vulkan, which works on AMD, Intel, Qualcomm and NVIDIA GPUs and is the mobile GPU path. | | **Apple Metal** | GPU acceleration on Apple silicon. | | **Qualcomm NPU** | NPU acceleration on Qualcomm Hexagon. | | **Google TPU** | Google TPU acceleration. | | **Intel NPU**, **AMD NPU**, **WebGPU** | Reserved for backends the runtime does not ship yet. They exist in the format so a credential issued today stays readable when they do. | ### Operating systems Which platforms the runtime may run on. The allowed set is exactly the entitlements set to true, and at least one has to be granted, so a credential that permits no operating system is refused at issue rather than shipped as an unusable artifact. **Linux**, **Windows**, **macOS**, **iOS**, **Android**, and **Web** (reserved for the WebAssembly target). A device whose operating system is not granted is refused with `PLATFORM_NOT_ALLOWED`. ### Execution environment | Entitlement | Grants | | --- | --- | | **Container allowed** | Running inside a container, for example under Docker or Kubernetes. | | **VM allowed** | Running inside a virtual machine. | Both are granted by default when a credential is issued. Turning one off is how a license that is meant for physical hardware stays on physical hardware. ### Capabilities | Entitlement | Grants | | --- | --- | | **Compression allowed** | The model-compression pipeline. Without it a compression call is refused with `COMPRESSION_NOT_LICENSED`. | | **State export allowed** | Exporting optimized state to a third-party format such as ONNX or LiteRT. Off unless granted, and a refusal reads `EXPORT_NOT_LICENSED`. | ### What is not an entitlement Whether a runtime can work air-gapped is not a grant. It follows from which kind of credential it holds: an Online key needs the platform, an Offline bundle does not. If you see `is_airgap`, `telemetry_level` or a requests-per-second field on an old credential, those are retired names the current format does not carry. `max_hardware` is not an entitlement either: it is the platform's own seat budget, described below, and the runtime never receives it. ## Entitlement profiles A profile is a reusable template for the grants above, scoped to your organization: a name, a description, the entitlement set, and a token lifetime in days (14 when you do not set one). It exists so a team issuing many credentials does not re-enter the same grant set each time, and so "what a lab device gets" is a thing with a name rather than a habit. The property that matters: **a profile is a prefill, not a live link.** An issued credential owns its own copy of the grants, so editing the profile afterwards never retroactively changes a credential already in the field. That is deliberate, because the alternative would mean a signed artifact whose meaning changes behind the holder's back. Profiles feed the enrollment path, where a device activates against a code that names one. Issuing a project credential directly takes its entitlements from the request instead, which is why the issue dialog shows the grants rather than only a profile name. What you ask for is capped by your organization's ceiling: a grant your organization does not hold cannot be issued to a project, and the issue response names any entitlement it had to narrow. The credential's lifetime is clamped the same way, to the shortest of what you asked for and what the issuing certificate itself outlives. ## Lifecycle - **Issue** mints a credential for a project. An Offline bundle's signed material appears in the issue response and nowhere else, so a lost one is replaced by rotating. An Online key can be read back after issue, which is recorded in the audit log. Treat both the way you treat any secret. - **Copy** on the credential's detail page hands an Online key back to a license manager of the project. The page shows the key's `clika_rk_` prefix, and **Copy** puts a freshly signed `CLIKA1-` bundle on the clipboard, which is the whole value a runtime takes; each one is recorded in the audit log. An Offline bundle cannot be read back: its private key is never stored, so a lost bundle is replaced by a rotation. - **Rotate** replaces a credential with a fresh one. For an Online key an acceptance grace can keep the old one working briefly; for an offline bundle it revokes and re-issues in one step. A runtime holding the old Online value is refused at its next check as revoked; a runtime holding the old offline bundle keeps working until that bundle expires, so redeploy it. The credential a rotation supersedes shows as rotated out, with the replacement beside it; a runtime still presenting a rotated-out Online key is refused with `LICENSE_STATE_BLOCKED: runtime key rotated out; the project issued a replacement`. - **Rename** is the one edit an offline bundle accepts, because the name is not signed. - **Revoke** ends a credential. An Online key is refused by the platform from that moment; a running runtime is cut off at its next check-in, within five minutes, and a new start is refused at once. The running runtime reports the cut-off as `LICENSE_STATE_BLOCKED`. An offline bundle stops being served immediately; a runtime already holding it is refused at its next start if the machine reaches the platform, and keeps verifying it until it expires if it does not. - **Expiry** ends a credential on its end date, with no grace period after it. A call made from then on is refused with the code name `LICENSE_EXPIRED`, where a call made with no credential at all, or an invalid one, is refused with `LICENSE_FAILED`. Re-issue before the date rather than after it. - **Last seen** is recorded per credential, which is the fastest way to find one nothing is using any more. ## Hardware seats A credential can carry a **hardware seat budget**: how many distinct machines may hold it. You set it when you issue or rotate a credential and change it in place on the credential's detail page. The web form takes a count of 1 or more and starts at 1 on every plan (one license takes one seat unless you type more); on a plan with a seat budget a request past the seats left is refused with `409 HARDWARE_SEAT_BUDGET_EXCEEDED`, which names the budget. A credential issued through the API, the CLI or the MCP server with no budget, or with `0`, has no seat limit, and the detail page shows it as Unlimited. The platform enforces it at every runtime check-in, for Online keys and Offline bundles alike: the first machines to present the credential take the seats, a machine beyond the budget is refused with a reason that says so, and a credential with a budget refuses a runtime that does not identify its hardware at all. The detail page lists every machine holding a seat with when it was first and last seen and a **Release** action per row, and the credential list shows seats as used / max. The budget is platform-side only: it is never part of the entitlements the runtime receives, so changing it does not re-issue anything. What identifies a machine is the runtime's hardware fingerprint, one digest over what the host lets the process read. On Linux, macOS and Windows that is the machine's own identity (the OS install id, the board and disk serials, the network hardware) beside its processor and accelerators. On Android it depends on the process: - **An app that loads the runtime through `ClikaRtAndroid.load`** identifies the device through the identity Android scopes to the app (`Settings.Secure.ANDROID_ID`), which survives a reinstall under the same signing key and changes on a factory reset. - **A command the app launches, such as the runtime's own command line run from a launcher app**, cannot read the device's serial (Android hides it from app processes) and has no app state of its own, so the runtime identifies the **installation** instead: an identity it mints once and keeps in a file under the cache root the launcher sets for the process (`XDG_CACHE_HOME`, or `HOME`). Two apps on one phone are two installations; clearing the app's cache or reinstalling it mints a new one, which takes a new seat. A launcher that sets neither variable leaves the process with no place to keep the identity, and the runtime then reports only the device model, which a budgeted credential refuses. - **A shell on the phone** (`adb shell`) reads the device serial and identifies the unit by it. The credentials the platform mints for its own benchmark runs are single-machine and are never counted as your organization's license usage. ## Where you see it
**Licenses** in the sidebar, inside a project. The **Licenses** tab lists the credentials with their kind, status, issue and expiry dates, key prefix and last-seen time; **Issue License** mints one; the **Profiles** tab holds the entitlement templates.
```bash clika-cli projects licenses ``` The CLI lists a project's credentials. Issuing, rotating, revealing and revoking take `project_licenses:write`, which no API key can carry, and the CLI signs in only with an API key, so under a key those commands answer `403 API_KEY_SCOPE_INSUFFICIENT`; do them in the web application. [Licensing](../cli/licensing.md) documents the commands.
Ask Claude: > List the runtime licenses in this project and tell me which ones have never been used. That reaches `get_projects_id_licenses`. Minting a credential is deliberately outside every served toolset, so issuing one stays a human action in the web application or the CLI.
## Platform licensing Everything on this page is about the runtime license, the credential a ClikaRT or Modelverse runtime presents when it starts inside your application. The platform that issues it is licensed separately, as a deployment, and that license is not something you issue from the web application. The platform itself can run where you need it. Clika operates it as a hosted service, and the same installation is available on-premise, either connected to the internet or fully air-gapped. An on-premise platform issues runtime licenses to your projects exactly as the hosted one does, and the runtime credential it produces is the same artifact described above. To ask about an on-premise installation of the platform, [contact Clika](https://clika.io/contact). ## Related pages - [Organization and project](organization-and-project.mdx): what a credential is issued against. - [ClikaRT](/clikart) and [Modelverse](/modelverse): the runtimes that present the credential. - [License the runtime](/clikart/how-to/license-the-runtime): where a ClikaRT or Modelverse runtime reads the credential, per language and packaging, and what a refused call looks like. --- # Service A process the platform keeps running on a device: its definition, desired state and reconciliation, and how it differs from a model deployment. Source: https://docs.clika.io/platform/concepts/service.md import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; A service is a process you tell the platform to keep running on a device. Unlike a [job](job.mdx), which runs to completion and produces a result, a service is meant to stay up, and the platform restarts it when it does not. This is the free-form version: you supply the command. When what you want to run is one of your own models behind an endpoint, the platform builds that service for you, and it has its own page ([Model deployment](model-deployment.mdx)). ## Why it exists Some work on a device is not a run with an end. A collector that samples a sensor, a helper daemon a benchmark needs, a container that has to be up while a test runs: all of it needs to be started, watched, restarted and eventually stopped, on a machine you may not be able to reach directly. A service is that, expressed once and then maintained by the platform. ## Desired state, not commands The platform stores a **desired state** per device: which services should be running, with what command, environment, artifacts and health check. A reconciler compares that against what the device reports in its heartbeat and issues the difference. Three useful behaviors fall out of that. - **A start survives an offline device.** Desired state persists while a device is unreachable, so starting a service on a device that is not connected is not refused: it is recorded, the service reads `pending`, and it starts when the device comes back. - **Crashes self-heal.** A service that dies falls out of the reported state, so the reconciler starts it again, with a backoff that gives up after repeated failures rather than restarting forever. - **The fleet is declarative.** You state what a device should run, and the platform converges toward it. Starting or stopping through the API or the web application updates desired state too, so the imperative path gets the same tracking. ## The service definition A definition is a reusable template. Its fields: | Field | What it declares | | --- | --- | | `command` | The process to run, as an argument array. | | `type` | `process` (a command) or `compose` (a Docker Compose project on a device that has Docker). | | `env` | Environment variables for it. | | `artifacts` | Files staged on the device before it starts, by artifact and tag. | | `health_check` | How the platform decides it is up: a URL to poll, an interval and a timeout. | | `auto_restart` | Whether to restart it when it exits. | | `port` | The port it listens on. | | `required_resources` | Minimum hardware the device must report. | | `resource_limits` | Bounds on the running process: CPU cores and memory. | [Write a service definition](../how-to/write-a-service-definition.md) covers the YAML in full, including how to start a service from a definition and override its defaults per device. ## States A service moves through `pending` (recorded, not started yet), `starting`, `running`, and then `failed` or `stopped`. A stop whose cleanup could not finish reads `cleanup_failed`. Stopping a service stops its whole process group, so a wrapper script that forked workers takes them with it. On Unix the stop is graded, a polite signal first and a kill once the grace period is spent; on Windows the process group is terminated outright, because a service there has no console to receive the polite signal. ## Where you see it
**Services** in the sidebar holds the definitions and the running services; a device's own **Services** tab shows what that device is running, with its state, uptime, restart count and a live log.
```bash clika-cli apply -f service-definition.yaml # kind: ServiceDefinition clika-cli apply -f start-service.yaml # kind: Service, naming its devices ```
Ask Claude: > What services are running on my devices, and has anything restarted more than once? That reaches `get_services` and the per-device service state on `get_devices_id`. Starting and stopping services stays outside the served toolsets, so those remain deliberate actions.
## Related pages - [Model deployment](model-deployment.mdx): the service the platform builds for you, for serving a model. - [Device](device.mdx): what a service runs on. - [Job](job.mdx): the other thing the platform runs on devices. --- # First steps New to the CLIKA Platform? Start here. What it is, the five-minute mental model, and a tutorial that ends with a benchmark result you can read. Source: https://docs.clika.io/platform/getting-started.md New to the platform? This section is where to start. It gives you enough of the model to hold the whole product in your head, then walks you from an empty fleet to the numbers of your first benchmark. Read it in order: 1. **[The Platform at a glance](overview.mdx)**: what the platform is and is not, who it is for, and the five ideas the rest is built on. 2. **Tutorial: your first benchmark**: six parts, each a complete step, with concepts explained where they first appear. [Take the tour](first-benchmark/01-overview.md), [add your first device](first-benchmark/02-add-your-first-device.md), [pick a device](first-benchmark/03-pick-a-device.md), [add a model](first-benchmark/04-add-a-model.md), [run a benchmark](first-benchmark/05-run-a-benchmark.md), and [read the results](first-benchmark/06-read-the-results.md). 3. **[What to read next](next-steps.md)**: where to go once you have a result. 4. **Quickstarts**, when you already know what you want and need one screen per way in: [the web application](quickstarts/web.md), [the CLI](quickstarts/cli.md), [an AI assistant over MCP](quickstarts/mcp.md) and [a licensed runtime on your own machine](quickstarts/licensing.md). Each ends in a first result and says how long each step takes. You need an account on a CLIKA Platform deployment, and a machine you can install software on. The tutorial's second part registers that machine and says what qualifies; if your team already runs a fleet you can skip it. An account comes from an invitation sent by an organization that already exists, or from a signup code that creates a new organization with you as its owner: [Sign up with a signup code](sign-up-with-a-signup-code.md) covers the second. Public self-service signup is closed during the alpha, so the sign-in page points an unregistered address to support. ## How the Platform docs are layered - **This section** orients: condensed, in reading order, concepts woven in. - **[Concepts](../concepts/index.md)**: one page per object (organization, project, device, model, artifact, benchmark, job, service, deployment), for when you need the full picture of one thing. - **[How-to guides](../how-to/index.md)**: problem-oriented recipes, from writing a job definition to managing API keys. - **[Using the CLI](../cli/index.md)**: the platform from a terminal, from the first command to every `clika-cli` command with its arguments and flags. - **[Using with AI (MCP)](../mcp/index.md)**: connecting Claude Desktop, Claude Code, Codex and other MCP clients to the platform. --- # Add a model Register a Hugging Face model with the platform, and understand why registration moves no weights and why the task matters. Source: https://docs.clika.io/platform/getting-started/first-benchmark/add-a-model.md A benchmark measures a model. This part registers one. It takes about ten seconds, and the reason it is that fast is worth understanding. ## Register the model Open **Models** and select **Add Model**. ![The Models list with the Add Model action](/img/platform/first-benchmark/03-models-list.png) Paste a Hugging Face model URL or its `owner/repo` identifier. An access token is only needed for a gated or private repository; public ones need nothing. ![The Add HuggingFace model dialog, with a model URL field and an optional access token](/img/platform/first-benchmark/03-add-model.png) For a first run, pick something small so the whole loop finishes in minutes rather than hours. `Qwen/Qwen2.5-0.5B-Instruct` is a reasonable first model on a workstation-class device. ## What registration did, and what it did not The platform read the repository's metadata and stored a reference: the canonical URL, the task, and the architecture and framework where the metadata names them. **No weights moved.** You can register a 70B model from a laptop, because nothing is downloaded until a run needs it, and then it is downloaded by the device that will run it. A Hugging Face repository has one record on the platform. When someone, in your organization or another, registered the same model before you, **Add Model** returns that record and lists it under your models as a platform catalog model, dated from the day it was added to your organization. Your benchmarks, results and deployments on it are yours and stay in your project; the record's own fields (name, task, architecture) are shared, so an edit to them is refused as shared. The **task** is the field that does real work. It is the model's Hugging Face `pipeline_tag`, for example `text-generation`, `automatic-speech-recognition` or `image-text-to-text`, and it decides which quality tests can run against the model in the next part. If the metadata does not carry a task, the model still registers and you can set the task by hand; the picker then offers the right tests. The list groups models into categories derived from that task: LLM, Vision, Audio, Image, Multimodal, Embedding. It is the same grouping the benchmark picker uses. ## The Hugging Face access token The dialog has an access-token field, and it is worth understanding what it is for before you need it. **What it is.** A personal access token from your Hugging Face account. It proves to the Hub that you are allowed to read a repository that is not public. **Why the platform needs one.** Public repositories need nothing at all, and most first benchmarks use one. A token becomes necessary when the model or the dataset behind a benchmark is **gated** (you accepted terms to get access, as with several Llama and GPQA releases) or **private** (it belongs to you or your organization). Three moments need it: reading the repository's metadata when you register the model, fetching the checkpoint on the device when a benchmark runs, and fetching it again when you [deploy the model](../../concepts/model-deployment.mdx) as an endpoint. **Where to get one.** In Hugging Face, under your account settings, at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens). A token with read permission is enough. A fine-grained token also works, as long as it grants access to the specific repositories you plan to benchmark. **Where it goes.** The token you type into the Add-model dialog is saved as your organization's Hugging Face credential, and the dialog says so: one token then serves the metadata read now, the checkpoint fetch when a benchmark runs, and the fetch when you deploy the model. You can see and replace it under **Settings**, then **Credentials**, as a HuggingFace credential. There a credential is either **Personal** (only your own runs use it) or **Organization** (anyone in the organization can), and the platform stores it encrypted. When both exist, your personal credential wins over the organization's. **What happens without one.** A public model behaves exactly as it did above. A gated model registers but the run fails on the device when the fetch is refused, and registering one from the dialog without a token reports that the metadata could not be fetched and to check the URL and, for gated or private models, the token. ## Pick one from the Modelverse The **Models** page has two views, switched at the left of its toolbar: **My Models**, the project's registered models, and **Modelverse**, the catalog of models CLIKA has confirmed to run well on ClikaRT. A link to the page with `?view=modelverse` opens the catalog directly, and the page remembers the view you left. Each catalog row is one model: the provider's mark and the `owner/repo` on the first line, and on the second its type, its parameter count, and **From**, the memory its smallest packaging needs to run. Clicking a row registers the model into My Models exactly as **Add Model** does, with the repository's public metadata; a row already registered in this project reads **Added**. Adding needs the same models write permission as Add Model, and a platform whose models come from its model store offers the catalog to browse but not to add from. The filter panel on the right narrows the rows. - **Type** lists the categories the catalog carries, each with the count it would leave under the other filters. - **Parameters** is a range with two handles over five stops (`<1B`, `3B`, `8B`, `14B`, `>30B`). The lower handle is the smallest size kept, the upper handle the largest, so `3B` to `14B` keeps models of 3 to 14 billion parameters; the two handles never share a stop. - **Fits my devices** lists your organization's devices with the memory the fit rule reads for each, the GPU's memory where the engine would use the card and the device's total memory otherwise. Selecting devices keeps only the models whose smallest packaging fits every one of them. A device whose memory the platform does not know rules nothing out, and the section is absent for a member who may not list devices. The name field in the toolbar searches by `owner/repo`; **Clear All** resets the filters and keeps the search. The catalog is curated by a platform administrator in the admin dashboard, on its **Modelverse** page: the entries, their providers and the provider logos the rows show. An empty catalog means nothing has been curated on this platform yet. [Modelverse](/modelverse) documents the model library itself. ## The other way in You do not have to register a model first. The benchmark flow accepts a Hugging Face URL directly, and registers it as a model on the way through. Registering up front is worth it when several people run against the same model, because everyone then picks it from a list instead of pasting a URL and hoping they pasted the same one. ## What you have now A model reference, a device, and a project holding both. The next part runs one against the other. - Full detail: [Artifact](../../concepts/artifact.mdx). - Next: [Run a benchmark](05-run-a-benchmark.md). --- # Add your first device Register a machine with the CLIKA agent: the install command, the activation code it carries, and what each of Linux, macOS, Windows and Android does differently. Source: https://docs.clika.io/platform/getting-started/first-benchmark/add-your-first-device.md A benchmark runs on hardware you own, so the platform needs a machine before it can measure anything. This part registers one, which is a single command on the machine itself. You need something you can install software on and that can reach the platform over the network: a Linux workstation or server (x86_64, arm64 or armv7), a Jetson or similar single-board computer, a Mac, a Windows machine, or an Android phone. The device dials out to the platform, so it needs no public address and no inbound port. A laptop behind a home router qualifies. :::note If your team already runs a fleet, you can skip this part: the devices are shared across every project of the organization, so they are already in your list. [Pick a device](03-pick-a-device.md) is where you choose one. ::: ## Start the registration Open **Devices** in the sidebar and select **Register Device**. The dialog asks which platform you are installing on, gives you a one-line command to run there, and lists what that platform does differently. ![The Add Device dialog on the Linux tab, with the platform tabs, the downloader and install-scope choices, and the install command](/img/platform/first-benchmark/02-register-linux.png) The command carries a short **activation code** (`?code=XXXX-XXXX` at the end of its address), which is how the machine proves it may join your organization without anyone typing credentials on it. The dialog mints the code when it opens; the code enrolls one device and expires 30 minutes later, so open the dialog when the machine is ready. And the agent connects on its own once installed, so there is no device-side configuration step. The device appears in the list within a minute of the installer finishing, and the dialog turns into the device list by itself when it does. Give the device one more minute before its first benchmark: it is listed Online as soon as it connects, and the channel a job needs follows shortly after. Treat the command as a secret while it is unused. Anyone holding it can put a device into your organization, which is exactly what it is for. ### Registering a batch **Reusable, many devices** is a switch at the top of the dialog, beside the copy buttons. Turning it on generates a new command and revokes the one shown before, so set the switch first and copy the command after. The batch command registers any number of devices and expires after ten minutes. That is the tool for provisioning a rack or a shelf of phones in one pass, and the short life is what keeps a copied command from being useful later. ### Registering with a token The dialog's command already carries an activation code, and [Register a device by code](../../how-to/register-a-device-by-code.md) covers codes. An **enrollment token** is the other way in: the CLI mints one ([below](#from-the-cli)), and a device it enrolls is the same kind of device as one enrolled with a code. ## Linux Choose `curl` or `wget` for the downloader, and whether to install for the current user or system-wide with `sudo`. - With `sudo`, the agent installs to `/usr/local/bin` and registers a system-level systemd service, so it runs whether or not anyone is signed in. - Without `sudo`, it installs to `~/.local/bin` and registers a user-level systemd service, with a `cron @reboot` fallback where user services are unavailable. - Architectures: x86_64 (amd64), arm64 (aarch64) and armv7, so a Jetson or a Raspberry Pi class board is the same command as a server. Appending `-s -- --help` after the `| sh` lists every installer option and environment override without installing anything, which is the safe way to look before you run. If the platform refuses the device, the installer stops at "Enrollment with the activation code failed", with the reason printed above that line. (A token install from the CLI ends the same way, `==> ERROR: The platform refused this device` and exit 1; the agent's journal keeps the detail: `journalctl --user -u clika-runtime-agent` for a current-user install, the system unit under `sudo`.) The three refusals a first install meets: `ACTIVATION_CODE_USED` (the command's code has already enrolled its device; generate a new one from the dialog), "this device name is already registered" (a second agent on the same machine; set `DEVICE_NAME=` before the command), and "Your organization has N devices enrolled, the most its plan allows" (remove a device you no longer use, or ask support to raise the limit). Running the same command twice on one machine keeps one device: a token command answers "already enrolled ... nothing to do", and the second run of a single-use code command stops with `ACTIVATION_CODE_USED` and exit 1 while the first agent keeps running. ### From the CLI The same registration from a terminal is two commands: mint an enrollment token (the body needs a `name`), then print the install script for the device's operating system. `install list --raw` prints the whole script, so on the device itself pipe it to `sh`: ```bash clika-cli enrollment-tokens create --body '{"name":"my-laptop","max_uses":1}' ``` ```bash clika-cli install list --os linux --token --raw | sh ``` `clika-cli devices list` shows the device once it has connected. [Devices](../../cli/devices.md) has the rest of the device commands. ## macOS ![The Add Device dialog on the macOS tab](/img/platform/first-benchmark/02-register-macos.png) macOS offers the same choice as Linux. As the current user, the agent installs to `~/.local/bin` and registers a launchd agent for your account; with `sudo`, it installs to `/usr/local/bin` and registers a system-level launchd daemon that runs whether or not anyone is signed in. Apple silicon (arm64) Macs are supported. Intel Macs are not. ## Windows ![The Add Device dialog on the Windows tab, with the current-user and Administrator scopes](/img/platform/first-benchmark/02-register-windows.png) The command is PowerShell rather than a shell one-liner. Installing as Administrator registers a Windows service that runs even when nobody is signed in. Installing as the current user needs no elevation: the agent lands under `%LOCALAPPDATA%\Clika\DeviceAgent\`, starts at your sign-in, and runs as you with your file permissions, which also means it stops when you sign out. Both amd64 and arm64 are supported. ## Android ![The Add Device dialog on the Android tab: download the APK, then scan the enrollment QR code or paste the token](/img/platform/first-benchmark/02-register-android.png) Android is an app rather than a script, in three steps: download the APK, open the Clika Runtime Agent on the phone, and scan the enrollment QR code the dialog shows. If the phone cannot scan, the same token is there to paste into the app by hand. It needs Android 8.0 or newer, and the browser doing the download needs the "install unknown apps" permission when it prompts. The QR code encodes the enrollment token; the same token is printed under it for pasting into the app when scanning is not an option. ## What you have now A machine in the fleet, connected and reporting. The next part finds it in the device list and reads what the platform knows about it. - Full detail: [Device](../../concepts/device.mdx). - Next: [Pick a device](03-pick-a-device.md). --- # Overview A tour of the signed-in platform: what each section of the navigation holds, and how the organization and project switchers decide what you see. Source: https://docs.clika.io/platform/getting-started/first-benchmark/overview.md This tutorial takes you from an empty project to a benchmark result you can read, in six parts. Each part is a complete step, and the concepts are explained where they first appear. This part is the tour: what the platform contains and where each thing lives, so the four parts that follow are navigation you already recognize. If you want the conceptual model first (what the platform is, and the five ideas the whole product rests on), read [The Platform at a glance](../overview.mdx). This page is about the application in front of you. ![The signed-in platform: the organization switcher top left, the project switcher top right, and the navigation down the left side](/img/platform/first-benchmark/01-dashboard.png) ## The two switchers decide everything else **The organization switcher, top left.** An organization is the tenant: your company or team, the accounts in it, and the devices it owns. If you belong to more than one, this is where you change which one you are in. **The project switcher, top right.** A project is one piece of work, and it holds the models, benchmark runs, deployments and licenses of that work. Every organization starts with a default project. Anything you create lands in the project shown here, so if you later cannot find a run you made, check that you are in the project you made it in. One exception is worth learning now, because it surprises people later: **devices belong to the organization, not to a project.** Every project sees the same fleet, because a physical machine is shared hardware rather than the property of one piece of work. ## What the navigation holds The sidebar is the whole product. In the order you meet these things: - **Dashboard** is the active project: its recent benchmark runs, its deployments, its models, its members and its licenses. It is also the fastest route to a new run. - **Models** is what you can benchmark. A model here is a reference to a Hugging Face repository plus the task it performs. - **Benchmarks** is your runs, newest first, each with its status. **New Benchmark** starts one. - **Deployments** is your models running as endpoints on your own devices, which is the step after a benchmark tells you which model to ship. - **Licenses** is the runtime credentials this project issues, which is what a ClikaRT runtime presents to prove it may run. - **Devices** is the fleet, shared across the organization's projects. A device's page is where you read its health and reach into it. Three things live under it: - **Jobs**, the individual executions behind your runs: one job per model and device pair, with its state and its log. - **Services**, anything else the platform keeps running on a device, where you supply the command yourself. - **Artifacts**, the file library: datasets, weights, scripts, engine bundles. Versioned, checksummed, and pushed to devices on demand. Two ways to run on hardware you do not own sit inside those pages rather than in the sidebar: **Cloud helpers**, a view on the Devices page for your own cloud instances and the gateways they reach the platform through, and **hosted devices**, phones the platform rents by the minute, offered as a device choice when you start a benchmark. Your account menu, top right, opens **Settings**: your profile, your API keys under **Developer access**, your organization's members under **Organization**, its plan and credits under **Subscription**, and its activity trail under **Activity**. A platform administrator also sees **Admin**, the deployment's own administration, which is not a tenant surface. The same menu holds **Download ClikaRT**, the runtime SDK and the license it needs ([Download the ClikaRT SDK](../../how-to/download-the-clikart-sdk.md)). Members live in two places on purpose: the organization's members under **Settings**, then **Organization**, and a project's members on that project's dashboard under **Members**; [Roles](../../concepts/roles.md) says which role is set where and what each may do. **Settings**, then **Credentials**, is where a Hugging Face token and a cloud provider credential are stored. ## What this tutorial uses Four of those: **Devices** to register and choose the hardware, **Models** to register what you are measuring, **Benchmarks** to run it, and the run's own page to read the result. The rest is there when you need it, and [Concepts](../../concepts/index.md) has a page on each. ## What you have now The map. The next part puts a machine of your own into the fleet. - Full detail: [Organization and project](../../concepts/organization-and-project.mdx). - Next: [Add your first device](02-add-your-first-device.md). --- # Pick a device Find your device in the fleet, read what the platform knows about it, and learn the device page tab by tab. Source: https://docs.clika.io/platform/getting-started/first-benchmark/pick-a-device.md Your device is registered. This part finds it in the fleet, reads what the platform knows about it, and walks the device page, which is where you will come back whenever a run behaves oddly. If your team already had devices before you arrived, everything here works the same. Pick any one that is online, rather than the one you added in the previous part. ## Read the fleet Open **Devices**. The list is the organization's whole fleet, shared by every project. ![The Devices list: the online and offline counts above a table of status, name, hardware, operating system and architecture, tags, who registered the device and when it was last seen](/img/platform/first-benchmark/02-devices-list.png) Each row carries what you need to choose one: whether it is online, what hardware it has (cores, memory, GPU and VRAM where there is one), its operating system and architecture, and when it was last seen. Pick one that is online and not busy. A workstation-class Linux machine finishes a first run fastest; a phone or a small board works too, and takes longer. The list shows whether a device is **Online** or **Offline** (no heartbeat for 60 seconds), and for a cloud device whether it is still **Provisioning** or its provisioning failed. What an online device is doing shows on its own page and in the benchmark's device picker: | State | What it means for your benchmark | | --- | --- | | Idle | Ready to take work. In the first minute after enrollment a device is listed Online before its job channel is up; a benchmark started then may be refused, so wait a minute and run it again. | | Busy | Running a job already. Yours will queue behind it. | | Suspended | Taken out of service by an operator. | | Dirty | A previous job's teardown failed. New benchmark jobs are blocked until it is cleared. | ## Read the device page Select the device you picked. ![A device page showing status, tags, device info and live health](/img/platform/first-benchmark/02-device-detail.png) The header carries the status and how long ago the device was seen, with **Terminal** and **Remote Desktop** beside it. **Device Info** is the hardware inventory the agent detected when it registered. **Health** is live, refreshed by a heartbeat every 15 seconds: CPU, memory and disk use, temperature, uptime, and history over the last day. The tabs are the rest of what you can do with the machine. - **Overview** is the page above: inventory, live health, health history, and the device's tags. - **Benchmarks** is the benchmark runs this device took part in, with a link into each result. - **Jobs** is everything the platform has run on this device, newest first, each with its state and its result. After the next two parts, your benchmark appears here. - **Services** is what the device keeps running: a process per row with its state, uptime and restart count, and a live log you can open per service. This is where a [model deployment](../../concepts/model-deployment.mdx) shows up on the device that serves it. - **Files** is a file manager for the device. Browse its filesystem, upload from your machine or straight from the artifact library, download a file, and save what you downloaded back into the library as a new artifact. Large uploads are chunked and resume after a dropped connection. - **Commands** runs a one-off command and records it with its exit code, its output and the time it ran, so a command you ran last week is still readable this week. Two buttons in the header do the interactive work. **Terminal** opens a shell in the browser, a real PTY on Linux, macOS and rooted Android. **Remote Desktop** opens the device's screen where a provider exists for that platform; when one cannot be started, the device says why and names the remedy rather than hiding the button. Both buttons need the `devices:exec` capability, which the organization's owner and admins hold. For other roles the Terminal button is still shown, and the page it opens reads "Disconnected" and "Connection lost" rather than naming the permission; a Reconnect there cannot help. Ask an admin to run the command for you, or use the **Commands** tab where your role allows it. All of it travels over the single connection the agent opened outward to the platform. Nothing connects inward to the device, which is why a machine behind a home router or a corporate firewall works with no ports opened. ## What you have now A device the platform can reach, and enough of its page to recognize a healthy one. The next part registers the model you want to measure. - Full detail: [Device](../../concepts/device.mdx). - Next: [Add a model](04-add-a-model.md). --- # Read the results Read a finished benchmark run: the headline cards, the quality and performance sections, the per-sample evidence behind a score, and how to share it. Source: https://docs.clika.io/platform/getting-started/first-benchmark/read-the-results.md The run is finished. This part reads it. ## The summary Open the run from **Benchmarks**. It opens on **Summary**. ![A finished benchmark run: headline cards, a quality comparison and a performance section with per-metric tabs](/img/platform/first-benchmark/05-results-summary.png) The screenshots on this page come from a finished run that compares two models on one device, so there is something to read in every section. A single-model run shows the same sections with one bar in each. The header states what the run was: its status, the models, the devices, when it started and finished, how long it took, whether it ran the quick set or the full dataset, and the benchmark's license with a link to its source. **The headline cards** answer the three questions people ask first. Best accuracy is the highest quality score in the run. Fastest is the highest throughput, with the device it was measured on. Most lightweight is the smallest peak memory, again with its device. Each names the model, so a two-model run tells you which one won each question at a glance. **Quality** compares every model in the run on the selected quality metric, with the direction stated ("higher is better") so a number is never ambiguous. **Performance** carries one tab per metric. A performance benchmark runs a concurrency sweep, three prompt shapes at one, four and eight parallel streams, and the sweep rows are the section's primary detail, so you see how the device holds up under load and not only its single-stream figure. | Metric | What it measures | Direction | | --- | --- | --- | | Tokens/sec | Decode throughput, the tokens the model produced per second. | Higher is better. | | TTFT | Time to first token, the wait before output starts appearing. | Lower is better. | | TPOT | Time per output token, the pace once output has started. | Lower is better. | | Total latency | End to end time for the request. | Lower is better. | | Peak Memory | The most memory the run held at once. | Lower is better. | Where a metric was measured over many samples, the result also carries its minimum, median and maximum, so you can see whether an average hides a long tail. An LLM Performance run also shows the sweep as a table of **cells**. A cell is written `prompt tokens:output tokens` (`128:128` is a 128-token prompt answered with 128 tokens), and `c1`, `c4` and `c8` are the number of parallel streams. A run with only LLM Performance shows its own cards (TTFT, Tokens/sec, Peak decode, TPOT and Peak Memory), and its TTFT and Tokens/sec cards are the median over every cell of the sweep, with the minimum and maximum beside it. That median mixes one, four and eight streams, so to compare devices, compare the single-stream (`c1`) rows of the table. **Peak decode** is the busiest cell's total, which rises with streams while time to first token grows. A Performance run carries no quality score, because it does not check the answers; add an accuracy test to the run for that. ## The evidence behind a score A score on its own is a claim. **View output data** opens the per-sample record behind it, in a dialog over the result, with a model switcher when the run compared several. ![The per-sample output view, with answer-outcome filters and one question per card](/img/platform/first-benchmark/05-output-data.png) The card fits the task: a text answer beside its reference, the predicted and expected label for image classification, the generated image or audio for a generation task, the transcript for speech recognition. Each card is one item the model was given: the input it saw, what it answered, what the expected answer was, and whether that counted as correct. The filters at the top narrow to correct, incorrect or unanswered items, which is the fastest way to understand a score that surprised you. A multiple-choice test shows the choices in the order the model saw them; a transcription test shows the audio reference, the predicted text and the word error rate; a coding test shows the program and its test outcome. This is the difference between "the model scored 33 percent" and knowing which kind of question it missed. ## Share it **Share Results** turns the run into something you can hand to someone who has no account on your deployment. ![The Share Results dialog: a result card preview, with Share, Copy link, Download Image and Make private](/img/platform/first-benchmark/05-share.png) The dialog previews the card the recipient sees: the run's name, its models and devices, and the three headline numbers. Under it are the ways to pass it on. - **Share** publishes the run and gives it a link. From then on the state line reads that anyone with the link can view it, and the run is readable without signing in. - **Copy link** puts that link on your clipboard, which is what you paste into a message or a ticket. - **Download Image** saves the preview card as a picture, for a slide or a chat where a link would not render. It is a snapshot: it does not update when the run does, and it carries no link back to the platform. - **Make private** revokes the link. Anyone holding it stops being able to open the run, immediately. There is one live link per run, so re-sharing does not accumulate links behind your back, and revoking is a single action rather than a hunt. ## Deploy it **Deploy** on the result creates a deployment of the model the run measured, which [Model deployment](../../concepts/model-deployment.mdx) covers. It lands in the project shown in the project switcher, so check the switcher before pressing it when you work in more than one project. The benchmark has already put the engine on the device, so the deployment reads "Starting" for about two to three minutes before it is running; on a device that has never run the engine, the first start takes five to nine minutes. ## Reading two runs against each other Inside one run, the platform draws the comparison for you. The quality section puts every model side by side, and the performance section groups the same metric per model or per device, so a two-model run answers "which one, on this hardware" without any work from you. ![The performance section grouped by model, with a bar per model and device for the selected metric](/img/platform/first-benchmark/05-comparison.png) The grouping control switches between all results, by model and by device. By model reads naturally when you are choosing between models on one device; by device when you are choosing where to run one model. Across two separate runs, you read the two pages against each other, and two habits make that reliable. Performance metrics compare cleanly across devices, which is the point of measuring on the real hardware. Quality scores are safest compared within one device, because a test's score depends on the engine and the settings that device ran; when two devices disagree on a quality score, the per-item view is what turns the disagreement into an explanation, since you can compare the actual answers rather than the totals. ## What you have now A measured model, on real hardware, with the evidence attached. - Full detail: [Benchmark](../../concepts/benchmark.mdx) and [Read benchmark results](../../how-to/read-benchmark-results.md). - Next: [What to read next](../next-steps.md). --- # Run a benchmark Create a benchmark run: pick the devices and the models, then the quality tests, launch it, and watch the platform fan it out into jobs. Source: https://docs.clika.io/platform/getting-started/first-benchmark/run-a-benchmark.md Everything so far was setup. This part runs the benchmark. ## Start a run Open **Benchmarks** and select **New Benchmark**. ![The Benchmarks list with the New Benchmark action](/img/platform/first-benchmark/04-benchmarks-list.png) A run is built in two steps. The first step is the selection: the devices to run on and the models to run, side by side. The second step is the evaluation: the quality tests to run over every model and device pair, and how much of each test's dataset to cover. Nothing else is asked of you. Give the run a name while you are here: the pencil next to "Untitled benchmark" at the top. A named run is much easier to find a week later. ### Step 1: devices and models The selection step shows two columns, **Devices** on the left and **My Models** on the right, and takes them in either order. **Devices** lists the devices that can take a run right now, the most recently seen first, with the same hardware facts as the device page. A device that is offline or suspended is not offered. Search, the tag filter and the hardware filter narrow the list, **Select all** takes everything listed, and **Register a device** opens the enrollment dialog without leaving the flow; a device registered from there comes back selected once it is online. Where the deployment hosts phones in the cloud, a **Cloud hosted devices** group sits under your own fleet. **My Models** lists the project's registered models, with a type filter and a search box. **Add Model** registers a Hugging Face model without leaving the dialog, and it comes back selected. A run uses one model type: once a model is selected, models of another type grey out until the selection is cleared. Each column reads the other. With a device selected, the model list sorts by fit (fits every selected device, then some, then none) and a model that would be skipped somewhere says on which device and why, memory most often. With a model selected, a device row that cannot take every selected model says how many it would skip. Nothing is hidden for its fit; a pair that does not fit is skipped at launch rather than refused. Pick one device and one model for a first run. Two models make the result a comparison, which is what the results view is built for, but one is enough to see the loop. The footer counts the pairs (`models × devices = runs`) and says how many would be skipped. **Continue** moves to the second step once there is at least one device, one model and one pair that can run. ### Step 2: quality tests and dataset The evaluation step opens with a recap of the selection: the arithmetic of the pairs, how many launch and how many are skipped, and a **Show skipped** disclosure that names the skipped pairs grouped by model. Under it, the **Quality tests** card offers the tests the selected models' tasks can run, grouped by category, each a chip you toggle. A chip reads the test's name, an info button with its description, the license its dataset is used under and where the dataset comes from. A selected test that takes only part of the selection says so on the chip ("runs 3 of 4"), because a test accepts only the models whose task it declares compatible and the pairs that fit it. For a first run, **LLM Performance**, under **Performance**, is the fastest way to a result: it is the one test that measures speed (throughput, time to first token, latency and peak memory), and it needs no dataset download. An accuracy test such as **MMLU** or **ARC-Challenge** is a good second choice, and produces per-question data you can browse in the next part. See [Benchmark](../../concepts/benchmark.mdx#the-catalogue) for the whole catalog: what each test is, what it tests, and how it is scored. **Dataset** is one choice for the whole run: **Quick** covers 10% of each selected test's dataset, at most 100 samples; **Full** covers the complete datasets. Quick is the default and the right choice for a first run. Selecting several tests runs several tests, so a run over two models, two tests and two devices is eight jobs. The footer says how many runs will start and how many are skipped, and names any selected device that reports less memory than a test needs; the run still proceeds there. When the selection needs more than one benchmark group (one group carries one test), the footer names the groups it will create. Then select **Run Benchmarks**. **Back** returns to the selection with everything kept; closing the dialog asks before it discards the draft. ## Watch it The run appears at the top of **Benchmarks** with a live status. Open it while it is going and the page shows what is happening rather than making you wait for a verdict. ![A running benchmark: the leg's status, the device's live health, the finished count and a live log](/img/platform/first-benchmark/04-running.png) Each leg of the run is a row: the model, the device, and the device's own health while it works, so a run that is slow because the machine is saturated says so. Underneath, a progress line counts the finished legs, which for a one-leg run stays at 0 of 1 until the leg completes, and the live log streams what the benchmark is printing on the device: the log is where a run's progress shows, one line per measured cell. How long it takes is mostly the device's business. A small model answering a few hundred questions on a workstation is minutes. Each leg begins by sending the engine bundle to the device, which the log reports as `pushing_artifacts`: on a small machine that is the first four to five minutes of a quick LLM Performance run, which takes about fifteen minutes in all on a two-core laptop. Where your organization's plan meters credits, that transfer is what spends them, and a cancelled run that has started its transfer is charged too; the **Subscription** card under Settings shows the balance. ## What you have now A run with results attached. The next part reads them. - Full detail: [Benchmark](../../concepts/benchmark.mdx). - Next: [Read the results](06-read-the-results.md). --- # What to read next Where the Platform documentation goes after your first benchmark: concepts, how-to guides, the CLI and the MCP surface. Source: https://docs.clika.io/platform/getting-started/next-steps.md You ran a benchmark on real hardware and read the result. The rest of the documentation, in a useful reading order: - **[Concepts](../concepts/index.md)**: one page per object, for the moment you need the whole picture of one thing rather than the next step. [Benchmark](../concepts/benchmark.mdx) and [Job](../concepts/job.mdx) explain what happened during the tutorial; [Runtime licenses](../concepts/runtime-licenses.mdx) separates the two artifacts that both get called a license. - **[How-to guides](../how-to/index.md)**: problem-shaped recipes. [Write a job definition](../how-to/write-a-job-definition.md) is where you go to add a benchmark the deployment does not carry yet, and [Manage members and API keys](../how-to/manage-members-and-api-keys.md) is the one every team needs on day two. - **[Using the CLI](../cli/index.md)**: the platform from a terminal, from [the first command](../cli/get-started.md) to every `clika-cli` command with its arguments and flags. The CLI is how a benchmark run becomes a step in CI, since `clika-cli benchmarks watch` exits non-zero when a leg fails. - **[Using with AI (MCP)](../mcp/index.md)**: hands the platform to an AI assistant such as Claude Desktop, Claude Code or Codex, as a set of tools it can call. - **[ClikaRT](/clikart)**: the inference runtime the platform issues credentials for, and the library you link when you ship the model you measured. [Download the ClikaRT SDK](../how-to/download-the-clikart-sdk.md) is how a licensed deployment hands its members the runtime's release archives. - **[Modelverse](/modelverse)**: the CLIKA model library, for packaged models that run on ClikaRT as they are. Two things worth doing early with a real team: 1. **Name your runs.** The results view is built for comparison, and a run named for the question it answers is worth ten named "Untitled Benchmark". 2. **Give integrations their own API key.** A key carries a subset of its owner's permissions, so a CI job can dispatch benchmarks without holding an account's full access. [Manage members and API keys](../how-to/manage-members-and-api-keys.md) covers it. --- # The Platform at a glance What the CLIKA Platform is and is not, and the five-minute mental model behind every page of these docs. Source: https://docs.clika.io/platform/getting-started/overview.md The CLIKA Platform runs benchmarks on real devices and keeps those devices under management afterwards. You register the hardware you care about, point the platform at a model, and it measures the model on each device and brings the numbers back with the per-sample evidence behind them. It is not a training system, not a model hub, and not a hosted inference service. Your devices stay yours, the deployment runs on your infrastructure, and model weights are fetched on the device that needs them. ## The mental model Five ideas carry the whole product, in the order you meet them. 1. **Everything sits in an organization and a project.** The organization is your tenant: users, roles, and the device fleet. A project is one piece of work, and it holds the models, benchmark runs, services and runtime credentials of that work. The switchers in the top bar decide which of each you are looking at. Devices are the deliberate exception, shared across every project of the organization, because a physical machine is shared hardware. 2. **Devices come to you.** A device runs the CLIKA agent, and the agent dials out to the platform and keeps one connection open. Nothing connects inward, so a device behind a home router or a corporate firewall works with no ports opened. Over that one connection the platform pushes files, runs jobs, opens a terminal, and reads a heartbeat every 15 seconds. 3. **The platform issues the licenses your own code runs on.** A ClikaRT runtime has to prove it is allowed to run, and the credential it presents is one your platform mints, per project, and can update or revoke afterwards. Managing those credentials is one of the platform's main purposes rather than a side feature: you issue one, ship it with your application, and keep control of what it permits and how long it lasts. [Runtime licenses](../concepts/runtime-licenses.mdx) covers the entitlements a credential carries. 4. **A benchmark is models times tests times devices, and the platform picks the runner.** You choose models, quality tests and devices. The platform maps each combination to the job definition confirmed to work on that platform and architecture, dispatches one job per combination, and refuses honestly (`DEVICE_NOT_RUNNABLE`) where no confirmed mapping exists rather than substituting something unproven. 5. **A result carries its evidence.** Every completed job returns headline metrics, per-sample inputs and outputs, and metadata about what actually ran. That is what makes two numbers comparable: you can open the questions a model got wrong, not only its score. ## What talks to what ```text you the deployment your device ┌───────────┐ ┌────────────────────┐ ┌──────────────────┐ │ web app │ │ the platform API │ │ CLIKA agent │ │ CLI │ ─ HTTPS ─► │ job orchestration │ ── gRPC ────► │ pushes files │ │ Claude │ │ artifact storage │ (agent dials │ runs the job │ │ (over MCP)│ ◄── JSON ─ │ licensing │ out, stays │ reports health │ └───────────┘ └────────────────────┘ connected) │ returns output │ │ └──────────────────┘ ▼ results, metrics, samples ``` The web application never talks to a device. Every terminal keystroke, file transfer and job dispatch goes through the platform, over the agent's own connection. ## Three surfaces, one platform | Surface | Use it for | | --- | --- | | Web application | Everyday work: fleet, models, runs, results, members, licenses. | | `clika-cli` CLI | Terminals and CI. Resource YAML (`apply` and `export`), transfers, watching runs. | | Claude, over MCP | Asking for the work in plain language. The deployment serves its own operations as tools at `/api/mcp`, and the CLI can serve them locally over stdio. | ## Where to go next - [The tutorial](first-benchmark/01-overview.md) turns these five ideas into a benchmark result. - [Concepts](../concepts/index.md) expands each of them, one page at a time. --- # Quickstart: CLI Install clika-cli, sign in with an API key, register a device, run and read a benchmark, and deploy, from a terminal. Source: https://docs.clika.io/platform/getting-started/quickstarts/cli.md Time to first result: about 10 minutes to a signed-in CLI with a device online, then 15 minutes for the run. [Get started with the CLI](../../cli/get-started.md) explains the install and the login; this page is the short form of the whole loop. 1. **Install.** In the web application, **Settings**, then **Developer access**, prints the install one-liner (its download token lasts 15 minutes). The binary is `clika-cli`, installed to `~/.local/bin` without `sudo`; put that directory on your `PATH`. 2. **Sign in with an API key.** On the same page, **Create key** (the **worker** template is enough), copy it once, then save it to a profile: ```bash clika-cli --base-url https:// login --api-key ``` Email and password sign-in is browser-only. The key acts in the project that was open in the browser when you created it. 3. **Register a device.** Mint a token (the body needs a `name`), then run the install script for the device's operating system on the device. `install list --raw` prints the whole script, so on the device itself pipe it to `sh`: ```bash clika-cli enrollment-tokens create --body '{"name":"my-laptop","max_uses":1}' ``` ```bash clika-cli install list --os linux --token --raw | sh ``` ```bash clika-cli devices list ``` 4. **Register a model and run the quick benchmark.** `benchmark_type` is required for an LLM; the speed test is `llm_performance`, and `benchmark-types list` prints the others: ```bash clika-cli models create --body '{"huggingface_url":"HuggingFaceTB/SmolLM2-135M-Instruct"}' ``` ```bash clika-cli benchmark-groups create --body '{"name":"first-run","model_ids":[""],"device_ids":[""],"benchmark_type":"llm_performance","dataset_mode":"quick"}' ``` ```bash clika-cli benchmarks watch first-run ``` `watch` prints one line per state change: about five minutes in `pushing_artifacts`, then `running`, then `completed`. `jobs execution-log ` shows the device's progress in between. 5. **Read it.** ```bash clika-cli benchmarks results first-run ``` ```bash clika-cli jobs benchmark-results-summary -o yaml ``` 6. **Deploy.** Create (the body needs a `name`), then start; create alone leaves the serving `pending`: ```bash clika-cli servings create --body '{"name":"first-deployment","model_id":"","device_ids":[""]}' ``` ```bash clika-cli servings start ``` ```bash clika-cli servings key ``` `running` takes about two to three minutes here, because the benchmark already put the engine on the device; on a device that has never run the engine it takes five to nine minutes. Call the proxy address from `servings get` with `curl` and the key, as in the [web quickstart](web.md); the CLI's proxy command sends no body. 7. **Licenses** are read-only from the CLI (`projects licenses `); issue, rotate and revoke in the web application. --- # Quickstart: licensing Issue a runtime license in a project, download ClikaRT from the platform, and run a licensed prompt on your machine in six steps. Source: https://docs.clika.io/platform/getting-started/quickstarts/licensing.md Time to first licensed run: about 6 minutes, plus the download. [Runtime licenses](../../concepts/runtime-licenses.mdx) explains the credential and [Download the ClikaRT SDK](../../how-to/download-the-clikart-sdk.md) the archive; this page is the short form. 1. **Issue.** In the project, **Licenses**, **Issue License**, **Online** (the machine can reach the platform), keep the default entitlements, **Issue license**. Copy the `CLIKA1-` text from the result. Press the button once; a double click issues two. 2. **Download.** Profile menu, **Download ClikaRT**, your platform, **Command line only**, CUDA by your driver (none without an NVIDIA GPU), **Download**; or `clika-cli runtime-sdk download linux-amd64-cli ./clikart` with a key that carries `runtime_sdk:download`. 3. **Verify and extract.** Copy the SHA-256 from the dialog, check it, then extract: ```bash echo " " | sha256sum -c ``` ```bash tar -xf ``` A Python build is a zip: `unzip` it, then `pip install` the wheels. 4. **License the runtime.** For this shell: ```bash export CLIKA_RT_LICENSE='' ``` or once for your account, which prints the kind and the project: ```bash ./bin/clikart-license-init '' ``` 5. **Run.** ```bash ./bin/clikart-cli HuggingFaceTB/SmolLM2-135M-Instruct prompt "Hello" ``` A licensed run prints the answer and nothing about the license. A refused one names the cause (`no ClikaRT license is configured`, `LICENSE_STATE_BLOCKED` after a revocation, `ENTITLEMENT_DENIED` when the license does not permit this machine, `PHONE_HOME_STALLED` when an Online license cannot reach the platform); [License the runtime](/clikart/how-to/license-the-runtime) is the contract. 6. **Watch it stop.** Revoke the license on the Licenses page (type the project's name). A running `serve` is refused at its next check with the platform, within about five minutes, as `LICENSE_STATE_BLOCKED: runtime key revoked`; a short prompt that starts and ends between two checks can still complete. An Offline bundle keeps working until it expires. Limits you may meet: your plan caps the active Offline bundles across the organization (the refusal names the limit); every license's expiry is bounded by the plan's end; an API key cannot issue, rotate or revoke, so the CLI and an assistant read licenses and do not mint them. --- # Quickstart: MCP Give Claude Code the platform's benchmark tools with a scoped key, and run the first benchmark from a prompt. Source: https://docs.clika.io/platform/getting-started/quickstarts/mcp.md Time to first result: about 10 minutes to a connected assistant, then 15 minutes for the run. [Using with AI (MCP)](../../mcp/index.md) is the full guide; this page is the short form. 1. **Create a key for the assistant.** **Settings**, then **Developer access**, **Create key**. Start from the **Worker** template, which is what [Using with AI (MCP)](../../mcp/index.md) and the **Connect an AI agent** box recommend: it carries the model, benchmark, cancel and serving-invoke permissions an assistant needs. Before you press Create, open **Advanced** and untick `hosted_devices:write`, `cloud_instances:write` and `cloud_credentials:write`, which spend money, and `orgs:write` unless the assistant should manage members: the template ticks them for an owner. Pick the project the key acts in with the dialog's **Project** field; every request the key makes acts there, whatever project the header shows later. 2. **Connect.** The same page's **Connect an AI agent** box prints this `claude mcp add` line with your platform's address filled in: ```bash claude mcp add --transport http clika-rt-remote-full https:///api/mcp --header "Authorization: Bearer " ``` `/api/mcp` serves every tool your key can reach. The box's other line, `clika-rt-remote`, connects the narrower `benchmark-ops` toolset instead; `GET /api/mcp/toolsets` lists every toolset. 3. **Register the device.** Use the web dialog or the CLI's two commands ([Quickstart: CLI](cli.md)). An assistant can do it too when its key holds `devices:write`, which only an organization owner's or admin's key carries: on `device-ops` or `/api/mcp`, `post_enrollment_tokens` mints a token and `get_install` with it returns the one-line install command to run on the machine. 4. **Ask.** Once connected: > Register HuggingFaceTB/SmolLM2-135M-Instruct, check which benchmark types it can run on my online Linux device, start a quick llm_performance run on it, poll once a minute until it finishes, and summarize tokens per second and time to first token. Start the first run a minute after the device enrolled; a run started earlier may be refused with `LICENSE_RECOVERY`. While it runs, `get_jobs_id` carries `run_progress` (cells done of the total for a performance run, items done for an accuracy run); the run takes about 15 minutes. 5. **What stays yours to do.** Issuing a runtime license, creating a key, and inviting or removing organization members are not tools; do them in the web application. A refusal says why: a tool outside the key's scope `requires the capability`, one that needs a signed-in session `takes a user session`, and `unknown tool` means no tool has that name. --- # Quickstart: web From a fresh account to a benchmark result and a running endpoint in the browser, with the minutes each step takes. Source: https://docs.clika.io/platform/getting-started/quickstarts/web.md Time to first result on a two-core Linux laptop: about 25 minutes, 15 of them waiting for the run. The [tutorial](../first-benchmark/01-overview.md) explains each step; this page is the short form. 1. **Sign in.** From a signup code or an invitation ([Sign up with a signup code](../sign-up-with-a-signup-code.md)). The verification link lands you on the **Default Project** dashboard. 2. **Register a device** (2 minutes). **Devices**, **Register Device**, the tab for the machine's operating system, copy the command, run it on the machine. Leave the **Reusable, many devices** switch alone for now: turning it on asks whether to keep the command you copied or make a new one, and a new one revokes the one you copied. The dialog turns into the device list when the device arrives, within a minute. Give the device one more minute before step 4. 3. **Add a model** (1 minute). **Models**, **Add Model**, paste a Hugging Face URL or `owner/repo`; `HuggingFaceTB/SmolLM2-135M-Instruct` is a good first model. A model somebody registered before you comes back as a shared record under your models; your results on it are yours. 4. **Run a benchmark** (15 minutes). **Benchmarks**, **New Benchmark**, pick the device and the model, **Continue**, then **LLM Performance** with **Quick** on, **Run Benchmarks**. The log shows the engine push for the first five minutes and then one line per measured cell; the progress bar stays at 0 of 1 until the leg completes. 5. **Read the result** (1 minute). The run's page: TTFT, Tokens/sec, Peak decode, TPOT and Peak Memory, each with median, minimum and maximum, and a table of cells (`128:128` is prompt tokens : output tokens, `c8` is eight parallel streams). **Share Results** makes a link. 6. **Deploy it** (2 to 3 minutes after a benchmark on the same device; 5 to 9 on a device that has never run the engine). **Deploy** on the result, keep the name, **Deploy**. The page reads "Starting" and "Waiting for logs" while the engine starts on the device, then **Running** with the platform address, a hidden key (**Reveal key**) and the device's own address. Send one request: ```bash curl -s -X POST "https:///api/v1/servings//devices//proxy/v1/chat/completions" -H "Authorization: Bearer " -H "Content-Type: application/json" -d '{"model":"","messages":[{"role":"user","content":"Hello"}]}' ``` 7. **Invite your team.** **Settings**, then **Organization**, **Add member**: the organization role is Admin or Member; give project roles (Developer, Viewer) on the project's **Members** tab. Only an owner or admin can open a device's terminal. [Roles](../../concepts/roles.md) has the table. What spends credits, where your plan meters them: every benchmark leg and every deployment start sends the engine to the device, and the **Subscription** card under Settings shows the balance. --- # Sign up with a signup code Create an organization and its owner account from a signup link or code, verify the address, and know what the plan behind the code allows. Source: https://docs.clika.io/platform/getting-started/sign-up-with-a-signup-code.md Public self-service signup is closed during the alpha. An account is created by an invitation from an organization that already exists, or by a **signup code**, which an event or a partner hands you as a link. The code creates a new organization with you as its owner, on the plan the code names. ## What the link shows The signup page opens on a card naming the code, the date the code closes and the date the plan ends, then an email field. What the plan allows is on **Settings**, then **Subscription**, once you are in. A plan is one of two kinds. On a usage-based plan (the plan name wears a **Usage-based** badge) no device, benchmark-run or artifact limit applies: the page shows the credit balance and the spend, the data transfer and storage paid from credits, the offline-license count and the output retention. On a cap-based plan the page shows each limit beside its usage. ## The five steps The wizard has five steps, and the account exists only after the last one: 1. Email, then a password typed twice. The button reads **Create account**, but nothing is created yet. 2. The terms and privacy consent boxes (tick the boxes themselves; the labels open the documents). 3. Your name. 4. Your organization's name. 5. Your role. Nothing is sent until you finish step 5, so "Check your email" is not true before it. Then verify: the mail arrives within a minute, and its link signs you in and lands on the organization's **Default Project** dashboard. ## If the page sends you to sign in instead An address the platform already knows goes to the password step. Three cases: - **You already have an account.** Sign in; the code is for new accounts. - **You were invited by an organization and never accepted.** Use the invitation link from that mail, or ask the inviter to send it again. An address in that state cannot use a signup code; for a new organization, use another address. - **You typed a wrong password.** The page asks for an authenticator code even when the account has none; read the small print, go back, and use **Forgot your password?**. Repeated wrong attempts lock the account for about ten minutes. An unregistered address at the ordinary sign-in page is told to contact support, with a **Have a signup code?** link back to this flow. ## What you have now An organization on the code's plan, with you as owner, and an empty default project. Invite your team under **Settings**, then **Organization** ([Roles](../concepts/roles.md) says what each role may do), then [add your first device](first-benchmark/02-add-your-first-device.md). --- # How-to guides Problem-oriented recipes for the CLIKA Platform, grouped by what you are trying to get done. Source: https://docs.clika.io/platform/how-to.md Practical guides, each answering one concrete "how do I X". They assume the ground the [tutorial](../getting-started/first-benchmark/01-overview.md) covers (a project, a device, a model, a run) and go deeper on one problem at a time, in any order. Entries marked as coming are planned and land here as they are written. ## Benchmarks and results - [Read benchmark results](read-benchmark-results.md): the metrics a run returns, the per-sample evidence behind a score, and how to pull both through the API and the CLI. - Compare two runs (coming): pick a baseline, read a regression, and share the comparison. ## Definitions - [Write a job definition](write-a-job-definition.md): the `config.yaml` shape, every key explained, a complete example, and how to apply it to a deployment. - [Write a service definition](write-a-service-definition.md): the same treatment for a process the platform keeps running on a device. ## Automation - [Using with AI (MCP)](../mcp/index.md): give Claude Desktop, Claude Code, Codex or another assistant the platform as a set of tools. - Run benchmarks from CI (coming): an API key, `clika-cli apply`, and `clika-cli benchmarks watch` as a gate. ## Administration - [Manage members and API keys](manage-members-and-api-keys.md): who is in the organization, who is in the project, and how an integration gets its own scoped credential. - [Register a device by code](register-a-device-by-code.md): enroll a device with a short activation code typed into the agent, when pasting the token command is not practical. - [Activate a deployment license](activate-a-deployment-license.md): activate a platform license, renew it, and issue the runtime credentials a ClikaRT project needs. - [Update an on-premise platform](update-an-on-premise-platform.md): learn that a release is published, apply it with the installer, and return to the previous version from the backup the upgrade takes. ## ClikaRT on your own machine - [Download the ClikaRT SDK](download-the-clikart-sdk.md): get the runtime, the `clikart-cli` model CLI and the C++ SDK for your workstation or target device from the deployment, through the web dialog, `clika-cli runtime-sdk` or MCP, and verify the archive. For the vocabulary these guides use, see [Concepts](../concepts/index.md). For the commands, see [Using the CLI](../cli/index.md). --- # Activate a deployment license Activate a platform license on a fresh deployment, keep it renewed, and issue the runtime credentials a ClikaRT project needs. Source: https://docs.clika.io/platform/how-to/activate-a-deployment-license.md Two operations use the word license, and this guide covers both in the order you meet them: **activating the deployment** with the platform license CLIKA issued for it, and **issuing runtime credentials** to the ClikaRT runtimes your projects ship. [Runtime licenses](../concepts/runtime-licenses.mdx) explains why they are separate artifacts. ## Activate the deployment A fresh deployment refuses ordinary API requests and shows an activation screen instead of the dashboard. Activating it takes the `CLIKA1-...` platform license key CLIKA issued for that specific deployment. Check the current state first. This route is public, so it answers before activation and from any client: ```bash curl https://platform.clika.io/api/license/state ``` Then paste the key into the activation screen, or submit it directly: ```bash curl -X POST https://platform.clika.io/api/license/activate \ -H "Content-Type: application/json" \ -d '{"license_key": "CLIKA1-..."}' ``` The platform verifies the whole artifact before it stores anything: the signature over the exact bundle bytes against the CLIKA root certificate, the certificate chain inside it, the key material, the entitlement grant, and the signed policy that binds this deployment, its expiry and its entitlement ceiling. A key issued for a different deployment, or for a different product, is refused, and nothing is written. Activation is rate limited, so a paste error costs you a moment rather than a lockout. ### Restart when the response asks for it If the deployment issues device certificates, the response reports that a restart is required. Restart every replica before issuing or renewing device credentials. The certificate authority the platform serves from is the one activation replaced, and a running process is still holding the previous one. After a restart the platform verifies the stored license again on startup and loads its key material from it. A deployment that cannot verify its own stored license refuses to start rather than falling back to unrelated key material, which is the behavior you want: an activated deployment issues credentials under exactly one authority. ## Keep it renewed A connected deployment reports to the CLIKA License Portal on its own schedule and receives a signed verdict in return. An extension bought or granted in the Portal reaches the deployment through that channel, so a renewal needs no re-activation and no downtime: the deployment picks up the new expiry on its next report. The same channel carries the list of published releases. A connected deployment asks it once a day, and the administration dashboard shows a banner when a newer version exists; [Update an on-premise platform](update-an-on-premise-platform.md) covers the check and the upgrade. An expired license moves the deployment into a grace state rather than switching it off at the instant it expires. Treat the grace window as the time to fix the renewal, not as extra runway. An air-gapped deployment reports to nothing, so renewing it means receiving a reissued bundle and activating it exactly as you activated the first one. ## Issue a runtime credential A ClikaRT runtime proves it may run by presenting a credential your platform issued for the project it belongs to. Open **Licenses** in the project. The banner shows the platform's own state (active, the device count, the expiry), the **Licenses** tab lists the credentials issued for this project, and **Issue License** mints one. Choose the kind by how the runtime authenticates. | Kind | Choose it when | Notes | | --- | --- | --- | | **Online** | The runtime can reach the platform | A `CLIKA1-...` license key presented at start and at intervals. Revocation takes effect at the runtime's next check with the platform, within about five minutes for a running server. A project can hold as many as it has runtimes. | | **Offline** | The runtime cannot reach anything | A signed `CLIKA1-...` license bundle the runtime verifies locally. Your organization's plan caps the active bundles across its projects. | Either kind hands you the same thing to distribute, the `CLIKA1-...` text the license shows as its license key; a bare `clika_rk_...` API key is not a credential. [License the runtime](/clikart/how-to/license-the-runtime) is the placement contract for both products. The list gives each credential a name, kind, status, issue and expiry dates, a key prefix, and when it was last seen, which is the fastest way to find a credential nothing is using any more. **Prefer Online unless the runtime genuinely cannot reach the platform.** The intuition that an offline bundle is the safer artifact is backwards: a runtime holding one keeps verifying it successfully until it expires, because there is no channel to tell it the credential was revoked. If revocation has to actually stop a running runtime, it has to be Online. Two behaviors follow from an offline bundle being signed material, and both are refusals rather than gaps: - **A bundle past your plan's cap is refused**, with `OFFLINE_LICENSE_LIMIT_REACHED` naming the limit and the count. Rotate an existing one instead, which replaces it atomically so the project is never without one. - **An offline bundle's entitlements cannot be edited.** They are inside the signature. Rotating re-signs, which is the only honest way to change them. The name is editable precisely because it is not signed. **Rotate** replaces a credential, **Revoke** ends it, and the **Profiles** tab holds reusable entitlement templates so a team issuing many credentials does not re-enter the same grant each time. ## Troubleshooting **The API answers with a licensing refusal after activation.** Restart the replicas if the activation response asked for it, then re-check `GET /api/license/state`. **Activation is refused with a signature or policy error.** The key does not belong to this deployment. Platform licenses are issued per deployment, so a key from another environment cannot be made to work here. **A runtime is refused after a credential was revoked, and another is not.** Revocation reaches an Online runtime at its next check with the platform, within about five minutes for a running server, and reaches an offline runtime only when its bundle expires. See the Online and Offline table above. ## Related pages - [Runtime licenses](../concepts/runtime-licenses.mdx): the two artifacts, and the ceiling relationship between them. - [Organization and project](../concepts/organization-and-project.mdx): what a runtime credential is issued against. - [License the runtime](/clikart/how-to/license-the-runtime): where a ClikaRT or Modelverse runtime reads the credential you issued here, per language and packaging. --- # Download the ClikaRT SDK Get the ClikaRT runtime, the clikart-cli model CLI and the C++ SDK for your own machine from a licensed platform deployment: the web dialog, the platform CLI and the MCP tools, and how to verify the archive. Source: https://docs.clika.io/platform/how-to/download-the-clikart-sdk.md A licensed deployment hands its members the ClikaRT release archives: the inference runtime, the `clikart-cli` model CLI and the C++ SDK headers and libraries, one archive per target platform, at the engine version the deployment's own benchmarks run with. The archive is CLIKA's release file, unmodified, so what you build and serve on your workstation behaves like what the platform measured. Every member of an organization on a deployment whose platform license is active may download; the capability is `runtime_sdk:download`, and every built-in organization role carries it. Each download is recorded in the audit log. The archive runs nothing on its own. The runtime starts only with the runtime license credential your project issues, which is why the download dialog shows the two side by side. [Runtime licenses](../concepts/runtime-licenses.mdx) explains the credential; this page covers the archive. ## Where the download is | Surface | Where | | --- | --- | | Web | Your name in the top bar, then **Download ClikaRT**. Or a project's **Overview**, where the Licenses card has **Download SDK**. Both open the same dialog; the Overview opens it on that project's license key. | | Platform CLI | `clika-cli runtime-sdk list` and `clika-cli runtime-sdk download [dest]`. | | MCP | The `get_runtime_sdk` and `post_runtime_sdk_dist_download_token` tools, in the `benchmark-ops` toolset and in `full`. | ## Pick the archive for where your program runs One archive is one target platform: the libraries in `lib/` and the model CLI in `bin/` are built for that operating system and CPU. The headers under `include/` are the same in every archive. | Archive | `dist` id | Runs on | Backends | | --- | --- | --- | --- | | Linux x86_64 | `linux-amd64` | Linux workstations and servers | CPU, CUDA, Vulkan | | Linux arm64 | `linux-arm64` | Linux on Arm, Jetson included | CPU, CUDA, Vulkan | | macOS Apple silicon | `darwin-arm64` | Apple silicon Macs | CPU, Metal | | Windows x86_64 | `windows-amd64` | Windows 10 and newer | CPU, Vulkan | | Android arm64 | `android-arm64` | Android devices; the CLI inside runs on the device, not on your workstation | CPU, Vulkan | The backends column is what each archive's own `dist.json` lists under `capabilities.backends`, and what the dialog and `clika-cli runtime-sdk list` print for the row. A release may ship an Android library add-on (`android-arm64-aar`) beside the Android archive; when it does, the dialog lists it indented under the Android row. A program cross-built for another platform takes that platform's archive to link against and ship with, plus your workstation's archive for the tools you run while developing: fetching and converting models, serving one locally. An Android app built on Linux takes Android arm64 and Linux x86_64. ## Download from the web dialog The dialog has three numbered sections, in the order a first run needs them. **1. Choose your target platform and build.** The dialog is a chooser: the platform first, then the language (Python, C and C++, command line only, Kotlin, or the examples), then the CUDA image where the platform offers one (CPU and Vulkan only, CUDA 12, or CUDA 13; the newest CUDA image is proposed on Linux, and CPU and Vulkan only is the choice without an NVIDIA GPU), then the Python version for a Python selection. Every choice ends on one of the release's own files, as the runtime vendor published it: the platform's whole archive for C and C++ and for the command line, one wheel for a Python version (with its CUDA variant where one exists), the Maven zip or the archive for Kotlin, the Android library, or the examples archive. Nothing is assembled on the platform. The row shows the file name, the backends inside, the size and the vendor's SHA-256 with a Copy control. The runtime's languages are C++ and Python (the full API) and Kotlin (the inference API); [Language bindings](/clikart/bindings) says what each carries. Download asks the platform for a short-lived link and then hands that link to the browser, so the file lands in your downloads folder like any other. The link is good for ten minutes and names one file; the audit record is written when the link is minted, not when the bytes move. The organization's credits are charged once per link, at the file's whole size and the plan's rate per GB, when the link is first followed: an interrupted download costs the whole file, and a resume or retry inside the link's ten minutes costs nothing more. **2. Your license key.** The newest active online license key of the project in context, with its name and key prefix. From the Overview that is the project you are looking at; from the top-bar menu it is your current project, with a project picker when you belong to several. **Reveal key** shows the key and offers a copy button; revealing takes the license-manager role in the project, and a member without it reads a hint instead and asks a license manager for the key. A project with no active online key says so, and offers **Issue License** to a member who may issue one; it opens the Licenses page with its issue dialog. Under the key, two lines hand it to the runtime, with the key filled in once revealed: ```bash export CLIKA_RT_LICENSE= ``` ```bash ./bin/clikart-license-init ``` The variable is for one shell; the init command stores the key under your user account so every later run finds it. **3. Start.** The first commands for the archive you downloaded (or, before any download, for the platform the page is open on): unpack it, `cd` into it, list the devices the runtime sees, serve a model on the CPU. Under them, the CMake line to build against the runtime and the platform-CLI command that fetches the same archive. ```bash tar -xJf ClikaRT_linux_x86_64-.tar.xz ``` ```bash cd ClikaRT_linux_x86_64- && ./bin/clikart-cli devices ``` ```bash ./bin/clikart-cli serve --device cpu --port 8129 ``` The first line follows the file: `tar -xJf` for a `.tar.xz` archive, `unzip` for the Windows `.zip` archive, and `pip install` for a Python wheel (the dialog's start block creates a virtual environment first). The Kotlin Maven zip is added to a build, not unpacked by hand. The release's `README-Modelverse.md` inside an archive describes what the archive carries. To build against the runtime from CMake (C++17 or newer): ```cmake find_package(ClikaRT CONFIG REQUIRED) ``` and link `ClikaRT::ClikaRT`. The archive's own `README.md` covers the build in detail, and its `examples/` directory holds a walkthrough per topic. [ClikaRT](/clikart) is the runtime's documentation. ## Download with the platform CLI `runtime-sdk` is a hand-written group on `clika-cli`, beside the generated resource groups. The API key it runs under needs `runtime_sdk:download`. ```bash clika-cli runtime-sdk list ``` One row per file: the `dist` id that `download` takes (the platform's id for the whole archive, as in `linux-amd64`; a language and Python-version suffix for a wheel, as in `linux-amd64-python-cp312`, with `-cu12` or `-cu13` for its CUDA variant; `android-arm64-aar` for the Android library; `examples` for the examples archive), the platform and architecture, the file kind, the backends inside, the size, and, in the JSON output, whether the file is staged on this deployment (`staged`; the table omits the column). In table mode the engine version prints as a notice line on standard error, so a piped table stays a table; `-o json` and `-o yaml` print the whole listing with the version inside it. ```bash clika-cli runtime-sdk download linux-amd64 ./ ``` `download [dest]` asks the platform for the download link, streams the archive to disk while hashing it, and compares the digest with the one the platform announced. `dest` defaults to the current directory; an existing directory keeps the release file name inside it, and any other path is the file to write. On success the command prints the path written and the verified SHA-256. On a mismatch it removes the file and exits non-zero naming both digests, so a corrupt archive is never left where the next step could unpack it. Two refusals have their own advice. A `dist` the deployment offers no archive for points you at `runtime-sdk list`; an archive that is listed but not yet staged says so and names your platform administrator as the person to ask. ## Download through MCP The two tools are the listing and the link mint. `get_runtime_sdk` answers the archives with their `dist` ids, file names, sizes, digests and backends; `post_runtime_sdk_dist_download_token` takes a `dist` and answers a link with the file name, size and SHA-256 to verify against. The `url` is a path on the platform (prefix the platform's address), and the link works only for the account or API key that minted it: send the same `Authorization` header the mint used, or the platform answers `403 RUNTIME_SDK_TOKEN_INVALID`. It honors `Range` requests so an interrupted download resumes, and stops working ten minutes after it was minted; mint another when it has expired. An assistant hands the link and the header to a download tool rather than reading the bytes itself. [Using with AI (MCP)](../mcp/index.md) covers connecting an assistant. ## Verify what you downloaded The dialog shows the file's SHA-256 beside its download button, with the line that verifies the file on your platform; the CLI and the MCP listing carry the same digest. Linux: ```bash echo " " | sha256sum -c ``` macOS: ```bash echo " " | shasum -a 256 -c ``` Windows (PowerShell): ```powershell (Get-FileHash ).Hash -eq "" ``` Every file's digest is the runtime vendor's own, the one published beside the file, because the platform serves the release's files as they are. The platform CLI runs this check itself and refuses a file that does not match. ## Troubleshooting **A row is listed with no size and a disabled Download.** The deployment's installer did not stage that archive. The listing comes from the installed benchmark data bundle, which names every archive of the release; the bytes are staged separately. Ask your platform administrator. **The link answers that it is invalid or has expired.** Links live ten minutes. Mint a new one from the dialog, the CLI or the tool. **The link is refused after the deployment was updated.** A link delivers exactly the bytes it was minted for. When the deployment has moved to another build in between, the old link is refused; list the archives again and mint a new one. **The model commands fail with `LICENSE_FAILED`.** The runtime found no valid runtime license credential. Set `CLIKA_RT_LICENSE` or run `clikart-license-init` with the key from section 2 of the dialog, or from the project's Licenses page. [License the runtime](/clikart/how-to/license-the-runtime) covers every placement. The message names the cause, and every one ends by naming `CLIKA_RT_LICENSE` and the init tool even when the variable is set: | Message | Meaning | | --- | --- | | `no ClikaRT license is configured for this process` | Neither the variable nor an installed license was found. | | `LICENSE_STATE_BLOCKED: runtime key revoked` | The license was revoked or rotated on the platform; copy the current one from the project's Licenses page. | | `ENTITLEMENT_DENIED: this license does not permit running on a virtual machine` | The license's entitlements exclude this machine; issue one with the grant. | | `PHONE_HOME_STALLED: cannot reach the license server` | An Online license could not reach the platform; check the network, or use an Offline license. | A licensed run prints nothing about its license. `clikart-license-init` prints the kind of credential it installed, and for an Offline bundle the project and the expiry with it; an Online key is validated by the platform on its first use, so the tool prints no project for one. ## Related pages - [Runtime licenses](../concepts/runtime-licenses.mdx): the credential the runtime needs, and the two kinds a project issues. - [Activate a deployment license](activate-a-deployment-license.md): issuing that credential from the project's Licenses page. - [Manage members and API keys](manage-members-and-api-keys.md): the `runtime_sdk:download` capability an integration's key needs. - [ClikaRT](/clikart) and [Modelverse](/modelverse): what is inside the archive. --- # Manage members and API keys Add people to an organization and a project, give them the right role, and mint a scoped API key for an integration. Source: https://docs.clika.io/platform/how-to/manage-members-and-api-keys.md Two kinds of access exist side by side: **people**, who sign in and hold roles, and **keys**, which authenticate an integration as the person who created them. This guide covers both, and the one rule that connects them. ## Add someone to the organization Open **Settings**, then **Organization**. The members list is there, along with your own role and the organization's identity-provider settings. **Add member** invites someone by email address. They receive a single-use invitation link, and following it activates their account. If the invitation is lost or expires, re-inviting the same person issues a fresh one and invalidates the old. An organization member holds one of three roles: `owner`, `admin` or `member`. Ownership transfers explicitly (**Make owner** in a member's row menu), and an organization always has one. Owners and admins create projects and may remove any of the organization's cloud credentials. Every member reads the device fleet by virtue of membership. A `member` is not read-only: in the projects they belong to, members register models, start and cancel their own benchmarks and deploy. Read-only access is a project role, `viewer`. [Roles](../concepts/roles.md) is the full table. Removing a member ends their sessions, API keys, and project and team memberships in that organization at once; demoting one narrows their project access the same way. An API key stops working while its owner's account is suspended or deleted, and is deleted with the account. ## Add someone to a project Organization membership does not grant access to a project's work. Open the project's dashboard and use the **Members** tab, then **Manage Members**. A project member holds one of four roles. | Role | Give it to | | --- | --- | | `owner` | Whoever is accountable for the project. The last owner cannot be removed or demoted. | | `admin` | Someone who manages the project's members and resources. | | `developer` | Someone who registers models and runs benchmarks. | | `viewer` | Someone who reads results. | The distinction that catches people out: [devices are organization-scoped](../concepts/organization-and-project.mdx#devices-are-the-exception), so every project sees the same fleet. Everything else (models, runs, services, artifacts, licenses) belongs to one project and is invisible from the others. Project roles say nothing about devices. A project admin cannot mint an enrollment token or change a device unless the organization grants that capability separately. Creating a project is likewise an organization owner's or admin's action, not something a project role confers. ## Mint an API key Open **Settings**, then **Developer access**. **Create key** takes a name (`CI pipeline`, `nightly sweep`) and the scope the key should carry. The key is shown once, at creation. Copy it then; the platform stores only its hash and a short display prefix, so a lost key is replaced rather than recovered. A key also acts in the **project** that is open in the project switcher when you create it: benchmarks and deployments made through the key land there, and a deployment of another project answers `404` to it by id. Create the key with the right project open, and name it after the project. Use it as a bearer token: ```bash curl -H "Authorization: Bearer clika_your_api_key" \ https://platform.clika.io/api/devices ``` Or save it to a CLI profile, which prompts without echoing so nothing lands in your shell history: ```bash clika-cli --base-url https://platform.clika.io login --api-key ``` ### Scopes, and why a key can never exceed you A key authenticates as the user who created it, and its scope is an allowlist of capabilities. The permissions a request actually gets are **the owner's capabilities intersected with the key's scope**, recomputed on every request. Two consequences follow. - A key can never do something its owner cannot, and demoting the owner narrows every key they own immediately, with no revocation sweep to run. - A key can do considerably less than its owner, which is the point. A key scoped to reading devices and running benchmarks is a safe thing to put in CI, or to hand to an [MCP client](../mcp/index.md), because no prompt and no script can make it edit members or licenses. Membership itself is out of every key's reach, whatever the scope: see [What a key can never do](#what-a-key-can-never-do). An empty scope grants nothing, which is a useful thing to create deliberately. It is not the same as the unrestricted scope some older keys carry. ### The three templates The creation dialog offers three ready-made scopes, expanded server-side from what you yourself hold. | Template | What it grants | | --- | --- | | **viewer** | Every read and list capability you hold. A safe default for dashboards, reporting and anything that only observes. | | **worker** | The viewer set plus write, cancel and invoke. This is the CI shape: it can register a model, launch a benchmark, cancel one, and call a model deployment. | | **admin** | Every capability you hold that a key may carry at all, including delete. | No template ever includes the sharper verbs: `delete` outside the admin template, and never `exec`, `tunnel`, `remote_desktop`, the cross-team `read_all` and `manage_all`, the storage `test` probes, or `manage`. If a key needs one of those, grant it explicitly and know why. `runtime_sdk:download` sits in the Admin template only (the dialog marks it "In Admin"), although every organization role holds it: tick it under **Advanced** when the key will run `clika-cli runtime-sdk`. A key created by a member is capped to what that member holds: the dialog greys out the boxes for capabilities the member lacks, and its counter shows how many permissions the key will carry. ### The capabilities a key can carry Capabilities are named `resource:action`. What the scope catalogue offers you is bounded by what you hold, so your own list may be shorter than this one. | Resource | Capabilities | Holding them admits | | --- | --- | --- | | `devices` | `read`, `write`, `delete`, `exec`, `tunnel`, `remote_desktop`, `read_all`, `manage_all` | Reading and managing the fleet. `exec` runs arbitrary commands on a device, `tunnel` opens a TCP tunnel to one, `remote_desktop` opens an interactive screen and keyboard session, and the two `_all` forms are the cross-team administrative view. | | `jobs` | `read`, `write`, `cancel` | Reading, submitting and cancelling jobs, which is what launching a benchmark needs. | | `artifacts` | `read`, `write`, `delete` | Reading the library, uploading and updating artifacts, and deleting them. | | `model_serving` | `read`, `write`, `invoke` | Seeing model deployments, managing them, and calling one through the platform proxy. `invoke` is separate on purpose: seeing a deployment is not the same as spending its GPU. | | `services` | `read`, `write` | Managing the device-side services subsystem. | | `metrics` | `read`, `write` | Reading and writing metric definitions and data points. | | `alerts` | `read`, `write` | Reading and managing alert rules. | | `events` | `read` | Reading the organization's event stream: device, job, command, transfer and alert activity. | | `projects` | `create`, `read`, `write`, `delete` | Creating projects, which the organization's owners and admins hold, and managing projects and their membership. | | `orgs` | `read`, `write` | Reading and updating the organization record (its name and settings). `write` also invites and removes members, transfers ownership and deletes the organization, so give it to a key only when that key must manage the organization. | | `teams` | `read`, `write` | Creating and renaming teams. Team membership and deleting a team take a signed-in session. | | `webhooks` | `read`, `write` | Managing webhooks. | | `hosted_devices` | `read`, `write`, `delete` | Renting and running on hosted phones. `write` is the capability that spends money, since hosted devices bill by the minute. | | `cloud_credentials`, `cloud_instances` | `read`, `write`, `delete` | Managing stored cloud provider credentials and the instances created with them. | | `credentials` | `read`, `write`, `delete` | Managing container-registry credentials, which external artifacts use to pull. | | `file_storage` | `read`, `write`, `test` | Viewing and changing the artifact storage backend. `test` makes live outbound probes against the configured bucket. | | `license` | `read` | Reading the organization's plan and activation codes. | | `license_profiles` | `read`, `write` | Reading and curating entitlement profile templates. | | `project_licenses` | `read` | Reading a project's runtime credentials. | | `runtime_deployments` | `read` | Reading the runtime deployment plane. | | `runtime_sdk` | `download` | Listing the ClikaRT release archives the deployment offers and minting a download link for one. Every built-in organization role holds it. See [Download the ClikaRT SDK](download-the-clikart-sdk.md). | | `vpn` | `read`, `write`, `delete` | Managing VPN configuration. | ### What a key can never do Some capabilities are refused to every key, however you scope it, because a credential that holds them can widen itself and the scope would be decorative. They stay available to a signed-in session with the right role. - **Identity and access**: creating or editing users, roles, and SSO configuration. - **Platform administration**: the `/api/admin` surface, platform settings, and support access. - **Minting credentials**: issuing a runtime credential or writing the organization's license. Both create new authority, which is exactly what a leaked key must not be able to do. - **Billing**: re-planning the organization or starting a payment flow. A second group is refused to every key by route rather than by capability. These are the organization's identity writes: inviting or re-inviting a member, changing a member's role, removing a member, issuing a member's password reset link, transferring ownership, deleting the organization, adding or removing a team member, deleting a team, and approving or denying a support-access request. A key carrying `orgs:write` or `teams:write` still cannot call them; neither can a key created before scopes existed. They answer `API_KEY_ROUTE_SESSION_ONLY`, and an MCP client authenticated with a key is not offered them as tools. Sign in to do these. ## Troubleshooting **The browser works, my API key gets 403 on one route.** Read the error code. `API_KEY_SCOPE_INSUFFICIENT` names the capability the route needs and your key does not carry, which you fix by re-scoping the key. `API_KEY_ROUTE_SESSION_ONLY` means the route is one of the identity writes listed above: no key reaches it whatever its scope, so sign in and do it from the web application; the CLI signs in with an API key only. `API_KEY_ROUTE_NOT_SCOPEABLE` means the route declares no capability at all, so no key can reach it while a signed-in session still can; that is a gap in the deployment's route declarations rather than something you can fix from the key, so report the route to whoever runs the deployment. **A member cannot see a project's resources.** Check project membership rather than organization membership. A resource in a project you are not a member of answers `404`, deliberately, so that probing ids reveals nothing. **An invitation link does not work.** Links are single-use and expire. Re-inviting issues a new one and invalidates the previous. ## Related pages - [Organization and project](../concepts/organization-and-project.mdx): what each container holds. - [Using with AI (MCP)](../mcp/index.md): the main consumer of a scoped key. --- # Read benchmark results Pull a benchmark run's numbers and its per-sample evidence: what each metric means, the API routes behind the result view, and how to diagnose a leg that failed. Source: https://docs.clika.io/platform/how-to/read-benchmark-results.md A finished run carries more than a score. This guide covers what the numbers mean, how to get them out of the platform programmatically, and what to do with a leg that did not finish. Reading a run in the web application is covered by the tutorial: [Read the results](../getting-started/first-benchmark/06-read-the-results.md). ## What the metrics mean | Metric | What it measures | Direction | | --- | --- | --- | | `tokens_per_sec` | Decode throughput: tokens produced per second. | Higher is better. | | `ttft_ms` | Time to first token: the wait before output starts. | Lower is better. | | `latency_ms` | End to end time for one request. | Lower is better. | | `tpot_ms` | Time per output token, derived from latency, time to first token and the token count. | Lower is better. | | `peak_memory_mb` | The most memory the run held at once. | Lower is better. | | `rtf` | Real-time factor, for speech models: audio seconds processed per second. | Higher is better. | | A quality score | What the test defines: accuracy, pass rate, ROUGE, BLEU, word error rate. | Stated per test. | Where a metric was measured per sample, the summary also carries `_min`, `_median` and `_max`. A metric with no per-sample spread (peak memory, a single accuracy score) stays a point value, and the result view shows the same number for all three rather than inventing a range. Not every run carries every key. A run whose engine reported no time to first token omits that family instead of reporting a zero. For a performance benchmark the headline `tokens_per_sec` is the mean over the run's concurrency sweep (three prompt shapes at one, four and eight parallel streams); the per-combination rows are in the result's details, with the device's peak beside the single-stream figure. ## Get the results from the API Start from the run (the API calls it a benchmark group) and walk down to a job. ```bash # the run, with its legs curl -H "Authorization: Bearer $CLIKA_API_KEY" \ https://platform.clika.io/api/benchmark-groups/$GROUP_ID # the results of every leg curl -H "Authorization: Bearer $CLIKA_API_KEY" \ https://platform.clika.io/api/benchmark-groups/$GROUP_ID/results # one leg's jobs, when you need their ids curl -H "Authorization: Bearer $CLIKA_API_KEY" \ https://platform.clika.io/api/benchmark-groups/$GROUP_ID/jobs ``` With a job id, four routes go progressively deeper. | Route | What it returns | | --- | --- | | `GET /api/jobs/{id}/benchmark-results` | The whole result: summary, metadata, and what the runner recorded about the run. | | `GET /api/jobs/{id}/benchmark-results/summary` | The headline metrics alone. | | `GET /api/jobs/{id}/benchmark-results/samples` | The per-sample rows, paginated. | | `GET /api/jobs/{id}/benchmark-results/io` | The per-sample inputs and outputs, with the `io_schema` that says how to read them. | The `io` route is the one behind **View output data**: each row carries what the model was given, what it answered, what was expected, and whether it counted as correct. The `io_schema` names the shape (multiple choice, transcription, code generation, and others), so a client can render the fields rather than guessing at them. The raw file the script wrote is stored as an artifact and stays downloadable from the job, whether or not it parsed. ## From the CLI ```bash # watch a run to completion; exits non-zero if any leg ends other than completed clika-cli benchmarks watch "nightly-llm-sweep" # the finished run's results as a table, or as JSON for a script clika-cli benchmarks results "nightly-llm-sweep" clika-cli benchmarks results "nightly-llm-sweep" -o json ``` The non-zero exit on failure is what makes `watch` usable as a CI gate: dispatch the run, watch it, and let the step fail when a leg does. ## From Claude With the `benchmark-ops` toolset connected ([Using with AI (MCP)](../mcp/index.md)), the same reading is a question: > Read the results of my last benchmark run and tell me which model was fastest and whether any leg failed. That reaches `get_benchmark_groups`, `get_benchmark_groups_id_results` and `get_jobs_id_benchmark_results`. The per-item evidence is `get_jobs_id_benchmark_results_samples`, which is worth asking for by name when you want to know what a model got wrong rather than only its score. ## Diagnose a leg that did not finish A run's **Details** tab (and `GET /api/benchmark-groups/{id}/jobs`) gives each leg's state. What the state tells you: | State | What happened, and what to look at | | --- | --- | | `failed` | A step failed, or the declared output file was never written. The job's error detail names the failing step; for a script failure it carries the tail of what the script printed. | | `partial` | The benchmark itself succeeded and the results are usable, but teardown failed. The device is flagged dirty and takes no further benchmark jobs until it is cleared. | | `interrupted` | The device disconnected mid-run. Teardown is retried when it reconnects, and a late result can still complete the job. | | Refused at dispatch | `DEVICE_NOT_RUNNABLE` means no confirmed mapping exists for that test on that device's platform and architecture. `RESOURCE_MISMATCH` names the hardware requirement the device did not meet. | A job that failed with an output-missing flag is saying something specific: collection ran and found nothing at the path the definition declared. That is a job-definition problem (the script wrote elsewhere, or exited before writing) rather than a platform one. The other case is a job that completed carrying a parse diagnostic, which means something was written and it was not a result envelope. [Write a job definition](write-a-job-definition.md) covers both contracts. ## Share a run **Share Results** on the run's page mints a link that shows it to someone with no account on the deployment. There is one live link per run, and revoking it invalidates that link immediately. Use it for a result you want a customer or a colleague to read, and revoke it when the conversation is over. ## Related pages - [Benchmark](../concepts/benchmark.mdx): what a run is and what a result contains. - [Job](../concepts/job.mdx): the states above, in full. - [Write a job definition](write-a-job-definition.md): the envelope your own benchmark has to emit. --- # Register a device by code Enroll a device with a short activation code: the code the Register Device dialog's command carries, and how to mint one from the CLI or an MCP client. Source: https://docs.clika.io/platform/how-to/register-a-device-by-code.md The install command the Register Device dialog gives carries a short **activation code** (`?code=XXXX-XXXX` at the end of its address) rather than a long enrollment token. Because the code is a few characters, it can also be read out or typed on a device where pasting is awkward, such as a phone, a kiosk or a machine reached over a serial console. ## Mint a code Open **Devices**, then **Register Device**. The dialog mints a code when it opens and shows the install command that uses it. The code enrolls one device and expires in 30 minutes; the **Reusable, many devices** switch gives a command many devices can use instead. From the CLI: ```bash clika-cli enrollment-codes create --body '{}' ``` From an MCP client, `post_enrollment_codes` with an empty body does the same, and `get_install` with the code and the operating system returns the install command that uses it. ## Use it Run the dialog's command on the device, or install the agent and give it the code when it asks. The device appears in the list within a minute, like a token enrollment, and is then an ordinary device of the organization. A code can name an entitlement profile, so devices enrolled with it start with that profile's grants; see [Runtime licenses](../concepts/runtime-licenses.mdx#entitlement-profiles). ## Codes and tokens An enrollment token, minted with `clika-cli enrollment-tokens create`, enrolls a device the same way ([Add your first device](../getting-started/first-benchmark/02-add-your-first-device.md#from-the-cli) shows the two commands). A device enrolled with a code runs benchmarks and deployments like one enrolled with a token. ## Related pages - [Add your first device](../getting-started/first-benchmark/02-add-your-first-device.md): the dialog, platform by platform. - [Device](../concepts/device.mdx): what a device is and how it joins. - [Devices](../cli/devices.md): registering and driving devices from the CLI. --- # Update an on-premise platform Learn when a newer platform release is published, apply it with the installer, and return to the previous version from the backup the upgrade takes. Source: https://docs.clika.io/platform/how-to/update-an-on-premise-platform.md An on-premise deployment is upgraded by running the installer, either the release archive Clika sends you or the one the installer downloads from the update channel. Nothing is applied on its own: the platform tells you that a release exists, and you run the command on the platform node. This guide assumes a platform installed by `clika-install`, the installer inside the release archive, and covers the notice, the upgrade and the way back. ## Learn that a release exists A deployment whose license reports to the CLIKA License Portal asks the Portal's update channel once a day which release is published for it. When a newer version exists, the administration dashboard (**Admin**) shows one banner with the version, the date it was published, a link to its release notes and the exact command to run on the platform node. **Dismiss for this version** hides the banner in your browser until the next version is published. The banner applies nothing. An air-gapped deployment never asks the Portal. Its dashboard shows a note that updates are delivered as archives, and a release reaches it as a file, as below. The same check runs from the node, with the installer you installed with: ```bash sudo ./clika-platform--install.run --check-updates ``` It prints the installed version, the newest release on the channel, the one it would apply next (one release line at a time), whether that release is signed by a key the installer trusts, whether its archive is already downloaded, which archive it would fetch, and the command to apply it. Nothing is downloaded or changed. ## Apply a release From the channel: ```bash sudo ./clika-platform--install.run --upgrade ``` The installer downloads the release into `/var/lib/clika-platform/downloads` (the last two releases are kept there), checks Clika's signature and the archive's checksum before anything of it runs, shows what the upgrade does and how long it takes, asks once, and then runs the downloaded installer. A download interrupted by a dropped connection resumes where it stopped when you run the command again. From a file, copy the newer archive to the platform node and run it as you ran the first one: ```bash sudo ./clika-platform--install.run ``` Either way the installer recognises the existing installation, keeps every answer you gave, asks once for the license passphrase, shows what changes before it proceeds, backs up the database (below), replaces the platform's container images and waits for them to roll out. Your data, your license, your certificates and your administrator account are kept, and registered devices stay connected. The web application is unavailable for about a minute while the services restart. Afterwards, `--verify` re-checks the installation without changing it. ### Which archive applies A release is published as the full archive, `clika-platform--install.run`. When the ClikaRT engine and the benchmark catalogue did not change since the previous release, a platform-only patch, `clika-platform--patch.run`, is published beside it with the same installer, images and agent payload and without the benchmark catalogue. `--check-updates` names the one `--upgrade` will fetch, the smallest that applies to your node. A patch carried by hand refuses a fresh install and a node whose catalogue is not the one it expects, naming the full archive to apply instead. Both archives are signed the same way. ### Versions the installer refuses An archive older than the installed version is refused; going back is the rollback below, not an install. An archive more than one release line ahead of the installed version is refused as well, naming the version to install first. `--upgrade` steps one release line at a time and says so when the channel's newest release lies beyond the next step. ### The release notes Every archive carries its release notes as `RELEASE_NOTES.md`, the text the dashboard banner and `--check-updates` link, so they can be read on the node without a browser. The notes say what changed, what the upgrade does, how long it takes and what is visible while it runs. ### The benchmark catalogue Before it applies the catalogue, the installer checks the archive's benchmark payload against the platform's store. A missing required artifact stops the catalogue step and is named in the summary; the platform itself is already upgraded at that point, and re-running the installer once the artifact is in place completes the catalogue. A missing optional artifact is reported together with the benchmark definitions it affects, and the catalogue is applied without them. ## Return to the previous version Before the upgrade writes anything, the installer backs the database up under `/var/lib/clika-platform/backups/-/` on the platform node, together with a record of the object-store volume and of the installer state as it was. The last three backups are kept; `--backup-dir` and `--backup-keep` change the directory and the count. The upgrade refuses to start when the disk cannot hold the backup, naming the free space, the database size and the flag. Artifacts stay on their own volume across upgrades and rollbacks and are not part of the backup. The closing lines of an upgrade name the backup and the exact rollback command. To return to the version the upgrade replaced, run the same installer file with `--rollback`: ```bash sudo ./clika-platform--install.run --rollback ``` It brings back the previous version's images (kept on the node since the upgrade), shows what it will change and asks again before it touches the database. Restoring the database from the backup discards everything written since the upgrade, so that question takes a typed answer, and `--restore-database` answers it for an unattended run. When the upgrade changed the database schema the restore is required; otherwise the database is kept by default. `--bundle ` supplies the previous images when the node dropped them. Secrets and certificates are untouched. ## Where the installed version shows The upgrade banner names the installed version beside the one it offers, and the platform's health endpoints (`/api/health`, `/api/ready` and `/api/livez`) carry it as `version`, so monitoring can tell which release a deployment runs without reading its image tag. ## Related pages - [Activate a deployment license](activate-a-deployment-license.md): the platform license whose channel carries the release list. - [Events, metrics and alerts](../cli/monitoring.md#deployment-health): the health probes and the `version` they report. --- # Write a job definition The job definition YAML: every key explained, a complete worked example, the environment the platform injects, and how to apply the file to a deployment. Source: https://docs.clika.io/platform/how-to/write-a-job-definition.md A job definition is the template a benchmark [job](../concepts/job.mdx) executes on a device: what to push, what to install, what to run, where the result lands, and what to clean up afterwards. You write one when you are adding a benchmark a deployment does not carry yet. For an ordinary run you never touch one, because the platform picks the definition itself. This guide covers the file, the contract your script has to honor, and how to get the file into a deployment. ## The document A definition is a YAML document with `kind: JobDefinition`. Field names are snake_case, matching the JSON API, and several documents can share a file separated by `---`. ```yaml kind: JobDefinition name: "Inference Benchmark (llama.cpp, CPU)" description: "Throughput on a GGUF model through llama-bench, into the structured envelope." tags: category: "inference" artifacts: - artifact_id: "llamacpp-runtime" # name or lineage UUID dest_path: "/tmp/llamacpp-runtime" tag: "latest" # resolved to a version at dispatch setup_steps: - type: exec command: "python3 -m venv /tmp/llamacpp-runtime/.venv && /tmp/llamacpp-runtime/.venv/bin/pip install -r /tmp/llamacpp-runtime/requirements.txt" script: command: - "sh" - "-lc" - 'mkdir -p /tmp/lc-out && exec /tmp/llamacpp-runtime/.venv/bin/python /tmp/llamacpp-runtime/runner.py bench --model "${CLIKA_MODEL_HF_URL#https://huggingface.co/}" --out "${CLIKA_OUTPUT_PATH}"' working_dir: "/tmp/llamacpp-runtime" timeout_sec: 1800 env: CLIKA_BENCH_DEVICE: "cpu" HF_HOME: "/tmp/clika-model-cache" # keep the weights cache inside a path cleanup can reach output_path: "/tmp/lc-out/results.json" result_type: "structured" teardown_steps: [] required_resources: min_disk_bytes: 10737418240 # 10 GiB cleanup_policy: remove_output: false custom_paths: - "/tmp/clika-model-cache" ``` That is the whole shape. The committed definitions a deployment ships live under `products/clika-runtime-platform/job-defs//config.yaml` in the platform repository, one directory per model and device variation, and they are the reference to copy from. ## Every key | Key | Required | What it does | | --- | --- | --- | | `kind` | yes | Must be `JobDefinition`. | | `name` | yes | The definition's identity. Applying a document twice with the same name updates the existing definition rather than creating a second one. | | `description` | no | What this benchmark measures. It is what a reader sees in the catalog. | | `tags` | no | Key and value pairs for organizing the catalog. Tags do **not** decide which definition a run uses. | | `artifacts[]` | no | Files pushed to the device before anything runs. | | `artifacts[].artifact_id` | yes | The artifact by name or by lineage UUID. | | `artifacts[].dest_path` | yes | Where it lands on the device. A `directory` artifact is extracted here; a `file` artifact is written here verbatim. | | `artifacts[].tag` | no | Which version the tag resolves to at dispatch (`latest` by default). A tag that does not exist fails the dispatch. | | `artifacts[].credential_name` | no | The registry credential to authenticate an external (Docker or Git) artifact with. | | `setup_steps[]` | no | Steps run in order after the files land, before the script. | | `script` | yes | The benchmark itself. | | `script.command` | yes | The command as an argument array. A shell one-liner is `["sh", "-lc", "..."]`. | | `script.working_dir` | no | Working directory on the device. | | `script.env` | no | Environment variables, merged with the ones the platform injects. | | `script.timeout_sec` | no | Wall-clock limit. When it elapses the platform stops the whole process group and reports the timeout as itself, not as an exit code. `0` means no limit. | | `output_path` | yes | Where the script must write its result. | | `result_type` | no | `raw` keeps the file as an artifact. `structured` parses it, which is what unlocks metrics, charts and comparison. Default `raw`. | | `teardown_steps[]` | no | Steps that always run at the end, successful or not. | | `required_resources` | no | Minimum hardware, checked before dispatch. | | `cleanup_policy` | no | What is removed after teardown. | ### Steps A step is a `type` plus a command. `exec` runs a command and waits for it to exit. `start_service` starts a managed service and waits for it to report running. `wait_healthy` polls a URL until it answers. Two spellings of a command exist in the wild, and both work: the argument array (`["bash", "-c", "..."]`) and a single shell string. The definitions currently deployed use the single-string form for setup and teardown, so a file captured from a live deployment will look like that. Teardown steps run after the script whatever happened, including after a cancel or a timeout. If a teardown step fails, the device is flagged dirty and takes no further benchmark jobs until it is cleared, so keep teardown simple and idempotent. ### Required resources ```yaml required_resources: min_cpu_cores: 4 min_memory_bytes: 8589934592 min_gpu_vram_bytes: 8589934592 min_gpu_count: 1 min_disk_bytes: 10737418240 accelerators: ["cuda"] custom: jetpack_version: "5.1" ``` Every numeric minimum must be met, every listed accelerator must be present, and every custom pair must match what the device reported. A dispatch that fails validation returns `RESOURCE_MISMATCH` and names each unmet requirement (`gpu_vram: need 8.0 GB, have 4.0 GB`). Passing `force` on the dispatch bypasses the numeric checks, for the case where you know better than auto-detection. ### Cleanup policy ```yaml cleanup_policy: remove_artifacts: true # default remove_output: false # keep it false custom_paths: - "/tmp/clika-model-cache" ``` Cleanup runs in a fixed order after the script: your teardown steps, then the pushed artifacts, then `custom_paths`, then the output path. Three rules are worth following exactly. - **Keep `remove_output` false.** The platform collects the output after teardown, and a cleanup that removes it first loses the result silently. - **Never sweep a path that contains your `output_path`.** Same failure, arrived at from the other direction. - **Anything downloaded at run time needs a path cleanup can reach.** A model cache under the default `$HOME/.cache/huggingface` sits outside the agent's writable roots, so point `HF_HOME` at a path you also list in `custom_paths`. Definitions that push a large engine bundle usually set `remove_artifacts: false`. The agent skips a transfer whose checksum already matches on the device, and that only helps if the file is still there. ## The contract your script honors The platform injects environment variables at dispatch, on top of `script.env`. | Variable | What it carries | | --- | --- | | `CLIKA_OUTPUT_PATH` | The absolute path the result must be written to (`output_path`, with `${JOB_ID}` expanded). | | `CLIKA_JOB_ID` | This job's id. Echo it into the result's metadata. | | `CLIKA_MODEL_HF_URL` | The canonical Hugging Face URL of the model being benchmarked. | | `CLIKA_MODEL_SOURCE` | `huggingface` when the run named a Hugging Face model. | | `CLIKA_MODEL_NAME`, `CLIKA_MODEL_PATH`, `CLIKA_MODEL_ARTIFACT_ID` | Set where the model reaches the device another way. | | `CLIKA_HF_TOKEN` | A Hugging Face token, when the organization has one configured. Absent for public repositories. | | `CLIKA_SERVER_URL` | The platform's externally reachable URL. | | `CLIKA_AGENT_BIN` | The agent's own executable, for definitions that run a benchmark the agent carries. | Writing `${JOB_ID}` into `output_path` is worth doing: it expands platform-side, so two jobs on the same device never write to the same results file. With `result_type: structured`, the file at `output_path` must be the result envelope: ```json { "summary": { "tokens_per_sec": 242.0, "ttft_ms": 118.4, "latency_ms": 940.2 }, "samples": [ { "input": "...", "output": "...", "correct": true } ], "metadata": { "job_id": "...", "benchmark": "llm_performance", "device_placed": "cpu" } } ``` `summary` is what the result view charts, `samples` is the per-item evidence, and `metadata` records what actually ran. Where a metric was measured per sample, add its `_min`, `_median` and `_max` on the same base name and the result view renders them. Two collection outcomes are worth designing for. A declared `output_path` with nothing at it **fails** the job, even when the script exited zero, because it is a broken contract with the platform. A file that is not a valid envelope **completes** the job with a diagnostic, and the raw file stays attached, because the format was your script's choice and its exit code was its own verdict. ## Apply it The CLI reads the YAML and applies it by name, which is create-or-update, so a committed file applied twice does not create a duplicate. ```bash clika-cli apply -f config.yaml ``` Applying the same document twice is safe, which is what makes a committed definition file the source of truth rather than a snapshot of one. A definition that exists is not yet reachable from the **New Benchmark** flow. A benchmark run resolves its definition from a table the deployment owns, keyed by benchmark type, device platform and device architecture, so a platform administrator has to add the row that maps your definition to the device shapes it is confirmed to work on. Until that row exists, a run of that type against that device shape refuses with `DEVICE_NOT_RUNNABLE`, which is the honest answer: nobody has verified the combination yet. The mapping is managed in the admin area's benchmark catalog. ## Two footguns - **Windows definitions take a single-element command.** The agent runs a script through `cmd /c` on Windows, and the platform quotes a multi-element array for a POSIX shell, which `cmd` cannot parse. Write the whole command as one string in `cmd.exe` syntax, with `%VAR%` expansion. - **Keep large files out of the output directory.** Output collection has failed in practice because a multi-hundred-megabyte engine library sat in the same directory as the result file. Push the engine somewhere else and point `output_path` at a small directory of its own. ## Related pages - [Job](../concepts/job.mdx): the states a job moves through and what each guarantees. - [Artifact](../concepts/artifact.mdx): versions, tags and checksum deduplication. - [Write a service definition](write-a-service-definition.md): the sibling document for long-running processes. --- # Write a service definition The service definition YAML: every key explained, a worked example, how a definition becomes a running service, and what is not shipped yet. Source: https://docs.clika.io/platform/how-to/write-a-service-definition.md A service definition is the template for a process the platform keeps running on a device: an inference server, a collector, a helper daemon. Unlike a [job](../concepts/job.mdx), which runs once and produces a result, a [service](../concepts/service.mdx) is meant to stay up, and the platform restarts it when it does not. ## The document ```yaml kind: ServiceDefinition name: "inference-server" description: "Serves the packaged model over HTTP on the device." type: process # "process" (default) or "compose" command: - "/bin/sh" - "-lc" - "exec python3 /opt/app/serve.py --port 8000" env: MODEL_PATH: "/opt/models/latest.bin" auto_restart: true port: 8000 health_check: type: http url: "http://127.0.0.1:8000/health" interval_sec: 10 timeout_sec: 5 artifacts: - artifact_id: "model-weights" dest_path: "/opt/models/latest.bin" tag: "latest" required_resources: min_gpu_count: 1 min_gpu_vram_bytes: 4294967296 resource_limits: max_cpu_cores: 4 max_memory_bytes: 8589934592 cleanup_steps: [] ``` The platform's own model-serving template lives at `products/clika-runtime-platform/service-defs/clika-modelverse-serve/config.yaml` in the platform repository, and is the reference for a definition that is hydrated per use. ## Every key | Key | Required | What it does | | --- | --- | --- | | `kind` | yes | Must be `ServiceDefinition`. | | `name` | yes | The definition's identity. Applying the same name twice updates rather than duplicates. | | `description` | no | What the service is for. | | `type` | no | `process` (default) runs a command. `compose` runs a Docker Compose project on a device that has Docker. | | `command` | yes for `process` | The command as an argument array. | | `env` | no | Environment variables for the process. | | `auto_restart` | no | Restart the process when it exits. Default false. | | `port` | no | The port the service listens on. Recorded so the platform and the UI can reach it. | | `health_check` | no | How the platform decides the service is up: `type`, `url`, `interval_sec`, `timeout_sec`. The agent polls it to move the service from starting to running. | | `artifacts[]` | no | Files staged on the device before the process starts. Same shape as a job definition's, including `tag` and `credential_name`. | | `required_resources` | no | Minimum hardware, validated before the service is placed. Same fields as a job definition's. | | `resource_limits` | no | Bounds applied to the running process: `max_cpu_cores`, `max_memory_bytes`. Enforced through cgroups where the platform supports it, and reported back with an out-of-memory flag when the limit is what killed it. | | `cleanup_steps` | no | Steps run when the service is removed. | ## Starting one A definition is a template. Starting a service from it takes a second document: ```yaml kind: Service service_definition: "inference-server" # name or UUID device: "jetson-orin-02" # or device_ids: [a, b] service_id: "inference-server" env: MODEL_PATH: "/opt/models/other.bin" # overrides the definition ``` `clika-cli apply -f service.yaml` resolves the definition, merges your overrides over its defaults, and starts the service on each device you named. Fields you can override are the ones the start request carries: `command`, `working_dir`, `env`, `artifacts`, `auto_restart`, `restart_delay_sec`, `log_buffer_lines`, `port`, `type`, `resource_limits`. A device that is offline does not refuse the start. The service and its desired state are persisted, the service reads `pending`, and the platform starts it when the device comes back. ## What happens on the device The platform stores a **desired state** per device and reconciles it against what the device reports in its heartbeat. That is why a crashed service comes back, why a start survives an offline device, and why removing a service is a state change rather than a signal you have to time. Stopping a service stops its whole process group, so a wrapper script that forked workers takes them with it. The stop is graded on Unix (a term signal first, then a kill once the grace period is spent); on Windows the process group is terminated outright, because a service has no console to receive the polite signal. A service moves through `pending`, `starting`, `running`, and then `failed` or `stopped`. A stop whose cleanup could not finish reads `cleanup_failed`. ## Model serving is a service you do not have to write Running one of your registered models on your devices as an OpenAI-compatible endpoint needs no definition of your own. The **Deployments** page (and `POST /api/v1/servings`) takes a model, the devices, and whether it runs on CPU or GPU, and the platform builds the service for you, stages its runner, watches its health, and exposes an authenticated proxy per device while it is running. Write a definition of your own when you are running something that is not a served model. ## Not shipped yet - **Declarative desired state through the CLI (coming).** The platform stores desired state per device and the API can set it (`PUT /api/devices/{id}/desired-state/services/{svc_id}`), but `clika-cli apply` does not accept a `DesiredState` document today. The kinds it accepts are `JobDefinition`, `ServiceDefinition`, `Job`, `Service`, `Model` and `Benchmark`. - **Device selectors in an applied document (coming).** A `Service` document names its devices explicitly (`device` or `device_ids`). Selecting by tag, platform or GPU model is refused with a message saying so. ## Related pages - [Service](../concepts/service.mdx): desired state, servings and the two meanings of "deployment". - [Write a job definition](write-a-job-definition.md): the sibling document, for work that finishes. --- # Using with AI (MCP) Connect Claude or ChatGPT to the CLIKA Platform and ask about your devices and benchmarks in plain words: which page to follow for your app, the API key it needs, what it can and cannot do, and the endpoint details for advanced users. Source: https://docs.clika.io/platform/mcp.md Connect an AI assistant such as Claude or ChatGPT to the CLIKA Platform, and you can ask it about your devices and benchmarks in plain words, or have it start a benchmark and summarize the results for you. The connection uses MCP (Model Context Protocol), the standard way an AI assistant uses tools outside itself. You do not need to know how it works to use it. ## Pick your app Desktop apps, no terminal needed: - [Claude Desktop](claude-desktop.md): add an extension you download from the platform. - [ChatGPT Desktop](chatgpt-desktop.md): add the platform in ChatGPT Desktop's Settings. Developer tools: - [Claude Code](claude-code.md): one `clika-cli` command, or `claude mcp add`. - [Codex](codex.md): one `clika-cli` command, `codex mcp add`, or an entry in `config.toml`. - [Other clients](other-clients.md): any MCP client that connects to a URL. Each page takes you from nothing to a first answer, and ends with how to remove the connection again. ## What you can ask Once connected, ask the way you would ask a colleague: > Which of my Clika Platform devices are online? > List the devices that are online, and the benchmark types that can run against Qwen/Qwen2.5-0.5B-Instruct. Then launch an LLM performance run of that model on the fastest online Linux device, wait for it to finish, and summarize tokens per second and time to first token. The assistant does what the [tutorial](../getting-started/first-benchmark/05-run-a-benchmark.md) walks through by hand. A quick LLM Performance run of a small model on a two-core device takes about fifteen minutes, five of them sending the engine to the device, so ask the assistant to check on the run once a minute, and start the first run a minute after the device enrolled. ## What it can and cannot do The assistant can read your devices and their health, run a command on a device, move files to and from it, browse your models and artifacts, start a benchmark, follow it to the end and read its results. It can only do what its API key allows, and every request it makes is checked exactly like one from the web application. A key made from the **Worker** template can look at devices and run benchmarks, but cannot change members or licenses. Some things stay yours to do, in the web application: - Create an API key, or invite or remove a member of the organization. - Issue, rotate or revoke a runtime license. - Register a new device (**Devices**, then **Register Device**), unless the key holds `devices:write`: see [Where the tools stop](#where-the-tools-stop). ## You need an API key Every app connects with an API key. Each app's page walks you through making one; in short: 1. Sign in to the CLIKA Platform and open **Settings**, then **Developer access**. 2. Press **Create key**. Give it a name that says what will use it (`claude-desktop`, `chatgpt-desktop`), choose the **Worker** template under **Start from a template**, and press **Create**. 3. Copy the key. It is shown only once, and it stops working 90 days after you created it. [Manage members and API keys](../how-to/manage-members-and-api-keys.md#mint-an-api-key) covers the templates and permissions in full. ## For advanced users ### The MCP endpoint and its toolsets Each deployment serves MCP over Streamable HTTP at `/api/mcp`, where `` is the address you open in the browser. A toolset is a named selection of the same tools for one kind of work, served at its own path below the endpoint. | Endpoint | What it serves | | --- | --- | | `/api/mcp` | Every documented operation outside the platform administration API (`/api/admin/**`), which no MCP surface serves. | | `/api/mcp/api-key` | The operations an API key can be granted. Sign-in flows, the capabilities only a browser session can hold (users, roles, SSO, license and billing changes) and multipart-only uploads are served on `/api/mcp` only. | | `/api/mcp/device-ops` | Day-to-day device work: list and inspect devices, read health, run a command, transfer files, browse artifacts. | | `/api/mcp/benchmark-ops` | The benchmark loop end to end: list benchmark types, check compatibility, list and register models, pick devices, launch a run, poll it, read its results and samples. | | `GET /api/mcp/toolsets` | The toolsets this deployment serves, with a description of each. | A deployment serves a toolset only when its API carries every operation the toolset names. Ask a deployment what it serves: ```bash curl -H "Authorization: Bearer $CLIKA_API_KEY" https://platform.clika.io/api/mcp/toolsets ``` A toolset decides which tools are listed. It is not a permission boundary: every call is checked against the key whichever endpoint it arrives on, and each caller's tool list holds the tools its key can call. To limit what an assistant may do, scope its [API key](../how-to/manage-members-and-api-keys.md#scopes-and-why-a-key-can-never-exceed-you). MCP requests authenticate like REST requests, with `Authorization: Bearer `. Do not give an assistant your browser session token: it expires, and it carries all of your permissions. ### Where the tools stop - Registering a device takes `devices:write`, which only an organization owner's or admin's key carries. With it, `device-ops` and the full endpoint have the two tools: `post_enrollment_tokens` mints a token (its body needs a `name`), and `get_install` with that token and the operating system returns the one-line install command to run on the machine. `benchmark-ops` has neither. - `get_jobs_id_benchmark_results_samples` carries a sample's index, latency and status; its question and answer text is in `get_jobs_id_benchmark_results_io`. `benchmark-ops` serves both. A run with only LLM Performance has no samples, so both answer empty lists for it. - A tool call is sent as JSON, so an operation that takes only a `multipart/form-data` upload cannot be called as a tool; its tool description says so. A file still reaches the platform through the chunked upload session (`POST /api/artifacts/uploads` with the size and SHA-256, then one base64-encoded chunk per call), through an external source the platform fetches itself (`POST /api/artifacts/external`), or through a JSON sibling tool that references an artifact already uploaded. - A refusal names its reason. A tool outside the key's scope answers that it `requires the capability, which this credential does not hold`; one that takes a signed-in session answers that it `takes a user session`; one outside the toolset you connected names the endpoint that serves it. `unknown tool` means no tool has that name. ### The CLIKA CLI The CLIKA CLI serves the same tools over standard input and output with [`clika-cli mcp`](../cli/mcp.md), the `api-key` toolset unless `--toolset` names another one. `clika-cli mcp install` writes a client's configuration for you and keeps the API key out of the client's files: the entry names a CLIKA CLI profile, and the CLIKA CLI hands the key to the client when it connects, so a new key needs a new `clika-cli login` and no change to any client. [Get started with the CLI](../cli/get-started.md) installs the CLIKA CLI, and its installer can register the clients in the same step. ## Related pages - [Troubleshooting](troubleshooting.md): a refused key, a self-signed certificate, a client that does not show the tools. - [MCP server and generic dispatch](../cli/mcp.md): the `mcp` command, its flags, and `mcp install`. - [Manage members and API keys](../how-to/manage-members-and-api-keys.md): scoping the key an assistant uses. --- # ChatGPT Desktop Add the CLIKA Platform to ChatGPT Desktop from its Settings, check that it works, and remove it. Source: https://docs.clika.io/platform/mcp/chatgpt-desktop.md Follow this page and you can ask ChatGPT things like "which of my devices are online?" or "run a benchmark of this model", and ChatGPT answers from the CLIKA Platform. It takes about five minutes. You need ChatGPT Desktop installed on your computer and an account on the CLIKA Platform. ## Step 1. Get a key 1. Sign in to the CLIKA Platform (for example `platform.clika.io`) and open **Settings**, then **Developer access**. 2. Press **Create key**. - **Key name**: `chatgpt-desktop`. - **Start from a template**: **Worker**. ChatGPT can then look at your devices and also run benchmarks for you. - Leave everything else as it is, and press **Create**. ![The Create API key dialog: Key name chatgpt-desktop, Expiry Platform default, the Worker template selected and its permissions ticked, and Create](/img/platform/mcp/chatgpt-desktop/create-key.png) 3. Press the copy button next to the key and paste the key somewhere for a moment, for example in a note. It is shown only once: after you press **Done**, you cannot see it again. ![The API key created dialog: the new key, which starts with clika_, the copy button next to it, and Done](/img/platform/mcp/chatgpt-desktop/key-created.png) ## Step 2. Add the platform to ChatGPT Desktop 1. In ChatGPT Desktop, open **Settings**. On Windows, press `Ctrl+,`. On macOS, press `⌘,`. 2. Under **Integrations**, choose **Plugins**, then the **MCPs** tab. 3. Press **Add**, then **Add MCP server**. ![ChatGPT Settings, Plugins, MCPs: the Add menu open with Create plugin, Add a marketplace and Add MCP server](/img/platform/mcp/chatgpt-desktop/01-add-menu.png) 4. On **Connect to a custom MCP**, fill in: - **Name**: `clika-platform` - **Type**: press **Streamable HTTP**. - **URL**: `https://platform.clika.io/api/mcp`. If your company signs in to the platform at a different address, use that address followed by `/api/mcp`. - **Headers**: in **Key**, type `Authorization`. In **Value**, type `Bearer`, one space, and then paste the key you copied in step 1. Leave everything else empty, and press **Save**. ![The Connect to a custom MCP form: Name clika-platform, Type Streamable HTTP, the URL, an Authorization header whose value is Bearer and the API key, and Save](/img/platform/mcp/chatgpt-desktop/02-custom-mcp-form.png) 5. **clika-platform** now appears under **Servers**, turned on. ![The Servers list with clika-platform and its toggle on](/img/platform/mcp/chatgpt-desktop/03-server-added.png) ## Step 3. Try it Start a new chat and ask: > Which of my Clika Platform devices are online? The first time, ChatGPT asks for permission to use the platform. Press **Always allow**, or **Allow once** to be asked again next time. When ChatGPT answers with your devices, you are ready. ### If it does not work - Check that **clika-platform** is turned on under **Servers** (Step 2, 5). - Press the gear next to **clika-platform** and check the header's value: `Bearer`, one space, and the whole key, which starts with `clika_`. - A key stops working 90 days after you created it. Create a new one (Step 1), press the gear next to **clika-platform**, replace the key in the header's value, and press **Save**. - Quit ChatGPT Desktop completely and open it again. If it still does not work, see [Troubleshooting](troubleshooting.md). ## Remove it Open **Settings**, then **Plugins** and the **MCPs** tab, press the gear next to **clika-platform**, then **Uninstall**. ChatGPT Desktop removes it at once, without asking to confirm. ![The clika-platform server's page in ChatGPT Settings: Back, Update Clika-platform MCP and the Uninstall button](/img/platform/mcp/chatgpt-desktop/04-uninstall.png) If nothing else uses the key, revoke it on the platform under **Settings**, then **Developer access**. ## Good to know - A key stops working 90 days after you created it. Create a new one and replace it with the gear next to **clika-platform**. - ChatGPT Desktop keeps the key in a settings file on your computer, which anyone who can open your files can read. Use a key made for ChatGPT Desktop only, and revoke it when you stop using it. --- # Claude Code Connect Claude Code to the platform with clika-cli mcp install, with claude mcp add, or with an entry in ~/.claude.json, then check the connection and remove it. Source: https://docs.clika.io/platform/mcp/claude-code.md Connect Claude Code to the CLIKA Platform and you can ask it about your devices and benchmarks while you work. There are three ways to connect; pick one. | Way | Use it when | Where the API key lives | | --- | --- | --- | | [`clika-cli mcp install`](#register-with-the-clika-cli-recommended) (recommended) | You use the CLIKA CLI. | In your CLIKA CLI profile. Claude Code's files hold no key. | | [`claude mcp add`](#add-it-with-claude-mcp-add) | You do not use the CLIKA CLI. | In `~/.claude.json`, in plain text. | | [An entry by hand](#add-the-entry-by-hand) | You manage the config file yourself. | In your CLIKA CLI profile. | Each needs an [API key](index.md#you-need-an-api-key). ## Register with the CLIKA CLI (recommended) With the CLIKA CLI installed and logged in ([Get started with the CLI](../cli/get-started.md)), one command registers the platform: ```bash clika-cli mcp install claude-code ``` If a Claude Code session is running, the command says so: restart it to load the server. What the command does: - It first checks the profile's key with one request to the deployment. If the profile has no key or the deployment refuses it, the command offers to log in when it runs at a terminal, and prints the `login` command to run otherwise. - It adds an entry named `clika-platform` to the user-scope `mcpServers` of `~/.claude.json` (`$CLAUDE_CONFIG_DIR/.claude.json` when `CLAUDE_CONFIG_DIR` is set), so every project on the machine sees the server. - The entry is an HTTP connection to `/api/mcp`. Its `Authorization` header comes from a `headersHelper`: a command Claude Code runs for each connection, here `clika-cli --profile default mcp headers`, which prints the header from the CLIKA CLI profile. The API key never lands in `~/.claude.json`, and a new key needs a new `clika-cli login` and no change to the entry. Useful flags: | Flag | What it does | | --- | --- | | `--profile staging` | Registers another deployment's profile. | | `--name clika-staging` | Picks the entry's name. | | `--dry-run` | Prints the entry without writing it. | | `--insecure-tls` | For a deployment with a self-signed certificate: the entry runs the CLIKA CLI as a local bridge (`mcp --remote`) instead of connecting over HTTP. [Troubleshooting](troubleshooting.md#a-self-signed-certificate-on-premise) explains why. | The [CLI reference](../cli/mcp.md#mcp-install-uninstall-status-and-headers) lists every flag. ## Add it with `claude mcp add` Without the CLIKA CLI, register the endpoint with Claude Code's own command: ```bash claude mcp add --transport http clika-platform https://platform.clika.io/api/mcp --header "Authorization: Bearer clika_your_api_key" --scope user ``` - `--scope user` makes the server available in every project. - Claude Code stores the key in `~/.claude.json` as part of the entry, readable by anything that can read that file. A new key means removing the entry and adding it again. - On a deployment with a self-signed certificate, trust its certificate authority in Claude Code (`NODE_EXTRA_CA_CERTS`) or register with `clika-cli mcp install claude-code --insecure-tls` instead ([Troubleshooting](troubleshooting.md#a-self-signed-certificate-on-premise)). - The **Connect an AI agent (MCP)** card under **Settings**, then **Developer access**, shows this command with your deployment's address filled in, twice: under the server name `clika-rt-remote-full` for the full endpoint `/api/mcp`, and under `clika-rt-remote` for the `benchmark-ops` toolset (`/api/mcp/benchmark-ops`). The name is yours to choose. ## Add the entry by hand Log in with the CLIKA CLI first (`clika-cli --base-url https://platform.clika.io login`). Then write the same entry `mcp install` writes, in the top-level `mcpServers` of `~/.claude.json` for every project, or in `.mcp.json` at a project's root for that project only: ```json { "mcpServers": { "clika-platform": { "type": "http", "url": "https://platform.clika.io/api/mcp", "headersHelper": "/usr/local/bin/clika-cli --profile default mcp headers" } } } ``` - Claude Code runs `headersHelper` through a shell, so put double quotes around a CLI path that contains a space. - Edit `~/.claude.json` while no Claude Code session is running, since a running session may write the file back. - An entry you write by hand is yours: `mcp install` and `mcp uninstall` never change or remove it. If it uses the name `clika-platform`, `mcp install` refuses that name and asks for `--name`. ## Check the connection `clika-cli mcp status` reports, for each client, whether `mcp install` registered the server, under which name and profile, and which config file it read. Then start Claude Code and ask: > Which of my Clika Platform devices are online? The answer comes from the device list tool, `get_devices`. If the tools do not appear, see [Troubleshooting](troubleshooting.md). ## Remove it Remove the entry the way you added it: | Added with | Remove with | | --- | --- | | `clika-cli mcp install` | `clika-cli mcp uninstall claude-code`. It deletes the entries it added for the profile and leaves every other entry alone. | | `claude mcp add` | `claude mcp remove clika-platform --scope user` | | By hand | Delete the entry from the file. | If a Claude Code session is running, restart it. If no other client uses the API key, revoke it under **Settings**, then **Developer access**. --- # Claude Desktop Add the CLIKA Platform to Claude Desktop with the extension from the platform's Settings page, check that it works, and remove it. Source: https://docs.clika.io/platform/mcp/claude-desktop.md Follow this page and you can ask Claude things like "which of my devices are online?" or "run a benchmark of this model", and Claude answers from the CLIKA Platform. It takes about five minutes. You need Claude Desktop installed on your computer and an account on the CLIKA Platform. ## Step 1. Get a key and the extension 1. Sign in to the CLIKA Platform (for example `platform.clika.io`) and open **Settings**, then **Developer access**. 2. Press **Create key**. - **Key name**: `claude-desktop`. - **Start from a template**: **Worker**. Claude can then look at your devices and also run benchmarks for you. - Leave everything else as it is, and press **Create**. ![The Create API key dialog: Key name claude-desktop, Expiry Platform default, the Worker template selected and its permissions ticked, and Create](/img/platform/mcp/claude-desktop/create-key.png) 3. Press the copy button next to the key and paste the key somewhere for a moment, for example in a note. It is shown only once: after you press **Done**, you cannot see it again. ![The API key created dialog: the new key, which starts with clika_, the copy button next to it, and Done](/img/platform/mcp/claude-desktop/key-created.png) 4. On the same page, find the **Claude Desktop** card. Check that the selected tab matches your computer, and pick the right one if it does not: - A Mac: **macOS (Apple Silicon)**. Macs with an Intel processor are not supported. - A Windows PC: **Windows (x86_64)**. Pick **Windows (arm64)** only on a Windows laptop with a Snapdragon processor. - A Linux PC: **Linux (x86_64)**, or **Linux (arm64)** on an Arm computer. 5. Press the download button under the tabs, for example **Download for Windows (x86_64)**. If the button no longer works, press **Refresh Links** and try again. ## Step 2. Add the extension to Claude Desktop 1. In Claude Desktop, open **Settings**. On Windows and Linux, open the menu at the top left and choose **File**, then **Settings**, or press `Ctrl+,`. On macOS, press `⌘,`. 2. Choose **Extensions**. The page ends with the line **Drag .MCPB or .DXT files here to install**. ![Claude Desktop Settings, Extensions, with no extensions installed: the Advanced settings button and the line Drag .MCPB or .DXT files here to install](/img/platform/mcp/claude-desktop/01-extensions-page.png) 3. Drag the file you downloaded in step 1 (its name ends in `.mcpb`) from your Downloads folder onto that page. **If that doesn't work,** add the file this way instead: 1. On the same page, press **Advanced settings**. 2. Under **Extension developer**, press **Install extension** and choose the file you downloaded. ![Claude Desktop Extensions, Advanced settings: the Extension developer section with the Install extension button](/img/platform/mcp/claude-desktop/02-advanced-settings.png) 4. The page for **Clika Platform** opens. Press **Install**. ![The Clika Platform extension page in Claude Desktop with the Install button](/img/platform/mcp/claude-desktop/03-install-preview.png) 5. When asked **Do you want to install Clika Platform?**, press **Install**. 6. In **Configure Clika Platform**, paste the key you copied in step 1 into **API key**. Leave **Platform URL** as it is, and press **Save**. ![The Configure Clika Platform dialog: API key above Platform URL, with Cancel and Save](/img/platform/mcp/claude-desktop/05-configure.png) 7. The extension is added turned off. Switch the toggle that reads **Disabled** so that it reads **Enabled**. ![The installed Clika Platform extension with its toggle reading Disabled](/img/platform/mcp/claude-desktop/06-installed-disabled.png) ## Step 3. Try it Close Settings, start a new conversation in **Chat**, and ask: > Which of my Clika Platform devices are online? The first time, Claude asks for permission to use the platform. Press **Always allow**, or **Allow once** to be asked again next time. When Claude answers with your devices, you are ready. ### If it does not work - Check that the extension's toggle reads **Enabled** (Step 2, 7). - Check that you pasted the whole key. It starts with `clika_`. - A key stops working 90 days after you created it. Create a new one (Step 1) and put it in with **Configure** next to **Clika Platform** in **Settings**, **Extensions**. - Quit Claude Desktop completely and open it again. If it still does not work, see [Troubleshooting](troubleshooting.md). ## Remove it Open **Settings**, then **Extensions** in Claude Desktop. Press **Configure** next to **Clika Platform**, then **Uninstall**, and confirm with **Uninstall**. Claude Desktop removes the extension and the key you entered. ![The Uninstall Clika Platform? confirmation in Claude Desktop](/img/platform/mcp/claude-desktop/07-uninstall-confirm.png) If nothing else uses the key, revoke it on the platform under **Settings**, then **Developer access**. ## Good to know - A key stops working 90 days after you created it. Create a new one and put it in with **Configure**. - When the CLIKA Platform is updated, download the extension again (Step 1) and add it the same way (Step 2). ## For advanced users The extension is a `.mcpb` file that carries the CLIKA CLI and the settings Claude Desktop needs to start it. Your deployment builds it from the CLIKA CLI it serves, so the extension's version is the deployment's CLI version, and it runs `clika-cli mcp` with the URL and the key you entered: it serves the `api-key` toolset, and of that the tools your key can call. It is offered for Linux and Windows on x86_64 and arm64, and for macOS on Apple silicon. If you already use the CLIKA CLI, you can register Claude Desktop with it instead of the extension, or write the entry by hand. ### A self-signed certificate The extension cannot skip certificate verification. On a deployment with a self-signed certificate, register with `clika-cli mcp install claude-desktop --insecure-tls` instead (next section). ### Register with the CLIKA CLI With the CLIKA CLI installed and logged in ([Get started with the CLI](../cli/get-started.md)), one command registers the platform: ```bash clika-cli mcp install claude-desktop ``` It adds an entry named `clika-platform` to `claude_desktop_config.json`. The entry runs the CLIKA CLI as a bridge to the deployment's MCP endpoint (`clika-cli --profile default mcp --remote`), so the tools follow the deployment with no extension to update. The entry holds no API key: the CLIKA CLI reads the key from its profile each time Claude Desktop starts it. The command finds the file Claude Desktop reads on each operating system, including the Microsoft Store install on Windows, and `clika-cli mcp status` shows which file it chose and why. Before writing anything it checks the profile's key with one request to the deployment. If the profile has no key or the deployment refuses it, the command offers to log in when it runs at a terminal, and prints the `login` command to run otherwise. If Claude Desktop is running, the command says so. Restart it to load the server. ```bash clika-cli mcp install claude-desktop --profile onprem --insecure-tls clika-cli mcp install claude-desktop --dry-run ``` `--profile` registers another deployment's profile, `--insecure-tls` adds the flag to the entry for a deployment with a self-signed certificate, and `--dry-run` prints the entry without writing it. The [CLI reference](../cli/mcp.md#mcp-install-uninstall-status-and-headers) lists every flag. ### Add the entry by hand Log in with the CLIKA CLI first (`clika-cli --base-url https://platform.clika.io login`), then add the entry to Claude Desktop's config file: | Operating system | File | | --- | --- | | macOS | `~/Library/Application Support/Claude/claude_desktop_config.json` | | Linux | `~/.config/Claude/claude_desktop_config.json`, or `$XDG_CONFIG_HOME/Claude/claude_desktop_config.json` when `XDG_CONFIG_HOME` is set | | Windows, Microsoft Store install | `%LOCALAPPDATA%\Packages\Claude_*\LocalCache\Roaming\Claude\claude_desktop_config.json` | | Windows, standalone installer | `%APPDATA%\Claude\claude_desktop_config.json` | The entry runs the CLIKA CLI by its absolute path, with the profile that holds your key: ```json { "mcpServers": { "clika-platform": { "command": "/usr/local/bin/clika-cli", "args": ["--profile", "default", "mcp", "--remote"] } } } ``` On Windows the path is the one the installer printed after `installed clika-cli to`, with each backslash doubled, for example `"C:\\Users\\you\\AppData\\Local\\Programs\\clika\\bin\\clika-cli.exe"`. For a deployment with a self-signed certificate, add `"--insecure-tls"` before `"mcp"`. Keep any other servers already in `mcpServers`. If Claude Desktop is running, restart it. An entry you write by hand is yours: `mcp install` and `mcp uninstall` never change or remove it. If it uses the name `clika-platform`, `mcp install` refuses that name and asks for `--name`. ### Remove a CLIKA CLI or hand-written entry An entry registered with the CLIKA CLI is removed with the CLIKA CLI, which deletes the entries `mcp install` added for the profile and leaves every other entry alone: ```bash clika-cli mcp uninstall claude-desktop ``` An entry written by hand is deleted from the config file by hand. In both cases, if Claude Desktop is running, restart it. If no other client uses the API key, revoke it under **Settings**, then **Developer access**. --- # Codex Connect the Codex CLI and the Codex IDE extension to the platform through config.toml, with clika-cli mcp install, with codex mcp add, or by hand, then check the connection and remove it. Source: https://docs.clika.io/platform/mcp/codex.md Connect the Codex CLI or the Codex IDE extension to the CLIKA Platform and you can ask it about your devices and benchmarks while you work. Both read one config file, `~/.codex/config.toml` (`$CODEX_HOME/config.toml` when `CODEX_HOME` is set, `%USERPROFILE%\.codex\config.toml` on Windows), and one entry there connects both. There are three ways to add it; pick one. | Way | Use it when | Where the API key lives | | --- | --- | --- | | [`clika-cli mcp install`](#register-with-the-clika-cli-recommended) (recommended) | You use the CLIKA CLI. | In your CLIKA CLI profile. `config.toml` holds no key. | | [`codex mcp add`](#add-it-with-codex-mcp-add) | You do not use the CLIKA CLI. | In an environment variable Codex reads when it starts. | | [An entry by hand](#add-the-entry-by-hand) | You manage the config file yourself. | In your CLIKA CLI profile. | Each needs an [API key](index.md#you-need-an-api-key). ChatGPT Desktop reads the same file. An entry added on this page shows up there too, and a server added in ChatGPT Desktop's Settings ([ChatGPT Desktop](chatgpt-desktop.md)) is one Codex uses as well. ## Register with the CLIKA CLI (recommended) With the CLIKA CLI installed and logged in ([Get started with the CLI](../cli/get-started.md)), one command registers the platform: ```bash clika-cli mcp install codex ``` If a Codex session or ChatGPT Desktop is running, the command says so: restart it to load the server. What the command does: - It first checks the profile's key with one request to the deployment. If the profile has no key or the deployment refuses it, the command offers to log in when it runs at a terminal, and prints the `login` command to run otherwise. - It adds a `[mcp_servers.clika-platform]` table to `config.toml`. Every other table and key in the file keeps its place and formatting. - The entry is an HTTP connection to `/api/mcp`. Its `Authorization` header comes from `http_headers_helper`: a command Codex runs to get the header, here `clika-cli --profile default mcp headers`, which prints it from the CLIKA CLI profile. The API key never lands in `config.toml`, and a new key needs a new `clika-cli login` and no change to the entry. Useful flags: | Flag | What it does | | --- | --- | | `--profile staging` | Registers another deployment's profile. | | `--name clika-staging` | Picks the entry's name. | | `--dry-run` | Prints the entry without writing it. | | `--insecure-tls` | For a deployment with a self-signed certificate: the entry runs the CLIKA CLI as a local bridge (`mcp --remote`) instead of connecting over HTTP. [Troubleshooting](troubleshooting.md#a-self-signed-certificate-on-premise) explains why. | The [CLI reference](../cli/mcp.md#mcp-install-uninstall-status-and-headers) lists every flag. ## Add it with `codex mcp add` Without the CLIKA CLI, register the endpoint with Codex's own command: ```bash codex mcp add clika-platform --url https://platform.clika.io/api/mcp --bearer-token-env-var CLIKA_API_KEY ``` - Codex has no option to write the header itself. `--bearer-token-env-var` names an environment variable, and Codex sends its value as the bearer token each time it starts the server. The entry in `config.toml` holds the variable's name, `bearer_token_env_var = "CLIKA_API_KEY"`, not the key. - Set `CLIKA_API_KEY` to your API key wherever Codex runs, for example with `export CLIKA_API_KEY=clika_your_api_key` in your shell's profile on macOS and Linux, or `setx CLIKA_API_KEY clika_your_api_key` on Windows and then a new terminal. The Codex IDE extension sees the variable only when the editor was started from an environment that has it. - A new key means changing the variable, with no change to the entry. - `codex mcp add` has no option to skip certificate verification. On a deployment with a self-signed certificate, register with `clika-cli mcp install codex --insecure-tls` instead ([Troubleshooting](troubleshooting.md#a-self-signed-certificate-on-premise)). ## Add the entry by hand Log in with the CLIKA CLI first (`clika-cli --base-url https://platform.clika.io login`). Then add the table to `config.toml`: ```toml [mcp_servers.clika-platform] url = "https://platform.clika.io/api/mcp" http_headers_helper = "/usr/local/bin/clika-cli --profile default mcp headers" ``` - Use the CLIKA CLI's absolute path. On Windows that is the path the installer printed after `installed clika-cli to`. A TOML literal string in single quotes keeps its backslashes as they are, for example `'C:\Users\you\AppData\Local\Programs\clika\bin\clika-cli.exe --profile default mcp headers'`. - If a Codex session or ChatGPT Desktop is running, restart it. - An entry you write by hand is yours: `mcp install` and `mcp uninstall` never change or remove it. If it uses the name `clika-platform`, `mcp install` refuses that name and asks for `--name`. ### A server added in ChatGPT Desktop ChatGPT Desktop saves the server as a `[mcp_servers.clika-platform]` table with the API key in its `http_headers`, in plain text. Like an entry written by hand, it is yours: `mcp install` and `mcp uninstall` never change or remove it, and `mcp install` refuses the name `clika-platform` while it exists and asks for `--name`. ## Check the connection `clika-cli mcp status` reports, for each client, whether `mcp install` registered the server, under which name and profile, and which config file it read. Then start a Codex session and ask: > Which of my Clika Platform devices are online? The answer comes from the device list tool, `get_devices`. If the tools do not appear, see [Troubleshooting](troubleshooting.md). ## Remove it Remove the entry the way you added it: | Added with | Remove with | | --- | --- | | `clika-cli mcp install` | `clika-cli mcp uninstall codex`. It deletes the tables it added for the profile and leaves the rest of the file as it was. | | `codex mcp add` | `codex mcp remove clika-platform`, then unset `CLIKA_API_KEY` if nothing else uses it. | | ChatGPT Desktop's Settings | ChatGPT Desktop ([Remove it](chatgpt-desktop.md#remove-it)). | | By hand | Delete the table from `config.toml`. | If a Codex session or ChatGPT Desktop is running, restart it. If no other client uses the API key, revoke it under **Settings**, then **Developer access**. --- # Other clients Connect any MCP client that speaks Streamable HTTP to the platform with the endpoint URL and an Authorization header, or run the CLIKA CLI as a local stdio server for a client that only starts commands. Source: https://docs.clika.io/platform/mcp/other-clients.md Any MCP client that speaks Streamable HTTP can connect to the platform. It needs two things: the endpoint URL, `/api/mcp` (or a [toolset](index.md#the-mcp-endpoint-and-its-toolsets) below it), and an `Authorization: Bearer` header carrying an [API key](index.md#you-need-an-api-key). The endpoint is stateless, and a request without a valid key is refused before any tool runs. Where a client takes a JSON server definition, the definition has this shape: ```json { "url": "https://platform.clika.io/api/mcp", "headers": { "Authorization": "Bearer clika_your_api_key" } } ``` The field names and the file they go in are the client's own, so check its documentation for both. Wherever the key is written, it is readable by anything that can read that file. A client that only starts local commands (stdio) can run the CLIKA CLI as its server instead. After `clika-cli login`, the command is `clika-cli mcp`, which serves the `api-key` toolset from the CLIKA CLI, or `clika-cli mcp --remote`, which bridges to the deployment's own endpoint. Give the command as an absolute path. In both forms the CLIKA CLI reads the key from its profile, so the client's config holds no key. The [CLI reference](../cli/mcp.md#mcp) describes both modes. Check the connection with a read-only question, such as which devices in your current project are online. If the client cannot connect, see [Troubleshooting](troubleshooting.md). --- # Troubleshooting What to do when an MCP client is refused with a 401, cannot verify a self-signed certificate, or does not show the platform's tools after a change. Source: https://docs.clika.io/platform/mcp/troubleshooting.md ## A 401, or "not logged in" `clika-cli mcp install` checks the profile's API key with one request to the deployment before it writes anything. It stops when the profile has no saved key, when the profile belongs to a different deployment than `--base-url`, or when the deployment answers 401. At a terminal it offers to log in on the spot. Without one, it prints the exact command to run, for example `clika-cli --base-url https://platform.clika.io login --profile default`. Run it, then run `mcp install` again. A client that connected before and is now refused is presenting a key the deployment refuses, because the key was revoked, its owner's account is suspended or was removed from the organization, or the organization was deleted; all four answer the same `401 invalid or expired token`. Create a new key and save it with `clika-cli login`. Entries written by `mcp install` read the key from the profile, so they need no change. A key written into a client's own config (`claude mcp add`, a hand-written `Authorization` header, the Desktop Extension) has to be replaced there. When a client reports that the server failed to start or to connect right after `mcp install`, run the header command the entry uses, `clika-cli --profile default mcp headers`. It prints the `Authorization` header from the profile, or says that the profile holds no key and which `login` command fixes it. ## A self-signed certificate (on-premise) A client that connects over HTTP has to trust the deployment's certificate. Trusting the deployment's certificate authority in the client is the preferred fix. Claude Code, for example, reads extra authorities from `NODE_EXTRA_CA_CERTS`. Where that is not possible, let the CLIKA CLI make the connection. With `--insecure-tls`, `mcp install` writes an entry that runs the CLIKA CLI as a local bridge to the deployment, `clika-cli --profile onprem --insecure-tls mcp --remote`, for every client: ```bash clika-cli mcp install claude-code --profile onprem --insecure-tls clika-cli mcp install codex --profile onprem --insecure-tls ``` The CLIKA CLI then skips certificate verification for its own connection to the deployment and for nothing else, which is why `mcp --remote` is the scoped exception: the alternative is switching off verification for the whole client process. A profile does not record `--insecure-tls`, so pass it to `mcp install` as well (or set `CLIKA_INSECURE_TLS=1`). Without it the key check fails with a certificate error and the message says to add the flag. The Desktop Extension has no option to skip verification. On such a deployment, register Claude Desktop with `clika-cli mcp install claude-desktop --insecure-tls` instead. ## The tools do not appear after a change ("restart it") A client reads its config when it starts. When `mcp install` or `mcp uninstall` finds the client running, it prints one line naming it, for example: ``` Claude Desktop is running; restart it to load "clika-platform". ``` Restart it, then start a new conversation. When the client is not running, nothing is printed and the change applies the next time it starts. ## Other refusals - `mcp install` refuses a name that an entry it did not write already uses, and changes nothing. Pass `--name` to register under another name. - A warning that the CLIKA CLI runs from a temporary location means the entry would point at a binary that is about to disappear, for example one built by `go run` or started from the system's temporary directory. Install the CLIKA CLI ([Get started with the CLI](../cli/get-started.md)) and run `mcp install` from the installed binary. - A toolset endpoint that answers with an error naming the servable toolsets is not carried by this deployment, because its API lacks one of the operations the toolset names. Use one of the names it lists, or `/api/mcp`.