Skip to main content

Use ClikaRT from Python

The clika-runtime wheel puts the runtime behind one import: import clika_runtime as crt loads libClikaRT.so and its backends from inside the wheel, with no library paths to set. The Python surface follows PyTorch's shapes: dtype objects such as crt.float32, a string spelling for every mode argument, nn.Module for models, and one exception class per failure kind. NumPy plays three roles: data entry, data exit, and the independent oracle you check results against. Compute runs in the runtime.

This guide assumes the wheel is installed; First steps covers getting it. Everything below is one script's worth of ground: the license credential, the NumPy boundary, operator chains, modes as strings, device placement, a model as nn.Module, and the error contract.

The credential goes in before the import​

The import is what loads the runtime, so CLIKA_RT_LICENSE has to hold the credential before import clika_runtime runs. Export it in the shell, or assign it above the import:

license.py
import os
os.environ["CLIKA_RT_LICENSE"] = "CLIKA1-..." # the credential text, or the path of a file holding it

import clika_runtime as crt

The alternative drops the variable: clikart-license-init <credential>, a console script the wheel installs, stores the credential once under your user account. clika_runtime.torch and the clika_modelverse package follow the same rule, because all three are the one runtime. License the runtime is the whole contract; without a valid credential a call raises crt.ClikaRTError with code_name LICENSE_FAILED.

NumPy in, NumPy out​

crt.tensor(array) copies the array in, crt.from_numpy(array) borrows its memory, and t.numpy() is the exit. Lists and scalars enter too, at the dtype NumPy would pick for them.

boundary.py
import numpy as np
import clika_runtime as crt

a = np.ones((2, 5), dtype=np.float32)
t = crt.tensor(a)

print(t.shape, t.dtype, t.device) # clika_runtime.Size([2, 5]) clika_runtime.float32 cpu
assert t.dtype == crt.float32 # one dtype object per storable dtype
assert isinstance(crt.float32, crt.dtype)
assert t.device == crt.Device("cpu")

a[0, 0] = 999.0 # entry copied: the tensor is unmoved
assert t.numpy()[0, 0] == 1.0
borrowed = crt.from_numpy(a) # from_numpy shares the array's memory
assert borrowed.numpy()[0, 0] == 999.0

assert crt.tensor([1, 2, 3]).dtype == crt.int64
half = t.to(crt.float16) # narrowing is explicit
assert half.dtype == crt.float16
print(half)
# tensor([[1., 1., 1., 1., 1.],
# [1., 1., 1., 1., 1.]], dtype=clika_runtime.float16)

Every NumPy-native dtype enters as itself (the float family, the signed ints, uint8, bool); a dtype with no tensor twin, complex64 for example, is refused with a TypeError that names the routes out. Payload dtypes NumPy cannot spell, bfloat16 among them, cross through bytes() and Tensor.from_bytes() instead of the array bridge. A tensor prints as tensor([...]): the dtype is named when it is not float32, the device when it is not the CPU, and a tensor above a thousand elements is abbreviated to its edge items.

A tensor prints as its values, wrapped in tensor(...), with no shape or device header:

print(2 * crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3)) + 3)
tensor([[ 3., 5., 7.],
[ 9., 11., 13.]])

Math that reads as math​

Operators compose the way the expression reads: Python numbers broadcast, @ is matmul, and method chains mirror the functional forms. Check anything against NumPy; that is what the oracle role means.

tensor_math.py
import numpy as np
import clika_runtime as crt

x = crt.tensor(np.arange(6, dtype=np.float32).reshape(2, 3))

y = 2.0 * x + 3.0 # scalars broadcast
z = (y - 3.0).abs().amax().item() # a method chain down to one float
print(z) # 10.0
assert np.allclose((x ** 2).numpy(), x.numpy() ** 2)

a = crt.tensor(np.ones((2, 3), dtype=np.float32))
b = crt.tensor(np.ones((3, 2), dtype=np.float32))
print((a @ b).numpy())
# [[3. 3.]
# [3. 3.]]

Modes are strings​

A mode argument takes its spelling as a string: approximate="tanh", mode="reflect", rounding_mode="floor", activation="relu". An unknown spelling raises a ValueError that lists the accepted ones.

modes.py
import numpy as np
import clika_runtime as crt

x = crt.tensor(np.linspace(-3.0, 3.0, 7, dtype=np.float32))

tanh_form = crt.gelu(x, approximate="tanh")
assert not np.array_equal(tanh_form.numpy(), crt.gelu(x).numpy())
assert np.array_equal(crt.gelu(x).numpy(), crt.gelu(x, approximate="none").numpy())

padded = crt.pad(x, [2, 2], mode="reflect")
assert np.array_equal(padded.numpy(), np.pad(x.numpy(), (2, 2), mode="reflect"))

try:
crt.div(x, 2.0, rounding_mode="ceil")
except ValueError as e:
print(e) # rounding_mode: 'ceil' is not a RoundingMode; choose one of 'none', 'trunc', 'floor'

Placement​

Placement is a constructor argument or a move: crt.tensor(arr, device=...) lands data where you say, .to("cpu") moves it, and crt.Device.gpu() names the machine's accelerator, or the CPU when it has none, so the same script runs everywhere. Each backend has a namespace: crt.cuda, crt.vulkan and crt.metal mirror crt.accelerator, with is_available(), device(index) and synchronize().

placement.py
import numpy as np
import clika_runtime as crt

gpu = crt.Device.gpu()
t = crt.zeros(2, device=gpu)
assert t.device == gpu
assert np.array_equal(crt.to(t, "cpu").numpy(), np.zeros(2, dtype=np.float32))

if crt.cuda.is_available():
x = crt.ones(2, 3, device=crt.cuda.device(0))
crt.cuda.synchronize()
print(crt.Device("cpu")) # cpu

A model is an nn.Module​

Assigning a layer in __init__ registers it, as in PyTorch: load_state_dict binds dotted names, named_parameters() enumerates them, state_dict() exports the same names back out, and the instance is callable. A Parameter is a Tensor.

model.py
import numpy as np
import clika_runtime as crt
import clika_runtime.nn as nn

class TinyMlp(nn.Module):
def __init__(self, d_in: int, d_hidden: int, d_out: int) -> None:
super().__init__()
# The first layer fuses its activation as an epilogue.
self.up = nn.Linear(d_in, d_hidden, activation="relu")
self.down = nn.Linear(d_hidden, d_out, bias=False)

def forward(self, x: crt.Tensor) -> crt.Tensor:
return self.down(self.up(x))

model = TinyMlp(4, 8, 2)
assert isinstance(model.up.weight, nn.Parameter) and isinstance(model.up.weight, crt.Tensor)
result = model.load_state_dict({
"up.weight": crt.tensor(np.full((8, 4), 0.1, dtype=np.float32)),
"up.bias": crt.tensor(np.zeros(8, dtype=np.float32)),
"down.weight": crt.tensor(np.full((2, 8), 0.1, dtype=np.float32)),
})
assert result.missing_keys == [] and result.unexpected_keys == []

y = model(crt.tensor(np.ones((3, 4), dtype=np.float32)))
print(y.shape) # clika_runtime.Size([3, 2])

exported = model.state_dict() # the same dotted names back out
print(sorted(exported)) # ['down.weight', 'up.bias', 'up.weight']

load_state_dict is strict by default: a missing or unexpected key raises a RuntimeError naming both sets; strict=False returns the report instead, with missing_keys and unexpected_keys. The state_dict() -> fresh load_state_dict() round trip reproduces the forward, which is the portable way to hand weights between processes. Author a model in Python builds a full decoder this way.

When it fails​

A runtime failure raises an exception typed by its kind: crt.InvalidArgumentError (also a ValueError) for a shape, dtype, device or option the call cannot accept, crt.NotFoundError (also a FileNotFoundError) for a missing file or entry, crt.UnsupportedError, crt.OutOfMemoryError and crt.UnavailableError for their kinds, all under crt.ClikaRTError. Every instance carries .code_name (the fine code, stable across builds) and .status (the coarse class); str(err) is the message alone. Branch on the class or the code name, never on the message text: an argument mistake reads as a sentence naming the operation and the values, while an E<digits> message is an internal fault code specific to the build that produced it (report it verbatim with the runtime version).

errors.py
import numpy as np
import clika_runtime as crt

try:
a = crt.tensor(np.ones((2, 3), dtype=np.float32))
b = crt.tensor(np.ones((4, 5), dtype=np.float32))
_ = a @ b # shape mismatch
except crt.InvalidArgumentError as e:
assert isinstance(e, ValueError)
print(f"failed with code {e.code_name!r}") # failed with code 'INVALID_ARGUMENT'

try:
crt.load("absent.safetensors")
except crt.NotFoundError as e:
assert isinstance(e, FileNotFoundError)
assert e.code_name == "NO_SUCHFILE"

From here, the rest of the Python surface follows the same grammar: GGUF and quantized weights and tokenizers have Python arms on their pages, ONNX models compile and run, tracing turns eager functions into graphs, pytrees carry structured inputs and outputs across those boundaries, and PyTorch interop covers the torch backend and tensor exchange. The wheel's own example programs double as a smoke suite for an installed wheel.