Coming from PyTorch or Hugging Face
ClikaRT's Python package reads like PyTorch and its model library loads a checkpoint by the name the Hugging Face hub gives it, so most of what you know carries over. This page lists the places where the runtime's convention differs from the one you expect, each with the one line that bridges it.
Tensors are channels-last, weights are OHWI
Activations are [N, spatial..., C] and a convolution weight is [out_channels, K..., in_channels / groups]
(output-channel-first, channels last). PyTorch stores an activation [N, C, H, W] and a
convolution weight [O, C / groups, K, K], and every checkpoint exported from it keeps that
order, so a weight is brought to OHWI once, at load, with one permute and a contiguous copy:
const Tensor w_ohwi = ops::contiguous(ops::permute(w_oihw, {0, 2, 3, 1})); // 2-D; {0, 2, 1} for 1-D, {0, 2, 3, 4, 1} for 3-D
w_ohwi = w_oihw.permute(0, 2, 3, 1).contiguous()
nn::Conv and nn.Conv declare their geometry in that layout, and
Load images and audio for inference decodes a picture straight into
[H, W, C].
A size is an integer at the operator's boundary
A split size, an index, a sequence length (max_seqlen_k of the attention operators) is an
integer slot. The C++ gateway type that takes a number or a tensor accepts a floating-point value
too, and refuses it at run time when the slot is integer-typed, so pass an int (or an
int64_t), never a double cast from one.
A read is the wait
An operator returns as soon as its work is queued; the kernels run behind it. A host read waits
for the value: item, numpy(), tolist(), printing a tensor, eval, to_string() and
item_as_vec<T>() in C++. Nothing you read is ever an unfinished value. The consequence for a
measurement: time a step at the read of its result, or after synchronize() on the producing
stream; the return of the dispatch is not the end of the work.
Control asynchronous execution walks the model in full.
JSON, audio and the tokenizer's files
- A JSON document reads an integer through
as_int64()and a floating-point value throughas_double().operator[]on a mutable document creates a missing key (it turns a null into an object);at(key)andat(index)read without creating. io::load_audioreturns the samples beside anAudioInfowhosesample_rate,channelsandframesdescribe them (frames / sample_rateis the duration in seconds).Tokenizer::from_huggingfacein C++,Tokenizer.from_filein Python andTokenizer.fromHuggingfacein Kotlin read a model directory: itstokenizer.jsonwith the special ids and the chat template; a SentencePiece model beside itstokenizer_config.json; or avocab.jsonfor a model no tokenizer class claims.
The thread count is read once
CLIKA_RT_NUM_THREADS sets the CPU worker count and is read before the first compute, so it goes
into the environment before the first operator runs: in the shell, or from the program before it
touches the runtime. An Android app sets it in its own process before it creates any tensor.
generate, side by side
The model library's Python surface mirrors the transformers idioms; the table the wheel's
examples/python/howto/generate_like_transformers/README.md carries is the whole map. The rows
a first program needs:
transformers | clika_runtime.modelverse |
|---|---|
AutoModelForCausalLM.from_pretrained(id) | AutoModelForCausalLM.from_pretrained(id, device="cpu") |
model.to("cuda") | from_pretrained(id, device="cuda:0"), chosen at load |
ids = tok(prompt).input_ids; out = model.generate(ids); tok.decode(out) | model.generate(prompt), which takes text and returns text |
generate(..., do_sample=False) | generate(..., temperature=0.0) |
tok.apply_chat_template(messages) | model.render(messages) |
apply_chat_template then generate | model.chat(messages) |
TextIteratorStreamer on a second thread | model.stream_generate(prompt), an iterator of text pieces |
pipeline("text-generation", model=id) | pipeline("text-generation", id) |
A model's weights download inside from_pretrained into the Hugging Face hub cache the other
tooling on the machine shares; mv.snapshot_download(id) downloads without loading, and a local
directory is a source everywhere a repository id is
(Run fully offline).
Where the same idea has another name
| You reach for | Here |
|---|---|
torch.compile(model) | crt.compile(model): capture on the first call, replay after (Trace eager code to graphs) |
torch.no_grad(), model.eval() | no gradients exist; model.eval() is accepted and changes nothing |
state_dict() / load_state_dict() | the same names and the same dotted keys; load_state_dict(state, assign=True) adopts the checkpoint's tensors as the parameters, one resident copy (Author a model in Python) |
torch.device("meta") | crt.device("meta"): shapes and dtypes without bytes, then to_empty(device=...) |
| DLPack exchange with torch | clika_runtime.torch (Use ClikaRT with PyTorch) |