Skip to main content

Coming from PyTorch or Hugging Face

ClikaRT's Python package reads like PyTorch and its model library loads a checkpoint by the name the Hugging Face hub gives it, so most of what you know carries over. This page lists the places where the runtime's convention differs from the one you expect, each with the one line that bridges it.

Tensors are channels-last, weights are OHWI​

Activations are [N, spatial..., C] and a convolution weight is [out_channels, K..., in_channels / groups] (output-channel-first, channels last). PyTorch stores an activation [N, C, H, W] and a convolution weight [O, C / groups, K, K], and every checkpoint exported from it keeps that order, so a weight is brought to OHWI once, at load, with one permute and a contiguous copy:

const Tensor w_ohwi = ops::contiguous(ops::permute(w_oihw, {0, 2, 3, 1})); // 2-D; {0, 2, 1} for 1-D, {0, 2, 3, 4, 1} for 3-D
w_ohwi = w_oihw.permute(0, 2, 3, 1).contiguous()

nn::Conv and nn.Conv declare their geometry in that layout, and Load images and audio for inference decodes a picture straight into [H, W, C].

A size is an integer at the operator's boundary​

A split size, an index, a sequence length (max_seqlen_k of the attention operators) is an integer slot. The C++ gateway type that takes a number or a tensor accepts a floating-point value too, and refuses it at run time when the slot is integer-typed, so pass an int (or an int64_t), never a double cast from one.

A read is the wait​

An operator returns as soon as its work is queued; the kernels run behind it. A host read waits for the value: item, numpy(), tolist(), printing a tensor, eval, to_string() and item_as_vec<T>() in C++. Nothing you read is ever an unfinished value. The consequence for a measurement: time a step at the read of its result, or after synchronize() on the producing stream; the return of the dispatch is not the end of the work. Control asynchronous execution walks the model in full.

JSON, audio and the tokenizer's files​

  • A JSON document reads an integer through as_int64() and a floating-point value through as_double(). operator[] on a mutable document creates a missing key (it turns a null into an object); at(key) and at(index) read without creating.
  • io::load_audio returns the samples beside an AudioInfo whose sample_rate, channels and frames describe them (frames / sample_rate is the duration in seconds).
  • Tokenizer::from_huggingface in C++, Tokenizer.from_file in Python and Tokenizer.fromHuggingface in Kotlin read a model directory: its tokenizer.json with the special ids and the chat template; a SentencePiece model beside its tokenizer_config.json; or a vocab.json for a model no tokenizer class claims.

The thread count is read once​

CLIKA_RT_NUM_THREADS sets the CPU worker count and is read before the first compute, so it goes into the environment before the first operator runs: in the shell, or from the program before it touches the runtime. An Android app sets it in its own process before it creates any tensor.

generate, side by side​

The model library's Python surface mirrors the transformers idioms; the table the wheel's examples/python/howto/generate_like_transformers/README.md carries is the whole map. The rows a first program needs:

transformersclika_runtime.modelverse
AutoModelForCausalLM.from_pretrained(id)AutoModelForCausalLM.from_pretrained(id, device="cpu")
model.to("cuda")from_pretrained(id, device="cuda:0"), chosen at load
ids = tok(prompt).input_ids; out = model.generate(ids); tok.decode(out)model.generate(prompt), which takes text and returns text
generate(..., do_sample=False)generate(..., temperature=0.0)
tok.apply_chat_template(messages)model.render(messages)
apply_chat_template then generatemodel.chat(messages)
TextIteratorStreamer on a second threadmodel.stream_generate(prompt), an iterator of text pieces
pipeline("text-generation", model=id)pipeline("text-generation", id)

A model's weights download inside from_pretrained into the Hugging Face hub cache the other tooling on the machine shares; mv.snapshot_download(id) downloads without loading, and a local directory is a source everywhere a repository id is (Run fully offline).

Where the same idea has another name​

You reach forHere
torch.compile(model)crt.compile(model): capture on the first call, replay after (Trace eager code to graphs)
torch.no_grad(), model.eval()no gradients exist; model.eval() is accepted and changes nothing
state_dict() / load_state_dict()the same names and the same dotted keys; load_state_dict(state, assign=True) adopts the checkpoint's tensors as the parameters, one resident copy (Author a model in Python)
torch.device("meta")crt.device("meta"): shapes and dtypes without bytes, then to_empty(device=...)
DLPack exchange with torchclika_runtime.torch (Use ClikaRT with PyTorch)