---
title: Coming from PyTorch or Hugging Face
description: "The conventions a reader who knows PyTorch and the Hugging Face libraries meets first: channels-last tensors and OHWI weights, integer slots, reads that wait, the JSON and audio accessors, the tokenizer's inputs, and the generate calls side by side."
---

{/* CERTIFICATION: every fact on this page is read from the public headers of the pinned release
     (nn/conv.h, compute/ops.h, json/json.h, io/io.h, tokenizer/tokenizer.h, threading/threading.h)
     and the wheel's generate_like_transformers README; the samples are excerpts, not a compiled program. */}

ClikaRT's Python package reads like PyTorch and its model library loads a checkpoint by the name
the Hugging Face hub gives it, so most of what you know carries over. This page lists the places
where the runtime's convention differs from the one you expect, each with the one line that
bridges it.

## Tensors are channels-last, weights are OHWI

Activations are `[N, spatial..., C]` and a convolution weight is `[out_channels, K..., in_channels / groups]`
(output-channel-first, channels last). PyTorch stores an activation `[N, C, H, W]` and a
convolution weight `[O, C / groups, K, K]`, and every checkpoint exported from it keeps that
order, so a weight is brought to OHWI once, at load, with one permute and a contiguous copy:

```cpp
const Tensor w_ohwi = ops::contiguous(ops::permute(w_oihw, {0, 2, 3, 1}));   // 2-D; {0, 2, 1} for 1-D, {0, 2, 3, 4, 1} for 3-D
```

```python
w_ohwi = w_oihw.permute(0, 2, 3, 1).contiguous()
```

`nn::Conv` and `nn.Conv` declare their geometry in that layout, and
[Load images and audio for inference](load-images-and-audio.mdx) decodes a picture straight into
`[H, W, C]`.

## A size is an integer at the operator's boundary

A split size, an index, a sequence length (`max_seqlen_k` of the attention operators) is an
integer slot. The C++ gateway type that takes a number or a tensor accepts a floating-point value
too, and refuses it at run time when the slot is integer-typed, so pass an `int` (or an
`int64_t`), never a `double` cast from one.

## A read is the wait

An operator returns as soon as its work is queued; the kernels run behind it. A host read waits
for the value: `item`, `numpy()`, `tolist()`, printing a tensor, `eval`, `to_string()` and
`item_as_vec<T>()` in C++. Nothing you read is ever an unfinished value. The consequence for a
measurement: time a step at the read of its result, or after `synchronize()` on the producing
stream; the return of the dispatch is not the end of the work.
[Control asynchronous execution](control-async-execution.mdx) walks the model in full.

## JSON, audio and the tokenizer's files

- A JSON document reads an integer through `as_int64()` and a floating-point value through
  `as_double()`. `operator[]` on a mutable document creates a missing key (it turns a null into an
  object); `at(key)` and `at(index)` read without creating.
- `io::load_audio` returns the samples beside an `AudioInfo` whose `sample_rate`, `channels` and
  `frames` describe them (`frames / sample_rate` is the duration in seconds).
- `Tokenizer::from_huggingface` in C++, `Tokenizer.from_file` in Python and `Tokenizer.fromHuggingface`
  in Kotlin read a model directory: its `tokenizer.json` with the special ids and the chat
  template; a SentencePiece model beside its `tokenizer_config.json`; or a `vocab.json` for a model
  no tokenizer class claims.

## The thread count is read once

`CLIKA_RT_NUM_THREADS` sets the CPU worker count and is read before the first compute, so it goes
into the environment before the first operator runs: in the shell, or from the program before it
touches the runtime. An Android app sets it in its own process before it creates any tensor.

## generate, side by side

The model library's Python surface mirrors the `transformers` idioms; the table the wheel's
`examples/python/howto/generate_like_transformers/README.md` carries is the whole map. The rows
a first program needs:

| `transformers` | `clika_runtime.modelverse` |
| --- | --- |
| `AutoModelForCausalLM.from_pretrained(id)` | `AutoModelForCausalLM.from_pretrained(id, device="cpu")` |
| `model.to("cuda")` | `from_pretrained(id, device="cuda:0")`, chosen at load |
| `ids = tok(prompt).input_ids; out = model.generate(ids); tok.decode(out)` | `model.generate(prompt)`, which takes text and returns text |
| `generate(..., do_sample=False)` | `generate(..., temperature=0.0)` |
| `tok.apply_chat_template(messages)` | `model.render(messages)` |
| `apply_chat_template` then `generate` | `model.chat(messages)` |
| `TextIteratorStreamer` on a second thread | `model.stream_generate(prompt)`, an iterator of text pieces |
| `pipeline("text-generation", model=id)` | `pipeline("text-generation", id)` |

A model's weights download inside `from_pretrained` into the Hugging Face hub cache the other
tooling on the machine shares; `mv.snapshot_download(id)` downloads without loading, and a local
directory is a source everywhere a repository id is
([Run fully offline](/modelverse/how-to/run-fully-offline)).

## Where the same idea has another name

| You reach for | Here |
| --- | --- |
| `torch.compile(model)` | `crt.compile(model)`: capture on the first call, replay after ([Trace eager code to graphs](trace-eager-code-to-graphs.mdx)) |
| `torch.no_grad()`, `model.eval()` | no gradients exist; `model.eval()` is accepted and changes nothing |
| `state_dict()` / `load_state_dict()` | the same names and the same dotted keys; `load_state_dict(state, assign=True)` adopts the checkpoint's tensors as the parameters, one resident copy ([Author a model in Python](author-a-model-in-python.mdx)) |
| `torch.device("meta")` | `crt.device("meta")`: shapes and dtypes without bytes, then `to_empty(device=...)` |
| DLPack exchange with torch | `clika_runtime.torch` ([Use ClikaRT with PyTorch](use-clikart-with-pytorch.mdx)) |
