Skip to main content

//clika-runtime/io.clika.runtime/Tokenizer

Tokenizer

[common]
class Tokenizer

A loaded tokenizer: text to ids and back, the tensor-shaped batch forms, the model's chat template, and a streaming decoder for generation.

The loaders are named by the format they read. fromHuggingface is the full entry: a model directory or one artifact file (tokenizer.json, tokenizer.model, tekken.json, vocab.json + merges.txt), with the sibling tokenizer_config.json read for the special ids and the chat template. fromFile detects the format for you. The bare loaders read one format and nothing beside it. Every failure is a ClikaRtException with its status and code name. AutoCloseable: close releases the tokenizer (a StreamingDecoder taken from it stays valid).

Tokenizer.fromHuggingface("/path/to/model").use { tokenizer ->
val ids = tokenizer.encode("Hello, world")
val text = tokenizer.decode(ids)
val prompt = tokenizer.applyChatTemplate("""[{"role": "user", "content": "Hi"}]""")
}

Types​

NameSummary
Companion[common]
object Companion

Properties​

NameSummary
bosId[common]
val bosId: Long
The begin-of-sequence id the model's config names, or -1 when it names none.
eosId[common]
val eosId: Long
The end-of-sequence id, or -1 when the config names none.
hasChatTemplate[common]
val hasChatTemplate: Boolean
True when the tokenizer carries a chat template (read beside tokenizer_config.json).
padId[common]
val padId: Long
The padding id, or -1 when the config names none.
vocabSize[common]
val vocabSize: Long
The number of entries in the vocabulary.

Functions​

NameSummary
applyChatTemplate[common]
fun applyChatTemplate(inputs: ChatTemplateInputs, options: ChatTemplateOptions? = null): String
Render with the full input set and, when given, per-render options.
[common]
fun applyChatTemplate(messages: String, addGenerationPrompt: Boolean = true): String
Render chat messages (a JSON array of {"role", "content"} turns, as text) into a prompt with the model's chat template; addGenerationPrompt appends the assistant-turn priming. A tokenizer with no template, malformed messages or a render error fails typed.
close[common]
open fun close()
decode[common]
fun decode(ids: IntArray, skipSpecialTokens: Boolean = true): String
Decode token ids back into text, the special tokens left out by default.
decodeBatch[common]
fun decodeBatch(ids: Tensor, skipSpecialTokens: Boolean = true): List<String>
Decode an ids tensor (Int32 or Int64): a [S] sequence yields one string, a padded [B, S] batch one string per row.
[common]
fun decodeBatch(sequences: List<IntArray>, skipSpecialTokens: Boolean = true): List<String>
Decode many id sequences, one string per sequence, in order.
[common]
fun decodeBatch(ids: Tensor, cuSeqlens: Tensor, skipSpecialTokens: Boolean = true): List<String>
Decode a flat varlen ids row sliced by its [B + 1] offset table, one string per sequence.
encode[common]
fun encode(text: String, addSpecialTokens: Boolean = true): IntArray
Encode text into token ids, wrapped with the model's special tokens by default.
encodeBatch[common]
fun encodeBatch(texts: List<String>, options: EncodeOptions = EncodeOptions()): Encoded
Encode a batch of texts into tensor form: a padded [B, S] matrix by default, or one flat row plus the offset table with EncodeOptions.varlen.
encodeChat[common]
fun encodeChat(inputs: ChatTemplateInputs, options: ChatTemplateOptions? = null, addSpecialTokens: Boolean = false): IntArray
Render and encode with the full input set and, when given, per-render options.
[common]
fun encodeChat(messages: String, addGenerationPrompt: Boolean = true, addSpecialTokens: Boolean = false): IntArray
Render chat messages with the template and encode the prompt in one step. The template writes its own bos and eos framing, so the ids are encoded without the special tokens by default; pass addSpecialTokens only for a template that writes no bos itself.
encodeText[common]
fun encodeText(text: String, options: EncodeOptions = EncodeOptions()): Encoded
Encode one text into tensor form, a one-element batch.
idToToken[common]
fun idToToken(id: Int): String
The surface text of a token id; an id out of range fails with NOT_FOUND.
streamingDecoder[common]
fun streamingDecoder(skipSpecialTokens: Boolean = true): StreamingDecoder
A streaming decoder over this tokenizer, for one generation stream at a time.
tokenize[common]
fun tokenize(text: String, addSpecialTokens: Boolean = true): List<Token>
Encode text into per-token records: each with its id, its surface text, its byte span in text, and whether the tokenizer injected it.
tokenToId[common]
fun tokenToId(token: String): Int
The id of a surface token; a token outside the vocabulary fails with NOT_FOUND.