//clika-runtime/io.clika.runtime/Tokenizer
Tokenizer
[common]
class Tokenizer
A loaded tokenizer: text to ids and back, the tensor-shaped batch forms, the model's chat template, and a streaming decoder for generation.
The loaders are named by the format they read. fromHuggingface is the full entry: a model directory or one artifact file (tokenizer.json, tokenizer.model, tekken.json, vocab.json + merges.txt), with the sibling tokenizer_config.json read for the special ids and the chat template. fromFile detects the format for you. The bare loaders read one format and nothing beside it. Every failure is a ClikaRtException with its status and code name. AutoCloseable: close releases the tokenizer (a StreamingDecoder taken from it stays valid).
Tokenizer.fromHuggingface("/path/to/model").use { tokenizer ->
val ids = tokenizer.encode("Hello, world")
val text = tokenizer.decode(ids)
val prompt = tokenizer.applyChatTemplate("""[{"role": "user", "content": "Hi"}]""")
}
Types
| Name | Summary |
|---|---|
| Companion | [common] object Companion |
Properties
| Name | Summary |
|---|---|
| bosId | [common] val bosId: Long The begin-of-sequence id the model's config names, or -1 when it names none. |
| eosId | [common] val eosId: Long The end-of-sequence id, or -1 when the config names none. |
| hasChatTemplate | [common] val hasChatTemplate: Boolean True when the tokenizer carries a chat template (read beside tokenizer_config.json). |
| padId | [common] val padId: Long The padding id, or -1 when the config names none. |
| vocabSize | [common] val vocabSize: Long The number of entries in the vocabulary. |
Functions
| Name | Summary |
|---|---|
| applyChatTemplate | [common] fun applyChatTemplate(inputs: ChatTemplateInputs, options: ChatTemplateOptions? = null): String Render with the full input set and, when given, per-render options. [common] fun applyChatTemplate(messages: String, addGenerationPrompt: Boolean = true): String Render chat messages (a JSON array of {"role", "content"} turns, as text) into a prompt with the model's chat template; addGenerationPrompt appends the assistant-turn priming. A tokenizer with no template, malformed messages or a render error fails typed. |
| close | [common] open fun close() |
| decode | [common] fun decode(ids: IntArray, skipSpecialTokens: Boolean = true): String Decode token ids back into text, the special tokens left out by default. |
| decodeBatch | [common] fun decodeBatch(ids: Tensor, skipSpecialTokens: Boolean = true): List<String> Decode an ids tensor (Int32 or Int64): a [S] sequence yields one string, a padded [B, S] batch one string per row.[common] fun decodeBatch(sequences: List<IntArray>, skipSpecialTokens: Boolean = true): List<String> Decode many id sequences, one string per sequence, in order. [common] fun decodeBatch(ids: Tensor, cuSeqlens: Tensor, skipSpecialTokens: Boolean = true): List<String> Decode a flat varlen ids row sliced by its [B + 1] offset table, one string per sequence. |
| encode | [common] fun encode(text: String, addSpecialTokens: Boolean = true): IntArray Encode text into token ids, wrapped with the model's special tokens by default. |
| encodeBatch | [common] fun encodeBatch(texts: List<String>, options: EncodeOptions = EncodeOptions()): Encoded Encode a batch of texts into tensor form: a padded [B, S] matrix by default, or one flat row plus the offset table with EncodeOptions.varlen. |
| encodeChat | [common] fun encodeChat(inputs: ChatTemplateInputs, options: ChatTemplateOptions? = null, addSpecialTokens: Boolean = false): IntArray Render and encode with the full input set and, when given, per-render options. [common] fun encodeChat(messages: String, addGenerationPrompt: Boolean = true, addSpecialTokens: Boolean = false): IntArray Render chat messages with the template and encode the prompt in one step. The template writes its own bos and eos framing, so the ids are encoded without the special tokens by default; pass addSpecialTokens only for a template that writes no bos itself. |
| encodeText | [common] fun encodeText(text: String, options: EncodeOptions = EncodeOptions()): Encoded Encode one text into tensor form, a one-element batch. |
| idToToken | [common] fun idToToken(id: Int): String The surface text of a token id; an id out of range fails with NOT_FOUND. |
| streamingDecoder | [common] fun streamingDecoder(skipSpecialTokens: Boolean = true): StreamingDecoder A streaming decoder over this tokenizer, for one generation stream at a time. |
| tokenize | [common] fun tokenize(text: String, addSpecialTokens: Boolean = true): List<Token> Encode text into per-token records: each with its id, its surface text, its byte span in text, and whether the tokenizer injected it. |
| tokenToId | [common] fun tokenToId(token: String): Int The id of a surface token; a token outside the vocabulary fails with NOT_FOUND. |