---
title: "Tokenizer"
sidebar_label: "Tokenizer"
description: "Kotlin binding reference: Tokenizer."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Tokenizer](index.md)

# Tokenizer

[common]\
class [Tokenizer](index.md)

A loaded tokenizer: text to ids and back, the tensor-shaped batch forms, the model's chat template, and a streaming decoder for generation.

The loaders are named by the format they read. [fromHuggingface](Companion/fromHuggingface.md) is the full entry: a model directory or one artifact file (`tokenizer.json`, `tokenizer.model`, `tekken.json`, `vocab.json` + `merges.txt`), with the sibling `tokenizer_config.json` read for the special ids and the chat template. [fromFile](Companion/fromFile.md) detects the format for you. The bare loaders read one format and nothing beside it. Every failure is a [ClikaRtException](../ClikaRtException/index.md) with its status and code name. `AutoCloseable`: [close](close.md) releases the tokenizer (a [StreamingDecoder](../StreamingDecoder/index.md) taken from it stays valid).

```kotlin
Tokenizer.fromHuggingface("/path/to/model").use { tokenizer ->
    val ids = tokenizer.encode("Hello, world")
    val text = tokenizer.decode(ids)
    val prompt = tokenizer.applyChatTemplate("""[{"role": "user", "content": "Hi"}]""")
}
```

## Types

| Name | Summary |
|---|---|
| [Companion](Companion/index.md) | [common]<br>object [Companion](Companion/index.md) |

## Properties

| Name | Summary |
|---|---|
| [bosId](bosId.md) | [common]<br>val [bosId](bosId.md): [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)<br>The begin-of-sequence id the model's config names, or -1 when it names none. |
| [eosId](eosId.md) | [common]<br>val [eosId](eosId.md): [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)<br>The end-of-sequence id, or -1 when the config names none. |
| [hasChatTemplate](hasChatTemplate.md) | [common]<br>val [hasChatTemplate](hasChatTemplate.md): [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html)<br>True when the tokenizer carries a chat template (read beside `tokenizer_config.json`). |
| [padId](padId.md) | [common]<br>val [padId](padId.md): [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)<br>The padding id, or -1 when the config names none. |
| [vocabSize](vocabSize.md) | [common]<br>val [vocabSize](vocabSize.md): [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)<br>The number of entries in the vocabulary. |

## Functions

| Name | Summary |
|---|---|
| [applyChatTemplate](applyChatTemplate.md) | [common]<br>fun [applyChatTemplate](applyChatTemplate.md)(inputs: [ChatTemplateInputs](../ChatTemplateInputs/index.md), options: [ChatTemplateOptions](../ChatTemplateOptions/index.md)? = null): [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)<br>Render with the full input set and, when given, per-render [options](applyChatTemplate.md).<br>[common]<br>fun [applyChatTemplate](applyChatTemplate.md)(messages: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html), addGenerationPrompt: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)<br>Render chat [messages](applyChatTemplate.md) (a JSON array of `{"role", "content"}` turns, as text) into a prompt with the model's chat template; [addGenerationPrompt](applyChatTemplate.md) appends the assistant-turn priming. A tokenizer with no template, malformed messages or a render error fails typed. |
| [close](close.md) | [common]<br>open fun [close](close.md)() |
| [decode](decode.md) | [common]<br>fun [decode](decode.md)(ids: [IntArray](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int-array/index.html), skipSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)<br>Decode token [ids](decode.md) back into text, the special tokens left out by default. |
| [decodeBatch](decodeBatch.md) | [common]<br>fun [decodeBatch](decodeBatch.md)(ids: [Tensor](../Tensor/index.md), skipSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)&gt;<br>Decode an ids tensor (Int32 or Int64): a `[S]` sequence yields one string, a padded `[B, S]` batch one string per row.<br>[common]<br>fun [decodeBatch](decodeBatch.md)(sequences: [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[IntArray](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int-array/index.html)&gt;, skipSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)&gt;<br>Decode many id sequences, one string per sequence, in order.<br>[common]<br>fun [decodeBatch](decodeBatch.md)(ids: [Tensor](../Tensor/index.md), cuSeqlens: [Tensor](../Tensor/index.md), skipSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)&gt;<br>Decode a flat varlen ids row sliced by its `[B + 1]` offset table, one string per sequence. |
| [encode](encode.md) | [common]<br>fun [encode](encode.md)(text: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html), addSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [IntArray](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int-array/index.html)<br>Encode [text](encode.md) into token ids, wrapped with the model's special tokens by default. |
| [encodeBatch](encodeBatch.md) | [common]<br>fun [encodeBatch](encodeBatch.md)(texts: [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)&gt;, options: [EncodeOptions](../EncodeOptions/index.md) = EncodeOptions()): [Encoded](../Encoded/index.md)<br>Encode a batch of texts into tensor form: a padded `[B, S]` matrix by default, or one flat row plus the offset table with [EncodeOptions.varlen](../EncodeOptions/varlen.md). |
| [encodeChat](encodeChat.md) | [common]<br>fun [encodeChat](encodeChat.md)(inputs: [ChatTemplateInputs](../ChatTemplateInputs/index.md), options: [ChatTemplateOptions](../ChatTemplateOptions/index.md)? = null, addSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = false): [IntArray](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int-array/index.html)<br>Render and encode with the full input set and, when given, per-render [options](encodeChat.md).<br>[common]<br>fun [encodeChat](encodeChat.md)(messages: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html), addGenerationPrompt: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true, addSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = false): [IntArray](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int-array/index.html)<br>Render chat [messages](encodeChat.md) with the template and encode the prompt in one step. The template writes its own bos and eos framing, so the ids are encoded without the special tokens by default; pass [addSpecialTokens](encodeChat.md) only for a template that writes no bos itself. |
| [encodeText](encodeText.md) | [common]<br>fun [encodeText](encodeText.md)(text: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html), options: [EncodeOptions](../EncodeOptions/index.md) = EncodeOptions()): [Encoded](../Encoded/index.md)<br>Encode one text into tensor form, a one-element batch. |
| [idToToken](idToToken.md) | [common]<br>fun [idToToken](idToToken.md)(id: [Int](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int/index.html)): [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)<br>The surface text of a token id; an id out of range fails with `NOT_FOUND`. |
| [streamingDecoder](streamingDecoder.md) | [common]<br>fun [streamingDecoder](streamingDecoder.md)(skipSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [StreamingDecoder](../StreamingDecoder/index.md)<br>A streaming decoder over this tokenizer, for one generation stream at a time. |
| [tokenize](tokenize.md) | [common]<br>fun [tokenize](tokenize.md)(text: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html), addSpecialTokens: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html) = true): [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[Token](../Token/index.md)&gt;<br>Encode [text](tokenize.md) into per-token records: each with its id, its surface text, its byte span in [text](tokenize.md), and whether the tokenizer injected it. |
| [tokenToId](tokenToId.md) | [common]<br>fun [tokenToId](tokenToId.md)(token: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Int](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-int/index.html)<br>The id of a surface token; a token outside the vocabulary fails with `NOT_FOUND`. |