---
title: "TokenizerBuilder"
sidebar_label: "TokenizerBuilder"
description: "The clika_runtime.tokenizer TokenizerBuilder class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

TokenizerBuilder(model, *, device=None) -> TokenizerBuilder

    Assemble a :class:`Tokenizer` from a :class:`Model` (the core stage, built
    from an explicit vocabulary) plus the optional :class:`Normalizer`,
    :class:`PreTokenizer`, :class:`PostProcessor` and :class:`Decoder` stages,
    added tokens, the well-known ids and a chat template, with no tokenizer
    file. Every setter takes its stage (a taken stage refuses its next use)
    and returns the builder, so a build reads as one chain ending in
    :meth:`build`, which checks the configuration and runs once::

        tokenizer = (
            TokenizerBuilder(Model.word_level({"hello": 0, "[UNK]": 1}, "[UNK]"))
            .pre_tokenizer(PreTokenizer.whitespace_split())
            .decoder(Decoder.fuse())
            .build()
        )

    ``device`` (a Device, a device string or a Stream) is the default
    placement of the tokenizer's tensor-shaped entries; None is the
    runtime's default.
    

## `__init__`

```python
__init__(self, model: 'Model', *, device: 'Device | Stream | str | None' = None) -> 'None'
```

Initialize self.  See help(type(self)) for accurate signature.

## `add_special_token`

```python
add_special_token(self, content: 'str', id: 'int') -> 'TokenizerBuilder'
```

add_special_token(content, id) -> TokenizerBuilder

Register a special added token under ``id``: matched before the model
runs, and skippable on decode.

## `add_token`

```python
add_token(self, content: 'str', id: 'int', *, single_word: 'bool' = False, lstrip: 'bool' = False, rstrip: 'bool' = False, normalized: 'bool' = True) -> 'TokenizerBuilder'
```

add_token(content, id, *, single_word=False, lstrip=False, rstrip=False, normalized=True) -> TokenizerBuilder

Register an ordinary added token: ``single_word`` matches at word
boundaries only, ``lstrip`` / ``rstrip`` consume the whitespace beside
it, ``normalized`` matches the normalized text.

## `bos_id`

```python
bos_id(self, id: 'int') -> 'TokenizerBuilder'
```

bos_id(id) -> TokenizerBuilder

Set the beginning-of-sequence token id.

## `build`

```python
build(self) -> 'Tokenizer'
```

build() -> Tokenizer

Assemble the tokenizer; raises on an inconsistent configuration (a
chat template that does not compile among it). Runs once: a second
call raises ``ValueError``.

## `chat_template`

```python
chat_template(self, jinja_source: 'str', bos_token: 'str', eos_token: 'str') -> 'TokenizerBuilder'
```

chat_template(jinja_source, bos_token, eos_token) -> TokenizerBuilder

Attach a chat template (Jinja source) for ``apply_chat_template`` and
``encode_chat``; ``bos_token`` and ``eos_token`` are the framing tokens
it references. The source compiles at :meth:`build`, which raises on
an invalid one.

## `decoder`

```python
decoder(self, decoder: 'Decoder') -> 'TokenizerBuilder'
```

decoder(decoder) -> TokenizerBuilder

Set the detokenization stage (taken).

## `eos_id`

```python
eos_id(self, id: 'int') -> 'TokenizerBuilder'
```

eos_id(id) -> TokenizerBuilder

Set the end-of-sequence token id.

## `normalizer`

```python
normalizer(self, normalizer: 'Normalizer') -> 'TokenizerBuilder'
```

normalizer(normalizer) -> TokenizerBuilder

Set the text-cleanup stage (taken).

## `pad_id`

```python
pad_id(self, id: 'int') -> 'TokenizerBuilder'
```

pad_id(id) -> TokenizerBuilder

Set the padding token id.

## `post_processor`

```python
post_processor(self, post_processor: 'PostProcessor') -> 'TokenizerBuilder'
```

post_processor(post_processor) -> TokenizerBuilder

Set the special-token stage (taken).

## `pre_tokenizer`

```python
pre_tokenizer(self, pre_tokenizer: 'PreTokenizer') -> 'TokenizerBuilder'
```

pre_tokenizer(pre_tokenizer) -> TokenizerBuilder

Set the pre-segmentation stage (taken).
