Skip to main content

TokenizerBuilder

TokenizerBuilder(model, *, device=None) -> TokenizerBuilder

Assemble a :class:`Tokenizer` from a :class:`Model` (the core stage, built
from an explicit vocabulary) plus the optional :class:`Normalizer`,
:class:`PreTokenizer`, :class:`PostProcessor` and :class:`Decoder` stages,
added tokens, the well-known ids and a chat template, with no tokenizer
file. Every setter takes its stage (a taken stage refuses its next use)
and returns the builder, so a build reads as one chain ending in
:meth:`build`, which checks the configuration and runs once::

tokenizer = (
TokenizerBuilder(Model.word_level({"hello": 0, "[UNK]": 1}, "[UNK]"))
.pre_tokenizer(PreTokenizer.whitespace_split())
.decoder(Decoder.fuse())
.build()
)

``device`` (a Device, a device string or a Stream) is the default
placement of the tokenizer's tensor-shaped entries; None is the
runtime's default.

__init__​

__init__(self, model: 'Model', *, device: 'Device | Stream | str | None' = None) -> 'None'

Initialize self. See help(type(self)) for accurate signature.

add_special_token​

add_special_token(self, content: 'str', id: 'int') -> 'TokenizerBuilder'

add_special_token(content, id) -> TokenizerBuilder

Register a special added token under id: matched before the model runs, and skippable on decode.

add_token​

add_token(self, content: 'str', id: 'int', *, single_word: 'bool' = False, lstrip: 'bool' = False, rstrip: 'bool' = False, normalized: 'bool' = True) -> 'TokenizerBuilder'

add_token(content, id, *, single_word=False, lstrip=False, rstrip=False, normalized=True) -> TokenizerBuilder

Register an ordinary added token: single_word matches at word boundaries only, lstrip / rstrip consume the whitespace beside it, normalized matches the normalized text.

bos_id​

bos_id(self, id: 'int') -> 'TokenizerBuilder'

bos_id(id) -> TokenizerBuilder

Set the beginning-of-sequence token id.

build​

build(self) -> 'Tokenizer'

build() -> Tokenizer

Assemble the tokenizer; raises on an inconsistent configuration (a chat template that does not compile among it). Runs once: a second call raises ValueError.

chat_template​

chat_template(self, jinja_source: 'str', bos_token: 'str', eos_token: 'str') -> 'TokenizerBuilder'

chat_template(jinja_source, bos_token, eos_token) -> TokenizerBuilder

Attach a chat template (Jinja source) for apply_chat_template and encode_chat; bos_token and eos_token are the framing tokens it references. The source compiles at :meth:build, which raises on an invalid one.

decoder​

decoder(self, decoder: 'Decoder') -> 'TokenizerBuilder'

decoder(decoder) -> TokenizerBuilder

Set the detokenization stage (taken).

eos_id​

eos_id(self, id: 'int') -> 'TokenizerBuilder'

eos_id(id) -> TokenizerBuilder

Set the end-of-sequence token id.

normalizer​

normalizer(self, normalizer: 'Normalizer') -> 'TokenizerBuilder'

normalizer(normalizer) -> TokenizerBuilder

Set the text-cleanup stage (taken).

pad_id​

pad_id(self, id: 'int') -> 'TokenizerBuilder'

pad_id(id) -> TokenizerBuilder

Set the padding token id.

post_processor​

post_processor(self, post_processor: 'PostProcessor') -> 'TokenizerBuilder'

post_processor(post_processor) -> TokenizerBuilder

Set the special-token stage (taken).

pre_tokenizer​

pre_tokenizer(self, pre_tokenizer: 'PreTokenizer') -> 'TokenizerBuilder'

pre_tokenizer(pre_tokenizer) -> TokenizerBuilder

Set the pre-segmentation stage (taken).