TokenizerBuilder
TokenizerBuilder(model, *, device=None) -> TokenizerBuilder
Assemble a :class:`Tokenizer` from a :class:`Model` (the core stage, built
from an explicit vocabulary) plus the optional :class:`Normalizer`,
:class:`PreTokenizer`, :class:`PostProcessor` and :class:`Decoder` stages,
added tokens, the well-known ids and a chat template, with no tokenizer
file. Every setter takes its stage (a taken stage refuses its next use)
and returns the builder, so a build reads as one chain ending in
:meth:`build`, which checks the configuration and runs once::
tokenizer = (
TokenizerBuilder(Model.word_level({"hello": 0, "[UNK]": 1}, "[UNK]"))
.pre_tokenizer(PreTokenizer.whitespace_split())
.decoder(Decoder.fuse())
.build()
)
``device`` (a Device, a device string or a Stream) is the default
placement of the tokenizer's tensor-shaped entries; None is the
runtime's default.
__init__
__init__(self, model: 'Model', *, device: 'Device | Stream | str | None' = None) -> 'None'
Initialize self. See help(type(self)) for accurate signature.
add_special_token
add_special_token(self, content: 'str', id: 'int') -> 'TokenizerBuilder'
add_special_token(content, id) -> TokenizerBuilder
Register a special added token under id: matched before the model
runs, and skippable on decode.
add_token
add_token(self, content: 'str', id: 'int', *, single_word: 'bool' = False, lstrip: 'bool' = False, rstrip: 'bool' = False, normalized: 'bool' = True) -> 'TokenizerBuilder'
add_token(content, id, *, single_word=False, lstrip=False, rstrip=False, normalized=True) -> TokenizerBuilder
Register an ordinary added token: single_word matches at word
boundaries only, lstrip / rstrip consume the whitespace beside
it, normalized matches the normalized text.
bos_id
bos_id(self, id: 'int') -> 'TokenizerBuilder'
bos_id(id) -> TokenizerBuilder
Set the beginning-of-sequence token id.
build
build(self) -> 'Tokenizer'
build() -> Tokenizer
Assemble the tokenizer; raises on an inconsistent configuration (a
chat template that does not compile among it). Runs once: a second
call raises ValueError.
chat_template
chat_template(self, jinja_source: 'str', bos_token: 'str', eos_token: 'str') -> 'TokenizerBuilder'
chat_template(jinja_source, bos_token, eos_token) -> TokenizerBuilder
Attach a chat template (Jinja source) for apply_chat_template and
encode_chat; bos_token and eos_token are the framing tokens
it references. The source compiles at :meth:build, which raises on
an invalid one.
decoder
decoder(self, decoder: 'Decoder') -> 'TokenizerBuilder'
decoder(decoder) -> TokenizerBuilder
Set the detokenization stage (taken).
eos_id
eos_id(self, id: 'int') -> 'TokenizerBuilder'
eos_id(id) -> TokenizerBuilder
Set the end-of-sequence token id.
normalizer
normalizer(self, normalizer: 'Normalizer') -> 'TokenizerBuilder'
normalizer(normalizer) -> TokenizerBuilder
Set the text-cleanup stage (taken).
pad_id
pad_id(self, id: 'int') -> 'TokenizerBuilder'
pad_id(id) -> TokenizerBuilder
Set the padding token id.
post_processor
post_processor(self, post_processor: 'PostProcessor') -> 'TokenizerBuilder'
post_processor(post_processor) -> TokenizerBuilder
Set the special-token stage (taken).
pre_tokenizer
pre_tokenizer(self, pre_tokenizer: 'PreTokenizer') -> 'TokenizerBuilder'
pre_tokenizer(pre_tokenizer) -> TokenizerBuilder
Set the pre-segmentation stage (taken).