Tokenizer
Tokenizer.from_file(path) / Tokenizer.from_huggingface(source) -> Tokenizer
A tokenizer over the runtime's tokenizer object: encode, decode,
encode_batch, decode_batch, token_to_id, id_to_token,
streaming_decoder and the bos_id / eos_id / pad_id /
vocab_size / has_chat_template properties answer as the runtime
object does; :meth:apply_chat_template and :meth:encode_chat also
take the conversation as a list of message dicts. Wrap a tokenizer a
model handed you with :meth:wrap; the runtime object itself is
:attr:bound.
bound (property)
The runtime tokenizer object this wrapper drives.
__init__
__init__(self, bound: '_core_tokenizer.Tokenizer') -> 'None'
Initialize self. See help(type(self)) for accurate signature.
apply_chat_template
apply_chat_template(self, messages: 'Messages', add_generation_prompt: 'bool' = True, *, now_epoch_seconds: 'int | None' = None) -> 'str'
apply_chat_template(messages, add_generation_prompt=True, *, now_epoch_seconds=None) -> str
Render the chat template over the conversation, given as a list of
{"role": ..., "content": ...} dicts or as the same conversation in
JSON text. Returns the prompt string. Templates that print the date
(Llama 3.2's system header) render it from now_epoch_seconds, a
Unix time formatted as UTC; the default renders at epoch 0, so a
rendering is reproducible. Pass int(time.time()) for the wall
clock.
encode_chat
encode_chat(self, messages: 'Messages', add_generation_prompt: 'bool' = True, add_special_tokens: 'bool' = False, *, now_epoch_seconds: 'int | None' = None) -> 'list[int]'
encode_chat(messages, add_generation_prompt=True, add_special_tokens=False, *, now_epoch_seconds=None) -> list[int]
Render the chat template and encode the result in one call.
messages takes the same two forms as :meth:apply_chat_template,
and now_epoch_seconds is the same clock (epoch 0 by default, so a
rendering is reproducible). add_special_tokens is False by
default: the template writes its own bos/eos framing, so the
tokenizer's post-processor must not add a second bos. Pass True
only for a template that writes no bos itself.
from_file
from_filefrom_file(path) -> Tokenizer
from_file(path) -> Tokenizer
Load a tokenizer from a tokenizer.json file.
from_huggingface
from_huggingfacefrom_huggingface(path) -> Tokenizer
from_huggingface(path) -> Tokenizer
Load a tokenizer from a tokenizer.json (or a snapshot directory
holding one), with the tokenizer_config.json overlay applied.
wrap
wrapwrap(bound) -> Tokenizer
wrap(bound) -> Tokenizer
The wrapper over a runtime tokenizer object (a wrapper passes through).