Model
The core tokenization stage of a TokenizerBuilder, built from an explicit vocabulary with no file: bpe, unigram, word_level or word_piece. A vocabulary is a dict of token to id (insertion order) or a sequence of (token, id) pairs. A builder takes the model; a taken model refuses its next use with ValueError.
__init__
__init__(self, /, *args, **kwargs)
Initialize self. See help(type(self)) for accurate signature.
bpe
bpebpe(vocab: object, merges: object, *, unk_token: str | None = None, continuing_subword_prefix: str | None = None, end_of_word_suffix: str | None = None, fuse_unk: bool = False, byte_fallback: bool = False, ignore_merges: bool = False) -> clika_runtime._core.tokenizer.Model
bpe(vocab: object, merges: object, *, unk_token: str | None = None, continuing_subword_prefix: str | None = None, end_of_word_suffix: str | None = None, fuse_unk: bool = False, byte_fallback: bool = False, ignore_merges: bool = False) -> clika_runtime._core.tokenizer.Model
bpe(vocab, merges, *, unk_token=None, continuing_subword_prefix=None, end_of_word_suffix=None, fuse_unk=False, byte_fallback=False, ignore_merges=False) -> Model
Byte-pair encoding over vocab and merges, a sequence of (left, right) pairs ranked by position. unk_token names the unknown token (None: unknown pieces are dropped); continuing_subword_prefix marks a continuing subword ('##' style) and end_of_word_suffix the end of a word, None for none; fuse_unk fuses consecutive unknown pieces into one; byte_fallback decomposes an unknown piece into byte tokens; ignore_merges skips the merge loop for a word already in the vocabulary. Raises on a merge naming a token the vocabulary lacks.
unigram
unigramunigram(vocab: object, *, unk_id: int | None = None, byte_fallback: bool = False) -> clika_runtime._core.tokenizer.Model
unigram(vocab: object, *, unk_id: int | None = None, byte_fallback: bool = False) -> clika_runtime._core.tokenizer.Model
unigram(vocab, *, unk_id=None, byte_fallback=False) -> Model
A unigram language model over vocab, a dict of piece to log-probability score or (piece, score) pairs; a higher score is preferred. unk_id is the unknown token's id (None: none); byte_fallback decomposes an unknown piece into byte tokens.
word_level
word_levelword_level(vocab: object, unk_token: str) -> clika_runtime._core.tokenizer.Model
word_level(vocab: object, unk_token: str) -> clika_runtime._core.tokenizer.Model
word_level(vocab, unk_token) -> Model
Whole-word lookup in vocab; a word outside it maps to unk_token.
word_piece
word_pieceword_piece(vocab: object, unk_token: str, continuing_subword_prefix: str = '##', max_input_chars_per_word: int = 100) -> clika_runtime._core.tokenizer.Model
word_piece(vocab: object, unk_token: str, continuing_subword_prefix: str = '##', max_input_chars_per_word: int = 100) -> clika_runtime._core.tokenizer.Model
word_piece(vocab, unk_token, continuing_subword_prefix='##', max_input_chars_per_word=100) -> Model
Greedy longest-match WordPiece over vocab; continuing_subword_prefix marks a non-initial subword, and a word longer than max_input_chars_per_word maps to unk_token.