---
title: "Model"
sidebar_label: "Model"
description: "The clika_runtime.tokenizer Model class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

The core tokenization stage of a TokenizerBuilder, built from an explicit vocabulary with no file: bpe, unigram, word_level or word_piece. A vocabulary is a dict of token to id (insertion order) or a sequence of (token, id) pairs. A builder takes the model; a taken model refuses its next use with ValueError.

## `__init__`

```python
__init__(self, /, *args, **kwargs)
```

Initialize self.  See help(type(self)) for accurate signature.

## `bpe`

```python
bpebpe(vocab: object, merges: object, *, unk_token: str | None = None, continuing_subword_prefix: str | None = None, end_of_word_suffix: str | None = None, fuse_unk: bool = False, byte_fallback: bool = False, ignore_merges: bool = False) -> clika_runtime._core.tokenizer.Model
```

bpe(vocab: object, merges: object, *, unk_token: str | None = None, continuing_subword_prefix: str | None = None, end_of_word_suffix: str | None = None, fuse_unk: bool = False, byte_fallback: bool = False, ignore_merges: bool = False) -> clika_runtime._core.tokenizer.Model

bpe(vocab, merges, *, unk_token=None, continuing_subword_prefix=None, end_of_word_suffix=None, fuse_unk=False, byte_fallback=False, ignore_merges=False) -> Model

Byte-pair encoding over `vocab` and `merges`, a sequence of (left, right) pairs ranked by position. `unk_token` names the unknown token (None: unknown pieces are dropped); `continuing_subword_prefix` marks a continuing subword ('##' style) and `end_of_word_suffix` the end of a word, None for none; fuse_unk fuses consecutive unknown pieces into one; byte_fallback decomposes an unknown piece into byte tokens; ignore_merges skips the merge loop for a word already in the vocabulary. Raises on a merge naming a token the vocabulary lacks.

## `unigram`

```python
unigramunigram(vocab: object, *, unk_id: int | None = None, byte_fallback: bool = False) -> clika_runtime._core.tokenizer.Model
```

unigram(vocab: object, *, unk_id: int | None = None, byte_fallback: bool = False) -> clika_runtime._core.tokenizer.Model

unigram(vocab, *, unk_id=None, byte_fallback=False) -> Model

A unigram language model over `vocab`, a dict of piece to log-probability score or (piece, score) pairs; a higher score is preferred. `unk_id` is the unknown token's id (None: none); byte_fallback decomposes an unknown piece into byte tokens.

## `word_level`

```python
word_levelword_level(vocab: object, unk_token: str) -> clika_runtime._core.tokenizer.Model
```

word_level(vocab: object, unk_token: str) -> clika_runtime._core.tokenizer.Model

word_level(vocab, unk_token) -> Model

Whole-word lookup in `vocab`; a word outside it maps to `unk_token`.

## `word_piece`

```python
word_pieceword_piece(vocab: object, unk_token: str, continuing_subword_prefix: str = '##', max_input_chars_per_word: int = 100) -> clika_runtime._core.tokenizer.Model
```

word_piece(vocab: object, unk_token: str, continuing_subword_prefix: str = '##', max_input_chars_per_word: int = 100) -> clika_runtime._core.tokenizer.Model

word_piece(vocab, unk_token, continuing_subword_prefix='##', max_input_chars_per_word=100) -> Model

Greedy longest-match WordPiece over `vocab`; `continuing_subword_prefix` marks a non-initial subword, and a word longer than `max_input_chars_per_word` maps to `unk_token`.
