Skip to main content

Decoder

The detokenization stage of a TokenizerBuilder: turns token strings back into text (the byte-level, metaspace and WordPiece reverses, fuse, strip, replace, byte fallback). sequence() composes several in order. A builder or a sequence takes the stage; a taken stage refuses its next use with ValueError.

__init__​

__init__(self, /, *args, **kwargs)

Initialize self. See help(type(self)) for accurate signature.

byte_fallback​

byte_fallbackbyte_fallback() -> clika_runtime._core.tokenizer.Decoder

byte_fallback() -> clika_runtime._core.tokenizer.Decoder

byte_fallback() -> Decoder

Decode a byte token ('<0x0A>') back into its byte.

byte_level​

byte_levelbyte_level() -> clika_runtime._core.tokenizer.Decoder

byte_level() -> clika_runtime._core.tokenizer.Decoder

byte_level() -> Decoder

Reverse the GPT-2 byte-level map.

fuse​

fusefuse() -> clika_runtime._core.tokenizer.Decoder

fuse() -> clika_runtime._core.tokenizer.Decoder

fuse() -> Decoder

Join every token into one with no separator.

metaspace​

metaspacemetaspace(replacement: object = '▁', prepend: bool = True) -> clika_runtime._core.tokenizer.Decoder

metaspace(replacement: object = '▁', prepend: bool = True) -> clika_runtime._core.tokenizer.Decoder

metaspace(replacement='▁', prepend=True) -> Decoder

Turn the one-character replacement marker back into a space; prepend drops a leading marker.

replace​

replacereplace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Decoder

replace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Decoder

replace(pattern, content, is_regex=False) -> Decoder

Replace pattern (a regular expression when is_regex, else an exact string) with content on each token; raises on a pattern that does not compile.

sequence​

sequencesequence(decoders: collections.abc.Sequence[clika_runtime._core.tokenizer.Decoder]) -> clika_runtime._core.tokenizer.Decoder

sequence(decoders: collections.abc.Sequence[clika_runtime._core.tokenizer.Decoder]) -> clika_runtime._core.tokenizer.Decoder

sequence(decoders) -> Decoder

Apply the given decoders in order; each is taken.

strip​

stripstrip(content: object, start: int, stop: int) -> clika_runtime._core.tokenizer.Decoder

strip(content: object, start: int, stop: int) -> clika_runtime._core.tokenizer.Decoder

strip(content, start, stop) -> Decoder

Trim up to start leading and stop trailing occurrences of the one-character content from each token.

wordpiece​

wordpiecewordpiece(prefix: str = '##', cleanup: bool = True) -> clika_runtime._core.tokenizer.Decoder

wordpiece(prefix: str = '##', cleanup: bool = True) -> clika_runtime._core.tokenizer.Decoder

wordpiece(prefix='##', cleanup=True) -> Decoder

Strip the continuing-subword prefix and join; cleanup fixes the spacing around punctuation.