Skip to main content

PostProcessor

The special-token stage of a TokenizerBuilder: wraps an encoded sequence with special tokens ('[CLS] $A [SEP]' templates, the RoBERTa framing) or re-encodes byte-level offsets. sequence() composes several in order. A builder or a sequence takes the stage; a taken stage refuses its next use with ValueError.

__init__​

__init__(self, /, *args, **kwargs)

Initialize self. See help(type(self)) for accurate signature.

byte_level​

byte_levelbyte_level(add_prefix_space: bool = True, trim_offsets: bool = True) -> clika_runtime._core.tokenizer.PostProcessor

byte_level(add_prefix_space: bool = True, trim_offsets: bool = True) -> clika_runtime._core.tokenizer.PostProcessor

byte_level(add_prefix_space=True, trim_offsets=True) -> PostProcessor

Re-encode byte-level offsets after the model (the GPT-2 and RoBERTa shape).

roberta​

robertaroberta(sep_token: str, sep_id: int, cls_token: str, cls_id: int, trim_offsets: bool = True, add_prefix_space: bool = True) -> clika_runtime._core.tokenizer.PostProcessor

roberta(sep_token: str, sep_id: int, cls_token: str, cls_id: int, trim_offsets: bool = True, add_prefix_space: bool = True) -> clika_runtime._core.tokenizer.PostProcessor

roberta(sep_token, sep_id, cls_token, cls_id, trim_offsets=True, add_prefix_space=True) -> PostProcessor

The RoBERTa ' ... ' framing from the separator and classifier tokens and their ids.

sequence​

sequencesequence(processors: collections.abc.Sequence[clika_runtime._core.tokenizer.PostProcessor]) -> clika_runtime._core.tokenizer.PostProcessor

sequence(processors: collections.abc.Sequence[clika_runtime._core.tokenizer.PostProcessor]) -> clika_runtime._core.tokenizer.PostProcessor

sequence(processors) -> PostProcessor

Apply the given post-processors in order; each is taken.

template_processing​

template_processingtemplate_processing(single: object, pair: object = (), special_tokens: object = ()) -> clika_runtime._core.tokenizer.PostProcessor

template_processing(single: object, pair: object = (), special_tokens: object = ()) -> clika_runtime._core.tokenizer.PostProcessor

template_processing(single, pair=(), special_tokens=()) -> PostProcessor

Assemble a sequence from a template: single for one sequence and pair for two, each a str of space-separated pieces ('[CLS] A[SEP]′)orasequenceofpieces,where′A [SEP]') or a sequence of pieces, where 'A' and '$B' are the sequences and any other piece a special token's key. special_tokens maps each key to its id (one token under its own key) or to a (tokens, ids) pair, as a dict or a sequence of pairs. A piece whose key the map lacks adds nothing to the sequence.