Skip to main content

Normalizer

The text-cleanup stage of a TokenizerBuilder: Unicode normalization, case folding, stripping, replacement. sequence() composes several in order. A builder or a sequence takes the stage; a taken stage refuses its next use with ValueError.

__init__​

__init__(self, /, *args, **kwargs)

Initialize self. See help(type(self)) for accurate signature.

bert​

bertbert(clean_text: bool = True, handle_chinese_chars: bool = True, strip_accents: bool | None = None, lowercase: bool = True) -> clika_runtime._core.tokenizer.Normalizer

bert(clean_text: bool = True, handle_chinese_chars: bool = True, strip_accents: bool | None = None, lowercase: bool = True) -> clika_runtime._core.tokenizer.Normalizer

bert(clean_text=True, handle_chinese_chars=True, strip_accents=None, lowercase=True) -> Normalizer

The BERT text cleaner; strip_accents=None follows lowercase.

lowercase​

lowercaselowercase() -> clika_runtime._core.tokenizer.Normalizer

lowercase() -> clika_runtime._core.tokenizer.Normalizer

lowercase() -> Normalizer

Lowercase the text.

nfc​

nfcnfc() -> clika_runtime._core.tokenizer.Normalizer

nfc() -> clika_runtime._core.tokenizer.Normalizer

nfc() -> Normalizer

Unicode NFC.

nfd​

nfdnfd() -> clika_runtime._core.tokenizer.Normalizer

nfd() -> clika_runtime._core.tokenizer.Normalizer

nfd() -> Normalizer

Unicode NFD.

nfkc​

nfkcnfkc() -> clika_runtime._core.tokenizer.Normalizer

nfkc() -> clika_runtime._core.tokenizer.Normalizer

nfkc() -> Normalizer

Unicode NFKC.

nfkd​

nfkdnfkd() -> clika_runtime._core.tokenizer.Normalizer

nfkd() -> clika_runtime._core.tokenizer.Normalizer

nfkd() -> Normalizer

Unicode NFKD.

nmt​

nmtnmt() -> clika_runtime._core.tokenizer.Normalizer

nmt() -> clika_runtime._core.tokenizer.Normalizer

nmt() -> Normalizer

The NMT text cleaner (control characters removed).

prepend​

prependprepend(prefix: str) -> clika_runtime._core.tokenizer.Normalizer

prepend(prefix: str) -> clika_runtime._core.tokenizer.Normalizer

prepend(prefix) -> Normalizer

Prepend prefix to a non-empty text.

replace​

replacereplace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Normalizer

replace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Normalizer

replace(pattern, content, is_regex=False) -> Normalizer

Replace every match of pattern (a regular expression when is_regex, else an exact string) with content; raises on a pattern that does not compile.

sequence​

sequencesequence(normalizers: collections.abc.Sequence[clika_runtime._core.tokenizer.Normalizer]) -> clika_runtime._core.tokenizer.Normalizer

sequence(normalizers: collections.abc.Sequence[clika_runtime._core.tokenizer.Normalizer]) -> clika_runtime._core.tokenizer.Normalizer

sequence(normalizers) -> Normalizer

Apply the given normalizers in order; each is taken.

strip​

stripstrip(left: bool = True, right: bool = True) -> clika_runtime._core.tokenizer.Normalizer

strip(left: bool = True, right: bool = True) -> clika_runtime._core.tokenizer.Normalizer

strip(left=True, right=True) -> Normalizer

Trim whitespace at the chosen ends.

strip_accents​

strip_accentsstrip_accents() -> clika_runtime._core.tokenizer.Normalizer

strip_accents() -> clika_runtime._core.tokenizer.Normalizer

strip_accents() -> Normalizer

Remove combining accents.