Normalizer
The text-cleanup stage of a TokenizerBuilder: Unicode normalization, case folding, stripping, replacement. sequence() composes several in order. A builder or a sequence takes the stage; a taken stage refuses its next use with ValueError.
__init__
__init__(self, /, *args, **kwargs)
Initialize self. See help(type(self)) for accurate signature.
bert
bertbert(clean_text: bool = True, handle_chinese_chars: bool = True, strip_accents: bool | None = None, lowercase: bool = True) -> clika_runtime._core.tokenizer.Normalizer
bert(clean_text: bool = True, handle_chinese_chars: bool = True, strip_accents: bool | None = None, lowercase: bool = True) -> clika_runtime._core.tokenizer.Normalizer
bert(clean_text=True, handle_chinese_chars=True, strip_accents=None, lowercase=True) -> Normalizer
The BERT text cleaner; strip_accents=None follows lowercase.
lowercase
lowercaselowercase() -> clika_runtime._core.tokenizer.Normalizer
lowercase() -> clika_runtime._core.tokenizer.Normalizer
lowercase() -> Normalizer
Lowercase the text.
nfc
nfcnfc() -> clika_runtime._core.tokenizer.Normalizer
nfc() -> clika_runtime._core.tokenizer.Normalizer
nfc() -> Normalizer
Unicode NFC.
nfd
nfdnfd() -> clika_runtime._core.tokenizer.Normalizer
nfd() -> clika_runtime._core.tokenizer.Normalizer
nfd() -> Normalizer
Unicode NFD.
nfkc
nfkcnfkc() -> clika_runtime._core.tokenizer.Normalizer
nfkc() -> clika_runtime._core.tokenizer.Normalizer
nfkc() -> Normalizer
Unicode NFKC.
nfkd
nfkdnfkd() -> clika_runtime._core.tokenizer.Normalizer
nfkd() -> clika_runtime._core.tokenizer.Normalizer
nfkd() -> Normalizer
Unicode NFKD.
nmt
nmtnmt() -> clika_runtime._core.tokenizer.Normalizer
nmt() -> clika_runtime._core.tokenizer.Normalizer
nmt() -> Normalizer
The NMT text cleaner (control characters removed).
prepend
prependprepend(prefix: str) -> clika_runtime._core.tokenizer.Normalizer
prepend(prefix: str) -> clika_runtime._core.tokenizer.Normalizer
prepend(prefix) -> Normalizer
Prepend prefix to a non-empty text.
replace
replacereplace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Normalizer
replace(pattern: str, content: str, is_regex: bool = False) -> clika_runtime._core.tokenizer.Normalizer
replace(pattern, content, is_regex=False) -> Normalizer
Replace every match of pattern (a regular expression when is_regex, else an exact string) with content; raises on a pattern that does not compile.
sequence
sequencesequence(normalizers: collections.abc.Sequence[clika_runtime._core.tokenizer.Normalizer]) -> clika_runtime._core.tokenizer.Normalizer
sequence(normalizers: collections.abc.Sequence[clika_runtime._core.tokenizer.Normalizer]) -> clika_runtime._core.tokenizer.Normalizer
sequence(normalizers) -> Normalizer
Apply the given normalizers in order; each is taken.
strip
stripstrip(left: bool = True, right: bool = True) -> clika_runtime._core.tokenizer.Normalizer
strip(left: bool = True, right: bool = True) -> clika_runtime._core.tokenizer.Normalizer
strip(left=True, right=True) -> Normalizer
Trim whitespace at the chosen ends.
strip_accents
strip_accentsstrip_accents() -> clika_runtime._core.tokenizer.Normalizer
strip_accents() -> clika_runtime._core.tokenizer.Normalizer
strip_accents() -> Normalizer
Remove combining accents.