//clika-runtime/io.clika.runtime/Tokenizer/Companion
Companion
[common]
object Companion
Functions
| Name | Summary |
|---|---|
| fromBertVocab | [common] fun fromBertVocab(directory: String): Tokenizer Load a BERT WordPiece tokenizer from a directory holding a bare vocab.txt. |
| fromFile | [common] fun fromFile(path: String): Tokenizer Load any supported tokenizer artifact, detecting its format: a directory is probed for its artifact, a file is read by its name or its content. A path no format claims fails naming every format tried. |
| fromGpt2 | [common] fun fromGpt2(directory: String): Tokenizer Load a GPT-2 style byte-level BPE tokenizer from a directory holding vocab.json and merges.txt. |
| fromHuggingface | [common] fun fromHuggingface(path: String): Tokenizer Load a Hugging Face tokenizer from a model directory or one artifact file, with the sibling tokenizer_config.json read for the special ids and the chat template when present. |
| fromSentencepiece | [common] fun fromSentencepiece(path: String): Tokenizer Load a bare SentencePiece model file: the pieces keep their own ids, nothing frames a sequence. |
| fromTekken | [common] fun fromTekken(path: String): Tokenizer Load a Mistral tekken.json tokenizer, which carries its own special tokens. |
| fromTiktoken | [common] fun fromTiktoken(path: String): Tokenizer Load an OpenAI .tiktoken ranks file; the special tokens of the family its table size names are added (a size outside the known families fails naming them). |