---
title: "Companion"
sidebar_label: "Companion"
description: "Kotlin binding reference: Companion."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../../index.md)/[io.clika.runtime](../../index.md)/[Tokenizer](../index.md)/[Companion](index.md)

# Companion

[common]\
object [Companion](index.md)

## Functions

| Name | Summary |
|---|---|
| [fromBertVocab](fromBertVocab.md) | [common]<br>fun [fromBertVocab](fromBertVocab.md)(directory: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load a BERT WordPiece tokenizer from a directory holding a bare `vocab.txt`. |
| [fromFile](fromFile.md) | [common]<br>fun [fromFile](fromFile.md)(path: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load any supported tokenizer artifact, detecting its format: a directory is probed for its artifact, a file is read by its name or its content. A path no format claims fails naming every format tried. |
| [fromGpt2](fromGpt2.md) | [common]<br>fun [fromGpt2](fromGpt2.md)(directory: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load a GPT-2 style byte-level BPE tokenizer from a directory holding `vocab.json` and `merges.txt`. |
| [fromHuggingface](fromHuggingface.md) | [common]<br>fun [fromHuggingface](fromHuggingface.md)(path: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load a Hugging Face tokenizer from a model directory or one artifact file, with the sibling `tokenizer_config.json` read for the special ids and the chat template when present. |
| [fromSentencepiece](fromSentencepiece.md) | [common]<br>fun [fromSentencepiece](fromSentencepiece.md)(path: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load a bare SentencePiece model file: the pieces keep their own ids, nothing frames a sequence. |
| [fromTekken](fromTekken.md) | [common]<br>fun [fromTekken](fromTekken.md)(path: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load a Mistral `tekken.json` tokenizer, which carries its own special tokens. |
| [fromTiktoken](fromTiktoken.md) | [common]<br>fun [fromTiktoken](fromTiktoken.md)(path: [String](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-string/index.html)): [Tokenizer](../index.md)<br>Load an OpenAI `.tiktoken` ranks file; the special tokens of the family its table size names are added (a size outside the known families fails naming them). |