Skip to main content

//clika-runtime/io.clika.runtime/EncodeOptions

EncodeOptions

[common]
data class EncodeOptions(val addSpecialTokens: Boolean = true, val returnTensors: Boolean = false, val varlen: Boolean = false, val padId: Int? = null, val device: Placement? = null, val maxLength: Long? = null, val truncationSide: TruncationSide = TruncationSide.RIGHT, val paddingSide: PaddingSide = PaddingSide.RIGHT, val padToMultipleOf: Long? = null)

The per-call controls of the tensor-shaped encode entry points (Tokenizer.encodeText, Tokenizer.encodeBatch).

Constructors​

EncodeOptions[common]
constructor(addSpecialTokens: Boolean = true, returnTensors: Boolean = false, varlen: Boolean = false, padId: Int? = null, device: Placement? = null, maxLength: Long? = null, truncationSide: TruncationSide = TruncationSide.RIGHT, paddingSide: PaddingSide = PaddingSide.RIGHT, padToMultipleOf: Long? = null)

Properties​

NameSummary
addSpecialTokens[common]
val addSpecialTokens: Boolean = true
wrap the ids with the model's special tokens (bos, eos, ...).
device[common]
val device: Placement? = null
where the tensors land: a Device, or a Stream so they land stream-ordered on it; null takes the tokenizer's default placement.
maxLength[common]
val maxLength: Long? = null
cap each sequence at this many tokens, trimming the text per truncationSide before the special tokens attach, so the total stays within the cap; null means no truncation.
paddingSide[common]
val paddingSide: PaddingSide
which end a padded batch fills; the lengths always carry the real counts.
padId[common]
val padId: Int? = null
the pad id of a padded batch; null takes the tokenizer's own.
padToMultipleOf[common]
val padToMultipleOf: Long? = null
round the padded row length up to a multiple of this (padded layout only; refused with varlen).
returnTensors[common]
val returnTensors: Boolean = false
build the tensor columns of Encoded beside the id rows.
truncationSide[common]
val truncationSide: TruncationSide
varlen[common]
val varlen: Boolean = false
lay inputIds out as one flat row with a cuSeqlens offset table instead of a padded [B, S] matrix; varlen never pads.