//clika-runtime/io.clika.runtime/Ops
Ops
[common]
object Ops
The generated operator surface; see the file banner for its laws.
Functions
| Name | Summary |
|---|---|
| abs | [common] fun abs(input: Tensor): Tensor abs(input: Tensor): the abs operator. Elementwise absolute value. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| absInPlace | [common] fun absInPlace(self: Tensor): Tensor absInPlace(self: Tensor): the abs_ operator. In-place abs: writes the result through self; same formula, arguments, and error conditions as abs(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| acos | [common] fun acos(input: Tensor): Tensor acos(input: Tensor): the acos operator. Elementwise arccosine. Inputs outside [-1, 1] produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| acosh | [common] fun acosh(input: Tensor): Tensor acosh(input: Tensor): the acosh operator. Elementwise inverse hyperbolic cosine. Inputs below 1 produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| acoshInPlace | [common] fun acoshInPlace(self: Tensor): Tensor acoshInPlace(self: Tensor): the acosh_ operator. In-place acosh: writes the result through self; same formula, arguments, and error conditions as acosh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| acosInPlace | [common] fun acosInPlace(self: Tensor): Tensor acosInPlace(self: Tensor): the acos_ operator. In-place acos: writes the result through self; same formula, arguments, and error conditions as acos(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| adaptiveAvgPool | [common] fun adaptiveAvgPool(input: Tensor, outputSize: LongArray): Tensor adaptiveAvgPool(input: Tensor, outputSize: LongArray): the adaptive_avg_pool operator. Rank-generic adaptive average pooling to a target output size, channels-last. The window geometry is DERIVED per output position so the spatial dims land exactly on output_size; no kernel/stride/padding to pick. |
| adaptiveAvgPool1d | [common] fun adaptiveAvgPool1d(input: Tensor, outputSize: LongArray): Tensor adaptiveAvgPool1d(input: Tensor, outputSize: LongArray): the adaptive_avg_pool1d operator. 1-D adaptive average pooling of [N, L, C] to [N, L', C]; the window geometry is derived from output_size. See adaptive_avg_pool. |
| adaptiveAvgPool2d | [common] fun adaptiveAvgPool2d(input: Tensor, outputSize: LongArray): Tensor adaptiveAvgPool2d(input: Tensor, outputSize: LongArray): the adaptive_avg_pool2d operator. 2-D adaptive average pooling of [N, H, W, C] to [N, H', W', C]; windows derived so the output lands exactly on output_size = {H', W'}. {1, 1} is global average pooling. |
| adaptiveAvgPool3d | [common] fun adaptiveAvgPool3d(input: Tensor, outputSize: LongArray): Tensor adaptiveAvgPool3d(input: Tensor, outputSize: LongArray): the adaptive_avg_pool3d operator. 3-D adaptive average pooling of [N, D, H, W, C] to [N, D', H', W', C]. See adaptive_avg_pool. |
| adaptiveMaxPool | [common] fun adaptiveMaxPool(input: Tensor, outputSize: LongArray): Tensor adaptiveMaxPool(input: Tensor, outputSize: LongArray): the adaptive_max_pool operator. Rank-generic adaptive MAX pooling to a target output size, channels-last, the max sibling of adaptive_avg_pool. A window holding any NaN element yields NaN. |
| adaptiveMaxPool1d | [common] fun adaptiveMaxPool1d(input: Tensor, outputSize: LongArray): Tensor adaptiveMaxPool1d(input: Tensor, outputSize: LongArray): the adaptive_max_pool1d operator. 1-D adaptive max pooling of [N, L, C] to [N, L', C]. See adaptive_max_pool. |
| adaptiveMaxPool2d | [common] fun adaptiveMaxPool2d(input: Tensor, outputSize: LongArray): Tensor adaptiveMaxPool2d(input: Tensor, outputSize: LongArray): the adaptive_max_pool2d operator. 2-D adaptive max pooling of [N, H, W, C] to [N, H', W', C]. {1, 1} is global max pooling. See adaptive_max_pool. |
| adaptiveMaxPool3d | [common] fun adaptiveMaxPool3d(input: Tensor, outputSize: LongArray): Tensor adaptiveMaxPool3d(input: Tensor, outputSize: LongArray): the adaptive_max_pool3d operator. 3-D adaptive max pooling of [N, D, H, W, C] to [N, D', H', W', C]. See adaptive_max_pool. |
| add | [common] fun add(input: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor add(input: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the add operator. Adds other (scaled) to input elementwise. Broadcasting and type promotion, the contract every binary op on this surface shares: operand shapes broadcast per the standard rules (trailing dims align; a 1 stretches), and dtypes promote to the dominant operand dtype per the promotion lattice (a scalar other keeps its weak kind; an integer literal with an integer tensor stays integral, a double promotes weak-float). The other ops below state "broadcasts and promotes as add" instead of restating this.[common] fun add(input: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor add(input: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the number form of add, other as a scalar.[common] fun add(input: Double, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor add(input: Double, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the add operator. Scalar-LHS . input keeps its kind: an integer scalar with an integer tensor stays integral (weak-int promotion); a double promotes weak-float. The parameter set mirrors the tensor-first form: alpha scales the TENSOR operand other, activation applies to the result. |
| addInPlace | [common] fun addInPlace(self: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor addInPlace(self: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the add_ operator. In-place add: writes through x; semantics as ops::add (which also documents broadcasting/promotion). Writes through self and returns it, so calls chain.[common] fun addInPlace(self: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor addInPlace(self: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the number form of add_, other as a scalar. |
| addLayerNorm | [common] fun addLayerNorm(input: Tensor, residual: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): List<Tensor> addLayerNorm(input: Tensor, residual: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): the add_layer_norm operator. Fused residual-add + normalization, on either side of the norm. Which data flow runs is inferred from which addends you supply; there is no mode flag: * residual only → {ACT(norm(x + residual)·w + b), x + residual} * post_residual only → {ACT(norm(x)·w + b) + post_residual, UNDEFINED} * both → {ACT(norm(x + residual)·w + b) + post_residual, x + residual} * neither → raises (that is a plain rms_norm/layer_norm) The second result is the PRE-norm sum, the residual stream the next sub-layer reads. Supplying only post_residual forms no such sum, so that element comes back UNDEFINED: read only the first one in that flow. The activation applies to the norm result, beforepost_residual is added; it is the norm's epilogue, not the sum's. Gated activations (SwiGlu / GeGlu / ReGlu) narrow their input and are rejected. |
| addRmsNorm | [common] fun addRmsNorm(input: Tensor, residual: Tensor? = null, residual2: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): List<Tensor> addRmsNorm(input: Tensor, residual: Tensor? = null, residual2: Tensor? = null, postResidual: Tensor? = null, normalizedShape: LongArray = longArrayOf(), skipBias: Tensor? = null, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): the add_rms_norm operator. Fused residual-add + normalization, on either side of the norm. Which data flow runs is inferred from which addends you supply; there is no mode flag: * residual only → {ACT(norm(x + residual)·w + b), x + residual} * post_residual only → {ACT(norm(x)·w + b) + post_residual, UNDEFINED} * both → {ACT(norm(x + residual)·w + b) + post_residual, x + residual} * neither → raises (that is a plain rms_norm/layer_norm) The second result is the PRE-norm sum, the residual stream the next sub-layer reads. Supplying only post_residual forms no such sum, so that element comes back UNDEFINED: read only the first one in that flow. The activation applies to the norm result, beforepost_residual is added; it is the norm's epilogue, not the sum's. Gated activations (SwiGlu / GeGlu / ReGlu) narrow their input and are rejected. |
| all | [common] fun all(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor all(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the all operator. True where EVERY element over dims is nonzero (logical AND reduce). Empty dims reduces every dimension. Output dtype is always Bool. An element is nonzero exactly when its Bool cast is true: +0 and -0 are zero; a subnormal, an infinity and a NaN are nonzero. |
| allclose | [common] fun allclose(input: Tensor, other: Tensor, rtol: Double = 1.0E-5, atol: Double = 1.0E-8, equalNan: Boolean = false): Tensor allclose(input: Tensor, other: Tensor, rtol: Double = 1e-5, atol: Double = 1e-8, equalNan: Boolean = false): the allclose operator. Whether EVERY elementwise pair is approximately equal, a reduction to one verdict. Shapes broadcast per the standard rules before the reduction. |
| amax | [common] fun amax(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor amax(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the amax operator. Maximum value of input over dims. Empty dims reduces every dimension. Returns VALUES (for the positions use argmax). |
| amin | [common] fun amin(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor amin(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the amin operator. Minimum value of input over dims. Empty dims reduces every dimension. Returns VALUES (for the positions use argmin). |
| aminmax | [common] fun aminmax(input: Tensor, dim: Long? = null, keepdim: Boolean = false): List<Tensor> aminmax(input: Tensor, dim: Long? = null, keepdim: Boolean = false): the aminmax operator. |
| any | [common] fun any(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor any(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the any operator. True where ANY element over dims is nonzero (logical OR reduce). Empty dims reduces every dimension. Output dtype is always Bool. An element is nonzero exactly when its Bool cast is true: +0 and -0 are zero; a subnormal, an infinity and a NaN are nonzero. |
| arange | [common] fun arange(start: Tensor, end: Tensor, step: Tensor? = null, dtype: DType = DType.INT64, device: Placement? = null): Tensor arange(start: Tensor, end: Tensor, step: Tensor? = null, dtype: DType = DType.INT64, device: Placement? = null): the arange operator. Evenly stepped 1-D range over the half-open interval [start, end). out[i] = start + i * step, for ceil((end - start) / step) elements; end itself is never included. Each bound is an int/float literal OR a 0-D Tensor (a tensor bound traces symbolically, never syncing to the host).[common] fun arange(start: Double, end: Double, step: Double = 1.0, dtype: DType = DType.INT64, device: Placement? = null): Tensor arange(start: Double, end: Double, step: Double = 1.0, dtype: DType = DType.INT64, device: Placement? = null): the number form of arange, start, end, step as scalars. |
| argmax | [common] fun argmax(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, indexDtype: DType = DType.INT64): Tensor argmax(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, indexDtype: DType = DType.INT64): the argmax operator. Index of the maximum of input over dims. Empty dims reduces every dimension (the index is then into the flattened tensor). The index dtype is Int64 by default; Int32 narrows it (an Int32 index addresses a reduced extent of at most INT32_MAX elements). |
| argmin | [common] fun argmin(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, indexDtype: DType = DType.INT64): Tensor argmin(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, indexDtype: DType = DType.INT64): the argmin operator. Index of the minimum of input over dims. Empty dims reduces every dimension (the index is then into the flattened tensor). The index dtype is Int64 by default; Int32 narrows it (an Int32 index addresses a reduced extent of at most INT32_MAX elements). |
| argsort | [common] fun argsort(input: Tensor, dim: Long = -1L, descending: Boolean = false, stable: Boolean = false): Tensor argsort(input: Tensor, dim: Long = -1L, descending: Boolean = false, stable: Boolean = false): the argsort operator. Indices that would sort input along dim (ascending unless descending). |
| asin | [common] fun asin(input: Tensor): Tensor asin(input: Tensor): the asin operator. Elementwise arcsine. Inputs outside [-1, 1] produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| asinh | [common] fun asinh(input: Tensor): Tensor asinh(input: Tensor): the asinh operator. Elementwise inverse hyperbolic sine. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| asinhInPlace | [common] fun asinhInPlace(self: Tensor): Tensor asinhInPlace(self: Tensor): the asinh_ operator. In-place asinh: writes the result through self; same formula, arguments, and error conditions as asinh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| asinInPlace | [common] fun asinInPlace(self: Tensor): Tensor asinInPlace(self: Tensor): the asin_ operator. In-place asin: writes the result through self; same formula, arguments, and error conditions as asin(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| atan | [common] fun atan(input: Tensor): Tensor atan(input: Tensor): the atan operator. Elementwise arctangent. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| atan2 | [common] fun atan2(input: Tensor, other: Tensor): Tensor atan2(input: Tensor, other: Tensor): the atan2 operator. Elementwise four-quadrant arc tangent of a/b (the angle of the point (b, a)), in radians. Broadcasts and promotes as add. |
| atan2InPlace | [common] fun atan2InPlace(self: Tensor, other: Tensor): Tensor atan2InPlace(self: Tensor, other: Tensor): the atan2_ operator. In-place atan2: writes the angles through x. Writes through self and returns it, so calls chain. |
| atanh | [common] fun atanh(input: Tensor): Tensor atanh(input: Tensor): the atanh operator. Elementwise inverse hyperbolic tangent. Inputs outside (-1, 1) produce NaN / ±inf. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| atanhInPlace | [common] fun atanhInPlace(self: Tensor): Tensor atanhInPlace(self: Tensor): the atanh_ operator. In-place atanh: writes the result through self; same formula, arguments, and error conditions as atanh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| atanInPlace | [common] fun atanInPlace(self: Tensor): Tensor atanInPlace(self: Tensor): the atan_ operator. In-place atan: writes the result through self; same formula, arguments, and error conditions as atan(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| atleast1d | [common] fun atleast1d(input: Tensor): Tensor atleast1d(input: Tensor): the atleast_1d operator. input with leading size-1 dims prepended until rank >= 1; a higher-rank input passes through unchanged. Returns a VIEW sharing input's storage (no copy). |
| atleast2d | [common] fun atleast2d(input: Tensor): Tensor atleast2d(input: Tensor): the atleast_2d operator. input with leading size-1 dims prepended until rank >= 2 (see atleast_1d). Returns a view (no copy). |
| atleast3d | [common] fun atleast3d(input: Tensor): Tensor atleast3d(input: Tensor): the atleast_3d operator. input with leading size-1 dims prepended until rank >= 3 (see atleast_1d). Returns a view (no copy). |
| attention | [common] fun attention(query: Tensor, key: Tensor, value: Tensor, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, kScale: Tensor? = null, vScale: Tensor? = null, keptPrefix: Tensor? = null): Tensor attention(query: Tensor, key: Tensor, value: Tensor, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, kScale: Tensor? = null, vScale: Tensor? = null, keptPrefix: Tensor? = null): the attention operator. Dense attention with the serving riders: per-head sink, logit soft-cap, sliding window, smoothed softmax. The core is scaled_dot_product_attention; each rider adjusts the softmax stage: - head_sink``[H_q]: a per-head virtual logit folded into the softmax denominator (attention that can "go nowhere"). - softcap: logits pass through cap * tanh(x / cap) before the softmax. - sliding_window: each query attends only the last N key positions. - smooth_softmax: adds one to the softmax denominator. Default false. |
| attentionOverCache | [common] fun attentionOverCache(query: Tensor, cacheKey: Tensor, cacheValue: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): Tensor attentionOverCache(query: Tensor, cacheKey: Tensor, cacheValue: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): the attention_over_cache operator. Attend a KV cache WITHOUT appending to it (a read-only re-attention). q is packed varlen [ΣS_q, H, D] (or hidden-folded [ΣS_q, num_heads*D] with num_heads set); cache_key/cache_value are a cache another attention call already appended: continuous head-major [max_seqs, H_kv, max_seq, D] with a rank-1 [B]``kvcache_start (the layout selector; its values are not read, and q's position derives from cu_seqlens_k), or a paged block pool [num_blocks, H_kv, block_size, D] with a rank-2 [B, max_blocks] block table (every entry -1 or in [0, num_blocks), no block index repeated within a row; an out-of-range or repeated entry refuses INVALID_ARGUMENT before any read). On the continuous cache slot_ids ([B] Int32, in range, pairwise distinct) names each batch row's cache row; absent, batch row b reads cache row b. cu_seqlens_k is each sequence's TOTAL cached length: per-seq [B] or cumulative [B+1]. The cache is never written. When rope_cos/rope_sin are bound, rotary applies to q only (the cached keys are already rotated). Use case: a q-only module re-attending a sibling layer's cache. head_sink is the per-head softmax sink [H_q], a virtual logit folded into the softmax denominator, the same contract as group_query_attention_varlen; it rides the parameter tail here. |
| attentionVarlen | [common] fun attentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, kScale: Tensor? = null, vScale: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): Tensor attentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, kScale: Tensor? = null, vScale: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): the attention_varlen operator. Variable-length (packed) form of attention: the serving riders over token-packed ragged batches. Tensors and offsets follow scaled_dot_product_attention_varlen; the riders (head_sink, softcap, sliding_window, smooth_softmax) follow attention. |
| avgPool | [common] fun avgPool(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, ceilMode: Boolean, countIncludePad: Boolean, divisorOverride: Long?): Tensor avgPool(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, ceilMode: Boolean, countIncludePad: Boolean, divisorOverride: Long?): the avg_pool operator. Rank-generic average pooling, channels-last. input is [N, D1..Dn, C]; the window rank is read from kernel_size's length. Each output element averages its window; count_include_pad decides whether padded positions count in the divisor, and divisor_override replaces the divisor outright. |
| avgPool1d | [common] fun avgPool1d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), ceilMode: Boolean = false, countIncludePad: Boolean = true): Tensor avgPool1d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), ceilMode: Boolean = false, countIncludePad: Boolean = true): the avg_pool1d operator. 1-D average pooling over [N, L, C] (channels-last). See avg_pool; stride empty = kernel_size. |
| avgPool2d | [common] fun avgPool2d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), ceilMode: Boolean = false, countIncludePad: Boolean = true, divisorOverride: Long? = null): Tensor avgPool2d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), ceilMode: Boolean = false, countIncludePad: Boolean = true, divisorOverride: Long? = null): the avg_pool2d operator. 2-D average pooling over [N, H, W, C] (channels-last). Averages each kernel_size window; count_include_pad includes the zero padding in the divisor; divisor_override fixes the divisor. stride empty = kernel_size. |
| avgPool3d | [common] fun avgPool3d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), ceilMode: Boolean = false, countIncludePad: Boolean = true, divisorOverride: Long? = null): Tensor avgPool3d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), ceilMode: Boolean = false, countIncludePad: Boolean = true, divisorOverride: Long? = null): the avg_pool3d operator. 3-D average pooling over [N, D, H, W, C] (channels-last). See avg_pool2d; parameters extend to {kD, kH, kW} etc. |
| batchNorm | [common] fun batchNorm(input: Tensor, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): Tensor batchNorm(input: Tensor, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): the batch_norm operator. Per-channel batch normalization (inference form), channels-last. input is [N, *spatial, C]; every operand is per-channel [C]. The supplied running_mean / running_var ARE the statistics (inference only; no training mode, no momentum). Optional fused activation applies to the result. |
| bernoulli | [common] fun bernoulli(probabilities: Tensor, device: Placement? = null): Tensor bernoulli(probabilities: Tensor, device: Placement? = null): the bernoulli operator. Independent Bernoulli draws from per-element success probabilities. Each output element is 1 with probability probabilities[i], else 0; shape and dtype mirror probabilities. |
| bernoulliInPlace | [common] fun bernoulliInPlace(self: Tensor, device: Placement? = null): Tensor bernoulliInPlace(self: Tensor, device: Placement? = null): the bernoulli_ operator. In-place: overwrite self, whose values are the per-element success probabilities, with the 0/1 draws (self[i] ~ Bernoulli(self[i])). Writes through self and returns it, so calls chain. |
| binaryCrossEntropy | [common] fun binaryCrossEntropy(input: Tensor, target: Tensor, weight: Tensor? = null, reduction: Reduction = Reduction.MEAN): Tensor binaryCrossEntropy(input: Tensor, target: Tensor, weight: Tensor? = null, reduction: Reduction = Reduction.MEAN): the binary_cross_entropy operator. Binary cross-entropy on element-wise PROBABILITIES. input must already be probabilities in [0, 1] (apply sigmoid first, or use binary_cross_entropy_with_logits for the fused, numerically safer form). weight re-weights each element's loss. |
| binaryCrossEntropyWithLogits | [common] fun binaryCrossEntropyWithLogits(input: Tensor, target: Tensor, weight: Tensor? = null, reduction: Reduction = Reduction.MEAN, posWeight: Tensor? = null): Tensor binaryCrossEntropyWithLogits(input: Tensor, target: Tensor, weight: Tensor? = null, reduction: Reduction = Reduction.MEAN, posWeight: Tensor? = null): the binary_cross_entropy_with_logits operator. Binary cross-entropy on RAW LOGITS (sigmoid fused, numerically stable). Computes binary_cross_entropy(sigmoid(input), target) in one pass without materializing the probabilities. pos_weight scales the positive-class term per element (class-imbalance correction). |
| bincount | [common] fun bincount(input: Tensor, weights: Tensor? = null, minlength: Long = 0): Tensor bincount(input: Tensor, weights: Tensor? = null, minlength: Long = 0L): the bincount operator. Occurrence count (or weight sum) of each non-negative integer value. The output is 1-D with length max(x) + 1, floored at minlength; entry i counts how often i occurs (with weights bound, it sums the weights at those positions instead). |
| bitwiseAnd | [common] fun bitwiseAnd(input: Tensor, other: Tensor): Tensor bitwiseAnd(input: Tensor, other: Tensor): the bitwise_and operator. Elementwise bitwise AND; the dtype law the whole bitwise family shares: integer and Bool dtypes only, and BOTH operands must carry ONE dtype (no promotion; a mixed pair is refused typed; a scalar other adopts input's dtype). Shapes broadcast per the standard rules; the output carries the shared dtype. The other bitwise ops state "dtype law as bitwise_and" instead of restating this.[common] fun bitwiseAnd(input: Tensor, other: Double): Tensor bitwiseAnd(input: Tensor, other: Double): the number form of bitwise_and, other as a scalar. |
| bitwiseAndInPlace | [common] fun bitwiseAndInPlace(self: Tensor, other: Tensor): Tensor bitwiseAndInPlace(self: Tensor, other: Tensor): the bitwise_and_ operator. In-place bitwise_and: writes the AND through x. Writes through self and returns it, so calls chain.[common] fun bitwiseAndInPlace(self: Tensor, other: Double): Tensor bitwiseAndInPlace(self: Tensor, other: Double): the number form of bitwise_and_, other as a scalar. |
| bitwiseLeftShift | [common] fun bitwiseLeftShift(input: Tensor, other: Tensor): Tensor bitwiseLeftShift(input: Tensor, other: Tensor): the bitwise_left_shift operator. Elementwise a << other. Integer dtypes only (Bool refuses; a shifted Bool byte has no meaning); otherwise dtype law as bitwise_and. The shift is a TOTAL function: a count outside [0, bit_width), negative included, yields the fully shifted-out value (zero fill), identically on every backend.[common] fun bitwiseLeftShift(input: Tensor, other: Double): Tensor bitwiseLeftShift(input: Tensor, other: Double): the number form of bitwise_left_shift, other as a scalar. |
| bitwiseLeftShiftInPlace | [common] fun bitwiseLeftShiftInPlace(self: Tensor, other: Tensor): Tensor bitwiseLeftShiftInPlace(self: Tensor, other: Tensor): the bitwise_left_shift_ operator. In-place bitwise_left_shift: writes the shifted values through x. Writes through self and returns it, so calls chain.[common] fun bitwiseLeftShiftInPlace(self: Tensor, other: Double): Tensor bitwiseLeftShiftInPlace(self: Tensor, other: Double): the number form of bitwise_left_shift_, other as a scalar. |
| bitwiseNot | [common] fun bitwiseNot(input: Tensor): Tensor bitwiseNot(input: Tensor): the bitwise_not operator. Elementwise bitwise NOT (~x; logical NOT for Bool). Integer and Bool dtypes; the output keeps input's dtype. |
| bitwiseNotInPlace | [common] fun bitwiseNotInPlace(self: Tensor): Tensor bitwiseNotInPlace(self: Tensor): the bitwise_not_ operator. In-place bitwise_not: complements x in place. Writes through self and returns it, so calls chain. |
| bitwiseOr | [common] fun bitwiseOr(input: Tensor, other: Tensor): Tensor bitwiseOr(input: Tensor, other: Tensor): the bitwise_or operator. Elementwise bitwise OR; dtype law as bitwise_and.[common] fun bitwiseOr(input: Tensor, other: Double): Tensor bitwiseOr(input: Tensor, other: Double): the number form of bitwise_or, other as a scalar. |
| bitwiseOrInPlace | [common] fun bitwiseOrInPlace(self: Tensor, other: Tensor): Tensor bitwiseOrInPlace(self: Tensor, other: Tensor): the bitwise_or_ operator. In-place bitwise_or: writes the OR through x. Writes through self and returns it, so calls chain.[common] fun bitwiseOrInPlace(self: Tensor, other: Double): Tensor bitwiseOrInPlace(self: Tensor, other: Double): the number form of bitwise_or_, other as a scalar. |
| bitwiseRightShift | [common] fun bitwiseRightShift(input: Tensor, other: Tensor): Tensor bitwiseRightShift(input: Tensor, other: Tensor): the bitwise_right_shift operator. Elementwise a >> other. Integer dtypes only (Bool refuses); otherwise dtype law as bitwise_and. ARITHMETIC (sign-propagating) for signed dtypes, logical for unsigned; the shift is a TOTAL function; a count outside [0, bit_width), negative included, yields the fully shifted-out value (sign fill 0/-1 for signed, 0 for unsigned), identically on every backend.[common] fun bitwiseRightShift(input: Tensor, other: Double): Tensor bitwiseRightShift(input: Tensor, other: Double): the number form of bitwise_right_shift, other as a scalar. |
| bitwiseRightShiftInPlace | [common] fun bitwiseRightShiftInPlace(self: Tensor, other: Tensor): Tensor bitwiseRightShiftInPlace(self: Tensor, other: Tensor): the bitwise_right_shift_ operator. In-place bitwise_right_shift: writes the shifted values through x. Writes through self and returns it, so calls chain.[common] fun bitwiseRightShiftInPlace(self: Tensor, other: Double): Tensor bitwiseRightShiftInPlace(self: Tensor, other: Double): the number form of bitwise_right_shift_, other as a scalar. |
| bitwiseXor | [common] fun bitwiseXor(input: Tensor, other: Tensor): Tensor bitwiseXor(input: Tensor, other: Tensor): the bitwise_xor operator. Elementwise bitwise XOR; dtype law as bitwise_and.[common] fun bitwiseXor(input: Tensor, other: Double): Tensor bitwiseXor(input: Tensor, other: Double): the number form of bitwise_xor, other as a scalar. |
| bitwiseXorInPlace | [common] fun bitwiseXorInPlace(self: Tensor, other: Tensor): Tensor bitwiseXorInPlace(self: Tensor, other: Tensor): the bitwise_xor_ operator. In-place bitwise_xor: writes the XOR through x. Writes through self and returns it, so calls chain.[common] fun bitwiseXorInPlace(self: Tensor, other: Double): Tensor bitwiseXorInPlace(self: Tensor, other: Double): the number form of bitwise_xor_, other as a scalar. |
| bmm | [common] fun bmm(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null): Tensor bmm(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null): the bmm operator. Batched matrix multiply of two rank-3 tensors, with an optional fused bias and activation epilogue. a [B, M, K] x b [B, K, N] -> [B, M, N], one independent matmul per batch index. |
| broadcastTensors | [common] fun broadcastTensors(tensors: List<Tensor>): List<Tensor> broadcastTensors(tensors: List<Tensor>): the broadcast_tensors operator. Broadcast every input to their common shape: dims are right-aligned, size-1 dims stretch, anything else must match. Returns VIEWS; each output shares its input's storage, with the stretched positions aliasing ONE stored element; treat the results as read-only (or contiguous one to materialize it). |
| broadcastTo | [common] fun broadcastTo(input: Tensor, shape: LongArray): Tensor broadcastTo(input: Tensor, shape: LongArray): the broadcast_to operator. Alias of expand under the NumPy name; same broadcasting rules, same view semantics. |
| bucketize | [common] fun bucketize(input: Tensor, boundaries: Tensor, outInt32: Boolean = false, right: Boolean = false): Tensor bucketize(input: Tensor, boundaries: Tensor, outInt32: Boolean = false, right: Boolean = false): the bucketize operator. Bucket index of each input element against a sorted 1-D boundaries. With right == false (default) an element in [b[i-1], b[i]) maps to bucket i; right == true uses (b[i-1], b[i]]. Output shape mirrors input. |
| cast | [common] fun cast(input: Tensor, target: DType, forceCopy: Boolean = false): Tensor cast(input: Tensor, target: DType, forceCopy: Boolean = false): the cast operator. Same-device dtype conversion. Converts per element to target. When input already has target's dtype and force_copy is false, the input passes through unchanged (no new storage); set force_copy = true to guarantee an owning copy. |
| castLike | [common] fun castLike(input: Tensor, reference: Tensor): Tensor castLike(input: Tensor, reference: Tensor): the cast_like operator. cast to another tensor's dtype: cast(x, reference.dtype()). |
| causalConvUpdate | [common] fun causalConvUpdate(input: Tensor, weight: Tensor, bias: Tensor?, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, activation: Activation? = null): Tensor causalConvUpdate(input: Tensor, weight: Tensor, bias: Tensor?, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, activation: Activation? = null): the causal_conv_update operator. Depthwise causal short-conv serving step over a rolling per-sequence window. x [B, S, dim]; weight [dim, W] (oldest tap first); optional bias [dim]; state [B, dim, W] (same dtype as x) is read AND updated IN PLACE in both modes: S 1 runs the prefill conv with each row's left context seeded from its window (a zero window is a fresh sequence, bit for bit; an S-token call over committed state equals S single-token steps exactly) and re-captures the window as the last W of (old window ++ the row's valid inputs; zero valid tokens leave it unchanged); S == 1 shift-inserts the new token and emits the tap dot. seq_lens ([B] Int32) bounds ragged prefill rows (their padding is zero post-activation). activation applies to the returned out [B, S, dim] only (Silu fuses in-kernel); the stored window stays pre-activation raw. slot_ids ([B] Int32, device-resident) addresses state as a SLAB [num_slots, dim, W]: batch row b reads/updates slab row slot_ids[b] in place (ids in range and DISTINCT per call, the caller's contract); absent keeps state row b. |
| cdist | [common] fun cdist(x1: Tensor, x2: Tensor, p: Double = 2.0): Tensor cdist(x1: Tensor, x2: Tensor, p: Double = 2.0): the cdist operator. Pairwise L_p distance between every ROW pair of two matrices. (B?, M, K) x (B?, N, K) -> (B?, M, N); the optional leading batch dims broadcast. |
| ceil | [common] fun ceil(input: Tensor): Tensor ceil(input: Tensor): the ceil operator. Elementwise ceiling: the smallest integer not below input. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| ceilInPlace | [common] fun ceilInPlace(self: Tensor): Tensor ceilInPlace(self: Tensor): the ceil_ operator. In-place ceil: writes the result through self; same formula, arguments, and error conditions as ceil(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| celu | [common] fun celu(input: Tensor, alpha: Double = 1.0): Tensor celu(input: Tensor, alpha: Double = 1.0): the celu operator. Continuously differentiable exponential linear unit. |
| celuInPlace | [common] fun celuInPlace(self: Tensor, alpha: Double = 1.0): Tensor celuInPlace(self: Tensor, alpha: Double = 1.0): the celu_ operator. In-place celu: writes the result through self; same formula, arguments, and error conditions as celu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| chunk | [common] fun chunk(input: Tensor, numChunks: Long, dim: Long = 0): List<Tensor> chunk(input: Tensor, numChunks: Long, dim: Long = 0L): the chunk operator. Split along dim into num_chunks near-equal parts (the last may be shorter). Returns VIEWS sharing the source's storage (one call, no copy); prefer it over repeated narrows. |
| circularPad | [common] fun circularPad(input: Tensor, pad: LongArray): Tensor circularPad(input: Tensor, pad: LongArray): the circular_pad operator. pad in wrap-around mode: the padding continues from the opposite edge. Same (lo, hi) pair layout as constant_pad. Copies. |
| clamp | [common] fun clamp(input: Tensor, min: Tensor? = null, max: Tensor? = null): Tensor clamp(input: Tensor, min: Tensor? = null, max: Tensor? = null): the clamp operator. Clamp input into min, max; an empty bound leaves that side unbounded. |
| clampInPlace | [common] fun clampInPlace(self: Tensor, min: Tensor? = null, max: Tensor? = null): Tensor clampInPlace(self: Tensor, min: Tensor? = null, max: Tensor? = null): the clamp_ operator. In-place clamp: writes the result through self; same formula, arguments, and error conditions as clamp(). Either bound may be absent (std::nullopt), a one-sided clamp; tensor bounds broadcast against self. A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| clampMax | [common] fun clampMax(input: Tensor, max: Tensor): Tensor clampMax(input: Tensor, max: Tensor): the clamp_max operator. Elementwise upper bound. max is a scalar or a tensor broadcast against input (numpy rules).[common] fun clampMax(input: Tensor, max: Double): Tensor clampMax(input: Tensor, max: Double): the number form of clamp_max, max as a scalar. |
| clampMaxInPlace | [common] fun clampMaxInPlace(self: Tensor, max: Tensor): Tensor clampMaxInPlace(self: Tensor, max: Tensor): the clamp_max_ operator. In-place clamp_max: writes the result through self; same formula, arguments, and error conditions as clamp_max(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun clampMaxInPlace(self: Tensor, max: Double): Tensor clampMaxInPlace(self: Tensor, max: Double): the number form of clamp_max_, max as a scalar. |
| clampMin | [common] fun clampMin(input: Tensor, min: Tensor): Tensor clampMin(input: Tensor, min: Tensor): the clamp_min operator. Elementwise lower bound. min is a scalar or a tensor broadcast against input (numpy rules).[common] fun clampMin(input: Tensor, min: Double): Tensor clampMin(input: Tensor, min: Double): the number form of clamp_min, min as a scalar. |
| clampMinInPlace | [common] fun clampMinInPlace(self: Tensor, min: Tensor): Tensor clampMinInPlace(self: Tensor, min: Tensor): the clamp_min_ operator. In-place clamp_min: writes the result through self; same formula, arguments, and error conditions as clamp_min(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun clampMinInPlace(self: Tensor, min: Double): Tensor clampMinInPlace(self: Tensor, min: Double): the number form of clamp_min_, min as a scalar. |
| clone | [common] fun clone(src: Tensor): Tensor clone(src: Tensor): the clone operator. Deep copy on the same device: fresh storage, identical shape, dtype, and values. Equivalent to copy(src) with the defaults. |
| concat | [common] fun concat(tensors: List<Tensor>, dim: Long = 0, activation: Activation = Activation.IDENTITY): Tensor concat(tensors: List<Tensor>, dim: Long = 0L, activation: Activation = Activation.IDENTITY): the concat operator. Join tensors along an EXISTING dim: all inputs share rank and off-dim extents; the dim extents add up. Copies into fresh storage. The optional activation fuses an elementwise epilogue into the single copy pass (act(concat(...))): it requires a float output dtype, every input already AT that dtype, and a non-gated kind; anything else raises. Activation::Identity (the default) is the plain join. |
| constantPad | [common] fun constantPad(input: Tensor, pad: LongArray, value: Tensor? = null): Tensor constantPad(input: Tensor, pad: LongArray, value: Tensor? = null): the constant_pad operator. pad with a constant fill. Pads each axis by interleaved (lo, hi) pairs in layout order: pair i pads axis i; fewer pairs than the rank pads only the leading axes; a negative width crops that side. Widths are literals or 0-D integer Tensors. Copies. |
| contiguous | [common] fun contiguous(input: Tensor, forceCopy: Boolean = false): Tensor contiguous(input: Tensor, forceCopy: Boolean = false): the contiguous operator. Packs the tensor into row-major contiguous storage. A no-op passthrough (no copy, same storage) when the input is already contiguous and force_copy is false. Kernels stride natively; call this only when YOUR host-side consumption needs dense memory, never "to be safe" before an op. |
| conv | [common] fun conv(input: Tensor, weight: Tensor, bias: Tensor?, stride: LongArray, padding: LongArray, dilation: LongArray, groups: Long, mode: PadMode = PadMode.CONSTANT, value: Double? = null, activation: Activation = Activation.IDENTITY): Tensor conv(input: Tensor, weight: Tensor, bias: Tensor?, stride: LongArray, padding: LongArray, dilation: LongArray, groups: Long, mode: PadMode = PadMode.CONSTANT, value: Double? = null, activation: Activation = Activation.IDENTITY): the conv operator. Convolution family, channels-last ([N, spatial..., C]), weights output-channel-first ([O, spatial..., Ig]). stride / dilation / output_padding are PER-AXIS window attributes everywhere: length 0 (defaulted), 1 (broadcast), or spatial-rank. padding carries TWO conventions; read the one that matches the op: - conv / conv_transpose*: interleaved (lo, hi) PAIRS in axis order (even length; pair i pads spatial axis i); asymmetric pads spell directly, e.g. 1-D {2, 3} = lo 2, hi 3. - conv1d/2d/3d: SYMMETRIC per-axis widths (length 0/1/spatial-rank), e.g. 1-D {2} = lo 2, hi 2; an asymmetric pad needs conv or an explicit ops::pad first. The inline defaults below encode the split: conv1d pads {0}, conv_transpose1d pads {0, 0}. |
| conv1d | [common] fun conv1d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), groups: Long = 1, activation: Activation = Activation.IDENTITY): Tensor conv1d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), groups: Long = 1L, activation: Activation = Activation.IDENTITY): the conv1d operator. 1-D convolution over a channels-last input. Layout is channels-last: input [N, spatial.., C], weight OHWI [O, K.., C/groups] (output channels first, kernel dims, then the per-group input channels), optional bias [O]. |
| conv2d | [common] fun conv2d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), groups: Long = 1, activation: Activation = Activation.IDENTITY): Tensor conv2d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), groups: Long = 1L, activation: Activation = Activation.IDENTITY): the conv2d operator. 2-D convolution over a channels-last input. Layout is channels-last: input [N, spatial.., C], weight OHWI [O, K.., C/groups] (output channels first, kernel dims, then the per-group input channels), optional bias [O]. |
| conv3d | [common] fun conv3d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1, 1), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), groups: Long = 1, activation: Activation = Activation.IDENTITY): Tensor conv3d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1, 1), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), groups: Long = 1L, activation: Activation = Activation.IDENTITY): the conv3d operator. 3-D convolution over a channels-last input. Layout is channels-last: input [N, spatial.., C], weight OHWI [O, K.., C/groups] (output channels first, kernel dims, then the per-group input channels), optional bias [O]. |
| convTranspose | [common] fun convTranspose(input: Tensor, weight: Tensor, bias: Tensor?, stride: LongArray, padding: LongArray, outputPadding: LongArray, groups: Long, dilation: LongArray, activation: Activation = Activation.IDENTITY): Tensor convTranspose(input: Tensor, weight: Tensor, bias: Tensor?, stride: LongArray, padding: LongArray, outputPadding: LongArray, groups: Long, dilation: LongArray, activation: Activation = Activation.IDENTITY): the conv_transpose operator. Rank-polymorphic (1-D/2-D/3-D) transposed (fractionally-strided) convolution (learnable upsampling). Layout is channels-last; the WEIGHT is the transpose-flip of conv's: input-channel-first [C_in, K.., O/groups]; C_in == weight[0] and C_out == weight[-1] * groups (the cuDNN-BackwardData / ONNX convention, NOT conv's OHWI). padding is per-side: two entries (lo, hi) per spatial dim. output_padding grows only the output's high side. All geometry spans are explicit here; the conv_transpose1d/2d/3d wrappers below carry the per-rank defaults. |
| convTranspose1d | [common] fun convTranspose1d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1), padding: LongArray = longArrayOf(0, 0), outputPadding: LongArray = longArrayOf(0), groups: Long = 1, dilation: LongArray = longArrayOf(1), activation: Activation = Activation.IDENTITY): Tensor convTranspose1d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1), padding: LongArray = longArrayOf(0, 0), outputPadding: LongArray = longArrayOf(0), groups: Long = 1L, dilation: LongArray = longArrayOf(1), activation: Activation = Activation.IDENTITY): the conv_transpose1d operator. 1-D transposed (fractionally-strided) convolution (learnable upsampling). Layout is channels-last; the WEIGHT is the transpose-flip of conv's: input-channel-first [C_in, K.., O/groups]; C_in == weight[0] and C_out == weight[-1] * groups (the cuDNN-BackwardData / ONNX convention, NOT conv's OHWI). padding is per-side: two entries (lo, hi) per spatial dim. output_padding grows only the output's high side. |
| convTranspose2d | [common] fun convTranspose2d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1), padding: LongArray = longArrayOf(0, 0, 0, 0), outputPadding: LongArray = longArrayOf(0, 0), groups: Long = 1, dilation: LongArray = longArrayOf(1, 1), activation: Activation = Activation.IDENTITY): Tensor convTranspose2d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1), padding: LongArray = longArrayOf(0, 0, 0, 0), outputPadding: LongArray = longArrayOf(0, 0), groups: Long = 1L, dilation: LongArray = longArrayOf(1, 1), activation: Activation = Activation.IDENTITY): the conv_transpose2d operator. 2-D transposed (fractionally-strided) convolution (learnable upsampling). Layout is channels-last; the WEIGHT is the transpose-flip of conv's: input-channel-first [C_in, K.., O/groups]; C_in == weight[0] and C_out == weight[-1] * groups (the cuDNN-BackwardData / ONNX convention, NOT conv's OHWI). padding is per-side: two entries (lo, hi) per spatial dim. output_padding grows only the output's high side. |
| convTranspose3d | [common] fun convTranspose3d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1, 1), padding: LongArray = longArrayOf(0, 0, 0, 0, 0, 0), outputPadding: LongArray = longArrayOf(0, 0, 0), groups: Long = 1, dilation: LongArray = longArrayOf(1, 1, 1), activation: Activation = Activation.IDENTITY): Tensor convTranspose3d(input: Tensor, weight: Tensor, bias: Tensor? = null, stride: LongArray = longArrayOf(1, 1, 1), padding: LongArray = longArrayOf(0, 0, 0, 0, 0, 0), outputPadding: LongArray = longArrayOf(0, 0, 0), groups: Long = 1L, dilation: LongArray = longArrayOf(1, 1, 1), activation: Activation = Activation.IDENTITY): the conv_transpose3d operator. 3-D transposed (fractionally-strided) convolution (learnable upsampling). Layout is channels-last; the WEIGHT is the transpose-flip of conv's: input-channel-first [C_in, K.., O/groups]; C_in == weight[0] and C_out == weight[-1] * groups (the cuDNN-BackwardData / ONNX convention, NOT conv's OHWI). padding is per-side: two entries (lo, hi) per spatial dim. output_padding grows only the output's high side. |
| copy | [common] fun copy(src: Tensor, target: Placement? = null, forceCopy: Boolean = true): Tensor copy(src: Tensor, target: Placement? = null, forceCopy: Boolean = true): the copy operator. Copies a tensor, optionally to another device or stream. With the default force_copy = true the result always owns fresh storage; with force_copy = false a same-device copy may pass the input through. target absent = src's own stream. |
| copyInPlace | [common] fun copyInPlace(self: Tensor, src: Tensor, target: Placement? = null): Tensor copyInPlace(self: Tensor, src: Tensor, target: Placement? = null): the copy_ operator. Copies src INTO self's existing storage, converting per element to self's dtype in the same single pass (never cast first; the copy IS the cast). Shapes must match after broadcasting src. Raises ClikaRT::Error where the shapes/dtypes cannot be served; returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| copyInto | [common] fun copyInto(self: Tensor, src: Tensor, target: Placement? = null): Tensor copyInto(self: Tensor, src: Tensor, target: Placement? = null): the copy_into operator. Write src into the caller-supplied out buffer (the destination-first spelling of the in-place copy; bind a persistent buffer once, write it every step with no allocation). Dtype conversion is interleaved with the write; a cross-device src is transferred first. Returns out. Writes through self and returns it, so calls chain. |
| copysign | [common] fun copysign(input: Tensor, other: Tensor): Tensor copysign(input: Tensor, other: Tensor): the copysign operator. Elementwise ` |
| copysignInPlace | [common] fun copysignInPlace(self: Tensor, other: Tensor): Tensor copysignInPlace(self: Tensor, other: Tensor): the copysign_ operator. In-place copysign: writes the re-signed magnitudes through x. Writes through self and returns it, so calls chain.[common] fun copysignInPlace(self: Tensor, other: Double): Tensor copysignInPlace(self: Tensor, other: Double): the number form of copysign_, other as a scalar. |
| copyToCpu | [common] fun copyToCpu(src: Tensor): Tensor copyToCpu(src: Tensor): the copy_to_cpu operator. Materializes a host-resident copy of the tensor (device-to-host transfer; an owning CPU tensor even when src is already on the CPU). |
| cos | [common] fun cos(input: Tensor): Tensor cos(input: Tensor): the cos operator. Elementwise cosine. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| cosh | [common] fun cosh(input: Tensor): Tensor cosh(input: Tensor): the cosh operator. Elementwise hyperbolic cosine. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| coshInPlace | [common] fun coshInPlace(self: Tensor): Tensor coshInPlace(self: Tensor): the cosh_ operator. In-place cosh: writes the result through self; same formula, arguments, and error conditions as cosh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| cosineSimilarity | [common] fun cosineSimilarity(x1: Tensor, x2: Tensor, dim: Long = 1, eps: Double? = null): Tensor cosineSimilarity(x1: Tensor, x2: Tensor, dim: Long = 1L, eps: Double? = null): the cosine_similarity operator. Cosine similarity along dim. |
| cosInPlace | [common] fun cosInPlace(self: Tensor): Tensor cosInPlace(self: Tensor): the cos_ operator. In-place cos: writes the result through self; same formula, arguments, and error conditions as cos(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| countNonzero | [common] fun countNonzero(input: Tensor, dims: LongArray = longArrayOf()): Tensor countNonzero(input: Tensor, dims: LongArray = longArrayOf()): the count_nonzero operator. Number of nonzero elements over dims. Empty dims counts across the whole tensor. Output dtype is always Int64. An element is nonzero exactly when its Bool cast is true: +0 and -0 are zero; a subnormal, an infinity and a NaN are nonzero. |
| cross | [common] fun cross(input: Tensor, other: Tensor, dim: Long? = null): Tensor cross(input: Tensor, other: Tensor, dim: Long? = null): the cross operator. 3-vector cross product along dim. Both operands share one shape; the dim axis (default: the LAST axis) must have size 3. The output mirrors the input shape. |
| crossEntropy | [common] fun crossEntropy(input: Tensor, target: Tensor, weight: Tensor? = null, ignoreIndex: Long? = null, reduction: Reduction = Reduction.MEAN): Tensor crossEntropy(input: Tensor, target: Tensor, weight: Tensor? = null, ignoreIndex: Long? = null, reduction: Reduction = Reduction.MEAN): the cross_entropy operator. Cross-entropy loss over class LOGITS; the class dim is LAST. Equivalent to log_softmax over the class axis followed by negative-log-likelihood: |
| cummax | [common] fun cummax(input: Tensor, dim: Long): List<Tensor> cummax(input: Tensor, dim: Long): the cummax operator. |
| cummin | [common] fun cummin(input: Tensor, dim: Long): List<Tensor> cummin(input: Tensor, dim: Long): the cummin operator. |
| cumprod | [common] fun cumprod(input: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): Tensor cumprod(input: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): the cumprod operator. Cumulative product along dim (inclusive scan). |
| cumprodInPlace | [common] fun cumprodInPlace(self: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): Tensor cumprodInPlace(self: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): the cumprod_ operator. In-place cumprod: rewrites self with its inclusive prefix products along dim and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| cumsum | [common] fun cumsum(input: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): Tensor cumsum(input: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): the cumsum operator. Cumulative sum along dim (inclusive scan). Output mirrors input's shape; dtype widens the accumulation/output when set (Undefined keeps input's dtype). |
| cumsumInPlace | [common] fun cumsumInPlace(self: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): Tensor cumsumInPlace(self: Tensor, dim: Long, dtype: DType = DType.UNDEFINED): the cumsum_ operator. In-place cumsum: rewrites self with its inclusive prefix sums along dim and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| deformConv | [common] fun deformConv(input: Tensor, weight: Tensor, offset: Tensor, mask: Tensor? = null, bias: Tensor? = null, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), dilation: LongArray = longArrayOf(), groups: Long = 1, offsetGroups: Long = 1, activation: Activation = Activation.IDENTITY): Tensor deformConv(input: Tensor, weight: Tensor, offset: Tensor, mask: Tensor? = null, bias: Tensor? = null, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), dilation: LongArray = longArrayOf(), groups: Long = 1L, offsetGroups: Long = 1L, activation: Activation = Activation.IDENTITY): the deform_conv operator. Deformable convolution, 1-D/2-D/3-D (rank derives from input), channels-last: x [N, D1..Dr, C], weight [O, K1..Kr, C/groups] (OHWI), offset [N, out-spatial..., offset_groups·∏K·r], mask [N, out-spatial..., offset_groups·∏K] (absent ⇒ unmodulated), bias [O]; returns [N, out-spatial..., O] at input's dtype. Each kernel tap samples at out·stride − pad_lo + tap·dilation + Δ (pixel units, bilinear; out-of-bounds reads 0); the offset channel for (group g, tap t, axis d) is (g·∏K + t)·r + d with taps row-major over the kernel and axes in layout order (2-D: Δh then Δw). padding is interleaved (lo, hi) pairs over the spatial axes; stride/dilation broadcast per spatial axis (empty ⇒ 1). All floating inputs must share input's dtype (f32/f64/f16/bf16; no silent promotion); activation is a fused elementwise epilogue applied after bias. |
| deg2rad | [common] fun deg2rad(input: Tensor): Tensor deg2rad(input: Tensor): the deg2rad operator. Converts degrees to radians elementwise: . |
| deg2radInPlace | [common] fun deg2radInPlace(self: Tensor): Tensor deg2radInPlace(self: Tensor): the deg2rad_ operator. In-place deg2rad: writes the radians through x. Writes through self and returns it, so calls chain. |
| diag | [common] fun diag(input: Tensor, diagonal: Long = 0): Tensor diag(input: Tensor, diagonal: Long = 0L): the diag operator. Rank-1 <-> rank-2 diagonal converter. A vector of length L becomes an `(L + |
| diagEmbed | [common] fun diagEmbed(input: Tensor, offset: Long = 0, dim1: Long = -2L, dim2: Long = -1L): Tensor diagEmbed(input: Tensor, offset: Long = 0L, dim1: Long = -2L, dim2: Long = -1L): the diag_embed operator. Embed the LAST dim along the diagonal of a fresh (dim1, dim2) plane: output rank = input rank + 1, both new dims sized `last + |
| diagonal | [common] fun diagonal(input: Tensor, offset: Long = 0, dim1: Long = 0, dim2: Long = 1): Tensor diagonal(input: Tensor, offset: Long = 0L, dim1: Long = 0L, dim2: Long = 1L): the diagonal operator. Extract the offset-th diagonal between dim1 and dim2: the two source dims are removed and a trailing dim of the diagonal's length is appended (rank - 1 total). Copies into fresh storage (deliberately not a view). |
| diff | [common] fun diff(input: Tensor, n: Long = 1, dim: Long = -1L, prepend: Tensor? = null, append: Tensor? = null): Tensor diff(input: Tensor, n: Long = 1L, dim: Long = -1L, prepend: Tensor? = null, append: Tensor? = null): the diff operator. n-th order finite difference along dim. Each application shortens the axis by one, so the output's dim extent is size(dim) - n. prepend / append are concatenated onto the axis before differencing (they must match input's shape outside dim). |
| div | [common] fun div(input: Tensor, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): Tensor div(input: Tensor, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): the div operator. Divides input by other elementwise, with an optional quotient rounding. Broadcasts and promotes as add.[common] fun div(input: Tensor, other: Double, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): Tensor div(input: Tensor, other: Double, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): the number form of div, other as a scalar.[common] fun div(input: Double, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): Tensor div(input: Double, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): the div operator. Scalar-LHS a / b. input keeps its kind (see Scalar). The parameter set mirrors the tensor-first form: rounding_mode picks true/trunc/floor division, activation applies to the result. |
| divInPlace | [common] fun divInPlace(self: Tensor, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): Tensor divInPlace(self: Tensor, other: Tensor, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): the div_ operator. In-place div: writes the (optionally rounded) quotient through x. Writes through self and returns it, so calls chain.[common] fun divInPlace(self: Tensor, other: Double, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): Tensor divInPlace(self: Tensor, other: Double, roundingMode: RoundingMode = RoundingMode.NONE, activation: Activation = Activation.IDENTITY): the number form of div_, other as a scalar. |
| dot | [common] fun dot(input: Tensor, other: Tensor): Tensor dot(input: Tensor, other: Tensor): the dot operator. Inner product of two 1-D tensors. Higher-rank inputs are rejected; use matmul (or inner) for batched contractions. |
| dynamicQuantize | [common] fun dynamicQuantize(input: Tensor): List<Tensor> dynamicQuantize(input: Tensor): the dynamic_quantize operator. Fused dynamic quantization: encode float input to uint8 affine codes with the scale and zero point derived from input's own runtime range. |
| einsum | [common] fun einsum(equation: String, operands: List<Tensor>): Tensor einsum(equation: String, operands: List<Tensor>): the einsum operator. Einstein-summation contraction from a notation string. equation names each operand's axes and (after ->) the output axes; omitting -> keeps the once-appearing labels in alphabetical order. One- and two-operand equations are served: one operand reduces and permutes; two operands contract through a single batched matmul. |
| elu | [common] fun elu(input: Tensor, alpha: Double = 1.0, scale: Double = 1.0, inputScale: Double = 1.0): Tensor elu(input: Tensor, alpha: Double = 1.0, scale: Double = 1.0, inputScale: Double = 1.0): the elu operator. Exponential linear unit. Elementwise; input_scale feeds only the exponential's argument. |
| eluInPlace | [common] fun eluInPlace(self: Tensor, alpha: Double = 1.0, scale: Double = 1.0, inputScale: Double = 1.0): Tensor eluInPlace(self: Tensor, alpha: Double = 1.0, scale: Double = 1.0, inputScale: Double = 1.0): the elu_ operator. In-place elu: writes the result through self; same formula, arguments, and error conditions as elu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| embedding | [common] fun embedding(indices: Tensor, weight: Tensor, bias: Tensor? = null, activation: Activation? = null): Tensor embedding(indices: Tensor, weight: Tensor, bias: Tensor? = null, activation: Activation? = null): the embedding operator. Embedding lookup: gathers rows of weight by indices, with an optional bias + activation epilogue. |
| empty | [common] fun empty(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): Tensor empty(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): the empty operator. Uninitialized tensor of the given shape; the contents are whatever the allocator hands back; write every element before reading any. Each shape extent is an int literal OR a 0-D integer Tensor (a tensor extent traces symbolically, never syncing to the host). |
| emptyLike | [common] fun emptyLike(reference: Tensor, device: Placement? = null): Tensor emptyLike(reference: Tensor, device: Placement? = null): the empty_like operator. Uninitialized tensor with reference's shape and dtype. |
| emptyStrided | [common] fun emptyStrided(size: LongArray, stride: LongArray, dtype: DType, device: Placement? = null): Tensor emptyStrided(size: LongArray, stride: LongArray, dtype: DType, device: Placement? = null): the empty_strided operator. Uninitialized tensor with caller-chosen sizes AND strides (element units). The storage is sized to the strided footprint. Extents and strides are host integers (no tensor-valued entries on this factory). |
| eq | [common] fun eq(input: Tensor, other: Tensor): Tensor eq(input: Tensor, other: Tensor): the eq operator. Elementwise equality a == other, the output contract the whole comparison family shares: the result is a Bool tensor at the broadcast shape (operands promote to their dominant dtype before the compare). A scalar other compares against every element. The other comparisons state "output as eq" instead of restating this.[common] fun eq(input: Tensor, other: Double): Tensor eq(input: Tensor, other: Double): the number form of eq, other as a scalar. |
| eqInPlace | [common] fun eqInPlace(self: Tensor, other: Tensor): Tensor eqInPlace(self: Tensor, other: Tensor): the eq_ operator. In-place eq: writes the comparison result through x (x keeps its own dtype; true/false land as one/zero). Writes through self and returns it, so calls chain.[common] fun eqInPlace(self: Tensor, other: Double): Tensor eqInPlace(self: Tensor, other: Double): the number form of eq_, other as a scalar. |
| erf | [common] fun erf(input: Tensor): Tensor erf(input: Tensor): the erf operator. Elementwise Gauss error function. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| erfc | [common] fun erfc(input: Tensor): Tensor erfc(input: Tensor): the erfc operator. Elementwise complementary error function. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| erfcInPlace | [common] fun erfcInPlace(self: Tensor): Tensor erfcInPlace(self: Tensor): the erfc_ operator. In-place erfc: writes the result through self; same formula, arguments, and error conditions as erfc(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| erfInPlace | [common] fun erfInPlace(self: Tensor): Tensor erfInPlace(self: Tensor): the erf_ operator. In-place erf: writes the result through self; same formula, arguments, and error conditions as erf(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| erfinv | [common] fun erfinv(input: Tensor): Tensor erfinv(input: Tensor): the erfinv operator. Elementwise inverse error function. Inputs outside (-1, 1) produce NaN; ±1 produce ±inf. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| erfinvInPlace | [common] fun erfinvInPlace(self: Tensor): Tensor erfinvInPlace(self: Tensor): the erfinv_ operator. In-place erfinv: writes the result through self; same formula, arguments, and error conditions as erfinv(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| exp | [common] fun exp(input: Tensor): Tensor exp(input: Tensor): the exp operator. Elementwise natural exponential. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| expand | [common] fun expand(input: Tensor, shape: LongArray, bidirectional: Boolean = false): Tensor expand(input: Tensor, shape: LongArray, bidirectional: Boolean = false): the expand operator. Broadcast input to a larger shape WITHOUT copying. Size-1 dims stretch to their target; leading target dims with no input counterpart broadcast; -1 keeps the input's extent at that position. Extents are literals or 0-D integer Tensors (trace symbolically). With bidirectional = true the aligned dims broadcast BOTH ways (the ONNX Expand rule): a target extent of 1 against a non-1 input dim keeps the input dim (the pairwise max) instead of raising. The default is the strict directed form (broadcast_to). Returns a VIEW: every stretched position aliases ONE stored element; treat the result as read-only (or contiguous it to materialize). |
| expandAs | [common] fun expandAs(input: Tensor, other: Tensor): Tensor expandAs(input: Tensor, other: Tensor): the expand_as operator. expand to other's shape; other supplies extents only; its data is never read. Same broadcasting rules and view semantics as expand. |
| expInPlace | [common] fun expInPlace(self: Tensor): Tensor expInPlace(self: Tensor): the exp_ operator. In-place exp: writes the result through self; same formula, arguments, and error conditions as exp(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| exponentialInPlace | [common] fun exponentialInPlace(self: Tensor, lambd: Double = 1.0, device: Placement? = null): Tensor exponentialInPlace(self: Tensor, lambd: Double = 1.0, device: Placement? = null): the exponential_ operator. In-place exponential fill with rate lambd: overwrites self with draws from p(x) = lambd * exp(-lambd * x) (x >= 0) and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| eye | [common] fun eye(n: Tensor, m: Tensor? = null, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor eye(n: Tensor, m: Tensor? = null, dtype: DType = DType.FLOAT32, device: Placement? = null): the eye operator. A 2-D identity matrix [n, m]: 1 on the diagonal, 0 elsewhere; m defaults to n (square).[common] fun eye(n: Double, m: Tensor? = null, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor eye(n: Double, m: Tensor? = null, dtype: DType = DType.FLOAT32, device: Placement? = null): the number form of eye, n as a scalar. |
| fastGelu | [common] fun fastGelu(input: Tensor): Tensor fastGelu(input: Tensor): the fast_gelu operator. FastGELU: the tanh GELU approximation as a standalone op. |
| fill | [common] fun fill(input: Tensor, value: Tensor, device: Placement? = null): Tensor fill(input: Tensor, value: Tensor, device: Placement? = null): the fill operator. A new tensor with input's shape and dtype, every element set to value. input supplies GEOMETRY only; its data is never read. value is a literal or a 0-D Tensor (a tensor value traces symbolically).[common] fun fill(input: Tensor, value: Double, device: Placement? = null): Tensor fill(input: Tensor, value: Double, device: Placement? = null): the number form of fill, value as a scalar. |
| fillDiagonal | [common] fun fillDiagonal(input: Tensor, fillValue: Double, wrap: Boolean = false, device: Placement? = null): Tensor fillDiagonal(input: Tensor, fillValue: Double, wrap: Boolean = false, device: Placement? = null): the fill_diagonal operator. A copy of the 2-D matrix input with fill_value written along its diagonal. Without wrap the diagonal stops at min(rows, cols); with wrap = true a TALL matrix continues the diagonal below the wrap row, restarting at the left edge (the classic tall-matrix wrap). |
| fillDiagonalInPlace | [common] fun fillDiagonalInPlace(self: Tensor, fillValue: Double, wrap: Boolean = false, device: Placement? = null): Tensor fillDiagonalInPlace(self: Tensor, fillValue: Double, wrap: Boolean = false, device: Placement? = null): the fill_diagonal_ operator. In-place fill_diagonal: writes the diagonal through self; same arguments and error conditions as fill_diagonal(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| fillInPlace | [common] fun fillInPlace(self: Tensor, value: Tensor, device: Placement? = null): Tensor fillInPlace(self: Tensor, value: Tensor, device: Placement? = null): the fill_ operator. In-place fill: overwrites every element of self with value; same arguments and error conditions as fill(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun fillInPlace(self: Tensor, value: Double, device: Placement? = null): Tensor fillInPlace(self: Tensor, value: Double, device: Placement? = null): the number form of fill_, value as a scalar. |
| flatten | [common] fun flatten(input: Tensor, startDim: Long = 0, endDim: Long = -1L): Tensor flatten(input: Tensor, startDim: Long = 0L, endDim: Long = -1L): the flatten operator. Collapse the INCLUSIVE dim range [start_dim, end_dim] into one dim; the defaults collapse everything to 1-D. Returns a view sharing storage when the input's stride layout factors into the merged shape (contiguous inputs always do); copies into a fresh dense tensor otherwise. |
| flip | [common] fun flip(input: Tensor, dims: LongArray): Tensor flip(input: Tensor, dims: LongArray): the flip operator. Reverse the element order along each dim in dims. Always copies; a reversed layout is not expressible as a view in this runtime, so this op never returns one. |
| fliplr | [common] fun fliplr(input: Tensor): Tensor fliplr(input: Tensor): the fliplr operator. flip on dim 1: reverse each row's column order. Input rank must be >= 2. Copies. |
| flipud | [common] fun flipud(input: Tensor): Tensor flipud(input: Tensor): the flipud operator. flip on dim 0: reverse the leading dim's order. Copies. |
| floor | [common] fun floor(input: Tensor): Tensor floor(input: Tensor): the floor operator. Elementwise floor: the largest integer not above input. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| floorDivide | [common] fun floorDivide(input: Tensor, other: Tensor): Tensor floorDivide(input: Tensor, other: Tensor): the floor_divide operator. Elementwise floor(a / other), i.e. ops::div with RoundingMode::Floor. Broadcasts and promotes as add.[common] fun floorDivide(input: Tensor, other: Double): Tensor floorDivide(input: Tensor, other: Double): the number form of floor_divide, other as a scalar. |
| floorDivideInPlace | [common] fun floorDivideInPlace(self: Tensor, other: Tensor): Tensor floorDivideInPlace(self: Tensor, other: Tensor): the floor_divide_ operator. In-place floor_divide: writes the floored quotient through x. Writes through self and returns it, so calls chain.[common] fun floorDivideInPlace(self: Tensor, other: Double): Tensor floorDivideInPlace(self: Tensor, other: Double): the number form of floor_divide_, other as a scalar. |
| floorInPlace | [common] fun floorInPlace(self: Tensor): Tensor floorInPlace(self: Tensor): the floor_ operator. In-place floor: writes the result through self; same formula, arguments, and error conditions as floor(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| fmax | [common] fun fmax(input: Tensor, other: Tensor): Tensor fmax(input: Tensor, other: Tensor): the fmax operator. Elementwise IEEE 754 maximum, NaN-IGNORING: where one operand is NaN the other value wins (both NaN gives NaN). For the NaN-propagating law use ops::maximum. Broadcasts and promotes as add. |
| fmin | [common] fun fmin(input: Tensor, other: Tensor): Tensor fmin(input: Tensor, other: Tensor): the fmin operator. Elementwise IEEE 754 minimum, NaN-IGNORING (the fmax dual). Broadcasts and promotes as add. |
| fmod | [common] fun fmod(input: Tensor, other: Tensor): Tensor fmod(input: Tensor, other: Tensor): the fmod operator. Elementwise remainder with the sign of the DIVIDEND, i.e. ops::mod with ModMode::C. Broadcasts and promotes as add.[common] fun fmod(input: Tensor, other: Double): Tensor fmod(input: Tensor, other: Double): the number form of fmod, other as a scalar. |
| fmodInPlace | [common] fun fmodInPlace(self: Tensor, other: Tensor): Tensor fmodInPlace(self: Tensor, other: Tensor): the fmod_ operator. In-place fmod: writes the remainders through x. Writes through self and returns it, so calls chain.[common] fun fmodInPlace(self: Tensor, other: Double): Tensor fmodInPlace(self: Tensor, other: Double): the number form of fmod_, other as a scalar. |
| fold | [common] fun fold(input: Tensor, outputSize: LongArray, kernelSize: LongArray, dilation: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), stride: LongArray = longArrayOf(), mode: PadMode = PadMode.CONSTANT, value: Double? = null): Tensor fold(input: Tensor, outputSize: LongArray, kernelSize: LongArray, dilation: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), stride: LongArray = longArrayOf(), mode: PadMode = PadMode.CONSTANT, value: Double? = null): the fold operator. col2im: the inverse of unfold; it sums overlapping windows back into a channels-last image. Input [N, L, prod(kernel_size), C]; overlapping window contributions ADD (so fold(unfold(x)) multiplies overlapped elements by their coverage count). |
| frac | [common] fun frac(input: Tensor): Tensor frac(input: Tensor): the frac operator. Elementwise fractional part. Keeps the sign of input (frac(-1.5) == -0.5). Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| fracInPlace | [common] fun fracInPlace(self: Tensor): Tensor fracInPlace(self: Tensor): the frac_ operator. In-place frac: writes the result through self; same formula, arguments, and error conditions as frac(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| frexp | [common] fun frexp(input: Tensor): List<Tensor> frexp(input: Tensor): the frexp operator. Decomposes each element into mantissa * 2^exponent with ` |
| full | [common] fun full(shape: LongArray, fillValue: Tensor, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): Tensor full(shape: LongArray, fillValue: Tensor, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): the full operator. A new tensor of the given shape with every element set to fill_value. Each shape extent is an int literal or a 0-D integer Tensor (traces symbolically). fill_value is a scalar literal or a tensor: a bool keeps Bool kind, an integer keeps its integer-ness across the boundary and stays exact past 2^53, and a tensor fills every element from its values (a 0-D tensor, or one broadcastable to shape), the same write as fill_ on an empty tensor in one kernel. Tensor::full is the literal-only form.[common] fun full(shape: LongArray, fillValue: Double, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): Tensor full(shape: LongArray, fillValue: Double, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): the number form of full, fillValue as a scalar. |
| fullLike | [common] fun fullLike(reference: Tensor, fillValue: Tensor, device: Placement? = null): Tensor fullLike(reference: Tensor, fillValue: Tensor, device: Placement? = null): the full_like operator. A new tensor with reference's shape and dtype, filled with fill_value. fill_value takes the same forms as full's: a bool, an integer, a double, or a tensor (0-D, or broadcastable to the reference's shape).[common] fun fullLike(reference: Tensor, fillValue: Double, device: Placement? = null): Tensor fullLike(reference: Tensor, fillValue: Double, device: Placement? = null): the number form of full_like, fillValue as a scalar. |
| gatedDeltaUpdate | [common] fun gatedDeltaUpdate(query: Tensor, key: Tensor, value: Tensor, beta: Tensor, gate: Tensor, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, scale: Double? = null, gateBias: Tensor? = null, gateScale: Tensor? = null): Tensor gatedDeltaUpdate(query: Tensor, key: Tensor, value: Tensor, beta: Tensor, gate: Tensor, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, scale: Double? = null, gateBias: Tensor? = null, gateScale: Tensor? = null): the gated_delta_update operator. Gated delta-rule recurrence step: per token, the [B, HV, K, V] Float32 state decays by exp(g), takes the delta-rule rank-1 update (beta·k) ⊗ (v − Sᵀk), and emits o = (scale·q)ᵀ S, updated IN PLACE. The gate's RANK picks the family: [B, T, HV] = one scalar per value head; [B, T, HV, K] = per key dim. q/k [B, T, H, K] (HV % H == 0, grouped heads), v [B, T, HV, V], beta [B, T, HV]; scale defaults to K^-1/2; seq_lens ([B] Int32) bounds ragged prefill rows. slot_ids ([B] Int32, device-resident) addresses state as a SLAB [num_slots, HV, K, V]: batch row b reads/updates slab row slot_ids[b] in place (ids in range and DISTINCT per call, the caller's contract); absent keeps state row b. TWO gate forms, told apart by which inputs are bound: with gate_bias and gate_scale absent, gate IS the log-space decay (Float32 or the activations' dtype); with both bound (Float32 gate_bias``[HV] beside a [B, T, HV] gate or [HV, K] beside a [B, T, HV, K] one, the checkpoint's dt_bias; Float32 gate_scale``[HV], the once-folded -exp(A_log)), gate is the RAW gate projection slice at the activations' dtype and the kernel forms the decay gate_scale * softplus(g + gate_bias) in fp32 registers, so no add / softplus / mul pass and no fp32 transient precede the call. One without the other rejects; a backend without the raw-gate arm declines it typed. |
| gatedRmsNorm | [common] fun gatedRmsNorm(input: Tensor, gate: Tensor, normalizedShape: LongArray, weight: Tensor? = null, eps: Double? = null): Tensor gatedRmsNorm(input: Tensor, gate: Tensor, normalizedShape: LongArray, weight: Tensor? = null, eps: Double? = null): the gated_rms_norm operator. Fused group RMS norm + post-norm silu(z) gate: out = rms_norm(x) · silu(z), the norm taken over the trailing normalized_shape group extent ({vd} for a per-head norm over [.., HV, vd]). weight is the optional [group] gamma applied after the normalization; an absent eps takes the runtime's rms-norm default. |
| gather | [common] fun gather(input: Tensor, dim: Long, index: Tensor): Tensor gather(input: Tensor, dim: Long, index: Tensor): the gather operator. Axis-wise gather: read input at positions given by index along dim. out[i][j][k] = x[index[i][j][k]][j][k] for dim = 0 (likewise for any other dim; only that axis's coordinate is replaced). The output takes index's shape and input's dtype. Index tensors are Int32 or Int64 on every backend; entries must lie in [0, x's dim extent). |
| ge | [common] fun ge(input: Tensor, other: Tensor): Tensor ge(input: Tensor, other: Tensor): the ge operator. Elementwise a >= other; output as eq.[common] fun ge(input: Tensor, other: Double): Tensor ge(input: Tensor, other: Double): the number form of ge, other as a scalar. |
| geglu | [common] fun geglu(input: Tensor, approximate: GeluMode = GeluMode.NONE): Tensor geglu(input: Tensor, approximate: GeluMode = GeluMode.NONE): the geglu operator. GeGLU gated activation over a concatenated gate‖up tensor. The input's LAST dimension must be even (2d); the output halves it (d). approximate selects the exact (erf) or tanh GELU for the gate. |
| geInPlace | [common] fun geInPlace(self: Tensor, other: Tensor): Tensor geInPlace(self: Tensor, other: Tensor): the ge_ operator. In-place ge: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain.[common] fun geInPlace(self: Tensor, other: Double): Tensor geInPlace(self: Tensor, other: Double): the number form of ge_, other as a scalar. |
| gelu | [common] fun gelu(input: Tensor, approximate: GeluMode = GeluMode.NONE): Tensor gelu(input: Tensor, approximate: GeluMode = GeluMode.NONE): the gelu operator. Gaussian error linear unit. approximate = GeluMode::Tanh selects the cheaper tanh form used by many transformer checkpoints; GeluMode::None is the exact erf form. |
| geluInPlace | [common] fun geluInPlace(self: Tensor, approximate: GeluMode = GeluMode.NONE): Tensor geluInPlace(self: Tensor, approximate: GeluMode = GeluMode.NONE): the gelu_ operator. In-place gelu: writes the result through self; same formula, arguments, and error conditions as gelu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| generateRotaryCache | [common] fun generateRotaryCache(rotaryDim: Long, maxPositions: Long, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null, device: Placement? = null, yarnTruncate: Boolean? = null): List<Tensor> generateRotaryCache(rotaryDim: Long, maxPositions: Long, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null, device: Placement? = null, yarnTruncate: Boolean? = null): the generate_rotary_cache operator. The rotary angle tables of one context-scaling family: cos and sin, each [max_positions, rotary_dim / 2] Float32. |
| glu | [common] fun glu(input: Tensor, dim: Long = -1L): Tensor glu(input: Tensor, dim: Long = -1L): the glu operator. Gated linear unit: splits input in half along dim and gates the first half with the sigmoid of the second. The size of dim must be even; the output halves it. |
| gridSample | [common] fun gridSample(input: Tensor, grid: Tensor, mode: GridSampleMode = GridSampleMode.BILINEAR, paddingMode: GridSamplePaddingMode = GridSamplePaddingMode.ZEROS, alignCorners: Boolean = false): Tensor gridSample(input: Tensor, grid: Tensor, mode: GridSampleMode = GridSampleMode.BILINEAR, paddingMode: GridSamplePaddingMode = GridSamplePaddingMode.ZEROS, alignCorners: Boolean = false): the grid_sample operator. Samples a channels-last input at arbitrary grid coordinates. input``[N, spatial_in.., C]; grid``[N, spatial_out.., S] with S the spatial rank, coordinates normalized to [-1, 1]. mode picks the interpolation (Bilinear / Nearest / …), padding_mode the out-of-range policy (Zeros / Border / Reflection), align_corners the corner convention. |
| groupNorm | [common] fun groupNorm(input: Tensor, numGroups: Long, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): Tensor groupNorm(input: Tensor, numGroups: Long, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): the group_norm operator. Group normalization: channels split into num_groups groups, normalized per group (channels-last). Statistics are computed per (sample, group) over the group's channels and the spatial dims; with running_mean / running_var present they are folded per channel instead (batch-norm style). Optional fused activation applies to the result. |
| groupQueryAttention | [common] fun groupQueryAttention(query: Tensor, key: Tensor, value: Tensor, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null): List<Tensor> groupQueryAttention(query: Tensor, key: Tensor, value: Tensor, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null): the group_query_attention operator. Fused GQA with RoPE + KV cache. Bind out_present_key/out_present_value to the same buffers as past_key/past_value (a KVCache::keys/values(layer) view) + pass kvcache_start to append the new post-RoPE K/V IN PLACE into the cache (the decode perf path). q/k/v are hidden-folded [ΣS, heads*head_dim]; num_heads/kv_num_heads drive the in-op head split (GQA). head_sink is the per-head softmax sink [H_q], a virtual logit folded into the softmax denominator (attention-sink models bind one per layer), the same contract as group_query_attention_varlen; it rides the parameter tail here. q_norm_gain/k_norm_gain engage the POST-rope per-head RMS norm: after the in-op rotation, every head's [head_dim] q (and new-k) vector is RMS-normalized with the gain BEFORE any cache append, so the cache holds rotated+normed keys. Each gain is rank-1 [head_dim] (one vector shared across heads; any other shape rejects), any float dtype, applied at its own dtype. The gains require the in-op rope planes (rope_cos/rope_sin), and they travel WITH qk_norm_eps: pass the model's own rms-norm epsilon alongside the gains, or neither (a gain without the epsilon, or an epsilon with no gain, rejects). Absent ⇒ the gain-less path, unchanged. A rope-free per-head norm composes qk_rms_norm instead. slot_ids ([B] Int32) names the cache row each batch row appends to and attends from on a continuous [max_seqs, H_kv, max_seq, D] cache: a sequence keeps its row while the batch composition changes around it. Absent, batch row b uses cache row b. Every entry must lie in [0, max_seqs) and no two rows may share one (each rejects). On that cache kvcache_start stays the rank-1 [B] layout selector whose values are not read: each row's write offset is cu_seqlens_k[b] - q_len[b]. A paged block table and the dense in-place form take no slot_ids (the block table is its own row map; the dense form addresses rows by batch index). A paged block table ([B, max_blocks] Int32) is validated whole before any read: every entry is -1 (the unused-tail pad) or a block index in [0, num_blocks), and no two entries of one row name the same physical block; an out-of-range entry and a repeated one alike refuse INVALID_ARGUMENT (a repeat would alias two positions onto one block and decode wrong values silently). A read-only attend (key and value absent) takes no present outputs: use attention_over_cache, or leave out_present_key/out_present_value unbound, and present_key/present_value come back undefined. A bound present refuses INVALID_ARGUMENT. kept_prefix keeps each sequence's first P keys attended beside the sliding window (a 0-D value for every sequence, or [B] one per sequence; Int32 or Int64): under the causal bound a key is attended when it is inside the first P keys or inside the window; absent, the window alone. A negative P refuses. |
| groupQueryAttentionVarlen | [common] fun groupQueryAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): List<Tensor> groupQueryAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): the group_query_attention_varlen operator. The varlen variant of the fused GQA above, with the same qk-norm tail: q_norm_gain/k_norm_gain (rank-1 [head_dim], post-rope, applied before the K/V append) travel with qk_norm_eps (the model's rms-norm epsilon) and require the in-op rope planes; absent ⇒ unchanged. The gains are the after-rotation order; a model that norms its heads before the rotation (the Qwen3 order) runs qk_rms_norm on its projections ahead of this op and passes no gains. The same slot_ids contract: on a continuous cache it names each batch row's cache row ([B] Int32, in range, pairwise distinct; absent = row b for batch row b), and kvcache_start's values stay unread there (the write offsets derive from cu_seqlens_k). The same block-table contract on a paged pool: every entry of the [B, max_blocks] table is -1 or a block index in [0, num_blocks), and no two entries of one row name the same physical block; an out-of-range or repeated entry refuses INVALID_ARGUMENT before any read or append. The same read-only rule: with key and value absent it takes no present outputs; use attention_over_cache, or leave out_present_key/out_present_value unbound. kept_prefix keeps each request's first P keys attended beside the sliding window (a 0-D value or [B], Int32 or Int64), the same law as attention's. kv_position_offset ([B], Int32 or Int64) is each paged row's served-key offset: when a windowed paged cache serves only a row's recent blocks, its served key at view position i sits at absolute position i + offset, so the in-op rotary and the window read true positions; pass KVCache::StepIndices::kv_position_offset through as-is (undefined on every other cache). |
| gt | [common] fun gt(input: Tensor, other: Tensor): Tensor gt(input: Tensor, other: Tensor): the gt operator. Elementwise a > other; output as eq.[common] fun gt(input: Tensor, other: Double): Tensor gt(input: Tensor, other: Double): the number form of gt, other as a scalar. |
| gtInPlace | [common] fun gtInPlace(self: Tensor, other: Tensor): Tensor gtInPlace(self: Tensor, other: Tensor): the gt_ operator. In-place gt: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain.[common] fun gtInPlace(self: Tensor, other: Double): Tensor gtInPlace(self: Tensor, other: Double): the number form of gt_, other as a scalar. |
| hardshrink | [common] fun hardshrink(input: Tensor, lambd: Double = 0.5): Tensor hardshrink(input: Tensor, lambd: Double = 0.5): the hardshrink operator. Hard shrinkage: zeroes every element within [-lambd, lambd]. |
| hardshrinkInPlace | [common] fun hardshrinkInPlace(self: Tensor, lambd: Double = 0.5): Tensor hardshrinkInPlace(self: Tensor, lambd: Double = 0.5): the hardshrink_ operator. In-place hardshrink: writes the result through self; same formula, arguments, and error conditions as hardshrink(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| hardsigmoid | [common] fun hardsigmoid(input: Tensor): Tensor hardsigmoid(input: Tensor): the hardsigmoid operator. Piecewise-linear sigmoid approximation. |
| hardsigmoidInPlace | [common] fun hardsigmoidInPlace(self: Tensor): Tensor hardsigmoidInPlace(self: Tensor): the hardsigmoid_ operator. In-place hardsigmoid: writes the result through self; same formula, arguments, and error conditions as hardsigmoid(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| hardswish | [common] fun hardswish(input: Tensor): Tensor hardswish(input: Tensor): the hardswish operator. Piecewise-linear swish approximation (the MobileNet-v3 form). |
| hardswishInPlace | [common] fun hardswishInPlace(self: Tensor): Tensor hardswishInPlace(self: Tensor): the hardswish_ operator. In-place hardswish: writes the result through self; same formula, arguments, and error conditions as hardswish(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| hardtanh | [common] fun hardtanh(input: Tensor, minVal: Double = -1.0, maxVal: Double = 1.0): Tensor hardtanh(input: Tensor, minVal: Double = -1.0, maxVal: Double = 1.0): the hardtanh operator. Clamps every element to [min_val, max_val]. |
| hardtanhInPlace | [common] fun hardtanhInPlace(self: Tensor, minVal: Double = -1.0, maxVal: Double = 1.0): Tensor hardtanhInPlace(self: Tensor, minVal: Double = -1.0, maxVal: Double = 1.0): the hardtanh_ operator. In-place hardtanh: writes the result through self; same formula, arguments, and error conditions as hardtanh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| hash128 | [common] fun hash128(input: Tensor, seed: Long = 0): Tensor hash128(input: Tensor, seed: Long = 0L): the hash_128 operator. 128-bit sibling of hash_64; same byte/portability contract. |
| hash256 | [common] fun hash256(input: Tensor, seed: Long = 0): Tensor hash256(input: Tensor, seed: Long = 0L): the hash_256 operator. 256-bit sibling of hash_64; same byte/portability contract. |
| hash64 | [common] fun hash64(input: Tensor, seed: Long = 0): Tensor hash64(input: Tensor, seed: Long = 0L): the hash_64 operator. 64-bit content hash of input's packed bytes (row-major logical order; strided inputs are materialized internally). A frozen function of (seed, bytes): identical on every platform, backend, and ISA tier, so digests are safe to persist and compare across devices. seed perturbs the key; 0 selects the runtime's fixed content-addressing key. |
| hashChain | [common] fun hashChain(rows: Tensor, parent: Tensor, seed: Long = 0): Tensor hashChain(rows: Tensor, parent: Tensor, seed: Long = 0L): the hash_chain operator. Rowwise 128-bit CHAINED hash: for rows [N, ..] and a [2] UInt64 parent (zeros = the chain's zero element), row i's digest hashes (digest i-1 ‖ row i's packed bytes); one dispatch computes a whole sequence's rolling content hashes. Same portability contract as hash_64. |
| hashLanes | [common] fun hashLanes(input: Tensor, key: Long = 0, algo: HashLaneAlgo = HashLaneAlgo.AUTO, maskBits: Long = 0): Tensor hashLanes(input: Tensor, key: Long = 0L, algo: HashLaneAlgo = HashLaneAlgo.AUTO, maskBits: Long = 0L): the hash_lanes operator. Elementwise per-lane BIJECTIVE hash, shape-preserving: every lane of an integer input maps through a permutation of its width, so distinct lane values stay distinct (a lane-wise content id, a bucket index, a deterministic shuffle key). A 32-bit lane lands as UInt32, a 64-bit lane as UInt64; algo selects the permutation within the input's width (HashLaneAlgo::Auto: Triple32 for 32-bit lanes, Moremur for 64-bit ones). key perturbs the permutation (0 is a fixed valid default, so the same key gives the same map on every platform and backend). mask_bits in (0, width] restricts the bijection to [0, 2^mask_bits): a lane below that bound maps to a lane below it, so a table of that size is permuted onto itself; 0 (and the lane width) keeps the full width. |
| hashTensor | [common] fun hashTensor(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, mode: HashTensorMode = HashTensorMode.XOR_SUM): Tensor hashTensor(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, mode: HashTensorMode = HashTensorMode.XOR_SUM): the hash_tensor operator. Bitwise hash-reduction over dims (empty = all): each element's bits are mixed and combined per mode into a UInt64 digest, standard reduction shape. Order-independent under HashTensorMode::XorSum, and identical on every backend, a cheap whole-tensor fingerprint for parity checks and cache keys. |
| histogram | [common] fun histogram(input: Tensor, bins: Long = 100, range: Tensor? = null, weight: Tensor? = null, density: Boolean = false): List<Tensor> histogram(input: Tensor, bins: Long = 100L, range: Tensor? = null, weight: Tensor? = null, density: Boolean = false): the histogram operator. |
| huberLoss | [common] fun huberLoss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, delta: Double = 1.0): Tensor huberLoss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, delta: Double = 1.0): the huber_loss operator. Huber loss: quadratic near zero, linear past the delta knee. |
| hypot | [common] fun hypot(input: Tensor, other: Tensor): Tensor hypot(input: Tensor, other: Tensor): the hypot operator. Elementwise hypotenuse . Broadcasts and promotes as add. |
| hypotInPlace | [common] fun hypotInPlace(self: Tensor, other: Tensor): Tensor hypotInPlace(self: Tensor, other: Tensor): the hypot_ operator. In-place hypot: writes the hypotenuses through x. Writes through self and returns it, so calls chain. |
| indexAdd | [common] fun indexAdd(input: Tensor, dim: Long, indices: Tensor, src: Tensor): Tensor indexAdd(input: Tensor, dim: Long, indices: Tensor, src: Tensor): the index_add operator. A copy of input with rows of src ACCUMULATED at indices along dim: out[.., indices[i], ..] += src[.., i, ..]; duplicate indices add up. |
| indexAddInPlace | [common] fun indexAddInPlace(self: Tensor, dim: Long, indices: Tensor, src: Tensor): Tensor indexAddInPlace(self: Tensor, dim: Long, indices: Tensor, src: Tensor): the index_add_ operator. In-place index_add: self[.., indices[i], ..] += src[.., i, ..]; same arguments and error conditions as index_add(). Accumulates through self's storage (a strided view reaches its base buffer; aliases observe the write). Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| indexCopy | [common] fun indexCopy(input: Tensor, dim: Long, indices: Tensor, src: Tensor): Tensor indexCopy(input: Tensor, dim: Long, indices: Tensor, src: Tensor): the index_copy operator. A copy of input with rows of src WRITTEN at indices along dim: out[.., indices[i], ..] = src[.., i, ..]. Prefer unique indices; a duplicated position is written more than once. Same shape/index-dtype rules as index_add. |
| indexCopyInPlace | [common] fun indexCopyInPlace(self: Tensor, dim: Long, indices: Tensor, src: Tensor): Tensor indexCopyInPlace(self: Tensor, dim: Long, indices: Tensor, src: Tensor): the index_copy_ operator. In-place index_copy: self[.., indices[i], ..] = src[.., i, ..]; same arguments and error conditions as index_copy(). Writes through self's storage (a strided view, e.g. a cache slice, reaches its base buffer; aliases observe the write). Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| indexFill | [common] fun indexFill(input: Tensor, dim: Long, indices: Tensor, value: Tensor): Tensor indexFill(input: Tensor, dim: Long, indices: Tensor, value: Tensor): the index_fill operator. A copy of input with the rows at indices along dim set to value: out[.., indices[i], ..] = value. indices is 1-D Int32/Int64; value is a literal or a 0-D Tensor.[common] fun indexFill(input: Tensor, dim: Long, indices: Tensor, value: Double): Tensor indexFill(input: Tensor, dim: Long, indices: Tensor, value: Double): the number form of index_fill, value as a scalar. |
| indexFillInPlace | [common] fun indexFillInPlace(self: Tensor, dim: Long, indices: Tensor, value: Tensor): Tensor indexFillInPlace(self: Tensor, dim: Long, indices: Tensor, value: Tensor): the index_fill_ operator. In-place index_fill: self[.., indices[i], ..] = value; same arguments and error conditions as index_fill(). Fills through self's storage (a strided view reaches its base buffer; aliases observe the write). Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun indexFillInPlace(self: Tensor, dim: Long, indices: Tensor, value: Double): Tensor indexFillInPlace(self: Tensor, dim: Long, indices: Tensor, value: Double): the number form of index_fill_, value as a scalar. |
| indexSelect | [common] fun indexSelect(input: Tensor, dim: Long, index: Tensor): Tensor indexSelect(input: Tensor, dim: Long, index: Tensor): the index_select operator. Select whole rows of input at index positions along dim. Output = input's shape with the dim extent replaced by index's length; rows may repeat and appear in any order. index is 1-D Int32/Int64; prefer Int32 where the extents allow (half the index bytes). |
| inner | [common] fun inner(input: Tensor, other: Tensor): Tensor inner(input: Tensor, other: Tensor): the inner operator. Last-axis contraction of two tensors. Sums over the LAST dim of each operand (which must agree); the output is the cartesian product of the leading shapes: [*a_lead, K] x [*b_lead, K] -> [*a_lead, *b_lead]. |
| instanceNorm | [common] fun instanceNorm(input: Tensor, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): Tensor instanceNorm(input: Tensor, weight: Tensor? = null, bias: Tensor? = null, runningMean: Tensor? = null, runningVar: Tensor? = null, eps: Double? = null, activation: Activation? = null): the instance_norm operator. Instance normalization: statistics per (sample, channel) over the spatial dims (channels-last). With running_mean / running_var present they are folded instead of computing per-instance statistics (inference form; no training mode). Optional fused activation applies to the result. |
| interpolate | [common] fun interpolate(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), mode: InterpMode = InterpMode.NEAREST, alignCorners: Boolean? = null, recomputeScaleFactor: Boolean = false, antialias: Boolean = false): Tensor interpolate(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), mode: InterpMode = InterpMode.NEAREST, alignCorners: Boolean? = null, recomputeScaleFactor: Boolean = false, antialias: Boolean = false): the interpolate operator. Resample the spatial dims to sizes OR by scale_factors (exactly one given; entries may be int literals or 0-D Tensors, so a dynamic output size flows without a host read). Rank picks the spatial variant. |
| isclose | [common] fun isclose(input: Tensor, other: Tensor, rtol: Double = 1.0E-5, atol: Double = 1.0E-8, equalNan: Boolean = false): Tensor isclose(input: Tensor, other: Tensor, rtol: Double = 1e-5, atol: Double = 1e-8, equalNan: Boolean = false): the isclose operator. Elementwise approximate equality, the elementwise map behind ops::allclose, same formula and defaults (rtol``1e-5, atol``1e-8, equal_nan``false). |
| isfinite | [common] fun isfinite(input: Tensor): Tensor isfinite(input: Tensor): the isfinite operator. Elementwise finiteness test (neither infinite nor NaN); output as eq. Integer inputs are finite everywhere. |
| isinf | [common] fun isinf(input: Tensor): Tensor isinf(input: Tensor): the isinf operator. Elementwise infinity test (either sign); output as eq. |
| isnan | [common] fun isnan(input: Tensor): Tensor isnan(input: Tensor): the isnan operator. Elementwise NaN test; output as eq. |
| isneginf | [common] fun isneginf(input: Tensor): Tensor isneginf(input: Tensor): the isneginf operator. Elementwise -inf test; output as eq. |
| isposinf | [common] fun isposinf(input: Tensor): Tensor isposinf(input: Tensor): the isposinf operator. Elementwise +inf test; output as eq. |
| klDiv | [common] fun klDiv(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, logTarget: Boolean = false): Tensor klDiv(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, logTarget: Boolean = false): the kl_div operator. Kullback–Leibler divergence loss. input is given in LOG space; target is probabilities unless log_target is set (then both are logs): |
| kron | [common] fun kron(input: Tensor, other: Tensor): Tensor kron(input: Tensor, other: Tensor): the kron operator. Kronecker product of two same-rank tensors. Output dim i is input.shape[i] * b.shape[i]; each input element scales a full copy of other into its block. |
| kthvalue | [common] fun kthvalue(input: Tensor, k: Long, dim: Long = -1L, keepdim: Boolean = false): List<Tensor> kthvalue(input: Tensor, k: Long, dim: Long = -1L, keepdim: Boolean = false): the kthvalue operator. |
| l1Loss | [common] fun l1Loss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN): Tensor l1Loss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN): the l1_loss operator. Mean-absolute-error loss between input and target. |
| layerNorm | [common] fun layerNorm(input: Tensor, normalizedShape: LongArray, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): Tensor layerNorm(input: Tensor, normalizedShape: LongArray, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): the layer_norm operator. Layer normalization over the trailing normalized_shape dims of input. Statistics are computed per position over the trailing normalized_shape dims (channels-last layout, [N, .., C]). The output keeps input's shape AND dtype. An optional activation is fused onto the post-affine value (gated kinds are not accepted here). |
| le | [common] fun le(input: Tensor, other: Tensor): Tensor le(input: Tensor, other: Tensor): the le operator. Elementwise a <= other; output as eq.[common] fun le(input: Tensor, other: Double): Tensor le(input: Tensor, other: Double): the number form of le, other as a scalar. |
| leakyRelu | [common] fun leakyRelu(input: Tensor, negativeSlope: Double = 0.01): Tensor leakyRelu(input: Tensor, negativeSlope: Double = 0.01): the leaky_relu operator. ReLU with a small slope on the negative side. |
| leakyReluInPlace | [common] fun leakyReluInPlace(self: Tensor, negativeSlope: Double = 0.01): Tensor leakyReluInPlace(self: Tensor, negativeSlope: Double = 0.01): the leaky_relu_ operator. In-place leaky_relu: writes the result through self; same formula, arguments, and error conditions as leaky_relu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| leInPlace | [common] fun leInPlace(self: Tensor, other: Tensor): Tensor leInPlace(self: Tensor, other: Tensor): the le_ operator. In-place le: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain.[common] fun leInPlace(self: Tensor, other: Double): Tensor leInPlace(self: Tensor, other: Double): the number form of le_, other as a scalar. |
| linear | [common] fun linear(input: Tensor, weight: Tensor, bias: Tensor? = null, activation: Activation? = null, situBeta: Double = 0.0, situLinearBeta: Double = 0.0): Tensor linear(input: Tensor, weight: Tensor, bias: Tensor? = null, activation: Activation? = null, situBeta: Double = 0.0, situLinearBeta: Double = 0.0): the linear operator. x @ weightᵀ (+ bias); weight in the [out, in] Linear layout. Same situ contract as matmul: Activation::Situ transforms both halves of the gate‖up projection under the two soft-caps (situ_beta = β, situ_linear_beta = lβ, both 0, from the model's config); either scalar with any other activation rejects. |
| linspace | [common] fun linspace(start: Tensor, end: Tensor, steps: Tensor, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor linspace(start: Tensor, end: Tensor, steps: Tensor, dtype: DType = DType.FLOAT32, device: Placement? = null): the linspace operator. steps evenly spaced values from start to end, ENDPOINTS INCLUDED. out[i] = start + i * (end - start) / (steps - 1); the last element is exactly end; steps = 1 yields just start. Contrast arange, whose upper bound is exclusive. Bounds and count are literals or 0-D Tensors (trace symbolically).[common] fun linspace(start: Double, end: Double, steps: Double, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor linspace(start: Double, end: Double, steps: Double, dtype: DType = DType.FLOAT32, device: Placement? = null): the number form of linspace, start, end, steps as scalars. |
| log | [common] fun log(input: Tensor): Tensor log(input: Tensor): the log operator. Elementwise natural logarithm. Zero produces -inf; negative inputs produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| log10 | [common] fun log10(input: Tensor): Tensor log10(input: Tensor): the log10 operator. Elementwise base-10 logarithm. Zero produces -inf; negative inputs produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| log10InPlace | [common] fun log10InPlace(self: Tensor): Tensor log10InPlace(self: Tensor): the log10_ operator. In-place log10: writes the result through self; same formula, arguments, and error conditions as log10(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| log1p | [common] fun log1p(input: Tensor): Tensor log1p(input: Tensor): the log1p operator. Elementwise log(1 + x), accurate for small input. Inputs below -1 produce NaN; exactly -1 produces -inf. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| log1pInPlace | [common] fun log1pInPlace(self: Tensor): Tensor log1pInPlace(self: Tensor): the log1p_ operator. In-place log1p: writes the result through self; same formula, arguments, and error conditions as log1p(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| log2 | [common] fun log2(input: Tensor): Tensor log2(input: Tensor): the log2 operator. Elementwise base-2 logarithm. Zero produces -inf; negative inputs produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| log2InPlace | [common] fun log2InPlace(self: Tensor): Tensor log2InPlace(self: Tensor): the log2_ operator. In-place log2: writes the result through self; same formula, arguments, and error conditions as log2(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| logaddexp | [common] fun logaddexp(input: Tensor, other: Tensor): Tensor logaddexp(input: Tensor, other: Tensor): the logaddexp operator. Elementwise , computed overflow-safely (the log-domain accumulation primitive). Broadcasts and promotes as add. |
| logaddexp2 | [common] fun logaddexp2(input: Tensor, other: Tensor): Tensor logaddexp2(input: Tensor, other: Tensor): the logaddexp2 operator. Elementwise , logaddexp in base 2. Broadcasts and promotes as add. |
| logicalAnd | [common] fun logicalAnd(input: Tensor, other: Tensor): Tensor logicalAnd(input: Tensor, other: Tensor): the logical_and operator. Elementwise logical AND; non-Bool inputs read as element != 0 before the logic, and the output is always Bool (output as eq). |
| logicalAndInPlace | [common] fun logicalAndInPlace(self: Tensor, other: Tensor): Tensor logicalAndInPlace(self: Tensor, other: Tensor): the logical_and_ operator. In-place logical_and: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain. |
| logicalNot | [common] fun logicalNot(input: Tensor): Tensor logicalNot(input: Tensor): the logical_not operator. Elementwise logical NOT (element == 0); output as eq. |
| logicalNotInPlace | [common] fun logicalNotInPlace(self: Tensor): Tensor logicalNotInPlace(self: Tensor): the logical_not_ operator. In-place logical_not: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain. |
| logicalOr | [common] fun logicalOr(input: Tensor, other: Tensor): Tensor logicalOr(input: Tensor, other: Tensor): the logical_or operator. Elementwise logical OR (non-Bool inputs read as != 0); output as eq. |
| logicalOrInPlace | [common] fun logicalOrInPlace(self: Tensor, other: Tensor): Tensor logicalOrInPlace(self: Tensor, other: Tensor): the logical_or_ operator. In-place logical_or: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain. |
| logicalXor | [common] fun logicalXor(input: Tensor, other: Tensor): Tensor logicalXor(input: Tensor, other: Tensor): the logical_xor operator. Elementwise logical XOR (non-Bool inputs read as != 0); output as eq. |
| logicalXorInPlace | [common] fun logicalXorInPlace(self: Tensor, other: Tensor): Tensor logicalXorInPlace(self: Tensor, other: Tensor): the logical_xor_ operator. In-place logical_xor: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain. |
| logInPlace | [common] fun logInPlace(self: Tensor): Tensor logInPlace(self: Tensor): the log_ operator. In-place log: writes the result through self; same formula, arguments, and error conditions as log(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| logit | [common] fun logit(input: Tensor, eps: Double? = null): Tensor logit(input: Tensor, eps: Double? = null): the logit operator. Elementwise log-odds. With eps set, the input is first clipped to [eps, 1 - eps]; absent means no clipping, so inputs outside (0, 1) produce NaN / ±inf per IEEE semantics. |
| logitInPlace | [common] fun logitInPlace(self: Tensor, eps: Double? = null): Tensor logitInPlace(self: Tensor, eps: Double? = null): the logit_ operator. In-place logit: writes the result through self; same formula, arguments, and error conditions as logit(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| logSigmoid | [common] fun logSigmoid(input: Tensor): Tensor logSigmoid(input: Tensor): the log_sigmoid operator. Logarithm of the sigmoid, computed stably. |
| logSigmoidInPlace | [common] fun logSigmoidInPlace(self: Tensor): Tensor logSigmoidInPlace(self: Tensor): the log_sigmoid_ operator. In-place log_sigmoid: writes the result through self; same formula, arguments, and error conditions as log_sigmoid(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| logSoftmax | [common] fun logSoftmax(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): Tensor logSoftmax(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): the log_softmax operator. Logarithm of the softmax along dim, computed stably (never log(softmax(x)) in two passes). |
| logSoftmaxInPlace | [common] fun logSoftmaxInPlace(self: Tensor, dim: Long = -1L): Tensor logSoftmaxInPlace(self: Tensor, dim: Long = -1L): the log_softmax_ operator. In-place log_softmax: writes the result through self; same formula, arguments, and error conditions as log_softmax(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| logsumexp | [common] fun logsumexp(input: Tensor, dims: LongArray, keepdim: Boolean = false): Tensor logsumexp(input: Tensor, dims: LongArray, keepdim: Boolean = false): the logsumexp operator. Numerically stable log(sum(exp(x))) over dims. Computed with the max-shift trick, so large magnitudes do not overflow. dims is required here (no reduce-all default). |
| lt | [common] fun lt(input: Tensor, other: Tensor): Tensor lt(input: Tensor, other: Tensor): the lt operator. Elementwise a < other; output as eq.[common] fun lt(input: Tensor, other: Double): Tensor lt(input: Tensor, other: Double): the number form of lt, other as a scalar. |
| ltInPlace | [common] fun ltInPlace(self: Tensor, other: Tensor): Tensor ltInPlace(self: Tensor, other: Tensor): the lt_ operator. In-place lt: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain.[common] fun ltInPlace(self: Tensor, other: Double): Tensor ltInPlace(self: Tensor, other: Double): the number form of lt_, other as a scalar. |
| maskedFill | [common] fun maskedFill(input: Tensor, mask: Tensor, value: Tensor): Tensor maskedFill(input: Tensor, mask: Tensor, value: Tensor): the masked_fill operator. Replace the elements of input where mask is true with value. mask is Bool and broadcasts to input's shape; value is a literal or a 0-D Tensor.[common] fun maskedFill(input: Tensor, mask: Tensor, value: Double): Tensor maskedFill(input: Tensor, mask: Tensor, value: Double): the number form of masked_fill, value as a scalar. |
| maskedFillInPlace | [common] fun maskedFillInPlace(self: Tensor, mask: Tensor, value: Tensor): Tensor maskedFillInPlace(self: Tensor, mask: Tensor, value: Tensor): the masked_fill_ operator. In-place masked_fill: writes value through self where mask is true; same arguments and error conditions as masked_fill(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun maskedFillInPlace(self: Tensor, mask: Tensor, value: Double): Tensor maskedFillInPlace(self: Tensor, mask: Tensor, value: Double): the number form of masked_fill_, value as a scalar. |
| maskedScatter | [common] fun maskedScatter(input: Tensor, mask: Tensor, source: Tensor): Tensor maskedScatter(input: Tensor, mask: Tensor, source: Tensor): the masked_scatter operator. A copy of input with the leading count(mask) elements of src (read row-major) written at the positions where mask is true. mask is Bool, broadcastable to input; src must supply at least count(mask) elements. |
| maskedScatterInPlace | [common] fun maskedScatterInPlace(self: Tensor, mask: Tensor, source: Tensor): Tensor maskedScatterInPlace(self: Tensor, mask: Tensor, source: Tensor): the masked_scatter_ operator. In-place masked_scatter: writes src's leading elements through self where mask is true; same arguments and error conditions as masked_scatter(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| maskedSelect | [common] fun maskedSelect(input: Tensor, mask: Tensor): Tensor maskedSelect(input: Tensor, mask: Tensor): the masked_select operator. The elements of input where mask is true, as a 1-D tensor. The output LENGTH is data-dependent (the number of true entries); under tracing it carries a data-dependent extent. mask is Bool and broadcasts to input's shape. |
| matmul | [common] fun matmul(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null, transposeA: Boolean = false, transposeB: Boolean = false, situBeta: Double = 0.0, situLinearBeta: Double = 0.0): Tensor matmul(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null, transposeA: Boolean = false, transposeB: Boolean = false, situBeta: Double = 0.0, situLinearBeta: Double = 0.0): the matmul operator. a @ b (+ bias) with an optional fused activation epilogue. A gated activation (SwiGlu/GeGlu/ReGlu/Situ) takes the product as a concatenated [*, 2d] gate‖up projection and emits [*, d] in one pass. situ_beta/situ_linear_beta are Activation::Situ's two soft-cap scalars (β and lβ in β·tanh(gate/β)·sigmoid(gate) · lβ·tanh(up/lβ)); pass the model's own config values. Situ requires BOTH 0, and either scalar with any other activation rejects (it would otherwise be silently ignored). The transform computes at fp32 end-to-end with one demote at the store for f16/bf16 outputs. |
| max | [common] fun max(input: Tensor, dim: Long, keepdim: Boolean = false): List<Tensor> max(input: Tensor, dim: Long, keepdim: Boolean = false): the max operator. |
| maximum | [common] fun maximum(input: Tensor, other: Tensor): Tensor maximum(input: Tensor, other: Tensor): the maximum operator. Elementwise maximum, NaN-PROPAGATING: a NaN in either operand yields NaN (use ops::fmax for the NaN-ignoring IEEE law). Broadcasts and promotes as add. |
| maximumInPlace | [common] fun maximumInPlace(self: Tensor, other: Tensor): Tensor maximumInPlace(self: Tensor, other: Tensor): the maximum_ operator. In-place maximum: writes the elementwise maxima through x. Writes through self and returns it, so calls chain. |
| maxPool | [common] fun maxPool(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, dilation: LongArray, ceilMode: Boolean): Tensor maxPool(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, dilation: LongArray, ceilMode: Boolean): the max_pool operator. Rank-generic max pooling, channels-last. input is [N, D1..Dn, C]; the window rank is read from kernel_size's length. Each output element is the maximum over its window; padding never wins. A window holding any NaN element yields NaN, and a window lying wholly in the padding yields -inf for a float dtype and the type minimum for an integer one. |
| maxPool1d | [common] fun maxPool1d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), ceilMode: Boolean = false): Tensor maxPool1d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), ceilMode: Boolean = false): the max_pool1d operator. 1-D max pooling over [N, L, C] (channels-last). See max_pool for the window semantics; stride empty = kernel_size. |
| maxPool1dWithIndices | [common] fun maxPool1dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), ceilMode: Boolean = false): List<Tensor> maxPool1dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0), dilation: LongArray = longArrayOf(1), ceilMode: Boolean = false): the max_pool1d_with_indices operator. |
| maxPool2d | [common] fun maxPool2d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), ceilMode: Boolean = false): Tensor maxPool2d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), ceilMode: Boolean = false): the max_pool2d operator. 2-D max pooling over [N, H, W, C] (channels-last). Each output element is the max over its kernel_size window; stride empty defaults to kernel_size (non-overlapping windows); dilation spaces the window's taps; ceil_mode rounds the output extents up. |
| maxPool2dWithIndices | [common] fun maxPool2dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), ceilMode: Boolean = false): List<Tensor> maxPool2dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0), dilation: LongArray = longArrayOf(1, 1), ceilMode: Boolean = false): the max_pool2d_with_indices operator. |
| maxPool3d | [common] fun maxPool3d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), ceilMode: Boolean = false): Tensor maxPool3d(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), ceilMode: Boolean = false): the max_pool3d operator. 3-D max pooling over [N, D, H, W, C] (channels-last). See max_pool2d; parameters extend to {kD, kH, kW} etc. |
| maxPool3dWithIndices | [common] fun maxPool3dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), ceilMode: Boolean = false): List<Tensor> maxPool3dWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray = longArrayOf(), padding: LongArray = longArrayOf(0, 0, 0), dilation: LongArray = longArrayOf(1, 1, 1), ceilMode: Boolean = false): the max_pool3d_with_indices operator. |
| maxPoolWithIndices | [common] fun maxPoolWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, dilation: LongArray, ceilMode: Boolean): List<Tensor> maxPoolWithIndices(input: Tensor, kernelSize: LongArray, stride: LongArray, padding: LongArray, dilation: LongArray, ceilMode: Boolean): the max_pool_with_indices operator. Max pooling that also returns each window's winning element as a flat index over the input's spatial extents (per plane, the same for every channel). The values follow max_pool. The index names the window's lowest-index NaN element when it holds one, else the first element holding the max (a tie keeps the lowest index); a window lying wholly in the padding reports -1. |
| mean | [common] fun mean(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): Tensor mean(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): the mean operator. Arithmetic mean of input over dims. Empty dims averages every element. dtype selects the output (and accumulation) dtype; the default keeps input's. |
| median | [common] fun median(input: Tensor, dim: Long? = null, keepdim: Boolean = false): List<Tensor> median(input: Tensor, dim: Long? = null, keepdim: Boolean = false): the median operator. |
| meshgrid | [common] fun meshgrid(tensors: List<Tensor>, indexing: MeshgridIndexing = MeshgridIndexing.IJ): List<Tensor> meshgrid(tensors: List<Tensor>, indexing: MeshgridIndexing = MeshgridIndexing.IJ): the meshgrid operator. Coordinate grids from 1-D axes: N inputs produce N N-D tensors, each input broadcast over every other axis, the NumPy meshgrid. With MeshgridIndexing::IJ (matrix indexing, the default) the output shapes follow the input order [len0, len1, ..]; XY (Cartesian) swaps the first two axes. |
| min | [common] fun min(input: Tensor, dim: Long, keepdim: Boolean = false): List<Tensor> min(input: Tensor, dim: Long, keepdim: Boolean = false): the min operator. |
| minimum | [common] fun minimum(input: Tensor, other: Tensor): Tensor minimum(input: Tensor, other: Tensor): the minimum operator. Elementwise minimum, NaN-PROPAGATING (the maximum dual; use ops::fmin for NaN-ignoring). Broadcasts and promotes as add. |
| minimumInPlace | [common] fun minimumInPlace(self: Tensor, other: Tensor): Tensor minimumInPlace(self: Tensor, other: Tensor): the minimum_ operator. In-place minimum: writes the elementwise minima through x. Writes through self and returns it, so calls chain. |
| mish | [common] fun mish(input: Tensor): Tensor mish(input: Tensor): the mish operator. Mish activation. |
| mishInPlace | [common] fun mishInPlace(self: Tensor): Tensor mishInPlace(self: Tensor): the mish_ operator. In-place mish: writes the result through self; same formula, arguments, and error conditions as mish(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| mlaAttention | [common] fun mlaAttention(qNope: Tensor, qPe: Tensor, newCkv: Tensor, newKpe: Tensor, ckvCache: Tensor, kpeCache: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, scale: Double, slotIds: Tensor? = null): Tensor mlaAttention(qNope: Tensor, qPe: Tensor, newCkv: Tensor, newKpe: Tensor, ckvCache: Tensor, kpeCache: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, scale: Double, slotIds: Tensor? = null): the mla_attention operator. Multi-head Latent Attention over a compressed KV cache, latent-space end to end: q_nope [ΣS, H, Dl] (W_UK-absorbed) + q_pe [ΣS, H, Dr] score against the per-token compressed rows; the caches ([max_seqs, max_seq, Dl/Dr]) take this step's new_ckv/new_kpe appends IN PLACE (write offsets = kvcache_start [B] Int32; cu_seqlens_q [B+1] locates each sequence's packed tokens); the returned [ΣS, H, Dl] output stays latent (apply the W_UV un-absorption after). scale is REQUIRED; the absorbed query's magnitude lives in the model's un-absorbed head dim. |
| mm | [common] fun mm(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null): Tensor mm(input: Tensor, other: Tensor, bias: Tensor? = null, activation: Activation? = null): the mm operator. Matrix multiply of two rank-2 tensors, with an optional fused bias and activation epilogue. a [M, K] x b [K, N] -> [M, N]. Inputs must be rank-2 (use bmm for batched operands, matmul for the broadcasting general form). |
| mod | [common] fun mod(input: Tensor, other: Tensor, mode: ModMode = ModMode.PYTHON): Tensor mod(input: Tensor, other: Tensor, mode: ModMode = ModMode.PYTHON): the mod operator. Elementwise modulo with a selectable sign convention. ModMode::Python (default) follows the DIVISOR's sign; ModMode::C follows the DIVIDEND's. ops::remainder is exactly the Python form, ops::fmod exactly the C form. Broadcasts and promotes as add.[common] fun mod(input: Tensor, other: Double, mode: ModMode = ModMode.PYTHON): Tensor mod(input: Tensor, other: Double, mode: ModMode = ModMode.PYTHON): the number form of mod, other as a scalar. |
| modInPlace | [common] fun modInPlace(self: Tensor, other: Tensor, mode: ModMode = ModMode.PYTHON): Tensor modInPlace(self: Tensor, other: Tensor, mode: ModMode = ModMode.PYTHON): the mod_ operator. In-place mod: writes the remainders through x. Writes through self and returns it, so calls chain.[common] fun modInPlace(self: Tensor, other: Double, mode: ModMode = ModMode.PYTHON): Tensor modInPlace(self: Tensor, other: Double, mode: ModMode = ModMode.PYTHON): the number form of mod_, other as a scalar. |
| moe | [common] fun moe(input: Tensor, routerLogits: Tensor, fc1Experts: Tensor, fc2Experts: Tensor, topK: Long, fc1Bias: Tensor? = null, fc2Bias: Tensor? = null, fc3Experts: Tensor? = null, fc3Bias: Tensor? = null, eScoreCorrectionBias: Tensor? = null, routerWeights: Tensor? = null, routingMode: MoeRouting? = null, renormalize: Boolean? = null, nGroup: Long? = null, topkGroup: Long? = null, routedScalingFactor: Double? = null, sparseMixerEps: Double? = null, applyRouterWeightOnInput: Boolean? = null, activation: Activation? = null, swigluFusion: SwigluFusion? = null, swigluAlpha: Double? = null, swigluBeta: Double? = null, swigluLimit: Double? = null, geluMode: GeluMode? = null, sharedOutput: Tensor? = null): Tensor moe(input: Tensor, routerLogits: Tensor, fc1Experts: Tensor, fc2Experts: Tensor, topK: Long, fc1Bias: Tensor? = null, fc2Bias: Tensor? = null, fc3Experts: Tensor? = null, fc3Bias: Tensor? = null, eScoreCorrectionBias: Tensor? = null, routerWeights: Tensor? = null, routingMode: MoeRouting? = null, renormalize: Boolean? = null, nGroup: Long? = null, topkGroup: Long? = null, routedScalingFactor: Double? = null, sparseMixerEps: Double? = null, applyRouterWeightOnInput: Boolean? = null, activation: Activation? = null, swigluFusion: SwigluFusion? = null, swigluAlpha: Double? = null, swigluBeta: Double? = null, swigluLimit: Double? = null, geluMode: GeluMode? = null, sharedOutput: Tensor? = null): the moe operator. Fused Mixture-of-Experts layer: route, run the top-k experts, combine in one call, with no per-expert dispatch from the caller. Per token, router_logits [T, E] select top_k experts under routing_mode; each selected expert applies its own MLP (fc1 [E, F·I, H] → activation → fc2 [E, H, I], with F = 2 for a gated activation, else 1, and an optional multiplicative fc3 [E, I, H] branch); the expert outputs combine under the routing weights. input is [T, H]; the result is [T, H] at input's dtype. Per-expert biases ride fc1_bias / fc2_bias / fc3_bias; e_score_correction_bias and n_group / topk_group / routed_scaling_factor serve the group-limited routing families; router_weights feeds MoeRouting::PreComputed (caller-supplied combine weights); sparse_mixer_eps tunes MoeRouting::SparseMixer. The gated-activation scalars (swiglu_*, gelu_mode) carry their swiglu / geglu meanings; shared_output [T, H] folds a shared-expert branch into the final combine. A config argument left std::nullopt takes the runtime default (SoftmaxTopK routing, renormalized top-k weights). |
| mropeRotaryEmbedding | [common] fun mropeRotaryEmbedding(input: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): Tensor mropeRotaryEmbedding(input: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): the mrope_rotary_embedding operator. Multi-section rotary position embedding over padded layouts. Multi-section RoPE (the multimodal form): the head dimension is split into mrope_sections, and each section takes its angles from its OWN row of position_ids (e.g. temporal/height/width position streams). interleaved_sections picks the section layout; theta overrides the frequency base. |
| mropeRotaryEmbeddingQkVarlen | [common] fun mropeRotaryEmbeddingQkVarlen(query: Tensor, key: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): List<Tensor> mropeRotaryEmbeddingQkVarlen(query: Tensor, key: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): the mrope_rotary_embedding_qk_varlen operator. Multi-axis RoPE over a packed q/k pair: ONE per-token angle table serves both projections in a single call, the media-prefill pre-rotation shape. query/key are head-exposed [n_tokens, heads, head_dim] (a flat [n_tokens, heads·head_dim] projection reshapes to it as a free view; the rotation spans the TRAILING dim, so a flat row would rotate across head boundaries); position_ids is [n_axes, n_tokens] with n_axes == mrope_sections.size() (strip any trailing zero sections a checkpoint's metadata pads; the axis count follows the section list); rotary_dim bounds the rotated span (dims beyond it pass through) and defaults to the trailing dim. Rows whose per-axis positions are all EQUAL rotate exactly as the plain ops do; serve pure-text calls through rotary_embedding* (cheaper: no per-token table), and use this form for packs whose rows carry genuinely multi-axis positions, then attend with the rope inputs ABSENT so the attention op consumes (and appends) q/k exactly as given. |
| mropeRotaryEmbeddingVarlen | [common] fun mropeRotaryEmbeddingVarlen(input: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): Tensor mropeRotaryEmbeddingVarlen(input: Tensor, positionIds: Tensor, mropeSections: LongArray, interleavedSections: Boolean? = null, theta: Double? = null, scaling: RopeScaling? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, scale: Double? = null): the mrope_rotary_embedding_varlen operator. Multi-section rotary position embedding over token-packed layouts (cu_seqlens``[B+1] Int32, as in rotary_embedding_varlen). Multi-section RoPE (the multimodal form): the head dimension is split into mrope_sections, and each section takes its angles from its OWN row of position_ids (e.g. temporal/height/width position streams). interleaved_sections picks the section layout; theta overrides the frequency base. |
| msDeformAttention | [common] fun msDeformAttention(value: Tensor, spatialShapes: Tensor, levelStartIndex: Tensor, samplingLocations: Tensor, attentionWeights: Tensor): Tensor msDeformAttention(value: Tensor, spatialShapes: Tensor, levelStartIndex: Tensor, samplingLocations: Tensor, attentionWeights: Tensor): the ms_deform_attention operator. Multi-scale deformable attention (2-D): per query and head, gather P bilinear samples from each of L flattened feature-map levels and combine them with the given weights. value [N, S, M, D] with S = Σ_l H_l·W_l; spatial_shapes [L, 2] = per-level (H_l, W_l) and level_start_index [L] (both Int32 or Int64); sampling_locations [N, Lq, M, L, P, 2]; last dim (x, y), normalized to [0, 1] per level, sampled at loc·size − 0.5 (bilinear; out-of-bounds reads 0); attention_weights [N, Lq, M, L, P] are consumed AS GIVEN (apply softmax beforehand if wanted). Returns [N, Lq, M, D] at value's dtype; accumulation is fp32. The float inputs must share value's dtype (f32/f64/f16/bf16; no silent promotion). |
| mseLoss | [common] fun mseLoss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN): Tensor mseLoss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN): the mse_loss operator. Mean-squared-error loss between input and target. Reduction::None returns the per-element losses (shape of input); Mean / Sum reduce to a 0-D scalar. |
| mul | [common] fun mul(input: Tensor, other: Tensor, activation: Activation = Activation.IDENTITY): Tensor mul(input: Tensor, other: Tensor, activation: Activation = Activation.IDENTITY): the mul operator. Multiplies input by other elementwise. Broadcasts and promotes as add; activation is the same fused float-only epilogue (default Activation::Identity).[common] fun mul(input: Tensor, other: Double, activation: Activation = Activation.IDENTITY): Tensor mul(input: Tensor, other: Double, activation: Activation = Activation.IDENTITY): the number form of mul, other as a scalar.[common] fun mul(input: Double, other: Tensor, activation: Activation = Activation.IDENTITY): Tensor mul(input: Double, other: Tensor, activation: Activation = Activation.IDENTITY): the mul operator. Scalar-LHS . input keeps its kind (see Scalar). activation applies to the result, mirroring the tensor-first form (mul carries no alpha; it would fold into the scalar). |
| mulInPlace | [common] fun mulInPlace(self: Tensor, other: Tensor, activation: Activation = Activation.IDENTITY): Tensor mulInPlace(self: Tensor, other: Tensor, activation: Activation = Activation.IDENTITY): the mul_ operator. In-place mul: writes through x. Writes through self and returns it, so calls chain.[common] fun mulInPlace(self: Tensor, other: Double, activation: Activation = Activation.IDENTITY): Tensor mulInPlace(self: Tensor, other: Double, activation: Activation = Activation.IDENTITY): the number form of mul_, other as a scalar. |
| multinomial | [common] fun multinomial(probabilities: Tensor, numSamples: Long, replacement: Boolean = false, device: Placement? = null): Tensor multinomial(probabilities: Tensor, numSamples: Long, replacement: Boolean = false, device: Placement? = null): the multinomial operator. Categorical sampling: draws num_samples category indices per row of a weight tensor. probabilities holds non-negative weights (they need not sum to 1), typically [batch, num_categories]; the output replaces the category axis with num_samples and is Int64. Without replacement each row samples distinct categories. |
| mv | [common] fun mv(input: Tensor, vec: Tensor, bias: Tensor? = null, activation: Activation? = null): Tensor mv(input: Tensor, vec: Tensor, bias: Tensor? = null, activation: Activation? = null): the mv operator. Matrix-vector multiply, with an optional fused bias and activation. a [M, K] x vec [K] -> [M]. |
| nanmean | [common] fun nanmean(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor nanmean(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the nanmean operator. Mean over dims, treating NaN entries as missing. Each output element divides by the count of NON-NaN contributors (an all-NaN slice yields NaN). Float dtypes only. |
| nanmedian | [common] fun nanmedian(input: Tensor, dim: Long? = null, keepdim: Boolean = false): List<Tensor> nanmedian(input: Tensor, dim: Long? = null, keepdim: Boolean = false): the nanmedian operator. |
| nanquantile | [common] fun nanquantile(input: Tensor, q: Tensor, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): Tensor nanquantile(input: Tensor, q: Tensor, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): the nanquantile operator. quantile that treats NaN entries as missing; each slice's quantile is computed over its non-NaN values (an all-NaN slice yields NaN). Parameters and shapes as quantile.[common] fun nanquantile(input: Tensor, q: Double, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): Tensor nanquantile(input: Tensor, q: Double, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): the number form of nanquantile, q as a scalar. |
| nansum | [common] fun nansum(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor nansum(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the nansum operator. Sum over dims, treating NaN entries as missing (contributing zero). Float dtypes only carry NaN; for integer inputs this is plain sum. |
| nanToNum | [common] fun nanToNum(input: Tensor, nan: Double? = null, posinf: Double? = null, neginf: Double? = null): Tensor nanToNum(input: Tensor, nan: Double? = null, posinf: Double? = null, neginf: Double? = null): the nan_to_num operator. Replaces NaN and infinities with finite values. Each absent replacement falls back to the dtype-specific default: 0 for NaN, the dtype's largest finite value for +inf, its lowest for -inf. |
| nanToNumInPlace | [common] fun nanToNumInPlace(self: Tensor, nan: Double? = null, posinf: Double? = null, neginf: Double? = null): Tensor nanToNumInPlace(self: Tensor, nan: Double? = null, posinf: Double? = null, neginf: Double? = null): the nan_to_num_ operator. In-place nan_to_num: writes the result through self; same formula, arguments, and error conditions as nan_to_num(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| narrow | [common] fun narrow(input: Tensor, dim: Long, start: Long, length: Long): Tensor narrow(input: Tensor, dim: Long, start: Long, length: Long): the narrow operator. length elements from start along dim. Both are int literals or 0-D integer Tensors (a tensor-valued window never syncs to the host). |
| ndim | [common] fun ndim(input: Tensor): Tensor ndim(input: Tensor): the ndim operator. The tensor's rank as a 0-D Int64 tensor. The value is produced ON DEVICE and is traceable; under tracing it carries the symbolic value. For a plain host integer use the *_host sibling instead. |
| ndimHost | [common] fun ndimHost(input: Tensor): Tensor ndimHost(input: Tensor): the ndim_host operator. The rank as a 0-D Int64 HOST tensor, written from metadata; the same no-sync contract as shape_host. |
| ne | [common] fun ne(input: Tensor, other: Tensor): Tensor ne(input: Tensor, other: Tensor): the ne operator. Elementwise a != other; output as eq.[common] fun ne(input: Tensor, other: Double): Tensor ne(input: Tensor, other: Double): the number form of ne, other as a scalar. |
| neg | [common] fun neg(input: Tensor): Tensor neg(input: Tensor): the neg operator. Elementwise negation. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| negInPlace | [common] fun negInPlace(self: Tensor): Tensor negInPlace(self: Tensor): the neg_ operator. In-place neg: writes the result through self; same formula, arguments, and error conditions as neg(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| neInPlace | [common] fun neInPlace(self: Tensor, other: Tensor): Tensor neInPlace(self: Tensor, other: Tensor): the ne_ operator. In-place ne: writes the result through x (one/zero at x's dtype). Writes through self and returns it, so calls chain.[common] fun neInPlace(self: Tensor, other: Double): Tensor neInPlace(self: Tensor, other: Double): the number form of ne_, other as a scalar. |
| nllLoss | [common] fun nllLoss(input: Tensor, target: Tensor, weight: Tensor? = null, ignoreIndex: Long? = null, reduction: Reduction = Reduction.MEAN): Tensor nllLoss(input: Tensor, target: Tensor, weight: Tensor? = null, ignoreIndex: Long? = null, reduction: Reduction = Reduction.MEAN): the nll_loss operator. Negative-log-likelihood loss over LOG-probabilities; class dim LAST. input is expected to already be log-probabilities (pair with log_softmax; cross_entropy is the fused form). |
| nms | [common] fun nms(boxes: Tensor, scores: Tensor, maxOutputBoxesPerClass: Tensor? = null, iouThreshold: Tensor? = null, scoreThreshold: Tensor? = null, centerPointBox: Boolean = false): Tensor nms(boxes: Tensor, scores: Tensor, maxOutputBoxesPerClass: Tensor? = null, iouThreshold: Tensor? = null, scoreThreshold: Tensor? = null, centerPointBox: Boolean = false): the nms operator. Batched, class-aware non-max suppression (argument order mirrors the ONNX operator). boxes is [batch, num_boxes, 4]; scores is [batch, classes, num_boxes]; returns the selected indices as an Int64 [num_selected, 3] of (batch, class, box) rows. Every threshold takes a literal OR a 0-D tensor (a tensor traces symbolically). An absent max_output_boxes_per_class selects NOTHING (the reference default); a negative cap clamps to 0. A literal iou_threshold outside 0, 1 raises. center_point_box picks the [cx, cy, w, h] box encoding over the corners form. |
| nonzero | [common] fun nonzero(input: Tensor): Tensor nonzero(input: Tensor): the nonzero operator. Coordinates of the nonzero elements: an [n, ndim] Int64 matrix, one row per nonzero element of input. An element is nonzero exactly when its Bool cast is true: +0 and -0 are zero; a subnormal, an infinity and a NaN are nonzero. The row count is data-dependent (see masked_select); a 1-arg where() call is sugar for this op. |
| norm | [common] fun norm(input: Tensor, p: Double = 2.0, dims: LongArray = longArrayOf(), keepdim: Boolean = false): Tensor norm(input: Tensor, p: Double = 2.0, dims: LongArray = longArrayOf(), keepdim: Boolean = false): the norm operator. The p-norm of input over dims. Empty dims reduces every dimension. |
| normal | [common] fun normal(mean: Tensor, stddev: Tensor, shape: LongArray, device: Placement? = null): Tensor normal(mean: Tensor, stddev: Tensor, shape: LongArray, device: Placement? = null): the normal operator. Normal draws with the given mean and standard deviation. mean and stddev are scalars or tensors broadcast over shape.[common] fun normal(mean: Double, stddev: Double, shape: LongArray, device: Placement? = null): Tensor normal(mean: Double, stddev: Double, shape: LongArray, device: Placement? = null): the number form of normal, mean, stddev as scalars. |
| normalInPlace | [common] fun normalInPlace(self: Tensor, mean: Tensor? = null, stddev: Tensor? = null, device: Placement? = null): Tensor normalInPlace(self: Tensor, mean: Tensor? = null, stddev: Tensor? = null, device: Placement? = null): the normal_ operator. In-place normal fill: overwrites self with N(mean, stddev^2) draws at self's shape/dtype and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun normalInPlace(self: Tensor, mean: Double = 0.0, stddev: Double = 1.0, device: Placement? = null): Tensor normalInPlace(self: Tensor, mean: Double = 0.0, stddev: Double = 1.0, device: Placement? = null): the number form of normal_, mean, stddev as scalars. |
| normalize | [common] fun normalize(input: Tensor, p: Double = 2.0, dim: Long = 1, eps: Double? = null): Tensor normalize(input: Tensor, p: Double = 2.0, dim: Long = 1L, eps: Double? = null): the normalize operator. L_p-normalizes input along dim: each slice is scaled to unit p-norm. |
| normalizeInPlace | [common] fun normalizeInPlace(self: Tensor, p: Double = 2.0, dim: Long = 1, eps: Double? = null): Tensor normalizeInPlace(self: Tensor, p: Double = 2.0, dim: Long = 1L, eps: Double? = null): the normalize_ operator. In-place normalize: rewrites self with its L_p-normalized value and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| numel | [common] fun numel(input: Tensor): Tensor numel(input: Tensor): the numel operator. The tensor's total element count as a 0-D Int64 tensor. The value is produced ON DEVICE and is traceable; under tracing it carries the symbolic value. For a plain host integer use the *_host sibling instead. |
| numelHost | [common] fun numelHost(input: Tensor): Tensor numelHost(input: Tensor): the numel_host operator. The element count as a 0-D Int64 HOST tensor, written from metadata; the same no-sync contract as shape_host. Composes with shape-consuming arguments, e.g. reshape(x, {numel_host(x)}) flattens without reading sizes to the host. |
| oneHot | [common] fun oneHot(indices: Tensor, numClasses: Long): Tensor oneHot(indices: Tensor, numClasses: Long): the one_hot operator. Expands an integer index tensor into a trailing one-hot dimension. |
| ones | [common] fun ones(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): Tensor ones(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): the ones operator. A new tensor of the given shape with every element set to one. Same shape-span and pinned_for contract as empty (extents are literals or 0-D integer Tensors). |
| onesLike | [common] fun onesLike(reference: Tensor, device: Placement? = null): Tensor onesLike(reference: Tensor, device: Placement? = null): the ones_like operator. A one-filled tensor with reference's shape and dtype (data never read); device absent = the reference's own device. |
| outer | [common] fun outer(input: Tensor, other: Tensor): Tensor outer(input: Tensor, other: Tensor): the outer operator. Outer product of two 1-D tensors: [m] x [n] -> [m, n]. |
| pad | [common] fun pad(input: Tensor, pad: LongArray, mode: PadMode = PadMode.CONSTANT, value: Tensor? = null): Tensor pad(input: Tensor, pad: LongArray, mode: PadMode = PadMode.CONSTANT, value: Tensor? = null): the pad operator. Pad each axis by interleaved (lo, hi) pairs in layout order, left-to-right: pair i pads axis i (the first pair is the leading axis). Provide fewer pairs than the rank to pad only the leading axes; a negative width crops that side. mode selects the fill; value is the Constant fill (default 0). |
| pairwiseDistance | [common] fun pairwiseDistance(x1: Tensor, x2: Tensor, p: Double = 2.0, eps: Double? = null, keepdim: Boolean = false): Tensor pairwiseDistance(x1: Tensor, x2: Tensor, p: Double = 2.0, eps: Double? = null, keepdim: Boolean = false): the pairwise_distance operator. Row-wise L_p distance between two batched vectors. Reduces the LAST axis; leading dims broadcast. |
| pdist | [common] fun pdist(input: Tensor, p: Double = 2.0): Tensor pdist(input: Tensor, p: Double = 2.0): the pdist operator. Condensed pairwise L_p distances within ONE 2-D input. Every unordered row pair of a [M, K] yields one entry; the output is 1-D of length M(M-1)/2 (upper-triangle order). |
| permute | [common] fun permute(input: Tensor, dims: LongArray): Tensor permute(input: Tensor, dims: LongArray): the permute operator. Reorder the dims by dims, a permutation of [0, rank). Returns a VIEW (metadata only, no copy). A kernel that later needs dense data materializes internally; no contiguous call is needed here. |
| pixelShuffle | [common] fun pixelShuffle(input: Tensor, upscaleFactor: Long, mode: PixelShuffleMode = PixelShuffleMode.CRD): Tensor pixelShuffle(input: Tensor, upscaleFactor: Long, mode: PixelShuffleMode = PixelShuffleMode.CRD): the pixel_shuffle operator. Rearranges channels into space (depth-to-space): channels-last [N, H, W, C] becomes [N, H*r, W*r, C/r^2]. |
| pixelUnshuffle | [common] fun pixelUnshuffle(input: Tensor, downscaleFactor: Long, mode: PixelShuffleMode = PixelShuffleMode.CRD): Tensor pixelUnshuffle(input: Tensor, downscaleFactor: Long, mode: PixelShuffleMode = PixelShuffleMode.CRD): the pixel_unshuffle operator. The inverse of pixel_shuffle (space-to-depth): channels-last [N, H, W, C] becomes [N, H/r, W/r, C*r^2]. |
| poisson | [common] fun poisson(rates: Tensor, device: Placement? = null): Tensor poisson(rates: Tensor, device: Placement? = null): the poisson operator. Independent Poisson draws from per-element rates. |
| pow | [common] fun pow(input: Tensor, exponent: Tensor): Tensor pow(input: Tensor, exponent: Tensor): the pow operator. Raises input to exponent elementwise: . Broadcasts and promotes as add; a scalar exponent keeps its weak kind (an integer exponent with an integer base stays integral).[common] fun pow(input: Tensor, exponent: Double): Tensor pow(input: Tensor, exponent: Double): the number form of pow, exponent as a scalar. |
| powInPlace | [common] fun powInPlace(self: Tensor, exponent: Tensor): Tensor powInPlace(self: Tensor, exponent: Tensor): the pow_ operator. In-place pow: writes the powers through x. Writes through self and returns it, so calls chain.[common] fun powInPlace(self: Tensor, exponent: Double): Tensor powInPlace(self: Tensor, exponent: Double): the number form of pow_, exponent as a scalar. |
| prelu | [common] fun prelu(input: Tensor, weight: Tensor? = null): Tensor prelu(input: Tensor, weight: Tensor? = null): the prelu operator. Parametric ReLU: a learned negative-side slope. weight is a scalar or a per-channel tensor broadcast against input. |
| preluInPlace | [common] fun preluInPlace(self: Tensor, weight: Tensor? = null): Tensor preluInPlace(self: Tensor, weight: Tensor? = null): the prelu_ operator. In-place prelu: writes the result through self; same formula, arguments, and error conditions as prelu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| prod | [common] fun prod(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): Tensor prod(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): the prod operator. Product of input's elements over dims. Empty dims multiplies every element. dtype widens the accumulation/output when set; the default keeps input's dtype. |
| put | [common] fun put(input: Tensor, index: Tensor, source: Tensor, accumulate: Boolean = false): Tensor put(input: Tensor, index: Tensor, source: Tensor, accumulate: Boolean = false): the put operator. A copy of input with source written at FLAT (row-major linearized) positions: out.flat[index[i]] = source.flat[i]. With accumulate = true duplicate positions ADD instead of overwrite. index and source carry the same element count; index dtype law as gather. |
| putInPlace | [common] fun putInPlace(self: Tensor, index: Tensor, source: Tensor, accumulate: Boolean = false): Tensor putInPlace(self: Tensor, index: Tensor, source: Tensor, accumulate: Boolean = false): the put_ operator. In-place put: writes (or accumulates) through self at flat positions; same arguments and error conditions as put(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| qkLayerNorm | [common] fun qkLayerNorm(query: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, queryBias: Tensor? = null, keyWeight: Tensor? = null, keyBias: Tensor? = null, headDim: Long = 0, eps: Double? = null): List<Tensor> qkLayerNorm(query: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, queryBias: Tensor? = null, keyWeight: Tensor? = null, keyBias: Tensor? = null, headDim: Long = 0L, eps: Double? = null): the qk_layer_norm operator. Per-head LAYER norm over packed attention projections, the mean-subtracting sibling of qk_rms_norm. Each contiguous head_dim run of query (and key, when present) is normalized as (x - mean) / sqrt(var + eps) * w + b; value, when given, is normalized the same way with no weight and no bias (it has no affine slots). |
| qkLayerNormInPlace | [common] fun qkLayerNormInPlace(self: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, queryBias: Tensor? = null, keyWeight: Tensor? = null, keyBias: Tensor? = null, headDim: Long = 0, eps: Double? = null): Tensor qkLayerNormInPlace(self: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, queryBias: Tensor? = null, keyWeight: Tensor? = null, keyBias: Tensor? = null, headDim: Long = 0L, eps: Double? = null): the qk_layer_norm_ operator. In-place qk_layer_norm: normalizes query (and key and value, when given) through their own storage (no output allocation). key and value are optional operands written through: pass a pointer to your tensor, or nullptr for absent. Every written-through handle is rebound to the op's output, so under a TracingScope the caller's key becomes the traced output exactly as query does. Otherwise the same arguments and error conditions as qk_layer_norm(). Returns query under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| qkRmsNorm | [common] fun qkRmsNorm(query: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, keyWeight: Tensor? = null, headDim: Long = 0, eps: Double? = null): List<Tensor> qkRmsNorm(query: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, keyWeight: Tensor? = null, headDim: Long = 0L, eps: Double? = null): the qk_rms_norm operator. Per-head RMS norm over packed attention projections: query, key and value in ONE call, no reshapes. Each contiguous head_dim run of query (and key, when present) is its own normalization group: out = x / sqrt(mean(x^2) + eps) * w. value passes through untouched (it rides along so one call serves the projection triplet). Equivalent to reshaping [S, heads*head_dim] to [S, heads, head_dim], applying rms_norm, and reshaping back, with none of those steps. |
| qkRmsNormInPlace | [common] fun qkRmsNormInPlace(self: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, keyWeight: Tensor? = null, headDim: Long = 0, eps: Double? = null): Tensor qkRmsNormInPlace(self: Tensor, key: Tensor? = null, value: Tensor? = null, queryWeight: Tensor? = null, keyWeight: Tensor? = null, headDim: Long = 0L, eps: Double? = null): the qk_rms_norm_ operator. In-place qk_rms_norm: normalizes query (and key and value, when given) through their own storage (no output allocation). key and value are optional operands written through: pass a pointer to your tensor, or nullptr for absent. Every written-through handle is rebound to the op's output, so under a TracingScope the caller's key becomes the traced output exactly as query does. Otherwise the same arguments and error conditions as qk_rms_norm(). Returns query under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| quantile | [common] fun quantile(input: Tensor, q: Tensor, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): Tensor quantile(input: Tensor, q: Tensor, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): the quantile operator. The q-th quantile of input along dim. q is a scalar or 1-D tensor of probabilities in [0, 1]. Absent dim flattens input first. The output prepends q's shape to the reduced shape; interpolation picks the between-ranks strategy (Linear by default; see QuantileInterp for the full set).[common] fun quantile(input: Tensor, q: Double, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): Tensor quantile(input: Tensor, q: Double, dim: Long? = null, keepdim: Boolean = false, interpolation: QuantileInterp = QuantileInterp.LINEAR): the number form of quantile, q as a scalar. |
| quantizeDequantize | [common] fun quantizeDequantize(input: Tensor, scale: Tensor, zeroPoint: Tensor? = null, quantAxis: Long = -1L, codeDtype: DType = DType.UNDEFINED, outDtype: DType = DType.UNDEFINED, blockSize: Long = 0): Tensor quantizeDequantize(input: Tensor, scale: Tensor, zeroPoint: Tensor? = null, quantAxis: Long = -1L, codeDtype: DType = DType.UNDEFINED, outDtype: DType = DType.UNDEFINED, blockSize: Long = 0L): the quantize_dequantize operator. Fake-quantize: encode input to affine codes and decode straight back (one op, no code tensor materialized), so the output is input snapped onto the quantization grid. scale / zero_point / quant_axis / block_size declare the scheme exactly as in quantize; code_dtype is the grid's code dtype (Undefined = the zero-point's dtype, else Int8); out_dtype picks the float output dtype (Undefined = input's own). The quantization-aware inspection/tooling entry. |
| quantizeInPlace | [common] fun quantizeInPlace(input: Tensor, out: Tensor, scale: Tensor, zeroPoint: Tensor? = null, quantAxis: Long = -1L, blockSize: Long = 0): Tensor quantizeInPlace(input: Tensor, out: Tensor, scale: Tensor, zeroPoint: Tensor? = null, quantAxis: Long = -1L, blockSize: Long = 0L): the quantize_ operator. In-place sibling taking a PLAIN pre-allocated code buffer + explicit affine params: encode input INTO out (no allocation) and attach the scheme, so on return out is a proper quantized tensor. Parameters as in quantize above. Returns out under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. |
| quickGelu | [common] fun quickGelu(input: Tensor): Tensor quickGelu(input: Tensor): the quick_gelu operator. QuickGELU: the sigmoid GELU approximation. |
| rad2deg | [common] fun rad2deg(input: Tensor): Tensor rad2deg(input: Tensor): the rad2deg operator. Converts radians to degrees elementwise: . |
| rad2degInPlace | [common] fun rad2degInPlace(self: Tensor): Tensor rad2degInPlace(self: Tensor): the rad2deg_ operator. In-place rad2deg: writes the degrees through x. Writes through self and returns it, so calls chain. |
| rand | [common] fun rand(shape: LongArray, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor rand(shape: LongArray, dtype: DType = DType.FLOAT32, device: Placement? = null): the rand operator. Uniform random tensor on [0, 1). |
| randint | [common] fun randint(low: Long, high: Long, shape: LongArray, dtype: DType = DType.INT64, device: Placement? = null): Tensor randint(low: Long, high: Long, shape: LongArray, dtype: DType = DType.INT64, device: Placement? = null): the randint operator. Uniform random integers in the half-open range [low, high). |
| randintLike | [common] fun randintLike(reference: Tensor, low: Long, high: Long, device: Placement? = null): Tensor randintLike(reference: Tensor, low: Long, high: Long, device: Placement? = null): the randint_like operator. Uniform integers in [low, high) shaped and typed like reference. |
| randLike | [common] fun randLike(reference: Tensor, device: Placement? = null): Tensor randLike(reference: Tensor, device: Placement? = null): the rand_like operator. Uniform [0, 1) draws shaped and typed like reference. |
| randn | [common] fun randn(shape: LongArray, dtype: DType = DType.FLOAT32, device: Placement? = null): Tensor randn(shape: LongArray, dtype: DType = DType.FLOAT32, device: Placement? = null): the randn operator. Standard-normal random tensor, N(0, 1). |
| randnLike | [common] fun randnLike(reference: Tensor, device: Placement? = null): Tensor randnLike(reference: Tensor, device: Placement? = null): the randn_like operator. N(0, 1) draws shaped and typed like reference. |
| randomInPlace | [common] fun randomInPlace(self: Tensor, low: Long? = null, high: Long? = null, device: Placement? = null): Tensor randomInPlace(self: Tensor, low: Long? = null, high: Long? = null, device: Placement? = null): the random_ operator. In-place uniform-INTEGER fill: overwrites self with draws from [low, high). low and high come both or neither; with both absent, the range is the dtype's full representable-integer span. A view input writes through its base storage. Returns self for chaining. Writes through self and returns it, so calls chain. |
| randperm | [common] fun randperm(n: Tensor, dtype: DType = DType.INT64, device: Placement? = null): Tensor randperm(n: Tensor, dtype: DType = DType.INT64, device: Placement? = null): the randperm operator. A random permutation of the integers 0 .. n-1.[common] fun randperm(n: Double, dtype: DType = DType.INT64, device: Placement? = null): Tensor randperm(n: Double, dtype: DType = DType.INT64, device: Placement? = null): the number form of randperm, n as a scalar. |
| reciprocal | [common] fun reciprocal(input: Tensor): Tensor reciprocal(input: Tensor): the reciprocal operator. Elementwise reciprocal. Zero produces ±inf. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| reciprocalInPlace | [common] fun reciprocalInPlace(self: Tensor): Tensor reciprocalInPlace(self: Tensor): the reciprocal_ operator. In-place reciprocal: writes the result through self; same formula, arguments, and error conditions as reciprocal(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| reflectPad | [common] fun reflectPad(input: Tensor, pad: LongArray): Tensor reflectPad(input: Tensor, pad: LongArray): the reflect_pad operator. pad in mirror mode: the padding reflects the tensor across each padded edge. Same (lo, hi) pair layout as constant_pad. Copies. |
| reglu | [common] fun reglu(input: Tensor): Tensor reglu(input: Tensor): the reglu operator. ReGLU gated activation over a concatenated gate‖up input. The relu-gated member of the GLU family, beside geglu / swiglu: [.., 2d] in, [.., d] out; the last dim must be even. |
| relu | [common] fun relu(input: Tensor): Tensor relu(input: Tensor): the relu operator. Rectified linear unit. |
| relu6 | [common] fun relu6(input: Tensor): Tensor relu6(input: Tensor): the relu6 operator. ReLU capped at 6. |
| relu6InPlace | [common] fun relu6InPlace(self: Tensor): Tensor relu6InPlace(self: Tensor): the relu6_ operator. In-place relu6: writes the result through self; same formula, arguments, and error conditions as relu6(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| reluInPlace | [common] fun reluInPlace(self: Tensor): Tensor reluInPlace(self: Tensor): the relu_ operator. In-place relu: writes the result through self; same formula, arguments, and error conditions as relu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| remainder | [common] fun remainder(input: Tensor, other: Tensor): Tensor remainder(input: Tensor, other: Tensor): the remainder operator. Elementwise remainder with the sign of the DIVISOR, i.e. ops::mod with ModMode::Python. Broadcasts and promotes as add.[common] fun remainder(input: Tensor, other: Double): Tensor remainder(input: Tensor, other: Double): the number form of remainder, other as a scalar. |
| remainderInPlace | [common] fun remainderInPlace(self: Tensor, other: Tensor): Tensor remainderInPlace(self: Tensor, other: Tensor): the remainder_ operator. In-place remainder: writes the remainders through x. Writes through self and returns it, so calls chain.[common] fun remainderInPlace(self: Tensor, other: Double): Tensor remainderInPlace(self: Tensor, other: Double): the number form of remainder_, other as a scalar. |
| renorm | [common] fun renorm(input: Tensor, p: Double, dim: Long, maxnorm: Double, eps: Double? = null): Tensor renorm(input: Tensor, p: Double, dim: Long, maxnorm: Double, eps: Double? = null): the renorm operator. Caps each sub-tensor's p-norm along dim at maxnorm. Every slice along dim whose p-norm exceeds maxnorm is rescaled to exactly maxnorm; slices already within the bound pass through unchanged. |
| renormInPlace | [common] fun renormInPlace(self: Tensor, p: Double, dim: Long, maxnorm: Double, eps: Double? = null): Tensor renormInPlace(self: Tensor, p: Double, dim: Long, maxnorm: Double, eps: Double? = null): the renorm_ operator. In-place renorm: rewrites self with each slice's norm capped at maxnorm and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| repeat | [common] fun repeat(input: Tensor, sizes: LongArray): Tensor repeat(input: Tensor, sizes: LongArray): the repeat operator. Tile input by per-dim repeat counts: sizes[i] copies along dim i. sizes must carry at least the input's rank; extra LEADING entries add fresh leading dims. Counts are literals or 0-D integer Tensors (a tensor count traces symbolically). Copies; numel scales by the product of the counts. |
| repeatInterleave | [common] fun repeatInterleave(input: Tensor, repeats: Tensor, dim: Long? = null, outputSize: Long? = null): Tensor repeatInterleave(input: Tensor, repeats: Tensor, dim: Long? = null, outputSize: Long? = null): the repeat_interleave operator. Repeat ELEMENTS, not blocks: each element along dim appears repeats times consecutively ([a, b] with repeats 2 ->[a, a, b, b]). repeats is a scalar (uniform count) or a 1-D tensor matching the dim's length (per-element counts, output length = their sum). With dim absent the input is flattened first. output_size is the known result length along the dim; pass it when the caller already has it, so a tensor-valued repeats skips the blocking host-side count. Copies.[common] fun repeatInterleave(input: Tensor, repeats: Double, dim: Long? = null, outputSize: Long? = null): Tensor repeatInterleave(input: Tensor, repeats: Double, dim: Long? = null, outputSize: Long? = null): the number form of repeat_interleave, repeats as a scalar. |
| replicatePad | [common] fun replicatePad(input: Tensor, pad: LongArray): Tensor replicatePad(input: Tensor, pad: LongArray): the replicate_pad operator. pad in edge-replicate mode: the padding repeats each padded edge's value. Same (lo, hi) pair layout as constant_pad. Copies. |
| resample | [common] fun resample(input: Tensor, origFreq: Long, newFreq: Long, lowpassFilterWidth: Long = 16, rolloff: Double = 0.945, beta: Double? = null): Tensor resample(input: Tensor, origFreq: Long, newFreq: Long, lowpassFilterWidth: Long = 16L, rolloff: Double = 0.945, beta: Double? = null): the resample operator. Audio-domain 1-D rational resample along the LAST axis: [.., S] at orig_freq → [.., ceil(S · L / M)] at new_freq, where L/M is the gcd-reduced rate pair. A kaiser-windowed-sinc polyphase FIR, the standard antialiased sample-rate converter; input outside the signal reads as zero, and equal rates pass the signal through exactly. Leading axes are batch/channels (transpose another samples axis to the back first, a view). The defaults are the kaiser preset: lowpass_filter_width sinc zero-crossings per side, rolloff of the target Nyquist, and (when beta is absent) the design beta 14.769656459379492 (~142.7 dB design stopband). Serves f32 natively and f16/bf16 through an f32 compute lane (output mirrors the input); integer and f64 signals refuse; cast to f32 first. CPU-served; device-resident inputs refuse until an accelerator kernel lands. |
| reshape | [common] fun reshape(input: Tensor, shape: LongArray): Tensor reshape(input: Tensor, shape: LongArray): the reshape operator. Reshape to shape; one entry may be -1 to infer it from the element count. A dim entry is an int literal OR a 0-D/1-D integer Tensor (IndexBound); e.g. reshape(x, {shape_host(x, 1), d}) composes without reading sizes to the host. |
| reshapeAs | [common] fun reshapeAs(input: Tensor, other: Tensor): Tensor reshapeAs(input: Tensor, other: Tensor): the reshape_as operator. Reshape input to other's shape (other supplies extents only; its data is never read). The element counts must match. Returns a view sharing storage when input's stride layout can express the new shape; copies into a fresh dense tensor otherwise. |
| rfft | [common] fun rfft(input: Tensor, nFft: Long? = null, normalized: Boolean = false): Tensor rfft(input: Tensor, nFft: Long? = null, normalized: Boolean = false): the rfft operator. One-sided real FFT along the last axis: real [..., n] → interleaved (re, im) pairs [..., n_fft/2 + 1, 2]. n_fft defaults to the last axis's extent; normalized scales the spectrum by 1/sqrt(n_fft). |
| rmsNorm | [common] fun rmsNorm(input: Tensor, normalizedShape: LongArray, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): Tensor rmsNorm(input: Tensor, normalizedShape: LongArray, weight: Tensor? = null, bias: Tensor? = null, eps: Double? = null, activation: Activation? = null): the rms_norm operator. Root-mean-square normalization over the trailing normalized_shape dims (no mean subtraction). The transformer-style norm: statistics are the mean SQUARE only, per position over the trailing dims. Optional fused activation applies to the post-affine value (gated kinds are not accepted). |
| roll | [common] fun roll(input: Tensor, shifts: LongArray, dims: LongArray = longArrayOf()): Tensor roll(input: Tensor, shifts: LongArray, dims: LongArray = longArrayOf()): the roll operator. Circularly shift elements: shifts[i] positions along dims[i]; elements that fall off one end re-enter at the other. With dims empty (the default) the tensor is treated as flattened row-major and shifted by shifts[0]. Copies. |
| rot90 | [common] fun rot90(input: Tensor, k: Long = 1, dims: LongArray = longArrayOf(0, 1)): Tensor rot90(input: Tensor, k: Long = 1L, dims: LongArray = longArrayOf(0, 1)): the rot90 operator. Rotate the plane spanned by dims by k x 90 degrees (k is taken mod 4; odd rotations swap the two dims' extents). Default plane {0, 1}. Copies. |
| rotaryEmbedding | [common] fun rotaryEmbedding(input: Tensor, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): Tensor rotaryEmbedding(input: Tensor, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): the rotary_embedding operator. Rotary position embedding over padded [.., S, D] layouts. Rotates the leading rotary_dim of each head vector by per-position angles (RoPE). cos / sin are the precomputed angle planes [max_pos, rotary_dim/2] Float32; position_ids (Int32/Int64) selects each token's row; absent, positions run 0, 1, 2, … per sequence. mode picks the pair layout (interleaved vs half-split); rotary_dim absent rotates the whole head dim. |
| rotaryEmbeddingQk | [common] fun rotaryEmbeddingQk(query: Tensor, key: Tensor, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): List<Tensor> rotaryEmbeddingQk(query: Tensor, key: Tensor, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): the rotary_embedding_qk operator. |
| rotaryEmbeddingQkVarlen | [common] fun rotaryEmbeddingQkVarlen(query: Tensor, key: Tensor, cuSeqlens: Tensor, seqlens: Tensor? = null, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): List<Tensor> rotaryEmbeddingQkVarlen(query: Tensor, key: Tensor, cuSeqlens: Tensor, seqlens: Tensor? = null, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): the rotary_embedding_qk_varlen operator. |
| rotaryEmbeddingVarlen | [common] fun rotaryEmbeddingVarlen(input: Tensor, cuSeqlens: Tensor, seqlens: Tensor? = null, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): Tensor rotaryEmbeddingVarlen(input: Tensor, cuSeqlens: Tensor, seqlens: Tensor? = null, positionIds: Tensor? = null, cos: Tensor? = null, sin: Tensor? = null, mode: RotaryMode? = null, rotaryDim: Long? = null, theta: Double? = null, scaling: RopeScaling? = null, scale: Double? = null, lowFreqFactor: Double? = null, highFreqFactor: Double? = null, originalMaxPos: Long? = null, betaFast: Double? = null, betaSlow: Double? = null, freqFactors: Tensor? = null): the rotary_embedding_varlen operator. Rotary position embedding over token-packed (variable-length) layouts. Packed [Sum(S), .., D] input with cu_seqlens``[B+1] Int32 prefix sums (per-sequence positions restart at 0 unless position_ids is given); seqlens optionally carries explicit per-sequence lengths. angles (RoPE). cos / sin are the precomputed angle planes [max_pos, rotary_dim/2] Float32; position_ids (Int32/Int64) selects each token's row; absent, positions run 0, 1, 2, … per sequence. mode picks the pair layout (interleaved vs half-split); rotary_dim absent rotates the whole head dim. |
| round | [common] fun round(input: Tensor, decimals: Long = 0): Tensor round(input: Tensor, decimals: Long = 0L): the round operator. Elementwise rounding to decimals fractional digits. decimals = 0 (default) rounds to integers; positive keeps that many fractional digits; negative rounds to tens, hundreds, …. |
| roundInPlace | [common] fun roundInPlace(self: Tensor, decimals: Long = 0): Tensor roundInPlace(self: Tensor, decimals: Long = 0L): the round_ operator. In-place round: writes the result through self; same formula, arguments, and error conditions as round(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| rsqrt | [common] fun rsqrt(input: Tensor): Tensor rsqrt(input: Tensor): the rsqrt operator. Elementwise reciprocal square root. Negative inputs produce NaN; zero produces +inf. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| rsqrtInPlace | [common] fun rsqrtInPlace(self: Tensor): Tensor rsqrtInPlace(self: Tensor): the rsqrt_ operator. In-place rsqrt: writes the result through self; same formula, arguments, and error conditions as rsqrt(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| scaledDotProductAttention | [common] fun scaledDotProductAttention(query: Tensor, key: Tensor, value: Tensor, attnMask: Tensor? = null, isCausal: Boolean = false, qScale: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null): Tensor scaledDotProductAttention(query: Tensor, key: Tensor, value: Tensor, attnMask: Tensor? = null, isCausal: Boolean = false, qScale: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null): the scaled_dot_product_attention operator. Scaled dot-product attention over dense head-major tensors. s defaults to (the q/k head size) and is replaced wholesale by q_scale when given. Layout is head-major, D innermost. |
| scaledDotProductAttentionVarlen | [common] fun scaledDotProductAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean = false, qScale: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null): Tensor scaledDotProductAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean = false, qScale: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null): the scaled_dot_product_attention_varlen operator. Variable-length (packed) scaled dot-product attention: ragged batches ride one token-packed tensor plus prefix-sum offsets, no padding. Same math as scaled_dot_product_attention; the batch structure moves into cu_seqlens_*. |
| scatter | [common] fun scatter(input: Tensor, dim: Long, index: Tensor, src: Tensor): Tensor scatter(input: Tensor, dim: Long, index: Tensor, src: Tensor): the scatter operator. A copy of input with src written at positions given by index along dim, the write mirror of gather: out[index[p]][j][k] = src[p] for dim = 0 (only that axis's coordinate is redirected). A scalar src broadcasts one value to every indexed position; a tensor src matches index's shape. On duplicate destinations one write wins; use scatter_add / scatter_reduce for well-defined accumulation. Index dtype law as gather.[common] fun scatter(input: Tensor, dim: Long, index: Tensor, src: Double): Tensor scatter(input: Tensor, dim: Long, index: Tensor, src: Double): the number form of scatter, src as a scalar. |
| scatterAdd | [common] fun scatterAdd(input: Tensor, dim: Long, index: Tensor, src: Tensor, deterministic: Boolean = false): Tensor scatterAdd(input: Tensor, dim: Long, index: Tensor, src: Tensor, deterministic: Boolean = false): the scatter_add operator. scatter with ACCUMULATION: out[.., index[p], ..] += src[p]; duplicate destinations sum. |
| scatterAddInPlace | [common] fun scatterAddInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor, deterministic: Boolean = false): Tensor scatterAddInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor, deterministic: Boolean = false): the scatter_add_ operator. In-place scatter_add: accumulates through self at the indexed positions; same arguments (incl. deterministic) and error conditions as scatter_add(). A strided view reaches its base buffer. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| scatterInPlace | [common] fun scatterInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor): Tensor scatterInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor): the scatter_ operator. In-place scatter: writes through self at the indexed positions; same arguments and error conditions as scatter(). Writes through self's storage (a strided view reaches its base buffer). Returns self for chaining. Writes through self and returns it, so calls chain.[common] fun scatterInPlace(self: Tensor, dim: Long, index: Tensor, src: Double): Tensor scatterInPlace(self: Tensor, dim: Long, index: Tensor, src: Double): the number form of scatter_, src as a scalar. |
| scatterReduce | [common] fun scatterReduce(input: Tensor, dim: Long, index: Tensor, src: Tensor, reduce: ScatterReduceMode, includeSelf: Boolean = true, deterministic: Boolean = false): Tensor scatterReduce(input: Tensor, dim: Long, index: Tensor, src: Tensor, reduce: ScatterReduceMode, includeSelf: Boolean = true, deterministic: Boolean = false): the scatter_reduce operator. scatter with a REDUCTION at each destination: Sum / Prod / Mean / AMax / AMin (ScatterReduceMode). |
| scatterReduceInPlace | [common] fun scatterReduceInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor, reduce: ScatterReduceMode, includeSelf: Boolean = true, deterministic: Boolean = false): Tensor scatterReduceInPlace(self: Tensor, dim: Long, index: Tensor, src: Tensor, reduce: ScatterReduceMode, includeSelf: Boolean = true, deterministic: Boolean = false): the scatter_reduce_ operator. In-place scatter_reduce: reduces into self at the indexed positions; same arguments and error conditions as scatter_reduce(). A strided view reaches its base buffer. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| searchsorted | [common] fun searchsorted(sortedSequence: Tensor, values: Tensor, outInt32: Boolean = false, right: Boolean = false, side: Boolean? = null, sorter: Tensor? = null): Tensor searchsorted(sortedSequence: Tensor, values: Tensor, outInt32: Boolean = false, right: Boolean = false, side: Boolean? = null, sorter: Tensor? = null): the searchsorted operator. Insertion points of values into a sorted sequence. For each value, the index in sorted_sequence's last axis where it would insert to keep the order: right == false gives the leftmost admissible slot, true the rightmost. side, when present, must AGREE with right (it is the same switch under its string-API name); sorter supplies indices that sort an unsorted sequence. |
| select | [common] fun select(input: Tensor, dim: Long, index: Long): Tensor select(input: Tensor, dim: Long, index: Long): the select operator. One position along dim (the dim is removed). index is an int literal OR a 0-D integer Tensor (e.g. a shape_host-derived index) so a data-dependent select never reads the value to the host. |
| selu | [common] fun selu(input: Tensor): Tensor selu(input: Tensor): the selu operator. Scaled exponential linear unit: elu with the fixed SELU constants (alpha ~= 1.6733, scale ~= 1.0507) from the self-normalizing-networks formulation. |
| seluInPlace | [common] fun seluInPlace(self: Tensor): Tensor seluInPlace(self: Tensor): the selu_ operator. In-place selu: writes the result through self; same formula, arguments, and error conditions as selu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sgn | [common] fun sgn(input: Tensor): Tensor sgn(input: Tensor): the sgn operator. Elementwise sign; identical to sign for real dtypes. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| sgnInPlace | [common] fun sgnInPlace(self: Tensor): Tensor sgnInPlace(self: Tensor): the sgn_ operator. In-place sgn: writes the result through self; same formula, arguments, and error conditions as sgn(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| shape | [common] fun shape(input: Tensor, dim: Long? = null): Tensor shape(input: Tensor, dim: Long? = null): the shape operator. The tensor's shape as a 1-D Int64 tensor (or one extent as 0-D). The value is produced ON DEVICE and is traceable; under tracing it carries the symbolic value. For a plain host integer use the *_host sibling instead. With dim set, returns that single extent as a 0-D Int64 tensor (negative dim counts from the end).[common] fun shape(input: Tensor, start: Long?, end: Long?): Tensor shape(input: Tensor, start: Long?, end: Long?): the shape operator. The [start, end) window of the shape as ONE 1-D Int64 tensor, a contiguous dim range in a single call (Python-slice bounds: negative values count from the end, std::nullopt leaves that side open). Same on-device, traceable contract as shape(x); the host-integer sibling is shape_host(x, start, end). |
| shapeHost | [common] fun shapeHost(input: Tensor): Tensor shapeHost(input: Tensor): the shape_host operator. The shape of input as a HOST-resident Int64 tensor, written from metadata: no device kernel, no readback, no stream synchronize, whatever device input lives on. The cheap way to feed shape values to shape-consuming arguments (a reshape dim) instead of Tensor::shape()'s concrete ints.[common] fun shapeHost(input: Tensor, dim: Long): Tensor shapeHost(input: Tensor, dim: Long): the shape_host operator. One dim of the shape as a 0-D Int64 host tensor (negative dim counts from the end). Same no-sync contract as shape_host(x).[common] fun shapeHost(input: Tensor, start: Long?, end: Long?): Tensor shapeHost(input: Tensor, start: Long?, end: Long?): the shape_host operator. The [start, end) window of the shape (Python-slice semantics: negatives count from the end, absent bounds are open) as a 1-D Int64 host tensor. |
| sigmoid | [common] fun sigmoid(input: Tensor): Tensor sigmoid(input: Tensor): the sigmoid operator. Logistic sigmoid. |
| sigmoidInPlace | [common] fun sigmoidInPlace(self: Tensor): Tensor sigmoidInPlace(self: Tensor): the sigmoid_ operator. In-place sigmoid: writes the result through self; same formula, arguments, and error conditions as sigmoid(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sign | [common] fun sign(input: Tensor): Tensor sign(input: Tensor): the sign operator. Elementwise sign: -1, 0, or 1. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| signInPlace | [common] fun signInPlace(self: Tensor): Tensor signInPlace(self: Tensor): the sign_ operator. In-place sign: writes the result through self; same formula, arguments, and error conditions as sign(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| silu | [common] fun silu(input: Tensor): Tensor silu(input: Tensor): the silu operator. Sigmoid linear unit (swish). |
| siluInPlace | [common] fun siluInPlace(self: Tensor): Tensor siluInPlace(self: Tensor): the silu_ operator. In-place silu: writes the result through self; same formula, arguments, and error conditions as silu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sin | [common] fun sin(input: Tensor): Tensor sin(input: Tensor): the sin operator. Elementwise sine. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| sinc | [common] fun sinc(input: Tensor): Tensor sinc(input: Tensor): the sinc operator. Elementwise normalized sinc. Normalized convention: zeros at nonzero integers, sinc(0) = 1 (matching torch / numpy). Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| sincInPlace | [common] fun sincInPlace(self: Tensor): Tensor sincInPlace(self: Tensor): the sinc_ operator. In-place sinc: writes the result through self; same formula, arguments, and error conditions as sinc(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sinh | [common] fun sinh(input: Tensor): Tensor sinh(input: Tensor): the sinh operator. Elementwise hyperbolic sine. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| sinhInPlace | [common] fun sinhInPlace(self: Tensor): Tensor sinhInPlace(self: Tensor): the sinh_ operator. In-place sinh: writes the result through self; same formula, arguments, and error conditions as sinh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sinInPlace | [common] fun sinInPlace(self: Tensor): Tensor sinInPlace(self: Tensor): the sin_ operator. In-place sin: writes the result through self; same formula, arguments, and error conditions as sin(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| slice | [common] fun slice(input: Tensor, dim: Long, start: Long? = null, end: Long? = null, step: Long = 1): Tensor slice(input: Tensor, dim: Long, start: Long? = null, end: Long? = null, step: Long = 1L): the slice operator. [start, end) with step along one dim. Each bound is an int literal, a 0-D integer Tensor, or absent ({}, an open bound); absent step is 1.[common] fun slice(input: Tensor, dim: LongArray, start: LongArray = longArrayOf(), end: LongArray = longArrayOf(), step: LongArray = longArrayOf()): Tensor slice(input: Tensor, dim: LongArray, start: LongArray = longArrayOf(), end: LongArray = longArrayOf(), step: LongArray = longArrayOf()): the slice operator. Strided slice: for each axis in dim, take [start, end) with step. An empty start/end/step means full-range / unit-step. |
| smoothL1Loss | [common] fun smoothL1Loss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, beta: Double = 1.0): Tensor smoothL1Loss(input: Tensor, target: Tensor, reduction: Reduction = Reduction.MEAN, beta: Double = 1.0): the smooth_l1_loss operator. Smooth-L1 loss: quadratic within beta of zero, L1 beyond it. beta == 0 degenerates to plain L1. |
| snake | [common] fun snake(input: Tensor, alpha: Tensor? = null, beta: Tensor? = null, eps: Double = 1.0E-9): Tensor snake(input: Tensor, alpha: Tensor? = null, beta: Tensor? = null, eps: Double = 1e-9): the snake operator. Periodic "snake" activation (neural vocoders), per channels-last channel. alpha (frequency) and beta (magnitude) are rank-1 [C] tensors over input's last dim, in the REAL domain. Apply exp once at setup for a log-scale checkpoint parameterization. beta absent selects the plain form (beta = alpha). One fused pass; the composed spelling costs five elementwise calls over the whole stream. |
| snakeInPlace | [common] fun snakeInPlace(self: Tensor, alpha: Tensor? = null, beta: Tensor? = null, eps: Double = 1.0E-9): Tensor snakeInPlace(self: Tensor, alpha: Tensor? = null, beta: Tensor? = null, eps: Double = 1e-9): the snake_ operator. In-place snake: writes the result through self; same formula, arguments, and error conditions as snake(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| softcapLogits | [common] fun softcapLogits(input: Tensor, cap: Tensor): Tensor softcapLogits(input: Tensor, cap: Tensor): the softcap_logits operator. Fused logit soft-cap: cap * tanh(x / cap) (true division), elementwise. A number cap must be 0 (0 or negative is rejected). A tensor cap broadcasts against input and every element must be non-zero; that precondition is the caller's contract and is not runtime-checked.[common] fun softcapLogits(input: Tensor, cap: Double): Tensor softcapLogits(input: Tensor, cap: Double): the number form of softcap_logits, cap as a scalar. |
| softcapLogitsInPlace | [common] fun softcapLogitsInPlace(self: Tensor, cap: Tensor): Tensor softcapLogitsInPlace(self: Tensor, cap: Tensor): the softcap_logits_ operator. In-place softcap_logits: writes the result through self; same formula, arguments, and error conditions as softcap_logits(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain.[common] fun softcapLogitsInPlace(self: Tensor, cap: Double): Tensor softcapLogitsInPlace(self: Tensor, cap: Double): the number form of softcap_logits_, cap as a scalar. |
| softmax | [common] fun softmax(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): Tensor softmax(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): the softmax operator. Softmax along dim, computed stably (max-subtracted). |
| softmaxInPlace | [common] fun softmaxInPlace(self: Tensor, dim: Long = -1L): Tensor softmaxInPlace(self: Tensor, dim: Long = -1L): the softmax_ operator. In-place softmax: writes the result through self; same formula, arguments, and error conditions as softmax(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| softmin | [common] fun softmin(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): Tensor softmin(input: Tensor, dim: Long = -1L, dtype: DType = DType.UNDEFINED): the softmin operator. Softmax of the negated input: weights small values highest. |
| softminInPlace | [common] fun softminInPlace(self: Tensor, dim: Long = -1L): Tensor softminInPlace(self: Tensor, dim: Long = -1L): the softmin_ operator. In-place softmin: writes the result through self; same formula, arguments, and error conditions as softmin(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| softplus | [common] fun softplus(input: Tensor, beta: Double = 1.0, threshold: Double = 20.0): Tensor softplus(input: Tensor, beta: Double = 1.0, threshold: Double = 20.0): the softplus operator. Smooth ReLU. For numerical stability the exact linear input is returned where beta * x > threshold. |
| softplusInPlace | [common] fun softplusInPlace(self: Tensor, beta: Double = 1.0, threshold: Double = 20.0): Tensor softplusInPlace(self: Tensor, beta: Double = 1.0, threshold: Double = 20.0): the softplus_ operator. In-place softplus: writes the result through self; same formula, arguments, and error conditions as softplus(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| softshrink | [common] fun softshrink(input: Tensor, lambd: Double = 0.5): Tensor softshrink(input: Tensor, lambd: Double = 0.5): the softshrink operator. Soft thresholding: shrinks every element toward zero by lambd. |
| softshrinkInPlace | [common] fun softshrinkInPlace(self: Tensor, lambd: Double = 0.5): Tensor softshrinkInPlace(self: Tensor, lambd: Double = 0.5): the softshrink_ operator. In-place softshrink: writes the result through self; same formula, arguments, and error conditions as softshrink(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| softsign | [common] fun softsign(input: Tensor): Tensor softsign(input: Tensor): the softsign operator. Softsign activation. |
| softsignInPlace | [common] fun softsignInPlace(self: Tensor): Tensor softsignInPlace(self: Tensor): the softsign_ operator. In-place softsign: writes the result through self; same formula, arguments, and error conditions as softsign(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| sort | [common] fun sort(input: Tensor, dim: Long = -1L, descending: Boolean = false, stable: Boolean = false): List<Tensor> sort(input: Tensor, dim: Long = -1L, descending: Boolean = false, stable: Boolean = false): the sort operator. |
| splitBySize | [common] fun splitBySize(input: Tensor, chunkSize: Long, dim: Long = 0): List<Tensor> splitBySize(input: Tensor, chunkSize: Long, dim: Long = 0L): the split_by_size operator. Split along dim into pieces of chunk_size, ceil(extent / chunk_size) of them, the last possibly shorter. Returns VIEWS sharing the source's storage (no copy). |
| splitWithSizes | [common] fun splitWithSizes(input: Tensor, sizes: LongArray, dim: Long = 0): List<Tensor> splitWithSizes(input: Tensor, sizes: LongArray, dim: Long = 0L): the split_with_sizes operator. Split along dim into chunks of the given lengths; sizes must sum to the dim's extent. Each length is a literal or a 0-D integer Tensor. Returns VIEWS; every chunk shares the source's storage (no copy). |
| sqrt | [common] fun sqrt(input: Tensor): Tensor sqrt(input: Tensor): the sqrt operator. Elementwise square root. Negative inputs produce NaN. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| sqrtInPlace | [common] fun sqrtInPlace(self: Tensor): Tensor sqrtInPlace(self: Tensor): the sqrt_ operator. In-place sqrt: writes the result through self; same formula, arguments, and error conditions as sqrt(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| square | [common] fun square(input: Tensor): Tensor square(input: Tensor): the square operator. Elementwise square. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| squareInPlace | [common] fun squareInPlace(self: Tensor): Tensor squareInPlace(self: Tensor): the square_ operator. In-place square: writes the result through self; same formula, arguments, and error conditions as square(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| squeeze | [common] fun squeeze(input: Tensor, dim: LongArray = longArrayOf()): Tensor squeeze(input: Tensor, dim: LongArray = longArrayOf()): the squeeze operator. Drop size-1 dims: the listed dims (each must be size 1; anything else raises), or EVERY size-1 dim when dim is empty (the default). Returns a view (metadata only, no copy). |
| squeezeInPlace | [common] fun squeezeInPlace(self: Tensor, dim: LongArray = longArrayOf()): Tensor squeezeInPlace(self: Tensor, dim: LongArray = longArrayOf()): the squeeze_ operator. In-place squeeze: reshapes self's handle in place; same rules and error conditions as squeeze(); the storage is untouched. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| ssdUpdate | [common] fun ssdUpdate(input: Tensor, dt: Tensor, aRate: Tensor, bMat: Tensor, cMat: Tensor, dSkip: Tensor?, dtBias: Tensor?, gate: Tensor?, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, dtSoftplus: Boolean = false): Tensor ssdUpdate(input: Tensor, dt: Tensor, aRate: Tensor, bMat: Tensor, cMat: Tensor, dSkip: Tensor?, dtBias: Tensor?, gate: Tensor?, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, dtSoftplus: Boolean = false): the ssd_update operator. Mamba2 / SSD selective-state serving step over a per-sequence [H, dim, dstate] Float32 state (updated IN PLACE; the returned tensor is out [B, T, H, dim], dtype following input). Per token: the state decays by exp(dt'*A[h]) (dt' = softplus(dt + dt_bias[h]) when dt_softplus), accumulates dt'*(x (outer) B), and emits S*C + D[h]*x (optionally silu-gated by gate). A/D/dt_bias are per-head Float32 parameters (A carries the NEGATIVE decay rate; fold -exp(A_log) at bind); B/C are [B, T, G, dstate] with H % G == 0. Decode is T == 1; a prefill runs the same sequential law. seq_lens ([B] Int32) bounds ragged rows. slot_ids ([B] Int32, device-resident) addresses state as a SLAB [num_slots, H, dim, dstate]: batch row b reads/updates slab row slot_ids[b] in place (ids in range and DISTINCT per call, the caller's contract); absent keeps state row b. |
| stack | [common] fun stack(tensors: List<Tensor>, dim: Long = 0): Tensor stack(tensors: List<Tensor>, dim: Long = 0L): the stack operator. Join tensors along a NEW dim at position dim: all inputs share one shape; output rank = input rank + 1, the new dim sized N. Copies. |
| std | [common] fun std(input: Tensor, dims: LongArray = longArrayOf(), correction: Long = 1, keepdim: Boolean = false): Tensor std(input: Tensor, dims: LongArray = longArrayOf(), correction: Long = 1L, keepdim: Boolean = false): the std operator. Standard deviation of input over dims, with Bessel correction. correction is the c above: 1 (the default) gives the sample standard deviation, 0 the population form. |
| stft | [common] fun stft(input: Tensor, nFft: Long, hopLength: Long? = null, winLength: Long? = null, window: Tensor? = null, center: Boolean = true, padMode: PadMode = PadMode.REFLECT, normalized: Boolean = false, onesided: Boolean = true): Tensor stft(input: Tensor, nFft: Long, hopLength: Long? = null, winLength: Long? = null, window: Tensor? = null, center: Boolean = true, padMode: PadMode = PadMode.REFLECT, normalized: Boolean = false, onesided: Boolean = true): the stft operator. Short-time Fourier transform of a real [L] / [B, L] signal: frames of win_length (default n_fft) at hop_length strides (default n_fft/4), windowed by window when given (a [win_length] tensor; absent = rectangular). Output [T, n_freq, 2] / [B, T, n_freq, 2] with n_freq = n_fft/2 + 1; frames along the time axis first, interleaved (re, im) pairs. center pads n_fft/2 per side in pad_mode before framing; normalized scales by 1/sqrt(n_fft). Only the one-sided form is served; onesided = false refuses. |
| sub | [common] fun sub(input: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor sub(input: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the sub operator. Subtracts other (scaled) from input elementwise. Broadcasts and promotes as add; alpha (default 1.0) scales other before the subtract, and activation is the same fused float-only epilogue (default Activation::Identity).[common] fun sub(input: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor sub(input: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the number form of sub, other as a scalar.[common] fun sub(input: Double, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor sub(input: Double, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the sub operator. Scalar-LHS . input keeps its kind (see Scalar). The parameter set mirrors the tensor-first form: alpha scales the TENSOR operand other, activation applies to the result. |
| subInPlace | [common] fun subInPlace(self: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor subInPlace(self: Tensor, other: Tensor, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the sub_ operator. In-place sub: writes through x. Writes through self and returns it, so calls chain.[common] fun subInPlace(self: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): Tensor subInPlace(self: Tensor, other: Double, alpha: Double = 1.0, activation: Activation = Activation.IDENTITY): the number form of sub_, other as a scalar. |
| sum | [common] fun sum(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): Tensor sum(input: Tensor, dims: LongArray = longArrayOf(), keepdim: Boolean = false, dtype: DType = DType.UNDEFINED): the sum operator. Sums input over dims. Empty dims reduces EVERY dimension (a 0-D result unless keepdim). Integer inputs accumulate and return at their own dtype unless dtype widens them explicitly. |
| swiglu | [common] fun swiglu(input: Tensor, alpha: Double = 1.0, beta: Double = 0.0, limit: Double = Double.POSITIVE_INFINITY): Tensor swiglu(input: Tensor, alpha: Double = 1.0, beta: Double = 0.0, limit: Double = Double.POSITIVE_INFINITY): the swiglu operator. SwiGLU over a concatenated [*, 2d] gate‖up input → [*, d], the clamped gated form: out = (clamp(up, ±limit) + beta) · G · sigmoid(alpha·G) with G = min(gate, limit). The defaults reduce exactly to the plain silu(gate) · up. The last dim must be even. |
| synchronizeAll | [common] fun synchronizeAll() synchronizeAll(): the synchronize_all operator. Block until every live stream in the runtime has finished, a barrier for "wait for all enqueued work to complete" (e.g. before reading a device result on the host, or timing a phase). Infallible: it never raises. |
| take | [common] fun take(self: Tensor, index: Tensor): Tensor take(self: Tensor, index: Tensor): the take operator. Read elements of self at FLAT (row-major linearized) positions. self is treated as flattened 1-D: out[..] = self.flat[index[..]]. The output takes index's shape and self's dtype. Index dtype law as gather. |
| takeAlongDim | [common] fun takeAlongDim(input: Tensor, index: Tensor, dim: Long? = null): Tensor takeAlongDim(input: Tensor, index: Tensor, dim: Long? = null): the take_along_dim operator. gather with broadcasting between input and index on the other dims. With dim absent both operands are treated as flattened 1-D. Same index dtype law as gather (Int32/Int64). |
| tan | [common] fun tan(input: Tensor): Tensor tan(input: Tensor): the tan operator. Elementwise tangent. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| tanh | [common] fun tanh(input: Tensor): Tensor tanh(input: Tensor): the tanh operator. Elementwise hyperbolic tangent. Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| tanhInPlace | [common] fun tanhInPlace(self: Tensor): Tensor tanhInPlace(self: Tensor): the tanh_ operator. In-place tanh: writes the result through self; same formula, arguments, and error conditions as tanh(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| tanInPlace | [common] fun tanInPlace(self: Tensor): Tensor tanInPlace(self: Tensor): the tan_ operator. In-place tan: writes the result through self; same formula, arguments, and error conditions as tan(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| tensordot | [common] fun tensordot(a: Tensor, b: Tensor, dimsA: LongArray, dimsB: LongArray): Tensor tensordot(a: Tensor, b: Tensor, dimsA: LongArray, dimsB: LongArray): the tensordot operator. Named-axis contraction: sums a over dims_a against b over dims_b, pairwise. dims_a[i] on a contracts with dims_b[i] on b (equal list lengths; per-pair sizes must agree). The output is a's non-contracted dims followed by b's. |
| threshold | [common] fun threshold(input: Tensor, threshold: Double, value: Double): Tensor threshold(input: Tensor, threshold: Double, value: Double): the threshold operator. Elementwise threshold: keeps values above threshold, replaces the rest with value. |
| thresholdInPlace | [common] fun thresholdInPlace(self: Tensor, threshold: Double, value: Double): Tensor thresholdInPlace(self: Tensor, threshold: Double, value: Double): the threshold_ operator. In-place threshold: writes the result through self; same formula, arguments, and error conditions as threshold(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| tile | [common] fun tile(input: Tensor, dims: LongArray): Tensor tile(input: Tensor, dims: LongArray): the tile operator. repeat that also accepts FEWER counts than the rank; missing leading entries default to 1. Counts are literals or 0-D integer Tensors (a tensor count traces symbolically). Copies. |
| to | [common] fun to(src: Tensor, target: Placement? = null): Tensor to(src: Tensor, target: Placement? = null): the to operator. Moves a tensor to a device or stream. Already-resident inputs pass through (no transfer). The result is produced on target's stream; downstream ops on that stream order after the transfer automatically.[common] fun to(src: Tensor, deviceStr: String): Tensor to(src: Tensor, deviceStr: String): the to operator. String-addressed to: the device is named by string: "cpu", "cuda", "cuda:1", ….[common] fun to(tensors: List<Tensor>, target: Placement? = null): List<Tensor> to(tensors: List<Tensor>, target: Placement? = null): the to operator. Multi-tensor to: moves a batch in ONE call; the batched transfer beats N single moves when several tensors cross together (one submission, one ordering point). |
| topk | [common] fun topk(input: Tensor, k: Long, dim: Long = -1L, largest: Boolean = true, sorted: Boolean = true): List<Tensor> topk(input: Tensor, k: Long, dim: Long = -1L, largest: Boolean = true, sorted: Boolean = true): the topk operator. |
| trace | [common] fun trace(input: Tensor): Tensor trace(input: Tensor): the trace operator. Sum of the main diagonal of a rank-2 tensor. |
| transpose | [common] fun transpose(input: Tensor, dim0: Long, dim1: Long): Tensor transpose(input: Tensor, dim0: Long, dim1: Long): the transpose operator. Swap dim0 and dim1, permute for exactly two dims. Returns a view (no copy). |
| tril | [common] fun tril(input: Tensor, diagonal: Long = 0): Tensor tril(input: Tensor, diagonal: Long = 0L): the tril operator. Zero out the entries ABOVE the chosen diagonal of the last two axes; the lower-triangular part survives. diagonal: 0 = main, +k above, -k below. Copies. |
| trilInPlace | [common] fun trilInPlace(self: Tensor, diagonal: Long = 0): Tensor trilInPlace(self: Tensor, diagonal: Long = 0L): the tril_ operator. In-place tril: zeroes the upper triangle through self; same arguments and error conditions as tril(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| triu | [common] fun triu(input: Tensor, diagonal: Long = 0): Tensor triu(input: Tensor, diagonal: Long = 0L): the triu operator. Zero out the entries BELOW the chosen diagonal of the last two axes; the upper-triangular part survives. diagonal: 0 = main, +k above, -k below. Copies. |
| triuInPlace | [common] fun triuInPlace(self: Tensor, diagonal: Long = 0): Tensor triuInPlace(self: Tensor, diagonal: Long = 0L): the triu_ operator. In-place triu: zeroes the lower triangle through self; same arguments and error conditions as triu(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| trunc | [common] fun trunc(input: Tensor): Tensor trunc(input: Tensor): the trunc operator. Elementwise truncation toward zero (drops the fraction). Value-domain violations follow IEEE semantics (NaN / ±inf in the result), never an error. |
| truncInPlace | [common] fun truncInPlace(self: Tensor): Tensor truncInPlace(self: Tensor): the trunc_ operator. In-place trunc: writes the result through self; same formula, arguments, and error conditions as trunc(). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| unflatten | [common] fun unflatten(input: Tensor, dim: Long, sizes: LongArray): Tensor unflatten(input: Tensor, dim: Long, sizes: LongArray): the unflatten operator. Split the dim at dim into sizes, the inverse of flatten. The product of sizes must equal that dim's extent. Returns a view sharing storage when the layout permits; copies otherwise. |
| unfold | [common] fun unfold(input: Tensor, kernelSize: LongArray, dilation: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), stride: LongArray = longArrayOf(), mode: PadMode = PadMode.CONSTANT, value: Double? = null): Tensor unfold(input: Tensor, kernelSize: LongArray, dilation: LongArray = longArrayOf(), padding: LongArray = longArrayOf(), stride: LongArray = longArrayOf(), mode: PadMode = PadMode.CONSTANT, value: Double? = null): the unfold operator. im2col: extracts sliding kernel windows from a channels-last input. Input [N, spatial.., C] (1-D/2-D/3-D); output [N, L, prod(kernel_size), C] where L is the number of window placements; channels-last throughout (window samples sit next to channels). |
| uniformInPlace | [common] fun uniformInPlace(self: Tensor, low: Double = 0.0, high: Double = 1.0, device: Placement? = null): Tensor uniformInPlace(self: Tensor, low: Double = 0.0, high: Double = 1.0, device: Placement? = null): the uniform_ operator. In-place uniform fill on [low, high): overwrites self at its own shape/dtype and returns it under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| unique | [common] fun unique(input: Tensor, sorted: Boolean = true, returnInverse: Boolean = false, returnCounts: Boolean = false, dim: Long? = null): List<Tensor> unique(input: Tensor, sorted: Boolean = true, returnInverse: Boolean = false, returnCounts: Boolean = false, dim: Long? = null): the unique operator. |
| uniqueConsecutive | [common] fun uniqueConsecutive(input: Tensor, returnInverse: Boolean = false, returnCounts: Boolean = false, dim: Long? = null): List<Tensor> uniqueConsecutive(input: Tensor, returnInverse: Boolean = false, returnCounts: Boolean = false, dim: Long? = null): the unique_consecutive operator. |
| unsqueeze | [common] fun unsqueeze(input: Tensor, dim: Long): Tensor unsqueeze(input: Tensor, dim: Long): the unsqueeze operator. Insert a size-1 dim at dim. The position is normalized against the OUTPUT rank, so -1 appends at the trailing end. Returns a view (no copy). |
| unsqueezeInPlace | [common] fun unsqueezeInPlace(self: Tensor, dim: Long): Tensor unsqueezeInPlace(self: Tensor, dim: Long): the unsqueeze_ operator. In-place unsqueeze: reshapes self's handle in place; same rules and error conditions as unsqueeze(); the storage is untouched. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| upsampleBicubic2d | [common] fun upsampleBicubic2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): Tensor upsampleBicubic2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): the upsample_bicubic2d operator. 2-D bicubic resize. Exactly one of sizes (target extents) / scale_factors (multipliers) is given; entries are literals or 0-D tensors (tensor entries trace symbolically). Input is channels-last. |
| upsampleBilinear2d | [common] fun upsampleBilinear2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): Tensor upsampleBilinear2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): the upsample_bilinear2d operator. 2-D bilinear resize. Exactly one of sizes (target extents) / scale_factors (multipliers) is given; entries are literals or 0-D tensors (tensor entries trace symbolically). Input is channels-last. |
| upsampleLinear1d | [common] fun upsampleLinear1d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): Tensor upsampleLinear1d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): the upsample_linear1d operator. interpolate with InterpMode::Linear (linear / bilinear / trilinear by input rank) or InterpMode::Bicubic. |
| upsampleNearest1d | [common] fun upsampleNearest1d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): Tensor upsampleNearest1d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): the upsample_nearest1d operator. interpolate with InterpMode::Nearest over a 1d/2d/3d input. |
| upsampleNearest2d | [common] fun upsampleNearest2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): Tensor upsampleNearest2d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): the upsample_nearest2d operator. 2-D nearest-neighbor resize. Exactly one of sizes (target extents) / scale_factors (multipliers) is given; entries are literals or 0-D tensors (tensor entries trace symbolically). Input is channels-last. |
| upsampleNearest3d | [common] fun upsampleNearest3d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): Tensor upsampleNearest3d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf()): the upsample_nearest3d operator. 3-D nearest-neighbor resize. Exactly one of sizes (target extents) / scale_factors (multipliers) is given; entries are literals or 0-D tensors (tensor entries trace symbolically). Input is channels-last. |
| upsampleTrilinear3d | [common] fun upsampleTrilinear3d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): Tensor upsampleTrilinear3d(input: Tensor, sizes: LongArray = longArrayOf(), scaleFactors: LongArray = longArrayOf(), alignCorners: Boolean = false): the upsample_trilinear3d operator. 3-D trilinear resize. Exactly one of sizes (target extents) / scale_factors (multipliers) is given; entries are literals or 0-D tensors (tensor entries trace symbolically). Input is channels-last. |
| validateRotaryDim | [common] fun validateRotaryDim(rotaryDim: Long, headDim: Long = -1L) validateRotaryDim(rotaryDim: Long, headDim: Long = -1L): the validate_rotary_dim operator. Checks a rotary-embedding configuration up front: rotary_dim must be positive, even, and (when head_dim is given) no larger than head_dim. |
| var | [common] fun var(input: Tensor, dims: LongArray = longArrayOf(), correction: Long = 1, keepdim: Boolean = false): Tensor var(input: Tensor, dims: LongArray = longArrayOf(), correction: Long = 1L, keepdim: Boolean = false): the var operator. Variance of input over dims, with Bessel correction. correction is the c above: 1 (the default) gives the sample variance, 0 the population form. |
| where | [common] fun where(condition: Tensor): Tensor where(condition: Tensor): the where operator. The coordinates where condition is nonzero, sugar for nonzero. Returns the same [n, ndim] Int64 coordinate matrix as ONE tensor (not a per-dim tuple); the row count is data-dependent.[common] fun where(condition: Tensor, input: Tensor, other: Tensor): Tensor where(condition: Tensor, input: Tensor, other: Tensor): the where operator. Elementwise select: condition ? x : y, with numpy-style broadcasting across all three operands. condition is Bool; the output takes the broadcast shape and the promoted common dtype of input and other. |
| xlog1py | [common] fun xlog1py(input: Tensor, other: Tensor): Tensor xlog1py(input: Tensor, other: Tensor): the xlog1py operator. Elementwise with the convention (entropy-style sums stay finite where input is zero). Broadcasts and promotes as add.[common] fun xlog1py(input: Tensor, other: Double): Tensor xlog1py(input: Tensor, other: Double): the number form of xlog1py, other as a scalar. |
| xlogy | [common] fun xlogy(input: Tensor, other: Tensor): Tensor xlogy(input: Tensor, other: Tensor): the xlogy operator. Elementwise with the convention. Broadcasts and promotes as add.[common] fun xlogy(input: Tensor, other: Double): Tensor xlogy(input: Tensor, other: Double): the number form of xlogy, other as a scalar. |
| xlogyInPlace | [common] fun xlogyInPlace(self: Tensor, other: Tensor): Tensor xlogyInPlace(self: Tensor, other: Tensor): the xlogy_ operator. In-place xlogy: writes the result through x. Writes through self and returns it, so calls chain.[common] fun xlogyInPlace(self: Tensor, other: Double): Tensor xlogyInPlace(self: Tensor, other: Double): the number form of xlogy_, other as a scalar. |
| zeroInPlace | [common] fun zeroInPlace(self: Tensor, device: Placement? = null): Tensor zeroInPlace(self: Tensor, device: Placement? = null): the zero_ operator. In-place zero fill: overwrites every element of self with zero (the error conditions of zeros()). A view input writes through its base storage. Returns self under the value surface, Result<void> under the Result surface; CLIKART_CHECK is the mode-stable spelling. Writes through self and returns it, so calls chain. |
| zeros | [common] fun zeros(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): Tensor zeros(shape: LongArray, dtype: DType, device: Placement? = null, pinnedFor: Placement? = null): the zeros operator. A new tensor of the given shape with every element set to zero. Same shape-span and pinned_for contract as empty (extents are literals or 0-D integer Tensors). |
| zerosLike | [common] fun zerosLike(reference: Tensor, device: Placement? = null): Tensor zerosLike(reference: Tensor, device: Placement? = null): the zeros_like operator. A zero-filled tensor with reference's shape and dtype (data never read); device absent = the reference's own device. |