Skip to main content

//clika-runtime/io.clika.runtime/Ops/groupQueryAttentionVarlen

groupQueryAttentionVarlen

[common]
fun groupQueryAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): List<Tensor>

groupQueryAttentionVarlen(query: Tensor, key: Tensor, value: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, headSink: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): the group_query_attention_varlen operator. The varlen variant of the fused GQA above, with the same qk-norm tail: q_norm_gain/k_norm_gain (rank-1 [head_dim], post-rope, applied before the K/V append) travel with qk_norm_eps (the model's rms-norm epsilon) and require the in-op rope planes; absent ⇒ unchanged. The gains are the after-rotation order; a model that norms its heads before the rotation (the Qwen3 order) runs qk_rms_norm on its projections ahead of this op and passes no gains. The same slot_ids contract: on a continuous cache it names each batch row's cache row ([B] Int32, in range, pairwise distinct; absent = row b for batch row b), and kvcache_start's values stay unread there (the write offsets derive from cu_seqlens_k). The same block-table contract on a paged pool: every entry of the [B, max_blocks] table is -1 or a block index in [0, num_blocks), and no two entries of one row name the same physical block; an out-of-range or repeated entry refuses INVALID_ARGUMENT before any read or append. The same read-only rule: with key and value absent it takes no present outputs; use attention_over_cache, or leave out_present_key/out_present_value unbound. kept_prefix keeps each request's first P keys attended beside the sliding window (a 0-D value or [B], Int32 or Int64), the same law as attention's. kv_position_offset ([B], Int32 or Int64) is each paged row's served-key offset: when a windowed paged cache serves only a row's recent blocks, its served key at view position i sits at absolute position i + offset, so the in-op rotary and the window read true positions; pass KVCache::StepIndices::kv_position_offset through as-is (undefined on every other cache).