Skip to main content

//clika-runtime/io.clika.runtime/Ops/attentionOverCache

attentionOverCache

[common]
fun attentionOverCache(query: Tensor, cacheKey: Tensor, cacheValue: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): Tensor

attentionOverCache(query: Tensor, cacheKey: Tensor, cacheValue: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, cuSeqlensK: Tensor, maxSeqlenQ: Tensor? = null, maxSeqlenK: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null, kvPositionOffset: Tensor? = null): the attention_over_cache operator. Attend a KV cache WITHOUT appending to it (a read-only re-attention). q is packed varlen [ΣS_q, H, D] (or hidden-folded [ΣS_q, num_heads*D] with num_heads set); cache_key/cache_value are a cache another attention call already appended: continuous head-major [max_seqs, H_kv, max_seq, D] with a rank-1 [B]``kvcache_start (the layout selector; its values are not read, and q's position derives from cu_seqlens_k), or a paged block pool [num_blocks, H_kv, block_size, D] with a rank-2 [B, max_blocks] block table (every entry -1 or in [0, num_blocks), no block index repeated within a row; an out-of-range or repeated entry refuses INVALID_ARGUMENT before any read). On the continuous cache slot_ids ([B] Int32, in range, pairwise distinct) names each batch row's cache row; absent, batch row b reads cache row b. cu_seqlens_k is each sequence's TOTAL cached length: per-seq [B] or cumulative [B+1]. The cache is never written. When rope_cos/rope_sin are bound, rotary applies to q only (the cached keys are already rotated). Use case: a q-only module re-attending a sibling layer's cache. head_sink is the per-head softmax sink [H_q], a virtual logit folded into the softmax denominator, the same contract as group_query_attention_varlen; it rides the parameter tail here.