//clika-runtime/io.clika.runtime/Ops/groupQueryAttention
groupQueryAttention
[common]
fun groupQueryAttention(query: Tensor, key: Tensor, value: Tensor, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null): List<Tensor>
groupQueryAttention(query: Tensor, key: Tensor, value: Tensor, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null): the group_query_attention operator. Fused GQA with RoPE + KV cache. Bind out_present_key/out_present_value to the same buffers as past_key/past_value (a KVCache::keys/values(layer) view) + pass kvcache_start to append the new post-RoPE K/V IN PLACE into the cache (the decode perf path). q/k/v are hidden-folded [ΣS, heads*head_dim]; num_heads/kv_num_heads drive the in-op head split (GQA). head_sink is the per-head softmax sink [H_q], a virtual logit folded into the softmax denominator (attention-sink models bind one per layer), the same contract as group_query_attention_varlen; it rides the parameter tail here. q_norm_gain/k_norm_gain engage the POST-rope per-head RMS norm: after the in-op rotation, every head's [head_dim] q (and new-k) vector is RMS-normalized with the gain BEFORE any cache append, so the cache holds rotated+normed keys. Each gain is rank-1 [head_dim] (one vector shared across heads; any other shape rejects), any float dtype, applied at its own dtype. The gains require the in-op rope planes (rope_cos/rope_sin), and they travel WITH qk_norm_eps: pass the model's own rms-norm epsilon alongside the gains, or neither (a gain without the epsilon, or an epsilon with no gain, rejects). Absent ⇒ the gain-less path, unchanged. A rope-free per-head norm composes qk_rms_norm instead. slot_ids ([B] Int32) names the cache row each batch row appends to and attends from on a continuous [max_seqs, H_kv, max_seq, D] cache: a sequence keeps its row while the batch composition changes around it. Absent, batch row b uses cache row b. Every entry must lie in [0, max_seqs) and no two rows may share one (each rejects). On that cache kvcache_start stays the rank-1 [B] layout selector whose values are not read: each row's write offset is cu_seqlens_k[b] - q_len[b]. A paged block table and the dense in-place form take no slot_ids (the block table is its own row map; the dense form addresses rows by batch index). A paged block table ([B, max_blocks] Int32) is validated whole before any read: every entry is -1 (the unused-tail pad) or a block index in [0, num_blocks), and no two entries of one row name the same physical block; an out-of-range entry and a repeated one alike refuse INVALID_ARGUMENT (a repeat would alias two positions onto one block and decode wrong values silently). A read-only attend (key and value absent) takes no present outputs: use attention_over_cache, or leave out_present_key/out_present_value unbound, and present_key/present_value come back undefined. A bound present refuses INVALID_ARGUMENT. kept_prefix keeps each sequence's first P keys attended beside the sliding window (a 0-D value for every sequence, or [B] one per sequence; Int32 or Int64): under the causal bound a key is attended when it is inside the first P keys or inside the window; absent, the window alone. A negative P refuses.