---
title: "groupQueryAttention"
sidebar_label: "groupQueryAttention"
description: "Kotlin binding reference: groupQueryAttention."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Ops](index.md)/[groupQueryAttention](groupQueryAttention.md)

# groupQueryAttention

[common]\
fun [groupQueryAttention](groupQueryAttention.md)(query: [Tensor](../Tensor/index.md), key: [Tensor](../Tensor/index.md), value: [Tensor](../Tensor/index.md), pastKey: [Tensor](../Tensor/index.md)? = null, pastValue: [Tensor](../Tensor/index.md)? = null, kvcacheStart: [Tensor](../Tensor/index.md)? = null, ropeCos: [Tensor](../Tensor/index.md)? = null, ropeSin: [Tensor](../Tensor/index.md)? = null, positionIds: [Tensor](../Tensor/index.md)? = null, attnMask: [Tensor](../Tensor/index.md)? = null, isCausal: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html)? = null, qScale: [Tensor](../Tensor/index.md)? = null, softcap: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, slidingWindow: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)? = null, smoothSoftmax: [Boolean](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-boolean/index.html)? = null, rotaryMode: [RotaryMode](../RotaryMode/index.md)? = null, numHeads: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)? = null, kvNumHeads: [Long](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-long/index.html)? = null, outPresentKey: [Tensor](../Tensor/index.md)? = null, outPresentValue: [Tensor](../Tensor/index.md)? = null, kScale: [Tensor](../Tensor/index.md)? = null, vScale: [Tensor](../Tensor/index.md)? = null, headSink: [Tensor](../Tensor/index.md)? = null, qNormGain: [Tensor](../Tensor/index.md)? = null, kNormGain: [Tensor](../Tensor/index.md)? = null, qkNormEps: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, slotIds: [Tensor](../Tensor/index.md)? = null, keptPrefix: [Tensor](../Tensor/index.md)? = null): [List](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.collections/-list/index.html)&lt;[Tensor](../Tensor/index.md)&gt;

`groupQueryAttention(query: Tensor, key: Tensor, value: Tensor, pastKey: Tensor? = null, pastValue: Tensor? = null, kvcacheStart: Tensor? = null, ropeCos: Tensor? = null, ropeSin: Tensor? = null, positionIds: Tensor? = null, attnMask: Tensor? = null, isCausal: Boolean? = null, qScale: Tensor? = null, softcap: Double? = null, slidingWindow: Long? = null, smoothSoftmax: Boolean? = null, rotaryMode: RotaryMode? = null, numHeads: Long? = null, kvNumHeads: Long? = null, outPresentKey: Tensor? = null, outPresentValue: Tensor? = null, kScale: Tensor? = null, vScale: Tensor? = null, headSink: Tensor? = null, qNormGain: Tensor? = null, kNormGain: Tensor? = null, qkNormEps: Double? = null, slotIds: Tensor? = null, keptPrefix: Tensor? = null)`: the `group_query_attention` operator. Fused GQA with RoPE + KV cache. Bind `out_present_key`/`out_present_value` to the same buffers as `past_key`/`past_value` (a `KVCache::keys`/`values(layer)` view) + pass `kvcache_start` to append the new post-RoPE K/V IN PLACE into the cache (the decode perf path). q/k/v are hidden-folded `[ΣS, heads*head_dim]`; `num_heads`/`kv_num_heads` drive the in-op head split (GQA). `head_sink` is the per-head softmax sink `[H_q]`, a virtual logit folded into the softmax denominator (attention-sink models bind one per layer), the same contract as `group_query_attention_varlen`; it rides the parameter tail here. `q_norm_gain`/`k_norm_gain` engage the POST-rope per-head RMS norm: after the in-op rotation, every head's `[head_dim]` q (and new-k) vector is RMS-normalized with the gain BEFORE any cache append, so the cache holds rotated+normed keys. Each gain is rank-1 `[head_dim]` (one vector shared across heads; any other shape rejects), any float dtype, applied at its own dtype. The gains require the in-op rope planes (`rope_cos`/`rope_sin`), and they travel WITH `qk_norm_eps`: pass the model's own rms-norm epsilon alongside the gains, or neither (a gain without the epsilon, or an epsilon with no gain, rejects). Absent ⇒ the gain-less path, unchanged. A rope-free per-head norm composes `qk_rms_norm` instead. `slot_ids` (`[B]` Int32) names the cache row each batch row appends to and attends from on a continuous `[max_seqs, H_kv, max_seq, D]` cache: a sequence keeps its row while the batch composition changes around it. Absent, batch row b uses cache row b. Every entry must lie in `[0, max_seqs)` and no two rows may share one (each rejects). On that cache `kvcache_start` stays the rank-1 `[B]` layout selector whose values are not read: each row's write offset is `cu_seqlens_k[b] - q_len[b]`. A paged block table and the dense in-place form take no `slot_ids` (the block table is its own row map; the dense form addresses rows by batch index). A paged block table (`[B, max_blocks]` Int32) is validated whole before any read: every entry is `-1` (the unused-tail pad) or a block index in `[0, num_blocks)`, and no two entries of one row name the same physical block; an out-of-range entry and a repeated one alike refuse `INVALID_ARGUMENT` (a repeat would alias two positions onto one block and decode wrong values silently). A read-only attend (`key` and `value` absent) takes no present outputs: use `attention_over_cache`, or leave `out_present_key`/`out_present_value` unbound, and `present_key`/`present_value` come back undefined. A bound present refuses `INVALID_ARGUMENT`. `kept_prefix` keeps each sequence's first P keys attended beside the sliding window (a 0-D value for every sequence, or `[B]` one per sequence; Int32 or Int64): under the causal bound a key is attended when it is inside the first P keys or inside the window; absent, the window alone. A negative P refuses.