//clika-runtime/io.clika.runtime/Ops/mlaAttention
mlaAttention
[common]
fun mlaAttention(qNope: Tensor, qPe: Tensor, newCkv: Tensor, newKpe: Tensor, ckvCache: Tensor, kpeCache: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, scale: Double, slotIds: Tensor? = null): Tensor
mlaAttention(qNope: Tensor, qPe: Tensor, newCkv: Tensor, newKpe: Tensor, ckvCache: Tensor, kpeCache: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, scale: Double, slotIds: Tensor? = null): the mla_attention operator. Multi-head Latent Attention over a compressed KV cache, latent-space end to end: q_nope [ΣS, H, Dl] (W_UK-absorbed) + q_pe [ΣS, H, Dr] score against the per-token compressed rows; the caches ([max_seqs, max_seq, Dl/Dr]) take this step's new_ckv/new_kpe appends IN PLACE (write offsets = kvcache_start [B] Int32; cu_seqlens_q [B+1] locates each sequence's packed tokens); the returned [ΣS, H, Dl] output stays latent (apply the W_UV un-absorption after). scale is REQUIRED; the absorbed query's magnitude lives in the model's un-absorbed head dim.