---
title: "mlaAttention"
sidebar_label: "mlaAttention"
description: "Kotlin binding reference: mlaAttention."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Ops](index.md)/[mlaAttention](mlaAttention.md)

# mlaAttention

[common]\
fun [mlaAttention](mlaAttention.md)(qNope: [Tensor](../Tensor/index.md), qPe: [Tensor](../Tensor/index.md), newCkv: [Tensor](../Tensor/index.md), newKpe: [Tensor](../Tensor/index.md), ckvCache: [Tensor](../Tensor/index.md), kpeCache: [Tensor](../Tensor/index.md), kvcacheStart: [Tensor](../Tensor/index.md), cuSeqlensQ: [Tensor](../Tensor/index.md), scale: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html), slotIds: [Tensor](../Tensor/index.md)? = null): [Tensor](../Tensor/index.md)

`mlaAttention(qNope: Tensor, qPe: Tensor, newCkv: Tensor, newKpe: Tensor, ckvCache: Tensor, kpeCache: Tensor, kvcacheStart: Tensor, cuSeqlensQ: Tensor, scale: Double, slotIds: Tensor? = null)`: the `mla_attention` operator. Multi-head Latent Attention over a compressed KV cache, latent-space end to end: q_nope `[ΣS, H, Dl]` (W_UK-absorbed) + q_pe `[ΣS, H, Dr]` score against the per-token compressed rows; the caches (`[max_seqs, max_seq, Dl/Dr]`) take this step's `new_ckv`/`new_kpe` appends IN PLACE (write offsets = `kvcache_start [B]` Int32; `cu_seqlens_q [B+1]` locates each sequence's packed tokens); the returned `[ΣS, H, Dl]` output stays latent (apply the W_UV un-absorption after). `scale` is REQUIRED; the absorbed query's magnitude lives in the model's un-absorbed head dim.