---
title: "gatedDeltaUpdate"
sidebar_label: "gatedDeltaUpdate"
description: "Kotlin binding reference: gatedDeltaUpdate."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Ops](index.md)/[gatedDeltaUpdate](gatedDeltaUpdate.md)

# gatedDeltaUpdate

[common]\
fun [gatedDeltaUpdate](gatedDeltaUpdate.md)(query: [Tensor](../Tensor/index.md), key: [Tensor](../Tensor/index.md), value: [Tensor](../Tensor/index.md), beta: [Tensor](../Tensor/index.md), gate: [Tensor](../Tensor/index.md), state: [Tensor](../Tensor/index.md), seqLens: [Tensor](../Tensor/index.md)? = null, slotIds: [Tensor](../Tensor/index.md)? = null, scale: [Double](https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-double/index.html)? = null, gateBias: [Tensor](../Tensor/index.md)? = null, gateScale: [Tensor](../Tensor/index.md)? = null): [Tensor](../Tensor/index.md)

`gatedDeltaUpdate(query: Tensor, key: Tensor, value: Tensor, beta: Tensor, gate: Tensor, state: Tensor, seqLens: Tensor? = null, slotIds: Tensor? = null, scale: Double? = null, gateBias: Tensor? = null, gateScale: Tensor? = null)`: the `gated_delta_update` operator. Gated delta-rule recurrence step: per token, the `[B, HV, K, V]` Float32 `state` decays by `exp(g)`, takes the delta-rule rank-1 update `(beta·k) ⊗ (v − Sᵀk)`, and emits `o = (scale·q)ᵀ S`, updated IN PLACE. The gate's RANK picks the family: `[B, T, HV]` = one scalar per value head; `[B, T, HV, K]` = per key dim. q/k `[B, T, H, K]` (HV % H == 0, grouped heads), v `[B, T, HV, V]`, beta `[B, T, HV]`; `scale` defaults to K^-1/2; `seq_lens` (`[B]` Int32) bounds ragged prefill rows. `slot_ids` (`[B]` Int32, device-resident) addresses `state` as a SLAB `[num_slots, HV, K, V]`: batch row `b` reads/updates slab row `slot_ids[b]` in place (ids in range and DISTINCT per call, the caller's contract); absent keeps state row `b`. TWO gate forms, told apart by which inputs are bound: with `gate_bias` and `gate_scale` absent, `gate` IS the log-space decay (Float32 or the activations' dtype); with both bound (Float32 `gate_bias``[HV]` beside a `[B, T, HV]` gate or `[HV, K]` beside a `[B, T, HV, K]` one, the checkpoint's `dt_bias`; Float32 `gate_scale``[HV]`, the once-folded `-exp(A_log)`), `gate` is the RAW gate projection slice at the activations' dtype and the kernel forms the decay `gate_scale * softplus(g + gate_bias)` in fp32 registers, so no add / softplus / mul pass and no fp32 transient precede the call. One without the other rejects; a backend without the raw-gate arm declines it typed.