---
title: "msDeformAttention"
sidebar_label: "msDeformAttention"
description: "Kotlin binding reference: msDeformAttention."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

//[clika-runtime](../../../index.md)/[io.clika.runtime](../index.md)/[Ops](index.md)/[msDeformAttention](msDeformAttention.md)

# msDeformAttention

[common]\
fun [msDeformAttention](msDeformAttention.md)(value: [Tensor](../Tensor/index.md), spatialShapes: [Tensor](../Tensor/index.md), levelStartIndex: [Tensor](../Tensor/index.md), samplingLocations: [Tensor](../Tensor/index.md), attentionWeights: [Tensor](../Tensor/index.md)): [Tensor](../Tensor/index.md)

`msDeformAttention(value: Tensor, spatialShapes: Tensor, levelStartIndex: Tensor, samplingLocations: Tensor, attentionWeights: Tensor)`: the `ms_deform_attention` operator. Multi-scale deformable attention (2-D): per query and head, gather `P` bilinear samples from each of `L` flattened feature-map levels and combine them with the given weights. `value [N, S, M, D]` with `S = Σ_l H_l·W_l`; `spatial_shapes [L, 2]` = per-level `(H_l, W_l)` and `level_start_index [L]` (both Int32 or Int64); `sampling_locations [N, Lq, M, L, P, 2]`; last dim `(x, y)`, normalized to `[0, 1]` per level, sampled at `loc·size − 0.5` (bilinear; out-of-bounds reads 0); `attention_weights [N, Lq, M, L, P]` are consumed AS GIVEN (apply softmax beforehand if wanted). Returns `[N, Lq, M, D]` at `value`'s dtype; accumulation is fp32. The float inputs must share `value`'s dtype (f32/f64/f16/bf16; no silent promotion).