Skip to main content

KVCacheLayerInfo

One attention layer's KV-cache wiring in a graph: which inputs carry its past keys and values in, which outputs carry the present ones out (any may be empty), and the layer's attention properties.

is_cross_attention (property)​

True when the layer attends the encoder stream.

is_self_attention (property)​

True when the layer attends its own sequence.

layer_idx (property)​

The sequential layer index.

mask_kind (property)​

The mask family: 'causal' or 'sliding_window_causal'.

num_heads (property)​

Query heads.

num_kv_heads (property)​

Key/value heads (grouped-query attention when fewer than num_heads).

op_name (property)​

The attention node's name.

past_key_input (property)​

The graph input carrying the past keys ('' when none).

past_value_input (property)​

The graph input carrying the past values ('' when none).

present_key_output (property)​

The graph output carrying the present keys ('' when none).

present_value_output (property)​

The graph output carrying the present values ('' when none).

sliding_window_size (property)​

The sliding window in tokens; 0 when the layer attends the whole past.

__init__​

__init__(self, /, *args, **kwargs)

Initialize self. See help(type(self)) for accurate signature.