KVCacheLayerInfo
One attention layer's KV-cache wiring in a graph: which inputs carry its past keys and values in, which outputs carry the present ones out (any may be empty), and the layer's attention properties.
is_cross_attention (property)
True when the layer attends the encoder stream.
is_self_attention (property)
True when the layer attends its own sequence.
layer_idx (property)
The sequential layer index.
mask_kind (property)
The mask family: 'causal' or 'sliding_window_causal'.
num_heads (property)
Query heads.
num_kv_heads (property)
Key/value heads (grouped-query attention when fewer than num_heads).
op_name (property)
The attention node's name.
past_key_input (property)
The graph input carrying the past keys ('' when none).
past_value_input (property)
The graph input carrying the past values ('' when none).
present_key_output (property)
The graph output carrying the present keys ('' when none).
present_value_output (property)
The graph output carrying the present values ('' when none).
sliding_window_size (property)
The sliding window in tokens; 0 when the layer attends the whole past.
__init__
__init__(self, /, *args, **kwargs)
Initialize self. See help(type(self)) for accurate signature.