Skip to main content

ClikaRT::graph::ModelGraph

class

Header: ClikaRT/graph/model_graph.h

An executable model graph. Move-only.

Member functions​

ModelGraph(ModelGraph)​

ModelGraph(ModelGraph&&) noexcept

Move-only: transfers the graph.

Declared in ClikaRT/graph/model_graph.h, line 80

operator=(ModelGraph)​

ModelGraph& operator=(ModelGraph&&) noexcept

Move-assign: transfers the graph.

Declared in ClikaRT/graph/model_graph.h, line 82

ModelGraph(ModelGraph)​

ModelGraph(const ModelGraph&) =delete

Declared in ClikaRT/graph/model_graph.h, line 83

operator=(ModelGraph)​

ModelGraph& operator=(const ModelGraph&) =delete

Declared in ClikaRT/graph/model_graph.h, line 84

~ModelGraph()​

~ModelGraph()

Releases the graph.

Declared in ClikaRT/graph/model_graph.h, line 86

input_names()​

std::vector<std::string> input_names() const

Runtime input names, in the model's declaration order, the order the positional run binds. Weights are not inputs; they ride the graph.

Declared in ClikaRT/graph/model_graph.h, line 91

output_names()​

std::vector<std::string> output_names() const

Output names, declaration order, the order every run returns.

Declared in ClikaRT/graph/model_graph.h, line 93

inputs()​

std::vector<spec::TensorSpec> inputs() const

Input / output slots as (name, dtype, dims) specs, same order as input_names() / output_names(). A dimension the graph left dynamic reads as spec::TensorSpec::kDynamicDim (-1).

Declared in ClikaRT/graph/model_graph.h, line 98

outputs()​

std::vector<spec::TensorSpec> outputs() const

The produced edges as specs, the output twin of inputs().

Declared in ClikaRT/graph/model_graph.h, line 100

run(vector<Tensor>)​

std::vector<Tensor> run(std::vector<Tensor> inputs) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 111

run(NamedTensors)​

std::vector<Tensor> run(const NamedTensors& inputs) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 118

run(vector<Tensor>, vector<Tensor>)​

std::vector<Tensor> run(std::vector<Tensor> inputs, std::vector<Tensor> bound_outputs) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 128

run(NamedTensors, NamedTensors)​

std::vector<Tensor> run(const NamedTensors& inputs, const NamedTensors& bound_outputs) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 139

attach_kv_cache(nn::KVCacheConfig)​

void attach_kv_cache(const nn::KVCacheConfig& config)

Build ONE nn::KVCache on the graph's device from kv_cache_info() and bind it to every self-attention layer. Fields of config left at 0 are taken from the graph: num_layers (the self-attention layer count), num_kv_heads and head_dim (each layer's own values fill its row, so a stack with mixed geometry binds per layer); a layer's sliding window becomes its row's window; kv_dtype is taken from the graph's KV planes (the config's value is not consulted). mode, paged, max_seqs, max_tokens_per_seq and the quantization fields come from config as given. After this call the past inputs leave input_names(); the present outputs stay in output_names() and run(inputs, step) returns each layer's cache plane at its present position, no copy. Every self-attention layer reports KVCacheLayerInfo::bound_to_cache, and the graph runs through run(inputs, const nn::KVCache::StepIndices&).

The present K/V of a bound layer ARE the cache planes: read them through kv_cache()->keys(l) / values(l) (a continuous cache's [max_seqs, kv_heads, max_tokens_per_seq, head_dim] plane; a paged cache's block pool, addressed through paged_block_table()). A graph with an interior consumer of a present output (a graph node reading it) refuses to attach (Unsupported): the in-place append needs the plane bound at the output port.

Refuses InvalidArgument on a graph with no KV layers, one already attached, or a config the cache cannot build (the cache's own message); Unsupported on a graph with a cross-attention layer. One cache, one graph run at a time: the caller serializes run calls on the graph (the cache's prepare_step already requires one driving thread), and attach_kv_cache refuses InvalidArgument while a run is in flight.

Declared in ClikaRT/graph/model_graph.h, line 175

attach_kv_cache(shared_ptr<nn::KVCache>)​

void attach_kv_cache(std::shared_ptr<nn::KVCache> cache)

Build ONE nn::KVCache on the graph's device from kv_cache_info() and bind it to every self-attention layer. Fields of config left at 0 are taken from the graph: num_layers (the self-attention layer count), num_kv_heads and head_dim (each layer's own values fill its row, so a stack with mixed geometry binds per layer); a layer's sliding window becomes its row's window; kv_dtype is taken from the graph's KV planes (the config's value is not consulted). mode, paged, max_seqs, max_tokens_per_seq and the quantization fields come from config as given. After this call the past inputs leave input_names(); the present outputs stay in output_names() and run(inputs, step) returns each layer's cache plane at its present position, no copy. Every self-attention layer reports KVCacheLayerInfo::bound_to_cache, and the graph runs through run(inputs, const nn::KVCache::StepIndices&).

The present K/V of a bound layer ARE the cache planes: read them through kv_cache()->keys(l) / values(l) (a continuous cache's [max_seqs, kv_heads, max_tokens_per_seq, head_dim] plane; a paged cache's block pool, addressed through paged_block_table()). A graph with an interior consumer of a present output (a graph node reading it) refuses to attach (Unsupported): the in-place append needs the plane bound at the output port.

Refuses InvalidArgument on a graph with no KV layers, one already attached, or a config the cache cannot build (the cache's own message); Unsupported on a graph with a cross-attention layer. One cache, one graph run at a time: the caller serializes run calls on the graph (the cache's prepare_step already requires one driving thread), and attach_kv_cache refuses InvalidArgument while a run is in flight.

Declared in ClikaRT/graph/model_graph.h, line 190

kv_cache()​

std::shared_ptr<nn::KVCache> kv_cache() const

The attached cache; null when none is attached. Owned or injected, it is the one nn::KVCache the serving loop drives.

Declared in ClikaRT/graph/model_graph.h, line 196

run(vector<Tensor>, nn::KVCache::StepIndices)​

std::vector<Tensor> run(std::vector<Tensor> inputs, const nn::KVCache::StepIndices& step) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 210

run(NamedTensors, nn::KVCache::StepIndices)​

std::vector<Tensor> run(const NamedTensors& inputs, const nn::KVCache::StepIndices& step) const

Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.

Declared in ClikaRT/graph/model_graph.h, line 217

kv_cache_info()​

std::vector<KVCacheLayerInfo> kv_cache_info() const

Per-attention-layer KV-cache descriptors, layer order. Empty = the graph carries no KV cache. After attach_kv_cache, every self-attention layer reports bound_to_cache with empty past_*_input (they left the input list; the present names stay), and generate_kv_input_specs contributes nothing for those layers.

Declared in ClikaRT/graph/model_graph.h, line 228

nodes()​

std::vector<Node> nodes() const

Every node in topological order: each node after the producers of its inputs, the Input nodes first. Fails when the graph holds a cycle (the error's code_name() is INVALID_GRAPH).

Declared in ClikaRT/graph/model_graph.h, line 243

node()​

std::optional<Node> node(std::string_view name) const

The node named name, or std::nullopt.

Declared in ClikaRT/graph/model_graph.h, line 246

find_nodes()​

std::vector<Node> find_nodes(OpCode code) const

Every node whose op_code() is code, in name order.

Declared in ClikaRT/graph/model_graph.h, line 249

find_pattern()​

std::vector<Match> find_pattern(const Pattern& pattern) const

Every occurrence of pattern (pattern.h): one Match per occurrence, the matched node per pattern node in the pattern's add_node order, occurrences never overlapping. Fails when a pattern edge names a node that was never added (Status::InvalidArgument) and with a predicate's own failure when one throws.

Declared in ClikaRT/graph/model_graph.h, line 257

constants()​

std::vector<Value> constants() const

The graph's edge constants: every tensor bound onto an edge with bytes fixed for the graph's lifetime (an ONNX initializer, a constant a trace captured), one Value each, in the graph's own order; Value::constant() reads the bytes. A compiled ONNX initializer is also its operator's bound tensor (Node::bound_tensors(), the same bytes); a traced module's weights are bound into the operators alone and are not edge constants.

Declared in ClikaRT/graph/model_graph.h, line 268

is_finalized()​

bool is_finalized() const

True once the operators are finalized for serving: weights packed into their kernel layouts and constants absorbed, the default finish of compile and trace. A finalized graph runs; graph rewrites belong before this point.

Declared in ClikaRT/graph/model_graph.h, line 274

visualize()​

void visualize(std::string_view path) const

Write the compiled graph to path as an ONNX model for viewing in Netron or any ONNX graph viewer. The file holds the executable topology as compile left it: the runtime's fused and absorbed operators under their own op types, every edge as a tensor wire carrying its shape and dtype, and the graph's inputs and outputs under their names. Structure only. No weights are written, so the file is a picture of the graph, not a runnable model. path is created or overwritten.

Declared in ClikaRT/graph/model_graph.h, line 286

shape_domains()​

std::vector<ShapeDomain> shape_domains() const

The shapes each input admits, input_names() order. A dimension with a fixed size is Fixed; a dynamic dimension compiled with a recorded profile (an ONNX compile given InputSpec ranges or named-dim bindings) is Range; a dynamic dimension without one is Free. See ShapeDomain (bench.h).

Declared in ClikaRT/graph/model_graph.h, line 297

input_shape_domain()​

ShapeDomain input_shape_domain(std::string_view name) const

The domain of the input named name. Raises ClikaRT::Error (NotFound) when no input carries the name.

Declared in ClikaRT/graph/model_graph.h, line 304

validate_shapes(vector<InputShape>)​

void validate_shapes(const std::vector<InputShape>& shapes) const

Check named shapes against the domains: every runtime input present once, every name an input, every rank equal to the input's, and every extent admitted (DimDomain::admits). Raises ClikaRT::Error (InvalidArgument) naming the input and the dimension that failed.

Declared in ClikaRT/graph/model_graph.h, line 313

validate_shapes(vector<vector<int64_t>>)​

void validate_shapes(const std::vector<std::vector<std::int64_t>>& shapes) const

Check named shapes against the domains: every runtime input present once, every name an input, every rank equal to the input's, and every extent admitted (DimDomain::admits). Raises ClikaRT::Error (InvalidArgument) naming the input and the dimension that failed.

Declared in ClikaRT/graph/model_graph.h, line 320

bench(vector<InputShape>, BenchOptions)​

BenchReport bench(const std::vector<InputShape>& shapes, BenchOptions options = {}) const

Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.

Declared in ClikaRT/graph/model_graph.h, line 334

bench(vector<vector<int64_t>>, BenchOptions)​

BenchReport bench(
    const std::vector<std::vector<std::int64_t>>& shapes,
    BenchOptions options = {}
) const

Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.

Declared in ClikaRT/graph/model_graph.h, line 343

bench(NamedTensors, BenchOptions)​

BenchReport bench(const NamedTensors& inputs, BenchOptions options = {}) const

Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.

Declared in ClikaRT/graph/model_graph.h, line 353