ClikaRT::graph::ModelGraph
class
Header: ClikaRT/graph/model_graph.h
An executable model graph. Move-only.
Member functions
ModelGraph(ModelGraph)
ModelGraph(ModelGraph&&) noexcept
Move-only: transfers the graph.
Declared in ClikaRT/graph/model_graph.h, line 80
operator=(ModelGraph)
ModelGraph& operator=(ModelGraph&&) noexcept
Move-assign: transfers the graph.
Declared in ClikaRT/graph/model_graph.h, line 82
ModelGraph(ModelGraph)
ModelGraph(const ModelGraph&) =delete
Declared in ClikaRT/graph/model_graph.h, line 83
operator=(ModelGraph)
ModelGraph& operator=(const ModelGraph&) =delete
Declared in ClikaRT/graph/model_graph.h, line 84
~ModelGraph()
~ModelGraph()
Releases the graph.
Declared in ClikaRT/graph/model_graph.h, line 86
input_names()
std::vector<std::string> input_names() const
Runtime input names, in the model's declaration order, the order the positional run binds. Weights are not inputs; they ride the graph.
Declared in ClikaRT/graph/model_graph.h, line 91
output_names()
std::vector<std::string> output_names() const
Output names, declaration order, the order every run returns.
Declared in ClikaRT/graph/model_graph.h, line 93
inputs()
std::vector<spec::TensorSpec> inputs() const
Input / output slots as (name, dtype, dims) specs, same order as input_names() / output_names(). A dimension the graph left dynamic reads as spec::TensorSpec::kDynamicDim (-1).
Declared in ClikaRT/graph/model_graph.h, line 98
outputs()
std::vector<spec::TensorSpec> outputs() const
The produced edges as specs, the output twin of inputs().
Declared in ClikaRT/graph/model_graph.h, line 100
run(vector<Tensor>)
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 111
run(NamedTensors)
std::vector<Tensor> run(const NamedTensors& inputs) const
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 118
run(vector<Tensor>, vector<Tensor>)
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 128
run(NamedTensors, NamedTensors)
std::vector<Tensor> run(const NamedTensors& inputs, const NamedTensors& bound_outputs) const
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 139
attach_kv_cache(nn::KVCacheConfig)
void attach_kv_cache(const nn::KVCacheConfig& config)
Build ONE nn::KVCache on the graph's device from kv_cache_info() and bind it to every self-attention layer. Fields of config left at 0 are taken from the graph: num_layers (the self-attention layer count), num_kv_heads and head_dim (each layer's own values fill its row, so a stack with mixed geometry binds per layer); a layer's sliding window becomes its row's window; kv_dtype is taken from the graph's KV planes (the config's value is not consulted). mode, paged, max_seqs, max_tokens_per_seq and the quantization fields come from config as given. After this call the past inputs leave input_names(); the present outputs stay in output_names() and run(inputs, step) returns each layer's cache plane at its present position, no copy. Every self-attention layer reports KVCacheLayerInfo::bound_to_cache, and the graph runs through run(inputs, const nn::KVCache::StepIndices&).
The present K/V of a bound layer ARE the cache planes: read them through kv_cache()->keys(l) / values(l) (a continuous cache's [max_seqs, kv_heads, max_tokens_per_seq, head_dim] plane; a paged cache's block pool, addressed through paged_block_table()). A graph with an interior consumer of a present output (a graph node reading it) refuses to attach (Unsupported): the in-place append needs the plane bound at the output port.
Refuses InvalidArgument on a graph with no KV layers, one already attached, or a config the cache cannot build (the cache's own message); Unsupported on a graph with a cross-attention layer. One cache, one graph run at a time: the caller serializes run calls on the graph (the cache's prepare_step already requires one driving thread), and attach_kv_cache refuses InvalidArgument while a run is in flight.
Declared in ClikaRT/graph/model_graph.h, line 175
attach_kv_cache(shared_ptr<nn::KVCache>)
void attach_kv_cache(std::shared_ptr<nn::KVCache> cache)
Build ONE nn::KVCache on the graph's device from kv_cache_info() and bind it to every self-attention layer. Fields of config left at 0 are taken from the graph: num_layers (the self-attention layer count), num_kv_heads and head_dim (each layer's own values fill its row, so a stack with mixed geometry binds per layer); a layer's sliding window becomes its row's window; kv_dtype is taken from the graph's KV planes (the config's value is not consulted). mode, paged, max_seqs, max_tokens_per_seq and the quantization fields come from config as given. After this call the past inputs leave input_names(); the present outputs stay in output_names() and run(inputs, step) returns each layer's cache plane at its present position, no copy. Every self-attention layer reports KVCacheLayerInfo::bound_to_cache, and the graph runs through run(inputs, const nn::KVCache::StepIndices&).
The present K/V of a bound layer ARE the cache planes: read them through kv_cache()->keys(l) / values(l) (a continuous cache's [max_seqs, kv_heads, max_tokens_per_seq, head_dim] plane; a paged cache's block pool, addressed through paged_block_table()). A graph with an interior consumer of a present output (a graph node reading it) refuses to attach (Unsupported): the in-place append needs the plane bound at the output port.
Refuses InvalidArgument on a graph with no KV layers, one already attached, or a config the cache cannot build (the cache's own message); Unsupported on a graph with a cross-attention layer. One cache, one graph run at a time: the caller serializes run calls on the graph (the cache's prepare_step already requires one driving thread), and attach_kv_cache refuses InvalidArgument while a run is in flight.
Declared in ClikaRT/graph/model_graph.h, line 190
kv_cache()
std::shared_ptr<nn::KVCache> kv_cache() const
The attached cache; null when none is attached. Owned or injected, it is the one nn::KVCache the serving loop drives.
Declared in ClikaRT/graph/model_graph.h, line 196
run(vector<Tensor>, nn::KVCache::StepIndices)
std::vector<Tensor> run(std::vector<Tensor> inputs, const nn::KVCache::StepIndices& step) const
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 210
run(NamedTensors, nn::KVCache::StepIndices)
std::vector<Tensor> run(const NamedTensors& inputs, const nn::KVCache::StepIndices& step) const
Run with positional inputs (one per input_names() entry, that order). Returns the outputs in output_names() order.
Declared in ClikaRT/graph/model_graph.h, line 217
kv_cache_info()
std::vector<KVCacheLayerInfo> kv_cache_info() const
Per-attention-layer KV-cache descriptors, layer order. Empty = the graph carries no KV cache. After attach_kv_cache, every self-attention layer reports bound_to_cache with empty past_*_input (they left the input list; the present names stay), and generate_kv_input_specs contributes nothing for those layers.
Declared in ClikaRT/graph/model_graph.h, line 228
nodes()
std::vector<Node> nodes() const
Every node in topological order: each node after the producers of its inputs, the Input nodes first. Fails when the graph holds a cycle (the error's code_name() is INVALID_GRAPH).
Declared in ClikaRT/graph/model_graph.h, line 243
node()
std::optional<Node> node(std::string_view name) const
The node named name, or std::nullopt.
Declared in ClikaRT/graph/model_graph.h, line 246
find_nodes()
Every node whose op_code() is code, in name order.
Declared in ClikaRT/graph/model_graph.h, line 249
find_pattern()
Every occurrence of pattern (pattern.h): one Match per occurrence, the matched node per pattern node in the pattern's add_node order, occurrences never overlapping. Fails when a pattern edge names a node that was never added (Status::InvalidArgument) and with a predicate's own failure when one throws.
Declared in ClikaRT/graph/model_graph.h, line 257
constants()
std::vector<Value> constants() const
The graph's edge constants: every tensor bound onto an edge with bytes fixed for the graph's lifetime (an ONNX initializer, a constant a trace captured), one Value each, in the graph's own order; Value::constant() reads the bytes. A compiled ONNX initializer is also its operator's bound tensor (Node::bound_tensors(), the same bytes); a traced module's weights are bound into the operators alone and are not edge constants.
Declared in ClikaRT/graph/model_graph.h, line 268
is_finalized()
bool is_finalized() const
True once the operators are finalized for serving: weights packed into their kernel layouts and constants absorbed, the default finish of compile and trace. A finalized graph runs; graph rewrites belong before this point.
Declared in ClikaRT/graph/model_graph.h, line 274
visualize()
void visualize(std::string_view path) const
Write the compiled graph to path as an ONNX model for viewing in Netron or any ONNX graph viewer. The file holds the executable topology as compile left it: the runtime's fused and absorbed operators under their own op types, every edge as a tensor wire carrying its shape and dtype, and the graph's inputs and outputs under their names. Structure only. No weights are written, so the file is a picture of the graph, not a runnable model. path is created or overwritten.
Declared in ClikaRT/graph/model_graph.h, line 286
shape_domains()
std::vector<ShapeDomain> shape_domains() const
The shapes each input admits, input_names() order. A dimension with a fixed size is Fixed; a dynamic dimension compiled with a recorded profile (an ONNX compile given InputSpec ranges or named-dim bindings) is Range; a dynamic dimension without one is Free. See ShapeDomain (bench.h).
Declared in ClikaRT/graph/model_graph.h, line 297
input_shape_domain()
ShapeDomain input_shape_domain(std::string_view name) const
The domain of the input named name. Raises ClikaRT::Error (NotFound) when no input carries the name.
Declared in ClikaRT/graph/model_graph.h, line 304
validate_shapes(vector<InputShape>)
void validate_shapes(const std::vector<InputShape>& shapes) const
Check named shapes against the domains: every runtime input present once, every name an input, every rank equal to the input's, and every extent admitted (DimDomain::admits). Raises ClikaRT::Error (InvalidArgument) naming the input and the dimension that failed.
Declared in ClikaRT/graph/model_graph.h, line 313
validate_shapes(vector<vector<int64_t>>)
void validate_shapes(const std::vector<std::vector<std::int64_t>>& shapes) const
Check named shapes against the domains: every runtime input present once, every name an input, every rank equal to the input's, and every extent admitted (DimDomain::admits). Raises ClikaRT::Error (InvalidArgument) naming the input and the dimension that failed.
Declared in ClikaRT/graph/model_graph.h, line 320
bench(vector<InputShape>, BenchOptions)
BenchReport bench(const std::vector<InputShape>& shapes, BenchOptions options = {}) const
Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.
Declared in ClikaRT/graph/model_graph.h, line 334
bench(vector<vector<int64_t>>, BenchOptions)
BenchReport bench(
const std::vector<std::vector<std::int64_t>>& shapes,
BenchOptions options = {}
) const
Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.
Declared in ClikaRT/graph/model_graph.h, line 343
bench(NamedTensors, BenchOptions)
BenchReport bench(const NamedTensors& inputs, BenchOptions options = {}) const
Time the graph on synthesized inputs of the given shapes. The shapes are validated as validate_shapes does, one tensor per input is filled per options.fill on the benchmark device, options.warmup untimed runs precede options.iterations timed ones, and the report carries every timed run's wall-clock milliseconds (each output synchronized before the clock stops), the summary statistics, the process's peak resident set and the device's memory counters. Raises ClikaRT::Error on a shape the domains reject or a run that fails.
Declared in ClikaRT/graph/model_graph.h, line 353