BenchReport
One benchmark's measurement. Times are wall-clock milliseconds around each run, with every output synchronized before the clock stops, so a run's time includes the device work its outputs wait on. peak_rss_bytes is the whole process's peak resident set at the end of the benchmark, not the graph's own footprint. device_memory is the runtime's memory-pool snapshot for the device after the runs; device_memory_valid is False where none was available. to_csv() and to_json() render the report under the table's column names; to_dict() is that JSON as a Python dict.
batch_axis (property)
The axis items were counted along, or -1.
device (property)
Where the inputs were placed.
device_memory (property)
The runtime's memory-pool counters for the device (a MemoryStats).
device_memory_valid (property)
device_memory holds a snapshot.
fill (property)
How synthesized inputs were filled; for caller-provided tensors the option's value, not applied.
first_call_ms (property)
The first run of the call, warmup or measured.
items_per_iteration (property)
The first input's extent on batch_axis, else 1.
iteration_ms (property)
Every measured run's milliseconds, in order.
iterations (property)
Measured runs.
max_ms (property)
The slowest measured run.
mean_ms (property)
The arithmetic mean of iteration_ms.
min_ms (property)
The fastest measured run.
outputs (property)
The output shapes produced, output_names() order.
outputs_on_device (property)
The outputs live on the device (False: on the host).
outputs_reused (property)
The measured runs wrote preallocated outputs.
p50_ms (property)
The median of iteration_ms.
p95_ms (property)
The 95th percentile of iteration_ms (linear interpolation).
p99_ms (property)
The 99th percentile of iteration_ms (linear interpolation).
peak_rss_bytes (property)
The process's peak resident set, in bytes.
shapes (property)
The input shapes fed, input_names() order.
stddev_ms (property)
The population standard deviation of iteration_ms.
threads (property)
The runtime's CPU thread count during the runs.
throughput_items_per_s (property)
items_per_iteration / (p50_ms / 1000).
warmup (property)
Untimed runs.
__init__
__init__(self, /, *args, **kwargs)
Initialize self. See help(type(self)) for accurate signature.
iterations_csv
iterations_csviterations_csv(self) -> str
iterations_csv(self) -> str
iterations_csv() -> str
One CSV row per measured run: columns iteration (0-based) and ms.
to_csv
to_csvto_csv(self) -> str
to_csv(self) -> str
to_csv() -> str
The report as one CSV row under a header: the shapes and outputs as name:d0xd1x... entries joined with ';', the device as API:index, the fill as zeros / uniform, then the timing, throughput and memory columns. The same columns for every report, so a sweep's rows append.
to_dict
to_dictto_dict(self) -> object
to_dict(self) -> object
to_dict() -> dict
The to_json() object as a Python dict.
to_json
to_jsonto_json(self) -> str
to_json(self) -> str
to_json() -> str
Every field as a JSON object under the table's column names: the shapes and outputs as arrays of {name, dims}, iteration_ms as an array, the device as API:index, the memory counters as the nested device_memory object.