Skip to main content

BenchReport

One benchmark's measurement. Times are wall-clock milliseconds around each run, with every output synchronized before the clock stops, so a run's time includes the device work its outputs wait on. peak_rss_bytes is the whole process's peak resident set at the end of the benchmark, not the graph's own footprint. device_memory is the runtime's memory-pool snapshot for the device after the runs; device_memory_valid is False where none was available. to_csv() and to_json() render the report under the table's column names; to_dict() is that JSON as a Python dict.

batch_axis (property)​

The axis items were counted along, or -1.

device (property)​

Where the inputs were placed.

device_memory (property)​

The runtime's memory-pool counters for the device (a MemoryStats).

device_memory_valid (property)​

device_memory holds a snapshot.

fill (property)​

How synthesized inputs were filled; for caller-provided tensors the option's value, not applied.

first_call_ms (property)​

The first run of the call, warmup or measured.

items_per_iteration (property)​

The first input's extent on batch_axis, else 1.

iteration_ms (property)​

Every measured run's milliseconds, in order.

iterations (property)​

Measured runs.

max_ms (property)​

The slowest measured run.

mean_ms (property)​

The arithmetic mean of iteration_ms.

min_ms (property)​

The fastest measured run.

outputs (property)​

The output shapes produced, output_names() order.

outputs_on_device (property)​

The outputs live on the device (False: on the host).

outputs_reused (property)​

The measured runs wrote preallocated outputs.

p50_ms (property)​

The median of iteration_ms.

p95_ms (property)​

The 95th percentile of iteration_ms (linear interpolation).

p99_ms (property)​

The 99th percentile of iteration_ms (linear interpolation).

peak_rss_bytes (property)​

The process's peak resident set, in bytes.

shapes (property)​

The input shapes fed, input_names() order.

stddev_ms (property)​

The population standard deviation of iteration_ms.

threads (property)​

The runtime's CPU thread count during the runs.

throughput_items_per_s (property)​

items_per_iteration / (p50_ms / 1000).

warmup (property)​

Untimed runs.

__init__​

__init__(self, /, *args, **kwargs)

Initialize self. See help(type(self)) for accurate signature.

iterations_csv​

iterations_csviterations_csv(self) -> str

iterations_csv(self) -> str

iterations_csv() -> str

One CSV row per measured run: columns iteration (0-based) and ms.

to_csv​

to_csvto_csv(self) -> str

to_csv(self) -> str

to_csv() -> str

The report as one CSV row under a header: the shapes and outputs as name:d0xd1x... entries joined with ';', the device as API:index, the fill as zeros / uniform, then the timing, throughput and memory columns. The same columns for every report, so a sweep's rows append.

to_dict​

to_dictto_dict(self) -> object

to_dict(self) -> object

to_dict() -> dict

The to_json() object as a Python dict.

to_json​

to_jsonto_json(self) -> str

to_json(self) -> str

to_json() -> str

Every field as a JSON object under the table's column names: the shapes and outputs as arrays of {name, dims}, iteration_ms as an array, the device as API:index, the memory counters as the nested device_memory object.