Skip to main content

ClikaRT::device

namespace

Classes​

NameDescription
DevicePropertiesA device's static capabilities. A plain value type; copy it freely.
MemoryStats

Enumerations​

enum ComputeAPI​

enum class ComputeAPI : std::int32_t

The compute backend a device dispatches through. Names the runtime/SDK, not the physical hardware (one GPU may be reachable via more than one).

int32-backed for a stable ABI. ABI note: new APIs append at the end; existing ones never reorder or drop.

EnumeratorValueDescription
CPU0Host CPU (SIMD kernels).
CUDANVIDIA CUDA.
VulkanCross-vendor GPU compute.
MetalApple Metal.
HTPQualcomm Hexagon DSP/NPU (HTP).
TPUGoogle Cloud TPU (PJRT).

Declared in ClikaRT/compute/device.h, line 20

enum DeviceKind​

enum class DeviceKind : std::int32_t

Where a device sits relative to the host: what an application reads to decide how much to place on it and whether it shares the host's memory.

int32-backed for a stable ABI. ABI note: new kinds append at the end; existing ones never reorder or drop.

EnumeratorValueDescription
Unknown0The backend reports no kind.
CpuThe host processor itself: the CPU backend, or a software GPU driver.
IntegratedGpuA GPU sharing the host's memory (phones, Jetson boards, Apple silicon).
DiscreteGpuA GPU with its own memory behind a bus (desktop and data-center cards).
VirtualGpuA GPU reached through a virtualization layer.
AcceleratorAn NPU or TPU class device.

Declared in ClikaRT/compute/device_properties.h, line 24

enum MemorySource​

enum class MemorySource : std::uint8_t

A point-in-time snapshot. Physical numbers come from the driver (GPU devices) or the OS's memory accounting (the CPU device: reclaimable+free as "free", MemAvailable on Linux). A device whose driver reports free memory as a count of free system pages (an integrated CUDA part) reports the larger of that count and the OS's MemAvailable, since the page count leaves out the reclaimable cache; a device whose driver already reports a budget (a unified-memory GPU's working-set headroom) reports that budget as is. Pool numbers come from the runtime's caching allocator for that device. Where a platform has no physical source the device_* fields are 0 ("unknown"); the pool counters still report. How the physical figures of a MemoryStats were obtained: read from the device's driver or the OS's accounting (Measured), derived from a declared budget or a clamped read (Heuristic), or not available (Unknown: the device_* fields are then the honest 0, or a figure whose source the device did not declare). A decision that must not misfire (refusing to load a model for lack of memory) keys on Measured alone; the other two carry at most a warning.

EnumeratorDescription
Unknown
Heuristic
Measured

Declared in ClikaRT/compute/memory_stats.h, line 38

Variables​

kWholeMachineCpuIndex​

std::int32_t kWholeMachineCpuIndex = -1

The whole-machine CPU: every core, over memory the OS places on first touch. CPU indexing is uniform around it: a NEGATIVE index is the whole machine, a NON-NEGATIVE index is that CPU socket (its own cores, its own local memory). On a single-socket machine the two describe the same hardware.

Declared in ClikaRT/compute/device.h, line 46

Functions​

is_backend_compiled()​

bool is_backend_compiled(ComputeAPI backend) noexcept

Was backend compiled into this build of ClikaRT? A pure build-time fact; it does not attempt to load anything. CPU is always compiled in.

Declared in ClikaRT/compute/backends.h, line 25

is_backend_available()​

bool is_backend_available(ComputeAPI backend)

Does backend actually load and initialize on this machine? Returns false if the backend was not compiled in, its shared library is absent, or it fails to bring up (e.g. no driver). Brings the backend up on first call, then caches the result; CPU is always available.

Declared in ClikaRT/compute/backends.h, line 31

is_backend_loaded()​

bool is_backend_loaded(ComputeAPI backend) noexcept

Is backend loaded and initialized RIGHT NOW, without bringing it up? A plain read of the runtime's registry, which publishes a backend only once it has initialized: false for a backend nothing has used yet, one that failed to come up, and one this build does not carry. The probe for code that must never start a backend as a side effect (a memory-pressure handler, a status line) where is_backend_available would load it first.

Declared in ClikaRT/compute/backends.h, line 39

is_cpu_available()​

bool is_cpu_available()

Convenience predicates equivalent to is_backend_available(<api>).

Declared in ClikaRT/compute/backends.h, line 42

is_cuda_available()​

bool is_cuda_available()

Declared in ClikaRT/compute/backends.h, line 43

is_vulkan_available()​

bool is_vulkan_available()

Declared in ClikaRT/compute/backends.h, line 44

is_metal_available()​

bool is_metal_available()

Declared in ClikaRT/compute/backends.h, line 45

is_tpu_available()​

bool is_tpu_available()

Declared in ClikaRT/compute/backends.h, line 46

backend_load_status()​

void backend_load_status(ComputeAPI backend)

Declared in ClikaRT/compute/backends.h, line 59

backend_device_status()​

void backend_device_status(ComputeAPI backend)

Declared in ClikaRT/compute/backends.h, line 72

device_count()​

std::int32_t device_count(ComputeAPI backend)

How many devices backend exposes on this machine (e.g. the number of CUDA GPUs, or Hexagon DSPs). 0 if the backend is unavailable. For CPU this is the number of exposed host compute domains (one per socket). Brings the backend up on first call.

Declared in ClikaRT/compute/backends.h, line 78

enumerate_devices()​

std::vector<Device> enumerate_devices(ComputeAPI backend)

The concrete Device handles backend exposes on this machine, in index order ({api, 0}, {api, 1}, …). Empty if the backend is unavailable; an empty list is a valid answer, not an error. Pass each Device to get_device_properties for its capabilities, or to Tensor::to to place work on it.

Declared in ClikaRT/compute/backends.h, line 85

compute_api_name()​

const char* compute_api_name(ComputeAPI api) noexcept

A short, human-readable backend name (e.g. "CUDA"). Never null.

Declared in ClikaRT/compute/device.h, line 30

parse_device()​

Device parse_device(std::string_view spec)

Declared in ClikaRT/compute/device.h, line 138

manual_seed()​

void manual_seed(Device device, std::uint64_t seed)

Seed device's random-number generator and reset its draw counter, making subsequent random ops (e.g. ops::multinomial) reproducible on that device. Seeding is device-scoped; every stream on the device shares one sequence. Do not call concurrently with in-flight random ops on the same device.

Declared in ClikaRT/compute/device.h, line 147

current_seed()​

std::uint64_t current_seed(Device device)

The current seed of device's generator. A device never explicitly seeded reports the default seed (0); results are reproducible out of the box.

Declared in ClikaRT/compute/device.h, line 151

synchronize()​

void synchronize(Device device)

Wait for every operation queued on device's lanes: the calling thread's current lane there, every Stream::create(device) lane, and the device's default stream. On return the device's queued work is complete and each of its lanes reports Stream::query_idle(). Only device's lanes are waited on: work queued on another device runs on, and a completion callback (Tensor::on_complete) posted by the drained work may still be running. A device with nothing queued returns at once.

Declared in ClikaRT/compute/device.h, line 161

get_device_properties()​

DeviceProperties get_device_properties(Device device)

Declared in ClikaRT/compute/device_properties.h, line 97

memory_stats()​

MemoryStats memory_stats(Device device)

Snapshot the memory state of device. Fails only when the device itself is invalid or its backend is unavailable.

Declared in ClikaRT/compute/memory_stats.h, line 92

release_cached_memory()​

void release_cached_memory(Device device)

Return the runtime pool's idle memory for device to the driver or the OS. The pool keeps memory it handed out and got back (cached_bytes) so the next allocation is cheap; after a model is unloaded, or before a second model must fit beside the first, that idle reserve is memory the device (or, on a shared-memory part, the host) cannot use for anything else until it is released. Memory still in use, and memory whose last use has not finished on the device, is never touched: the call releases what is provably idle at the moment it runs and leaves the rest cached, so a later call can release more once that work has completed. Fails only when the device itself is invalid or its backend is unavailable.

Declared in ClikaRT/compute/memory_stats.h, line 106

reset_peak_memory_stats()​

void reset_peak_memory_stats(Device device)

Restart the high-water marks of the runtime pool for device from its present occupancy. MemoryStats::peak_active_bytes and peak_reserved_bytes only ever rise, so after a load or a warm-up they record the largest footprint the process has held so far and say nothing about the step that follows. This call sets the active mark to the pool's present active_bytes (and the reserved mark to reserved_bytes where the pool keeps one apart from the active mark); both track upward from there. The measuring idiom: reset, read active_bytes, run one step, snapshot again; peak_active_bytes less the active_bytes read before the step is the largest transient the step needed on top of what was already resident, even after the step's temporaries are gone. Nothing is allocated, freed, or waited for: the marks move, the memory does not. Fails only when the device itself is invalid or its backend is unavailable.

Declared in ClikaRT/compute/memory_stats.h, line 124

ClikaRT/compute/backends.h​

#include <ClikaRT/compute/backends.h>

Backend + device discovery: which compute backends this build was compiled with, which actually load on this machine, how many devices each exposes, and the concrete Device handles to address them.

get_device_properties (device_properties.h) describes ONE device; these functions let you find out which devices exist in the first place; the starting point for a program that adapts to whatever hardware it runs on (e.g. shard work across every available GPU).