Benchmarks
A benchmark measures how a model performs on real hardware: throughput, latency, memory, accuracy. You point the platform at one or more models and one or more devices, and it creates a benchmark group: a parent record with one leg per model and device pair. Each leg is a job that runs on its target and reports back.
Two command groups cover this:
benchmarksis a small hand-written workflow: launch, thenwatchorresults. It is what you use in a script.benchmark-groupsis the generated resource group, with the full create, read, update and delete surface plus sharing.
The remaining benchmark-* groups are read-only catalogs describing what the platform can measure.
The workflow
clika-cli apply -f benchmark.yaml
clika-cli benchmarks watch nightly-llm-sweep
clika-cli benchmarks results nightly-llm-sweep -o json
Launching through apply keeps the definition of the run in a file you can commit. You can equally create the group directly with benchmark-groups create --body '<json>'; the two reach the same endpoint.
benchmarks watch
clika-cli benchmarks watch <group-name-or-id> [flags]
Polls the group's legs, both child jobs on registered devices and hosted-device legs, prints a line to standard error whenever a leg changes state, and then renders the final table on standard output in whatever -o format you chose.
It exits non-zero when any leg ends in a state other than completed, which is what makes it usable as a CI gate.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--interval | duration | 5s | How often to poll. Accepts Go duration syntax, for example 30s or 2m. |
$ clika-cli benchmarks watch nightly-llm-sweep
jetson-01 / Qwen2.5-0.5B-Instruct: queued -> running
orin-02 / Qwen2.5-0.5B-Instruct: queued -> running
jetson-01 / Qwen2.5-0.5B-Instruct: running -> completed
orin-02 / Qwen2.5-0.5B-Instruct: running -> completed
DEVICE MODEL STATUS DURATION
jetson-01 Qwen2.5-0.5B-Instruct completed 4m12s
orin-02 Qwen2.5-0.5B-Instruct completed 3m48s
Gate a pipeline on it:
clika-cli apply -f benchmark.yaml && clika-cli benchmarks watch nightly-llm-sweep
Poll less often on a long run:
clika-cli benchmarks watch nightly-llm-sweep --interval 30s
benchmarks results
clika-cli benchmarks results <group-name-or-id>
Fetches the same legs together with their measured metrics, at any time, whether or not the run has finished. Unlike watch, it does not block and does not set a failure exit code.
$ clika-cli benchmarks results nightly-llm-sweep
DEVICE MODEL STATUS TOKENS/S TTFT PEAK MEM
jetson-01 Qwen2.5-0.5B-Instruct completed 142.6 210ms 1.8 GB
orin-02 Qwen2.5-0.5B-Instruct completed 118.3 265ms 1.8 GB
clika-cli benchmarks results nightly-llm-sweep -o json | jq '.[] | {device, tokens_per_second}'
Creating a group directly
clika-cli benchmark-groups create --body '<json>'
The body says what to run, on what, and where. Two constraints are enforced by the platform: at least one of device_ids or hosted_device_arns must be set, and exactly one of model_ids or huggingface_url.
| Body field | Type | Meaning |
|---|---|---|
name | string | Display name. Generated for you when omitted, but naming it is what lets you say benchmarks watch nightly-llm-sweep later. |
model_ids | array of strings | One or more registered model ids. Mutually exclusive with huggingface_url. |
huggingface_url | string | A convenience that registers and benchmarks a single Hugging Face model in one step. Mutually exclusive with model_ids. |
device_ids | array of strings | Registered devices to run on. Each must be online and benchmark capable. |
hosted_device_arns | array of strings | Rented phones from hosted-devices list instead of, or alongside, your own devices. 1 to 5 identifiers. One hosted run is created per selected model and linked to the group; its results appear marked execution: "hosted". |
hosted_backend | string | Which hosted cloud the identifiers belong to, device_farm or test_lab. Only needed when the deployment has more than one configured. |
hosted_android_version | string | The Android version a hosted run targets. Required by a backend that schedules a device model plus an OS version; ignored by one whose identifier already encodes it. |
benchmark_type | string | Which benchmark to run. Required only when the platform cannot infer it, which happens when the model's task maps to more than one compatible type. |
job_definition_id | string | A power-user override that pins the job definition used. Normally omitted: the platform maps the benchmark type and the device to the right definition through its compatibility matrix. Applies to registered devices only. |
force | boolean | Skip the device resource compatibility check. Use it when you know a device can take the model and the check disagrees. |
transmit_detailed_outputs | boolean | Whether each leg sends back the per-sample input, model output and expected answer that power the detailed-outputs view. Defaults to true; set it to false to keep results lightweight. |
clika-cli benchmark-groups create --body '{
"name": "nightly-llm-sweep",
"huggingface_url": "https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct",
"device_ids": ["1f0a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8"]
}'
The same run as a YAML file you can commit, applied with apply, which additionally lets you name devices and models instead of pasting UUIDs:
kind: Benchmark
name: nightly-llm-sweep
benchmark_type: performance
models:
- Qwen2.5-0.5B-Instruct
devices:
- jetson-01
- orin-02
Reading and managing a group
| Command | What it does |
|---|---|
benchmark-groups list | Every group, newest first. Takes --page and --page_size. |
benchmark-groups get <name-or-id> | One group with its status summary. |
benchmark-groups jobs <name-or-id> | The group's legs as jobs, which is what benchmarks watch polls. |
benchmark-groups results <name-or-id> | The group's per-leg results, the same data benchmarks results renders. |
benchmark-groups update <name-or-id> --body '{"name":"..."}' | Rename the group. |
benchmark-groups delete <name-or-id> | Delete the group. |
Sharing results
A benchmark group can be published behind an opaque link so someone without an account can read it.
| Command | What it does |
|---|---|
benchmark-groups share-mark <name-or-id> | Create the share link. Only one live link exists per group. |
benchmark-groups share <name-or-id> | Show the current link. |
benchmark-groups share-card <name-or-id> | Upload the preview card image shown when the link is unfurled. |
benchmark-groups share-mark-share_id <name-or-id> <share_id> | Revoke the share link. |
public benchmark-shares <token> | The read side of a share link, which is what an anonymous reader hits. |
public benchmark-shares-card.png <token> | The card image behind that link. |
The benchmark catalogs
These read-only lists describe what this deployment can measure. They are worth checking before you build a request by hand, because they tell you the exact slugs the API expects.
| Command | What it lists |
|---|---|
benchmark-types list | The benchmark types available, by slug. This is where a benchmark_type value comes from. |
benchmark-categories list | How those types are grouped for display. |
benchmark-metrics list | Every metric the platform can report, with its unit and meaning. |
benchmark-io-schemas list | The input and output schemas benchmark results conform to. |
benchmark-task-categories list | The map from a model task to its benchmark category. |
benchmark-compatibility list --model_id <uuid> | Which benchmark types are compatible with one model, and on what hardware. Run this when a group creation is refused as incompatible. |
model-recommendations list | Models the platform suggests for a given target. |
huggingface-tasks list | The Hugging Face task names the platform understands, which is what maps a model to its compatible benchmark types. |
Comparing runs
clika-cli comparisons list
clika-cli comparisons create --body '<json>'
clika-cli comparisons get <id>
clika-cli comparisons delete <id>
A comparison is a saved side-by-side of several benchmark results, the thing you keep when you want to show that a change helped.
Rescoring
Some benchmarks are scored on the server after the run. If the scoring logic changed, or a scorer failed, you can score an existing job again without re-running it on the device:
clika-cli jobs rescore <name-or-id>
Full command reference
Every command below is generated from the deployment's own API description, so
one subcommand is exactly one platform operation. Each entry names the method,
the endpoint and the MCP tool name, so the same operation is
identifiable whichever surface you drive it from. Path parameters are positional
arguments, query parameters are flags, and a request body is --body or
--body-file. The hand-written commands, the ones that stream, propagate an
exit code, or hand your terminal to ssh, carry no operation line.
The prose above covers the commands most people reach for. This section is the complete surface, for when you need the flag you have not used before.
clika-cli benchmarks
benchmarks has 2 subcommands.
clika-cli benchmarks results
Fetches the group's legs, child jobs and cloud hosted device legs, with their measured metrics.
clika-cli benchmarks results <group-name-or-id> [flags]
Positional arguments: required <group-name-or-id>.
clika-cli benchmarks watch
Polls a benchmark group's legs, child jobs on registered devices and cloud hosted device legs, and prints a line whenever one changes state, then renders the final table. Exits non-zero when any leg ends in a state other than completed, so it gates CI.
clika-cli benchmarks watch <group-name-or-id> [flags]
Positional arguments: required <group-name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--interval | duration | 5s | poll interval |
clika-cli benchmark-groups
benchmark-groups has 13 subcommands.
clika-cli benchmark-groups cancel
Cancel benchmark group
clika-cli benchmark-groups cancel <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: POST /api/benchmark-groups/{id}/cancel. MCP tool name: post_benchmark_groups_id_cancel.
clika-cli benchmark-groups create
Create benchmark group
clika-cli benchmark-groups create [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--body | string | none | request body (inline JSON) |
--body-file | string | none | request body (path to a JSON file) |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: POST /api/benchmark-groups. MCP tool name: post_benchmark_groups.
clika-cli benchmark-groups delete
Delete benchmark group
clika-cli benchmark-groups delete <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--delete_outputs | string | none | Also delete the artifacts the child jobs' pipelines collected (each job's output file and execution log). Default false: the outputs stay in the artifact library, tagged with their job id, for the record. |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: DELETE /api/benchmark-groups/{id}. MCP tool name: delete_benchmark_groups_id.
clika-cli benchmark-groups get
Get benchmark group
clika-cli benchmark-groups get <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups/{id}. MCP tool name: get_benchmark_groups_id.
clika-cli benchmark-groups jobs
List benchmark group jobs
clika-cli benchmark-groups jobs <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups/{id}/jobs. MCP tool name: get_benchmark_groups_id_jobs.
clika-cli benchmark-groups list
List benchmark groups
clika-cli benchmark-groups list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--page | string | none | Page number (default: 1) |
--page_size | string | none | Items per page |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups. MCP tool name: get_benchmark_groups.
clika-cli benchmark-groups outputs
List benchmark group output artifacts
clika-cli benchmark-groups outputs <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups/{id}/outputs. MCP tool name: get_benchmark_groups_id_outputs.
clika-cli benchmark-groups results
List benchmark group results
clika-cli benchmark-groups results <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups/{id}/results. MCP tool name: get_benchmark_groups_id_results.
clika-cli benchmark-groups share
Get benchmark share link
clika-cli benchmark-groups share <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-groups/{id}/share. MCP tool name: get_benchmark_groups_id_share.
clika-cli benchmark-groups share-card
Upload benchmark share card
clika-cli benchmark-groups share-card <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: PUT /api/benchmark-groups/{id}/share/card. MCP tool name: put_benchmark_groups_id_share_card.
clika-cli benchmark-groups share-mark
Create benchmark share link
clika-cli benchmark-groups share-mark <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: POST /api/benchmark-groups/{id}/share. MCP tool name: post_benchmark_groups_id_share.
clika-cli benchmark-groups share-mark-share_id
Revoke benchmark share link
clika-cli benchmark-groups share-mark-share_id <name-or-id> <share_id> [flags]
Positional arguments: required <name-or-id>, <share_id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: DELETE /api/benchmark-groups/{id}/share/{share_id}. MCP tool name: delete_benchmark_groups_id_share_share_id.
clika-cli benchmark-groups update
Rename benchmark group
clika-cli benchmark-groups update <name-or-id> [flags]
Positional arguments: required <name-or-id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--body | string | none | request body (inline JSON) |
--body-file | string | none | request body (path to a JSON file) |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: PATCH /api/benchmark-groups/{id}. MCP tool name: patch_benchmark_groups_id.
clika-cli benchmark-categories
benchmark-categories has 1 subcommands.
clika-cli benchmark-categories list
List benchmark categories
clika-cli benchmark-categories list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-categories. MCP tool name: get_benchmark_categories.
clika-cli benchmark-compatibility
benchmark-compatibility has 1 subcommands.
clika-cli benchmark-compatibility list
Get benchmark compatibility for a model
clika-cli benchmark-compatibility list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--model_id | string | none | Model ID (UUID) |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-compatibility. MCP tool name: get_benchmark_compatibility.
clika-cli benchmark-io-schemas
benchmark-io-schemas has 1 subcommands.
clika-cli benchmark-io-schemas list
List benchmark I/O schemas
clika-cli benchmark-io-schemas list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-io-schemas. MCP tool name: get_benchmark_io_schemas.
clika-cli benchmark-metrics
benchmark-metrics has 1 subcommands.
clika-cli benchmark-metrics list
List benchmark metrics
clika-cli benchmark-metrics list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-metrics. MCP tool name: get_benchmark_metrics.
clika-cli benchmark-task-categories
benchmark-task-categories has 1 subcommands.
clika-cli benchmark-task-categories list
List benchmark task-category map
clika-cli benchmark-task-categories list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-task-categories. MCP tool name: get_benchmark_task_categories.
clika-cli benchmark-types
benchmark-types has 1 subcommands.
clika-cli benchmark-types list
List benchmark types
clika-cli benchmark-types list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/benchmark-types. MCP tool name: get_benchmark_types.
clika-cli comparisons
comparisons has 4 subcommands.
clika-cli comparisons create
Create job comparison
clika-cli comparisons create [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--body | string | none | request body (inline JSON) |
--body-file | string | none | request body (path to a JSON file) |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: POST /api/comparisons. MCP tool name: post_comparisons.
clika-cli comparisons delete
Delete comparison
clika-cli comparisons delete <id> [flags]
Positional arguments: required <id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: DELETE /api/comparisons/{id}. MCP tool name: delete_comparisons_id.
clika-cli comparisons get
Get comparison
clika-cli comparisons get <id> [flags]
Positional arguments: required <id>.
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/comparisons/{id}. MCP tool name: get_comparisons_id.
clika-cli comparisons list
List comparisons
clika-cli comparisons list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/comparisons. MCP tool name: get_comparisons.
clika-cli huggingface-tasks
huggingface-tasks has 1 subcommands.
clika-cli huggingface-tasks list
List HuggingFace pipeline tasks
clika-cli huggingface-tasks list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/huggingface-tasks. MCP tool name: get_huggingface_tasks.
clika-cli model-recommendations
model-recommendations has 1 subcommands.
clika-cli model-recommendations list
Model recommendations for a benchmark type
clika-cli model-recommendations list [flags]
| Flag | Type | Default | Meaning |
|---|---|---|---|
--benchmark_type | string | none | Benchmark type slug |
--raw | bool | false | print raw response without pretty-printing |
Endpoint: GET /api/model-recommendations. MCP tool name: get_model_recommendations.
Related
- Jobs: a benchmark leg is a job, and the job commands read its logs, outputs and per-sample results.
- Devices: choosing and preparing the hardware to benchmark on.
- apply and export: keeping a benchmark definition in git.
- Artifacts: the models and datasets a benchmark consumes.
- Benchmark concept: what the platform measures, and the datasets it ships.