Benchmark a model on this machine
Whether a model meets your latency and throughput targets is a property of this machine, this quantization and this serving shape, and bench measures exactly that. Every runnable text family provides it, the flow drives the same serving pipeline prompt and serve run on, and the result is a structured table, never ad-hoc prints.
The default sweep
clika-modelverse meta-llama/Llama-3.2-1B-Instruct bench --device cuda
The default sweep runs three serving-shaped cells (balanced 128:128, prefill-heavy 2048:128, decode-heavy 128:1024) over the default concurrency ladder (1, 4, 8), each cell strictly one at a time so cells never contend with each other. One CSV row per cell and concurrency, the last column naming what the row measured:
isl,osl,concurrency,first_call_ms,ttft_ms,prefill_tok_s,itl_p50_ms,itl_p99_ms,prefill_chunk_ms_p50,prefill_chunk_ms_max,decode_tok_s,speedup_vs_c1,batching_efficiency,admitted_tok,retires,peak_active_bytes,num_allocs,num_inflight_park_waits,status,reason,notes
128,128,1,685.7974,14.375232,8904.204,4.709812,11.978596,0,0,194.13823,1,1,112,3,3422640128,59180,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,128,4,937.06464,25.81384,4969.912,6.5161915,15.847513,0,0,552.3547,2.845162,0.7112905,448,12,3746307584,60982,6,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,128,8,745.5794,36.65184,3492.3213,4.9885244,10.634405,0,0,1434.3369,7.388225,0.92352813,896,24,4083747328,93121,12,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,1,820.59033,67.9156,30155.074,5.575619,8.472811,0,0,159.87682,1,1,2032,3,4083747328,65325,20,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,4,1258.1548,245.00293,8360.729,7.7572827,15.318566,0,0,414.64108,2.5935037,0.6483759,8128,12,4192593408,67393,41,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
2048,128,8,1308.3411,464.48145,4409.434,6.383371,11.002449,0,0,787.79083,4.927487,0.61593586,16256,24,4744088576,99717,63,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,1,4980.8936,10.374589,12337.838,4.823417,8.758105,0,0,190.86696,1,1,112,3,4744088576,498596,0,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,4,7377.311,26.746096,4836.729,6.6391444,13.775986,0,0,592.14624,3.102403,0.77560073,448,12,4744088576,508606,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
128,1024,8,5360.8525,39.64144,3232.0154,5.407993,9.28407,0,0,1437.6843,7.5323896,0.9415487,896,24,4744088576,767141,3,ok,"",prompts=cold (a fresh prompt every iteration; prefill_tok_s prices prompt ingest)
The numbers are one machine's run (one workstation GPU, this model at its checkpoint dtype, bf16, cold prompts, the default); yours are the point. Stdout carries the CSV, ready for a spreadsheet or a script; diagnostics ride stderr. A failed cell reports error in status without invalidating the rows around it, and the reason column carries the model's own refusal or the cell's first failure in words (empty on an ok row). With --output, the report streams to disk as rows complete, so a run killed at its deadline keeps every finished cell.
Shape the sweep to your workload
--cells 512:256,4096:64: your ownisl:osllist, replacing the default three.--isl/--oslset a single cell directly.--concurrency 1,4,16: the parallel-session ladder;speedup_vs_c1andbatching_efficiencyin the report tell you what the added concurrency actually bought.--shared-prefix 256: how many prompt tokens every session shares (a system prompt, a RAG preamble); with it theadmitted_tokandretirescolumns show the paged prefix cache engaging. Every LLM cell runs cold prompts by default (a fresh random prompt per session every iteration, so a paged prefix cache serves no repeat;--unique-promptsnames that default and stays accepted).--repeat-promptsis the warm protocol: the warm-up and every iteration reuse one prompt set, soprefill_tok_sreads cache admits, and the row'snotessay so.--warmupand--iters: untimed full passes of each cell before it measures, then timed passes averaged. The default is two warm-up passes and one timed iteration;--warmup 0times a cold cell. A cell still cold after the default passes is a defect worth reporting, not a reason to raise the count.- The load knobs apply unchanged:
--kvcache paged|continuous(each verb has its own default andbenchruns paged),--prefill-chunk,--step-token-budget, a GGUF weight selector on the source,--max-seq. --memory-stages FILEappends the runtime's memory-pool counters to a JSON file at each stage of the model's life (after the weights bind, after the one-forward warm-up, after the first generation), one row per device, with the checkpoint's size. It is the footprint diagnostic behind the figures on Model requirements.--output PREFIXwrites the report to files beside the terminal view;--profileadds the per-op profile on stderr (the summary opens withProfiling Report Summary:, itstop_ops:block ranks op groups by self time, and the per-op table carries anAvgSelf(us)column, the per-op average that excludes nested spans),--profile-pipelinethe per-node one.
Reading the columns
ttft_ms is the enqueue-to-first-token wall, the number an interactive user feels; first_call_ms is the cell's very first call, which pays the kernel loads a fresh process owes and is kept out of ttft_ms for that reason. prefill_tok_s is prompt ingest; decode_tok_s is the aggregate generation rate across sessions. The inter-token percentiles itl_p50_ms/itl_p99_ms sample decode-only gaps; on a chunked-prefill run the chunk boundary walls report separately (prefill_chunk_ms_p50/_max), so a decode tail is never polluted by an ingest wall. When the p99 sits far above the p50, the deployment story is a batching or cache-pressure story, and the concurrency ladder narrows down which. Three columns carry the memory story: peak_active_bytes is the runtime pool's high-water mark for the cell, num_allocs the allocation count behind it, and num_inflight_park_waits the number of times a request waited for a slot rather than proceeding, a direct read on admission pressure. The -v diagnostics carry the load and peak-memory figures the Model requirements page's estimates are checked against.
Audio and media axes
A speech-to-text family's bench sweeps audio length instead of token shapes (--audio-seconds 5,30,120, with the same --concurrency ladder), reporting ingest as encoder frames over the end-to-end transcribe wall. Multimodal families add mm_bench, the media sweep over the serving path: --images per request, --image-resolution, --video-frames, with effective_isl reporting what the prefill actually ingested once media expansions counted, so a media-heavy row's ingest rate stays comparable to a text row's.
Benchmark the quantization you intend to ship (Run a specific GGUF quantization); a Q4_K_M and a Q8_0 of the same model can sit on different sides of a latency target. The serving flags that shaped a good bench row carry directly onto serve (Serve an OpenAI-compatible endpoint).