//clika-runtime/io.clika.modelverse/BenchOptions
BenchOptions
[common]
data class BenchOptions(val cells: List<BenchCell>? = null, val concurrency: List<Int>? = null, val warmup: Int? = null, val iters: Int? = null, val cacheMode: BenchCacheMode? = null, val seed: Long? = null, val kvQuant: String? = null, val prefillChunkTokens: Long? = null, val stepTokenBudget: Long? = null, val sharedPrefixTokens: Long? = null, val promptLookupNumTokens: Long? = null, val uniquePrompts: Boolean? = null, val profile: Boolean? = null, val profilePipeline: Boolean? = null, val profileDir: String? = null, val reportPrefix: String? = null, val maxKvCacheBytes: Long? = null)
The knobs of Modelverse.bench, the command line's bench verb's and the Python modelverse.bench.run's: every field null keeps the library's own default (the one the command line prints in its --help), so a BenchOptions() runs the library's single-cell sweep at concurrency 1.
The sweep runs cells times concurrency, one (cell, width) at a time: each row decodes concurrency sessions of the cell at once through the serving pipeline over the model, warmup unmeasured passes first (any non-zero count also runs one global one-token warm-up before the first cell, so the first row measures the packed steady state; 0 measures cold) and iters measured ones, the metrics averaged over them. cacheMode is the key-value cache every row's decoder runs on, seed draws the synthetic prompts, kvQuant names the cache scheme the model was loaded with for the record (the model binds it; this field reports it). prefillChunkTokens (prompt tokens one session ingests per step; 0 = the whole prompt) and stepTokenBudget (the cap on tokens launched per step across sessions; 0 = unbounded) are the decoder's pacing knobs, null the platform's default. sharedPrefixTokens makes a cell's prompts share a prefix (a paged cache then serves it; the admittedTok and retires columns show the reuse), promptLookupNumTokens turns prompt lookup on for every request (0 = off; the *PredictionTokens and speculativeSteps columns tally it), and uniquePrompts false reuses one prompt set across the iterations so the cache serves the repeats (the warm protocol; the default regenerates every iteration's prompts, the cold one). profile adds one profiled iteration per row over the decoder's phases, profilePipeline over the whole request pipeline instead; profileDir receives their artifacts (empty keeps the summary in the report alone). reportPrefix streams the report to <prefix>.json (the configuration, before the first row) and <prefix>.csv (every row as it lands), so a run killed before its end keeps its finished rows. maxKvCacheBytes refuses a row whose cache alone would exceed it (an ERROR row, never an out-of-memory); 0 disables the check, null keeps the load's own ceiling.
Constructors
| BenchOptions | [common] constructor(cells: List<BenchCell>? = null, concurrency: List<Int>? = null, warmup: Int? = null, iters: Int? = null, cacheMode: BenchCacheMode? = null, seed: Long? = null, kvQuant: String? = null, prefillChunkTokens: Long? = null, stepTokenBudget: Long? = null, sharedPrefixTokens: Long? = null, promptLookupNumTokens: Long? = null, uniquePrompts: Boolean? = null, profile: Boolean? = null, profilePipeline: Boolean? = null, profileDir: String? = null, reportPrefix: String? = null, maxKvCacheBytes: Long? = null) |
Properties
| Name | Summary |
|---|---|
| cacheMode | [common] val cacheMode: BenchCacheMode? = null |
| cells | [common] val cells: List<BenchCell>? = null |
| concurrency | [common] val concurrency: List<Int>? = null |
| iters | [common] val iters: Int? = null |
| kvQuant | [common] val kvQuant: String? = null |
| maxKvCacheBytes | [common] val maxKvCacheBytes: Long? = null |
| prefillChunkTokens | [common] val prefillChunkTokens: Long? = null |
| profile | [common] val profile: Boolean? = null |
| profileDir | [common] val profileDir: String? = null |
| profilePipeline | [common] val profilePipeline: Boolean? = null |
| promptLookupNumTokens | [common] val promptLookupNumTokens: Long? = null |
| reportPrefix | [common] val reportPrefix: String? = null |
| seed | [common] val seed: Long? = null |
| sharedPrefixTokens | [common] val sharedPrefixTokens: Long? = null |
| stepTokenBudget | [common] val stepTokenBudget: Long? = null |
| uniquePrompts | [common] val uniquePrompts: Boolean? = null |
| warmup | [common] val warmup: Int? = null |