Skip to main content

//clika-runtime/io.clika.modelverse/BenchOptions

BenchOptions

[common]
data class BenchOptions(val cells: List<BenchCell>? = null, val concurrency: List<Int>? = null, val warmup: Int? = null, val iters: Int? = null, val cacheMode: BenchCacheMode? = null, val seed: Long? = null, val kvQuant: String? = null, val prefillChunkTokens: Long? = null, val stepTokenBudget: Long? = null, val sharedPrefixTokens: Long? = null, val promptLookupNumTokens: Long? = null, val uniquePrompts: Boolean? = null, val profile: Boolean? = null, val profilePipeline: Boolean? = null, val profileDir: String? = null, val reportPrefix: String? = null, val maxKvCacheBytes: Long? = null)

The knobs of Modelverse.bench, the command line's bench verb's and the Python modelverse.bench.run's: every field null keeps the library's own default (the one the command line prints in its --help), so a BenchOptions() runs the library's single-cell sweep at concurrency 1.

The sweep runs cells times concurrency, one (cell, width) at a time: each row decodes concurrency sessions of the cell at once through the serving pipeline over the model, warmup unmeasured passes first (any non-zero count also runs one global one-token warm-up before the first cell, so the first row measures the packed steady state; 0 measures cold) and iters measured ones, the metrics averaged over them. cacheMode is the key-value cache every row's decoder runs on, seed draws the synthetic prompts, kvQuant names the cache scheme the model was loaded with for the record (the model binds it; this field reports it). prefillChunkTokens (prompt tokens one session ingests per step; 0 = the whole prompt) and stepTokenBudget (the cap on tokens launched per step across sessions; 0 = unbounded) are the decoder's pacing knobs, null the platform's default. sharedPrefixTokens makes a cell's prompts share a prefix (a paged cache then serves it; the admittedTok and retires columns show the reuse), promptLookupNumTokens turns prompt lookup on for every request (0 = off; the *PredictionTokens and speculativeSteps columns tally it), and uniquePrompts false reuses one prompt set across the iterations so the cache serves the repeats (the warm protocol; the default regenerates every iteration's prompts, the cold one). profile adds one profiled iteration per row over the decoder's phases, profilePipeline over the whole request pipeline instead; profileDir receives their artifacts (empty keeps the summary in the report alone). reportPrefix streams the report to <prefix>.json (the configuration, before the first row) and <prefix>.csv (every row as it lands), so a run killed before its end keeps its finished rows. maxKvCacheBytes refuses a row whose cache alone would exceed it (an ERROR row, never an out-of-memory); 0 disables the check, null keeps the load's own ceiling.

Constructors​

BenchOptions[common]
constructor(cells: List<BenchCell>? = null, concurrency: List<Int>? = null, warmup: Int? = null, iters: Int? = null, cacheMode: BenchCacheMode? = null, seed: Long? = null, kvQuant: String? = null, prefillChunkTokens: Long? = null, stepTokenBudget: Long? = null, sharedPrefixTokens: Long? = null, promptLookupNumTokens: Long? = null, uniquePrompts: Boolean? = null, profile: Boolean? = null, profilePipeline: Boolean? = null, profileDir: String? = null, reportPrefix: String? = null, maxKvCacheBytes: Long? = null)

Properties​

NameSummary
cacheMode[common]
val cacheMode: BenchCacheMode? = null
cells[common]
val cells: List<BenchCell>? = null
concurrency[common]
val concurrency: List<Int>? = null
iters[common]
val iters: Int? = null
kvQuant[common]
val kvQuant: String? = null
maxKvCacheBytes[common]
val maxKvCacheBytes: Long? = null
prefillChunkTokens[common]
val prefillChunkTokens: Long? = null
profile[common]
val profile: Boolean? = null
profileDir[common]
val profileDir: String? = null
profilePipeline[common]
val profilePipeline: Boolean? = null
promptLookupNumTokens[common]
val promptLookupNumTokens: Long? = null
reportPrefix[common]
val reportPrefix: String? = null
seed[common]
val seed: Long? = null
sharedPrefixTokens[common]
val sharedPrefixTokens: Long? = null
stepTokenBudget[common]
val stepTokenBudget: Long? = null
uniquePrompts[common]
val uniquePrompts: Boolean? = null
warmup[common]
val warmup: Int? = null