Skip to main content

//clika-runtime/io.clika.modelverse/GenerationReport

GenerationReport

[common]
data class GenerationReport(val finish: FinishReason, val promptTokens: Int, val completionTokens: Int, val timeToFirstTokenMs: Float, val prefillTokensPerSecond: Float, val decodeTokensPerSecond: Float, val latencyMs: Float, val accelerator: String, val cachedPromptTokens: Int = 0, val text: String = "", val reasoning: String = "", val visibleText: String = "", val toolCalls: List<ToolCall> = emptyList(), val toolsCalled: Boolean = false)

The counts and timings of one generation, delivered once through GenerationListener.onDone.

Constructors​

GenerationReport[common]
constructor(finish: FinishReason, promptTokens: Int, completionTokens: Int, timeToFirstTokenMs: Float, prefillTokensPerSecond: Float, decodeTokensPerSecond: Float, latencyMs: Float, accelerator: String, cachedPromptTokens: Int = 0, text: String = "", reasoning: String = "", visibleText: String = "", toolCalls: List<ToolCall> = emptyList(), toolsCalled: Boolean = false)

Properties​

NameSummary
accelerator[common]
val accelerator: String
The compute the request ran on, as the compute API's name (CPU; an accelerator among several carries its ordinal, CUDA:1).
cachedPromptTokens[common]
val cachedPromptTokens: Int = 0
Of promptTokens, the tokens the model's cache served from an earlier turn instead of computing them again: a chat sends the whole conversation on every turn, and the text model keeps a finished turn's cache for the next; 0 when nothing was served.
completionTokens[common]
val completionTokens: Int
The tokens the model produced.
decodeTokensPerSecond[common]
val decodeTokensPerSecond: Float
Reply tokens produced per second, over the decode span.
finish[common]
val finish: FinishReason
latencyMs[common]
val latencyMs: Float
The whole request, in milliseconds.
prefillTokensPerSecond[common]
val prefillTokensPerSecond: Float
Prompt tokens ingested per second.
promptTokens[common]
val promptTokens: Int
The prompt tokens the model ingested.
reasoning[common]
val reasoning: String
The reasoning channel's text (GenerationListener.onThinking joined); empty when the request had none.
text[common]
val text: String
The whole reply as the model wrote it, the reasoning channel's text included where the request had one; the text streamed so far on a canceled reply.
timeToFirstTokenMs[common]
val timeToFirstTokenMs: Float
From the request to the first streamed token, in milliseconds.
toolCalls[common]
val toolCalls: List<ToolCall>
The calls the reply made, as the model's call format wrote them (an id only where the format wrote one).
toolsCalled[common]
val toolsCalled: Boolean = false
True when the reply called a tool; finish then reads FinishReason.TOOL_CALLS.
visibleText[common]
val visibleText: String
The reply's user-facing text (GenerationListener.onToken joined): text without the reasoning channel.