//clika-runtime/io.clika.modelverse/GenerationReport
GenerationReport
[common]
data class GenerationReport(val finish: FinishReason, val promptTokens: Int, val completionTokens: Int, val timeToFirstTokenMs: Float, val prefillTokensPerSecond: Float, val decodeTokensPerSecond: Float, val latencyMs: Float, val accelerator: String, val cachedPromptTokens: Int = 0, val text: String = "", val reasoning: String = "", val visibleText: String = "", val toolCalls: List<ToolCall> = emptyList(), val toolsCalled: Boolean = false)
The counts and timings of one generation, delivered once through GenerationListener.onDone.
Constructors
| GenerationReport | [common] constructor(finish: FinishReason, promptTokens: Int, completionTokens: Int, timeToFirstTokenMs: Float, prefillTokensPerSecond: Float, decodeTokensPerSecond: Float, latencyMs: Float, accelerator: String, cachedPromptTokens: Int = 0, text: String = "", reasoning: String = "", visibleText: String = "", toolCalls: List<ToolCall> = emptyList(), toolsCalled: Boolean = false) |
Properties
| Name | Summary |
|---|---|
| accelerator | [common] val accelerator: String The compute the request ran on, as the compute API's name ( CPU; an accelerator among several carries its ordinal, CUDA:1). |
| cachedPromptTokens | [common] val cachedPromptTokens: Int = 0 Of promptTokens, the tokens the model's cache served from an earlier turn instead of computing them again: a chat sends the whole conversation on every turn, and the text model keeps a finished turn's cache for the next; 0 when nothing was served. |
| completionTokens | [common] val completionTokens: Int The tokens the model produced. |
| decodeTokensPerSecond | [common] val decodeTokensPerSecond: Float Reply tokens produced per second, over the decode span. |
| finish | [common] val finish: FinishReason |
| latencyMs | [common] val latencyMs: Float The whole request, in milliseconds. |
| prefillTokensPerSecond | [common] val prefillTokensPerSecond: Float Prompt tokens ingested per second. |
| promptTokens | [common] val promptTokens: Int The prompt tokens the model ingested. |
| reasoning | [common] val reasoning: String The reasoning channel's text (GenerationListener.onThinking joined); empty when the request had none. |
| text | [common] val text: String The whole reply as the model wrote it, the reasoning channel's text included where the request had one; the text streamed so far on a canceled reply. |
| timeToFirstTokenMs | [common] val timeToFirstTokenMs: Float From the request to the first streamed token, in milliseconds. |
| toolCalls | [common] val toolCalls: List<ToolCall> The calls the reply made, as the model's call format wrote them (an id only where the format wrote one). |
| toolsCalled | [common] val toolsCalled: Boolean = false True when the reply called a tool; finish then reads FinishReason.TOOL_CALLS. |
| visibleText | [common] val visibleText: String The reply's user-facing text (GenerationListener.onToken joined): text without the reasoning channel. |