---
title: "GenerationConfig"
sidebar_label: "GenerationConfig"
description: "The clika_runtime.modelverse.generation GenerationConfig class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

The per-request decode policy.

    ``max_new_tokens`` caps the reply (0 keeps the model's default budget).
    ``temperature`` at or below 0 decodes greedily; above 0 the logits are
    scaled by ``1 / temperature`` before sampling. ``top_k`` keeps the k
    most likely candidates (0 = no cut); ``top_p`` keeps the smallest set
    whose probability mass reaches it (1.0 = no cut). ``seed`` makes a
    sampled request reproducible when it samples alone on its device.
    ``stop`` lists extra stop strings and ``stop_token_ids`` extra terminal
    ids; ``min_tokens`` is a floor no stop condition ends the reply under;
    ``ignore_eos`` runs the full budget past the end token.
    ``reasoning_effort`` selects one of the family's declared reasoning
    levels ("" = the family default, "none" disables a toggleable channel,
    "on" turns an opt-in channel on). ``enable_thinking`` is the boolean
    spelling of that switch (True = "on", False = "none"; None leaves the
    family default); an explicit ``reasoning_effort`` wins over it, and a
    model without a switch refuses a False the way it refuses "none"
    (``GenerativeModel.supports_thinking`` and ``reasoning_mode`` say which).
    ``no_repeat_ngram_size`` removes, before each draw, every token that
    would repeat an n-gram of that many tokens the sequence already holds
    (0 = no rule); ``no_repeat_ngram_window`` limits the search to the
    latest that many tokens (0 = the whole sequence, prompt included).
    ``repetition_penalty`` divides, before each draw, the logit of every
    token the sequence already holds (prompt included) by it when positive
    and multiplies it when negative (1.0 = no penalty; a checkpoint's own
    value is the default).
    ``prompt_lookup_num_tokens`` turns on prompt lookup (speculative decoding
    with no second model): at each step the trailing n-gram of the sequence
    is looked up earlier in it and up to that many of the ids that followed
    are drafted, scored in one step and accepted up to the first
    disagreement, so the tokens equal the plain decode's and the step count
    shrinks (0 = off); ``max_matching_ngram_size`` is the longest n-gram the
    lookup tries first (1 or more). A draft runs alone: a repetition penalty,
    a no-repeat n-gram rule, a completion floor or log-probabilities beside
    it are refused, and ``GenerativeModel.supports_speculative_decoding``
    says whether the model serves one. The reply's ``report`` carries the
    tally (``accepted_prediction_tokens``, ``rejected_prediction_tokens``,
    ``speculative_steps``).
    ``num_assistant_tokens`` is the draft length per step when an assistant
    model is attached (``load(..., assistant_model=)`` or
    ``GenerativeModel.set_assistant``; 0 turns it off for the request, and it
    is inert without one; a request with both an assistant and
    ``prompt_lookup_num_tokens`` is refused: one draft source);
    ``num_assistant_tokens_schedule`` moves that length between steps
    ("constant", "heuristic": +2 while every drafted token held, else -1,
    never under 1; "heuristic_transient": the same per request);
    ``assistant_confidence_threshold`` stops a draft at the first token the
    assistant is less sure of than it (0 = no stop; in [0, 1]).
    

## `is_sampling` (property)

True when stochastic sampling is engaged: the temperature is at
least the sampling floor, 1e-5 (greedy otherwise).

## `__init__`

```python
__init__(self, max_new_tokens: 'int' = 0, temperature: 'float' = 0.0, top_k: 'int' = 0, top_p: 'float' = 1.0, seed: 'int | None' = None, ignore_eos: 'bool' = False, stop: 'list[str]' = <factory>, stop_token_ids: 'list[int]' = <factory>, min_tokens: 'int' = 0, reasoning_effort: 'str' = '', enable_thinking: 'bool | None' = None, no_repeat_ngram_size: 'int' = 0, no_repeat_ngram_window: 'int' = 0, repetition_penalty: 'float' = 1.0, prompt_lookup_num_tokens: 'int' = 0, max_matching_ngram_size: 'int' = 2, num_assistant_tokens: 'int' = 20, num_assistant_tokens_schedule: 'str' = 'constant', assistant_confidence_threshold: 'float' = 0.4) -> None
```

Initialize self.  See help(type(self)) for accurate signature.

## `from_bound`

```python
from_boundA config from the model library's own config object.
```

A config from the model library's own config object.

## `merged`

```python
merged(self, **overrides: 'Any') -> "'GenerationConfig'"
```

A copy with ``overrides`` applied; an unknown name raises TypeError
naming it and the accepted fields.

## `to_bound`

```python
to_bound(self) -> 'Any'
```

The model library's own config object carrying these values.
