Skip to main content

GenerationConfig

The per-request decode policy.

``max_new_tokens`` caps the reply (0 keeps the model's default budget).
``temperature`` at or below 0 decodes greedily; above 0 the logits are
scaled by ``1 / temperature`` before sampling. ``top_k`` keeps the k
most likely candidates (0 = no cut); ``top_p`` keeps the smallest set
whose probability mass reaches it (1.0 = no cut). ``seed`` makes a
sampled request reproducible when it samples alone on its device.
``stop`` lists extra stop strings and ``stop_token_ids`` extra terminal
ids; ``min_tokens`` is a floor no stop condition ends the reply under;
``ignore_eos`` runs the full budget past the end token.
``reasoning_effort`` selects one of the family's declared reasoning
levels ("" = the family default, "none" disables a toggleable channel,
"on" turns an opt-in channel on). ``enable_thinking`` is the boolean
spelling of that switch (True = "on", False = "none"; None leaves the
family default); an explicit ``reasoning_effort`` wins over it, and a
model without a switch refuses a False the way it refuses "none"
(``GenerativeModel.supports_thinking`` and ``reasoning_mode`` say which).
``no_repeat_ngram_size`` removes, before each draw, every token that
would repeat an n-gram of that many tokens the sequence already holds
(0 = no rule); ``no_repeat_ngram_window`` limits the search to the
latest that many tokens (0 = the whole sequence, prompt included).
``repetition_penalty`` divides, before each draw, the logit of every
token the sequence already holds (prompt included) by it when positive
and multiplies it when negative (1.0 = no penalty; a checkpoint's own
value is the default).
``prompt_lookup_num_tokens`` turns on prompt lookup (speculative decoding
with no second model): at each step the trailing n-gram of the sequence
is looked up earlier in it and up to that many of the ids that followed
are drafted, scored in one step and accepted up to the first
disagreement, so the tokens equal the plain decode's and the step count
shrinks (0 = off); ``max_matching_ngram_size`` is the longest n-gram the
lookup tries first (1 or more). A draft runs alone: a repetition penalty,
a no-repeat n-gram rule, a completion floor or log-probabilities beside
it are refused, and ``GenerativeModel.supports_speculative_decoding``
says whether the model serves one. The reply's ``report`` carries the
tally (``accepted_prediction_tokens``, ``rejected_prediction_tokens``,
``speculative_steps``).
``num_assistant_tokens`` is the draft length per step when an assistant
model is attached (``load(..., assistant_model=)`` or
``GenerativeModel.set_assistant``; 0 turns it off for the request, and it
is inert without one; a request with both an assistant and
``prompt_lookup_num_tokens`` is refused: one draft source);
``num_assistant_tokens_schedule`` moves that length between steps
("constant", "heuristic": +2 while every drafted token held, else -1,
never under 1; "heuristic_transient": the same per request);
``assistant_confidence_threshold`` stops a draft at the first token the
assistant is less sure of than it (0 = no stop; in [0, 1]).

is_sampling (property)​

True when stochastic sampling is engaged: the temperature is at least the sampling floor, 1e-5 (greedy otherwise).

__init__​

__init__(self, max_new_tokens: 'int' = 0, temperature: 'float' = 0.0, top_k: 'int' = 0, top_p: 'float' = 1.0, seed: 'int | None' = None, ignore_eos: 'bool' = False, stop: 'list[str]' = <factory>, stop_token_ids: 'list[int]' = <factory>, min_tokens: 'int' = 0, reasoning_effort: 'str' = '', enable_thinking: 'bool | None' = None, no_repeat_ngram_size: 'int' = 0, no_repeat_ngram_window: 'int' = 0, repetition_penalty: 'float' = 1.0, prompt_lookup_num_tokens: 'int' = 0, max_matching_ngram_size: 'int' = 2, num_assistant_tokens: 'int' = 20, num_assistant_tokens_schedule: 'str' = 'constant', assistant_confidence_threshold: 'float' = 0.4) -> None

Initialize self. See help(type(self)) for accurate signature.

from_bound​

from_boundA config from the model library's own config object.

A config from the model library's own config object.

merged​

merged(self, **overrides: 'Any') -> "'GenerationConfig'"

A copy with overrides applied; an unknown name raises TypeError naming it and the accepted fields.

to_bound​

to_bound(self) -> 'Any'

The model library's own config object carrying these values.