---
title: "SttModel"
sidebar_label: "SttModel"
description: "The clika_runtime.modelverse SttModel class."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

A loaded speech-to-text model.

## `audio_processor` (property)



## `max_active` (property)



## `tokenizer` (property)



## `features`

```python
features(self, audio: 'bytes') -> 'clika_runtime.Tensor'
```

The input features of an encoded audio payload.

## `set_stage_observer`

```python
set_stage_observer(self, callback: 'StageObserver | None') -> 'None'
```

Report every transcription's stages (``features``, ``encode``,
``decode``, the final text and the language) to ``callback``, called
with a :class:`StageEvent` on the model's own thread; it must return
at once. ``None`` detaches.

## `transcribe`

```python
transcribe(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'str'
```

Transcribe an audio file path, an encoded audio payload (bytes), or
already-computed input features (a tensor). ``language`` names the
spoken language as an ISO code (empty keeps the model's default);
``target_language`` names the text's language: empty keeps the spoken
one (a transcription), another code asks for speech-to-text
translation into it on a model whose prompt carries a target
language, and a model that cannot translate refuses it, naming what
it lacks. ``prompt`` is prior text
the decode conditions on (names, spellings, the style of what came
before); a model whose vocabulary carries no previous-text token
refuses it, naming the token. ``temperature`` is the sampling
temperature of each window's first decode: 0 (the default) picks the
most probable token at every step and climbs the fallback ladder,
``temperatures`` (the model's own rungs when ``None``: 0, 0.2, 0.4,
0.6, 0.8 and 1.0), decoding the window again at the next rung while
its mean token log-probability is under ``logprob_threshold`` (-1.0
when ``None``) or its text's compression ratio (its length over its
deflated length) is above ``compression_ratio_threshold`` (2.4 when
``None``: a decode that fell into repetition), and the window is not
silence (its no-speech probability above ``no_speech_threshold``,
0.6 when ``None``); a positive ``temperature`` samples, and starts
the ladder there. A
sampled decode draws from a generator seeded with ``seed``, so a call
repeats exactly. ``stop_requested`` is a cooperative stop for a clip
that spans several windows: the model asks it on its own thread at
every window boundary with the windows done and the total (0 on the
long-form walk), ``True`` ends the transcription there with the text
of the windows done; a single-window clip never asks. ``stage_observer``
receives this call's stage events (see ``set_stage_observer``) and is
detached after it.

## `transcribe_stream`

```python
transcribe_stream(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, word_timestamps: 'bool' = False, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'Iterator[TranscriptEvent]'
```

Transcribe as ``transcribe`` does (``prompt``, ``temperature`` and
``seed`` included), yielding the decode's events as
they happen: a ``Partial`` after each token (the open segment's words
so far), a ``Segment`` when one closes (its text, ids and span), and
the ``Done`` event last (the whole transcript, every id, the language
and the duration). ``word_timestamps=True`` places the text on the
audio word by word: the ``Segment`` and ``Done`` events carry
``words``, each a :class:`TranscriptWord` with its span and the mean
probability of its tokens; the alignment costs one more decoder pass
per window, and a model that cannot place its words (a checkpoint
naming no alignment heads, a vocabulary with no time grid) refuses,
naming what it lacks. A stream decodes each window once, at the fallback
ladder's first rung: its partials are never revised, so a window
``transcribe`` would decode again at a higher temperature streams as
its first decode. The model decodes on its own thread and this
iterator drains its events, so the calling thread never blocks the
model. Leaving the loop early (a ``break``, an exception) stops the
decode within one token; the next call starts a fresh one.
``stop_requested`` is ``transcribe``'s window-boundary stop.
