SttModel
A loaded speech-to-text model.
audio_processor (property)
max_active (property)
tokenizer (property)
features
features(self, audio: 'bytes') -> 'clika_runtime.Tensor'
The input features of an encoded audio payload.
set_stage_observer
set_stage_observer(self, callback: 'StageObserver | None') -> 'None'
Report every transcription's stages (features, encode,
decode, the final text and the language) to callback, called
with a :class:StageEvent on the model's own thread; it must return
at once. None detaches.
transcribe
transcribe(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'str'
Transcribe an audio file path, an encoded audio payload (bytes), or
already-computed input features (a tensor). language names the
spoken language as an ISO code (empty keeps the model's default);
target_language names the text's language: empty keeps the spoken
one (a transcription), another code asks for speech-to-text
translation into it on a model whose prompt carries a target
language, and a model that cannot translate refuses it, naming what
it lacks. prompt is prior text
the decode conditions on (names, spellings, the style of what came
before); a model whose vocabulary carries no previous-text token
refuses it, naming the token. temperature is the sampling
temperature of each window's first decode: 0 (the default) picks the
most probable token at every step and climbs the fallback ladder,
temperatures (the model's own rungs when None: 0, 0.2, 0.4,
0.6, 0.8 and 1.0), decoding the window again at the next rung while
its mean token log-probability is under logprob_threshold (-1.0
when None) or its text's compression ratio (its length over its
deflated length) is above compression_ratio_threshold (2.4 when
None: a decode that fell into repetition), and the window is not
silence (its no-speech probability above no_speech_threshold,
0.6 when None); a positive temperature samples, and starts
the ladder there. A
sampled decode draws from a generator seeded with seed, so a call
repeats exactly. stop_requested is a cooperative stop for a clip
that spans several windows: the model asks it on its own thread at
every window boundary with the windows done and the total (0 on the
long-form walk), True ends the transcription there with the text
of the windows done; a single-window clip never asks. stage_observer
receives this call's stage events (see set_stage_observer) and is
detached after it.
transcribe_stream
transcribe_stream(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, word_timestamps: 'bool' = False, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'Iterator[TranscriptEvent]'
Transcribe as transcribe does (prompt, temperature and
seed included), yielding the decode's events as
they happen: a Partial after each token (the open segment's words
so far), a Segment when one closes (its text, ids and span), and
the Done event last (the whole transcript, every id, the language
and the duration). word_timestamps=True places the text on the
audio word by word: the Segment and Done events carry
words, each a :class:TranscriptWord with its span and the mean
probability of its tokens; the alignment costs one more decoder pass
per window, and a model that cannot place its words (a checkpoint
naming no alignment heads, a vocabulary with no time grid) refuses,
naming what it lacks. A stream decodes each window once, at the fallback
ladder's first rung: its partials are never revised, so a window
transcribe would decode again at a higher temperature streams as
its first decode. The model decodes on its own thread and this
iterator drains its events, so the calling thread never blocks the
model. Leaving the loop early (a break, an exception) stops the
decode within one token; the next call starts a fresh one.
stop_requested is transcribe's window-boundary stop.