Skip to main content

SttModel

A loaded speech-to-text model.

audio_processor (property)​

max_active (property)​

tokenizer (property)​

features​

features(self, audio: 'bytes') -> 'clika_runtime.Tensor'

The input features of an encoded audio payload.

set_stage_observer​

set_stage_observer(self, callback: 'StageObserver | None') -> 'None'

Report every transcription's stages (features, encode, decode, the final text and the language) to callback, called with a :class:StageEvent on the model's own thread; it must return at once. None detaches.

transcribe​

transcribe(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'str'

Transcribe an audio file path, an encoded audio payload (bytes), or already-computed input features (a tensor). language names the spoken language as an ISO code (empty keeps the model's default); target_language names the text's language: empty keeps the spoken one (a transcription), another code asks for speech-to-text translation into it on a model whose prompt carries a target language, and a model that cannot translate refuses it, naming what it lacks. prompt is prior text the decode conditions on (names, spellings, the style of what came before); a model whose vocabulary carries no previous-text token refuses it, naming the token. temperature is the sampling temperature of each window's first decode: 0 (the default) picks the most probable token at every step and climbs the fallback ladder, temperatures (the model's own rungs when None: 0, 0.2, 0.4, 0.6, 0.8 and 1.0), decoding the window again at the next rung while its mean token log-probability is under logprob_threshold (-1.0 when None) or its text's compression ratio (its length over its deflated length) is above compression_ratio_threshold (2.4 when None: a decode that fell into repetition), and the window is not silence (its no-speech probability above no_speech_threshold, 0.6 when None); a positive temperature samples, and starts the ladder there. A sampled decode draws from a generator seeded with seed, so a call repeats exactly. stop_requested is a cooperative stop for a clip that spans several windows: the model asks it on its own thread at every window boundary with the windows done and the total (0 on the long-form walk), True ends the transcription there with the text of the windows done; a single-window clip never asks. stage_observer receives this call's stage events (see set_stage_observer) and is detached after it.

transcribe_stream​

transcribe_stream(self, audio: 'str | bytes | clika_runtime.Tensor', *, max_new_tokens: 'int' = 0, language: 'str' = '', target_language: 'str' = '', prompt: 'str' = '', temperature: 'float' = 0.0, temperatures: 'Sequence[float] | None' = None, logprob_threshold: 'float | None' = None, no_speech_threshold: 'float | None' = None, compression_ratio_threshold: 'float | None' = None, seed: 'int' = 0, word_timestamps: 'bool' = False, stop_requested: 'Callable[[int, int], bool | None] | None' = None, stage_observer: 'StageObserver | None' = None) -> 'Iterator[TranscriptEvent]'

Transcribe as transcribe does (prompt, temperature and seed included), yielding the decode's events as they happen: a Partial after each token (the open segment's words so far), a Segment when one closes (its text, ids and span), and the Done event last (the whole transcript, every id, the language and the duration). word_timestamps=True places the text on the audio word by word: the Segment and Done events carry words, each a :class:TranscriptWord with its span and the mean probability of its tokens; the alignment costs one more decoder pass per window, and a model that cannot place its words (a checkpoint naming no alignment heads, a vocabulary with no time grid) refuses, naming what it lacks. A stream decodes each window once, at the fallback ladder's first rung: its partials are never revised, so a window transcribe would decode again at a higher temperature streams as its first decode. The model decodes on its own thread and this iterator drains its events, so the calling thread never blocks the model. Leaving the loop early (a break, an exception) stops the decode within one token; the next call starts a fresh one. stop_requested is transcribe's window-boundary stop.