Skip to main content

TtsModel

A loaded text-to-speech model.

accepts_voice_reference (property)​

defaults (property)​

The family's trained baseline of the request knobs, as a dict of synthesize's keywords (a knob the family leaves to its own default reads None); a request starts from them where it sets none.

tokenizer (property)​

set_stage_observer​

set_stage_observer(self, callback: 'StageObserver | None') -> 'None'

Report every synthesis's stages (text, talker, codec and the audio produced) to callback, called with a :class:StageEvent on the model's own thread; it must return at once. None detaches.

synthesize​

synthesize(self, text: 'str', **options: 'Any') -> 'Any'

Synthesize text; returns an object with samples (a Float32 [S] tensor) and sample_rate. Options: max_frames, temperature, top_k, seed, speaker, voice, reference_audio, reference_sample_rate, exaggeration, cfg_scale, steps, extra (a dict of family knobs), stop_requested (a callable the model asks on its own thread between decode blocks: True stops the synthesis there, and the audio produced so far is the answer), and stage_observer (this call's stage events; detached after it).

synthesize_stream​

synthesize_stream(self, text: 'str', **options: 'Any') -> 'Iterator[AudioChunk]'

Synthesize as synthesize does, yielding the audio as the model produces it: one AudioChunk (samples, a Float32 [S] tensor; sample_rate; index; is_last) per block of codec frames where the family's vocoder is causal, else the utterance as one chunk; the chunks concatenate to the utterance. synthesize's options apply, plus chunk_frames (the frames per chunk on a family that vocodes per block; 0 keeps the family's own block). The model runs on its own thread and this iterator drains its chunks; leaving the loop early stops the synthesis after the current block, and stop_requested stops it between decode blocks as synthesize's does.