Skip to main content

DiarizationModel

A loaded speaker-diarization model: a 16 kHz mono clip in, who spoke when out.

sample_rate (property)​

The sample rate every waveform is read at, 16000 Hz: a path or an encoded payload is decoded and resampled to it; a tensor is taken as already at it.

diarize​

diarize(self, audio: 'str | bytes | clika_runtime.Tensor', *, streaming: 'bool' = False, streaming_mode: 'str' = '', threshold: 'float' = 0.5) -> 'list[Any]'

Who spoke when, in time order, as :class:SpeakerSegment rows (speaker, start and end in seconds). audio is an audio file path, an encoded payload (bytes), or a [S] Float32 tensor at 16 kHz. streaming runs chunk by chunk with the speaker cache, as a live stream does, streaming_mode naming the latency mode (the checkpoint's default when empty); threshold in (0, 1) is the activity a frame needs to count for a speaker.