DiarizationModel
A loaded speaker-diarization model: a 16 kHz mono clip in, who spoke when out.
sample_rate (property)
The sample rate every waveform is read at, 16000 Hz: a path or an encoded payload is decoded and resampled to it; a tensor is taken as already at it.
diarize
diarize(self, audio: 'str | bytes | clika_runtime.Tensor', *, streaming: 'bool' = False, streaming_mode: 'str' = '', threshold: 'float' = 0.5) -> 'list[Any]'
Who spoke when, in time order, as :class:SpeakerSegment rows
(speaker, start and end in seconds). audio is an audio
file path, an encoded payload (bytes), or a [S] Float32 tensor at
16 kHz. streaming runs chunk by chunk with the speaker cache, as a
live stream does, streaming_mode naming the latency mode (the
checkpoint's default when empty); threshold in (0, 1) is the
activity a frame needs to count for a speaker.