Skip to main content

Transcribe audio

A speech-to-text family provides transcribe, and the workflow is the same one the tutorial used for text: resolve the model, run its command, get the payload on stdout. This guide transcribes a file locally, then serves transcription over HTTP.

From the command line​

clika-modelverse openai/whisper-large-v3-turbo transcribe meeting.wav
Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week.

The transcript is the entire stdout payload, ready to redirect into a file. Decoding progress and timing ride stderr. The audio loader decodes WAV, FLAC, MP3 and OGG (Vorbis) and resamples to the model's expected rate, so a file in any of those containers transcribes as is. A container outside that set (OGG Opus, M4A/AAC, WMA) is refused by name, audio decode: <file>: OGG (Opus) is not supported; this build decodes WAV, FLAC, MP3, OGG (Vorbis), with exit code 1; a damaged file, or a file that is not audio at all, is refused the same way, with the file named and the first bytes it actually found quoted back. The web UI is not bound by this set: it decodes a recording or an upload in the browser and sends 16 kHz WAV, so anything the browser plays transcribes.

The load knobs from the tutorial apply unchanged: --device cuda places the encoder and decoder, --cache-dir controls where the snapshot lives, and a local directory works as the source for offline machines. Whisper checkpoints come in sizes from whisper-tiny (fits anywhere, fastest, roughest) to whisper-large-v3 (the accuracy reference); Model requirements has the figures. A clip longer than one 30 s window is windowed and transcribed whole.

Over HTTP​

The same model serves the OpenAI transcription route:

clika-modelverse openai/whisper-large-v3-turbo serve --port 8000

Call it as a multipart upload, the OpenAI shape:

curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F file=@meeting.wav \
-F model=whisper
{"text":" Good morning, everyone. Before we start, two quick announcements about the release schedule and the on-call rotation for next week."}

OpenAI SDK clients use their audio.transcriptions.create(...) call against the server's base URL, unchanged. The server-side flags and operations story (binding, health, admission control) is the common one in Serve an OpenAI-compatible endpoint.

For concurrent callers, --concurrent-sessions N admits N requests at once: Whisper keeps one encoder and mints decoder-only replicas over the same weights, encodes every window on the shared encoder with no lock held, and locks one replica for the decode alone, so N uploads run without a whole-request lock. The library form is LoadOptions::concurrent_sessions.

For transcription inside your own process, the library's SttModel and the transcription engine mount on the same OpenAiServer the 01_serve example demonstrates (Additional examples).