Skip to main content

ASR Models

Batch Transcription

TDT models process audio in chunks (~15s with overlap). Fast enough for dictation-style workflows.

Streaming Transcription

Custom Vocabulary

VAD Models

Diarization Models

TTS Models

HuggingFace Sources