WhisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Filament plugins, Laravel packages and open source starter kits actively maintained. All MIT, with releases for v3, v4 and v5 where applicable.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
ASR Configurator, Essentials and Atomic Testing
All packages have CI, Pest tests and open issues tagged good first issue.
New version available.