WhisperX transcription and diarization
Transcription and speaker separation in a single GPU pass: captions in 99 languages, word-timed in 41.
Transcription and speaker separation used to be two services. They now run in one GPU pass, which removed a second transcription we were paying for and improved the timestamps at the same time.
Captions cover 99 languages. Word level timing, the kind that highlights one word at a time, covers 41 of them; the rest get sentence level timing. Japanese and Thai are in the second group because they are written without spaces, so a word boundary is a judgement rather than a fact.
Multi-speaker clips split more cleanly, because the speaker labels and the word timings now come from the same pass instead of being matched afterwards.
2 videos a month, no card. What is still being built is on the roadmap.