← Changelog

WhisperX transcription and diarization

Transcription and speaker separation in a single GPU pass: captions in 99 languages, word-timed in 41.

Transcription and speaker separation used to be two services. They now run in one GPU pass, which removed a second transcription we were paying for and improved the timestamps at the same time.

Captions cover 99 languages. Word level timing, the kind that highlights one word at a time, covers 41 of them; the rest get sentence level timing. Japanese and Thai are in the second group because they are written without spaces, so a word boundary is a judgement rather than a fact.

Multi-speaker clips split more cleanly, because the speaker labels and the word timings now come from the same pass instead of being matched afterwards.

Try it free

2 videos a month, no card. What is still being built is on the roadmap.