Dubbing

Your clip, speaking a language you do not

A finished clip can be dubbed into 19 languages without re-recording anything and without spending your monthly video quota. Four different things get sold under the words AI dubbing, though, and they are not interchangeable. Here is the comparison first, so you can tell whether this is the one you wanted.

Four approaches, side by side

Each row is a real product category with real trade-offs. Two of them we do not offer, and the reasons are worth more to you than the feature would be.

ApproachWhat you getThe trade-offKatto
Translated subtitlesText in the target languageThe viewer still hears a language they do not speak. Fine for a clip that works muted, weak for anything carried by delivery.Also available
Synthetic voice dubbingA new spoken track in the target languageThe voice is not the speaker's. It reads naturally, but a viewer who knows the original will notice a stranger talking.What Katto does
Voice cloningThe speaker's own voice, in another languageNeeds the speaker's consent to be legitimate, and the same capability produces convincing impersonation. The exposure is not technical, it is what someone else does with it.Deliberately not shipped
Lip-sync dubbingThe mouth re-rendered to match the new audioHeaviest by far, and it falls apart on anything that is not a clean frontal head shot. On clipped podcast footage that is most of it.Not shipped

The hard part is the clock, not the voice

Calling a text-to-speech API is the easy half. The difficulty is that translation changes length: the same sentence is routinely a third longer in one language than another, and a dub that ignores this drifts further out of sync with every phrase until the speaker is answering a question they have not been asked yet.

So the original timing is treated as data rather than as something to overwrite. Each phrase is translated with its spoken duration as a constraint, synthesised on its own, and fitted back into the slot it came from, stretched only within limits that keep it sounding like a person. Captions are generated against the same clock, so the text and the new audio cannot part company.

What it costs you

Dubbing does not touch your monthly video quota, because that quota counts source videos you turn into clips and a dub is work on a clip you already made. It has its own cap instead: 20 dubs per rolling 30 days, up to 3 languages per request, with the oldest falling off the count as they age. The editor shows how many you have left before you pick, rather than after.

The cap exists because speech synthesis is the one step in the pipeline whose cost tracks usage rather than video length. We would rather state a number than sell an unlimited promise that gets quietly throttled later.

The one we are not shipping

Voice cloning is the feature everyone asks for, and it is the reason this page exists in this shape. Reproducing a speaker's voice from a clip is the same capability as putting words in their mouth, and shipping it openly means shipping it to people whose intentions we cannot check. If it ever arrives it will be gated behind evidence that the voice belongs to the person asking. Until then the dub uses a synthetic voice and does not pretend to be anyone.

Questions

How many languages?

19. The engine underneath can speak many more, and the picker will open up when the quality has been checked language by language rather than assumed. Announcing a number we have not verified is how a feature becomes a support queue.

Does dubbing use my monthly video quota?

No. The quota counts source videos you turn into clips. Dubbing is capped separately: 20 dubs per rolling 30 days, up to 3 languages in one go. The oldest ones drop off the count as they age.

Why cap it at all?

Because speech synthesis is the one step whose cost scales with how much you use it rather than with how long your video is. A cap is the honest version of that constraint. The alternative is an unlimited promise that quietly degrades.

Does the dub stay in sync?

That is the actual engineering problem, and it is not solved by calling a text-to-speech API. The original timing is kept as data, each phrase is translated for spoken duration rather than only for meaning, synthesised on its own, then fitted back into its slot. Captions are put on the same clock so they do not drift away from the new audio.

Why no voice cloning?

The technology works. The problem is that a tool which reproduces anyone's voice from a clip is also a tool for putting words in their mouth, and we would be shipping it to strangers. If it arrives it will arrive gated behind proof of consent, not as a toggle.

Can I dub through the API or an agent?

Yes. The same operation is available over the REST API and as an MCP tool, so an agent can dub a finished clip without a browser. It returns a job to poll, not a blocking call.

Reach an audience that does not speak your language

Dubbing is a Creator feature. The clip you already made is the input, so there is nothing to re-record and no quota to spend.