Two tools, one job: making one video work everywhere
If you sell to more than one language market, you already know the old way of localizing video is slow: re-record voiceover with a studio or freelancer per language, then either live with mismatched lip movement or pay for a separate re-shoot. This platform handles both halves of that problem with two dedicated tools that are designed to work together.
Dubbing: translate and re-voice
Dubbing takes your original video, transcribes and translates the spoken audio, and generates a new voice track in the target language β while preserving the original speaker's tone, pacing and emotion rather than producing a flat, robotic read. It supports 30+ languages, and you can choose to keep the original speaker's voice character or select a different voice for the new track.
Workflow:
- Upload your video. The original audio is transcribed and translated automatically.
- Choose target languages. Pick from 30+ options; keep the source voice character or swap it for a new one.
- Generate and download. One upload produces a localized audio version for every market you selected.
Lip Sync: match the mouth to the audio
Lip Sync solves the problem dubbing alone leaves behind: once you have new audio in a different language, the original speaker's mouth is still moving to the old language, which reads as obviously fake. Lip Sync maps any audio track onto any face with frame-accurate mouth movement, so the result looks genuinely spoken rather than dubbed-over. It's built to handle more than straight-on talking heads β head turns, partial occlusion and natural expression changes are all accounted for, not just a face staring directly at the camera.
Workflow:
- Upload your video β any clip with a clearly visible face: an interview, an ad, UGC-style content.
- Add the new audio β a translated dub, a re-record, or an entirely different voice.
- Generate and download β frame-accurate sync that holds up even in a close-up.
Combining them: the full localization pipeline
Used together, the two tools cover the complete job:
- Run your source video through Dubbing to get translated, re-voiced audio in each target language.
- Feed each translated audio track back through Lip Sync against the original video.
- Download a version per market where both the voice and the mouth movement match the local language.
One upload of your source footage becomes a full set of market-native videos, without re-shooting, re-casting or booking a voice actor per language.
Where this gets used
- Global ad campaigns β one hero video, localized voice and lip movement per market instead of separate shoots
- Multilingual product explainers and demos β especially useful for SaaS and e-commerce sellers running the same explainer across regions
- Repurposing creator content β turn one piece of long-form or social content into versions for audiences that don't share a language
- Onboarding and support videos β internal or customer-facing training content that needs to read as native, not subtitled
How this is different from generative video models
It's worth being clear about what Dubbing and Lip Sync are not: they don't generate new video from a text prompt the way Kling, Veo 3.1, Sora 2 and Seedance do. Those models create video (and in most cases audio) from nothing, based on a description. Dubbing and Lip Sync do the opposite β they start from footage and a performance that already exist and adapt the audio and mouth movement on top of it. If you need brand-new footage, you want a generative video model; if you need your existing footage to speak a different language convincingly, Dubbing and Lip Sync are the right tools for the job.
What it costs
Both tools are priced the same way as the rest of the video lineup β by output duration, at the platform's standard per-second video rate. See the credits guide for the exact numbers and worked examples across different clip lengths.