Multilingual dubbing

AI dubbing for audio and video

Translate speech, review each sentence, and export a dubbed video or aligned audio.

ASR + voice on our servers Context-aware translation
  1. Add mediaChoose a file and languages
  2. Review the scriptListen and edit each sentence
  3. Generate & exportAligned audio or video

Add a video or audio file

Audio or video up to 2 hours. Download aligned WAV audio, or export a dubbed video.

Drop audio or video here

MP3 · WAV · M4A · MP4 · MOV · WebM · Up to 2 hours / 2 GB

This tool sends extracted audio and text

Extracted audio and text are processed by server models. This implementation adds no server file storage or training use.

Processing & privacy
Workflow

Processing route

Transcribe, translate with context, and generate speech. Local draft playback is available when supported.

  1. ASR
    TranscribeServer
  2. LLM
    TranslateTranslation uses server
  3. TTS
    Create voiceFinal voice uses server

What leaves the browser

Audio is extracted on your device and sent for transcription and audio-based text checking. Text and neighboring context go to translation; reviewed and, when needed, shortened text goes to speech generation. Video pictures stay on your device.

The latest draft is saved on this device for 7 days when storage is available. Reselect the same file after refreshing to resume.

Three steps, with a review in the middle

ASR creates timed text. Translation keeps those passages aligned. You can fix both columns before TTS turns the approved translation into a WAV track.

Timing and voices

Each sentence starts at its original time. Long sentences may be made more concise or spoken faster; silent gaps are kept. Timing follows the transcript, which you can edit. Voices are not cloned.

AI dubbing guide

How to dub audio or video online

AI dubbing online for audio and video: transcribe, translate and review each passage, then export a dubbed video or aligned WAV voice track.

  1. Choose the source media

    Upload an audio or video file, then choose the spoken language and the language for the new voice track. The browser prepares the media in short timed chunks so longer recordings can be reviewed as a script.

  2. Review the translated script

    The tool transcribes speech, translates each timed passage and keeps the original beside the translation. Listen to the source, fix names or phrasing, and mark unclear passages before creating the voice.

  3. Create and export the dub

    Generate the reviewed speech track, listen to it against the original timing, then download the aligned WAV or export a dubbed video when the browser supports the source video and codec.

Audio and video dubbing formats

MP3 · WAV · M4A · AAC

Use a browser-readable audio file as the source. The generated voice is delivered as an aligned WAV track.

MP4 · MOV

The tool can preview the source picture and may export a dubbed video when the video track and browser codecs are supported.

WebM

Use when the browser can decode it. Download the aligned voice track if video export is unavailable.

What this AI dubbing tool does and does not do

The current workflow preserves passage order and timing, but it does not clone a speaker, create a live microphone, or guarantee lip sync. Audio is extracted on your device and sent with transcript text to server models; video pictures stay on your device. Exporting a dubbed video replaces the original soundtrack, including music and background sounds.

FAQ

AI dubbing questions

Can I dub a video into another language online?

Yes. Upload a supported video, choose the source and target languages, review the translated script, and generate a new voice track. Video export depends on the browser and source codec; the aligned WAV remains available.

Can I translate an audio recording and keep its timing?

Yes. Each transcript passage keeps its original start and end times. You can edit the translation before the speech model creates an aligned WAV track.

Is this voice cloning?

No. The generated track uses a selected synthetic speech model. It does not copy the identity or voiceprint of the original speaker.

Are files uploaded for AI dubbing?

Extracted audio and speech text are sent to server models for transcription, translation and voice generation. Video pictures stay on your device. The page explains this before processing; this implementation does not add file storage or training use.

Does dubbing include lip sync?

No. The workflow aligns the new speech to the original passage timing and can export a dubbed video, but it does not adjust mouth movement or guarantee lip synchronization.

Can I edit the translation before generating the voice?

Yes. The editor keeps the original transcript and translation in separate editable columns. Search the script, filter passages that need review, and translate again when the source language changes.

Continue with the same media