Audio to text

Convert audio and video to text

Turn recordings into editable text for free. Listen, correct each passage and download a transcript or subtitle file. No account required.

On-device or cloud transcription Video stays on your device
Audio to text
  1. 1Add file
  2. 2Transcribe
  3. 3Review & export

Audio stays on this device. First use downloads about 100 MB; cached when browser storage is available. Computers are recommended for long recordings.

Model not fully cached

Select the language spoken in the recording. Automatic detection is available in cloud mode.

Vocabulary hints are available in cloud mode. You can edit the local transcript after recognition.

Add a recording

Choose an audio or video file up to two hours long.

Drop audio or video here

MP4 · MOV · M4A · MP3 · WAV · AAC · FLAC · WebM · Up to 2 hours / 2 GB

On-device mode: your audio is not uploaded

Recognition runs in your browser with Whisper base. Model files are downloaded from this site. Switching to cloud requires choosing cloud mode and starting transcription yourself.

The latest draft is saved on this device for 7 days when storage is available. Reselect the same file after refreshing to resume.

How to convert audio to text online

  1. Choose your recording

    Select an audio or video file, or try the prepared example. Choose on-device or cloud mode and the spoken language. In cloud mode, add the correct spelling of names or specialist terms in the optional vocabulary field.

  2. Listen and edit

    Start transcription, then click a timestamp to replay that passage. Adjust playback speed, search the transcript and correct names, numbers or unclear words directly in the editor.

  3. Copy or export

    Copy the full text or download TXT, SRT, VTT or JSON. Edits are included in the export; searching filters the editor without removing other passages from the downloaded transcript.

Choose an audio or video file

TXT, SRT, VTT or JSON: choose your export

Input: MP3, M4A, WAV, AAC, FLAC, MP4, MOV and WebM, subject to codec support. Up to 2 hours and 2 GB per file; browser memory and available storage can impose lower practical limits. Video transcription recognizes the soundtrack, not words printed in the picture.

FormatWhen to use it
TXTPlain text for notes, interview quotes or pasting into a document. Timestamps are omitted.
SRTNumbered subtitle cues with start and end times, for video editors and players that support SubRip.
VTTWebVTT captions for compatible web players. Like SRT, this is a separate subtitle file, not a captioned video.
JSONStructured text, language and segment timings for your own workflow. Word timings are included when available and are removed from manually edited passages.

Get a more useful transcript

Set the language and vocabulary
Specify the spoken language. In cloud mode, add the correct spelling of unfamiliar terms. Hints can guide recognition but cannot guarantee the right words or recover inaudible speech.
Review difficult passages against the recording
Listen closely to names, dates, amounts and overlapping speakers. Long model paragraphs are split using real word timings when available; no new timing is guessed just to shorten a line.
Recover without losing completed work
Cloud recognition retries transient connection failures once. If processing stops, download completed text or retry the unfinished segments. When local storage is available, the latest draft stays on this device for seven days; reselect the same file after refreshing to restore it.

Audio to text questions

Is this audio to text converter free?

Yes. The tool currently requires no account or payment. Files are limited to 2 hours and 2 GB, and model availability and browser resources can affect processing. It is not an unlimited-capacity service.

Can I transcribe an MP3 or an iPhone voice memo?

Yes. Select a supported MP3 or M4A file saved on your device. Export a voice memo to Files first if needed, then choose it here. The tool does not access your recorder or messaging account.

Can I convert a video to text or SRT subtitles?

Yes, when the video contains a readable speech track. Select MP4, MOV or another supported file and download SRT or VTT after reviewing. This page does not read on-screen text or burn subtitles into a video.

Which transcription languages can I choose?

In cloud mode, use automatic detection, or select English, Chinese, Spanish, Portuguese, Japanese, Korean or French. Choosing a language guides recognition; it does not translate the recording.

Does it identify speakers or write meeting summaries?

It produces a timed speech transcript. Automatic speaker labels, meeting summaries and live microphone dictation are not provided. Review conversations against the original audio, particularly when people speak at the same time.

Are audio, vocabulary hints and transcripts kept locally?

In on-device mode, audio stays on your device. In cloud mode, audio segments and vocabulary hints are sent to Cloudflare Whisper. Video pictures stay local. When browser storage is available, the latest transcript, edits and settings are stored on this device for seven days. Start over or clear this site’s storage to remove the local draft; no copy of the original media is saved in it.

Why did transcription stop or miss some words?

Quiet audio, background noise, unfamiliar terms and overlapping voices can reduce recognition quality. An unsupported codec or limited memory can stop preparation. Try a shorter file or another supported format; retry a failed segment and compare the output with the recording.

Can I transcribe audio without uploading it?

Yes. Choose On this device and the spoken language. The first use downloads about 100 MB of model and runtime files; audio recognition then runs in the browser. Cache reuse depends on storage. Cloud mode is a separate choice and sends audio and vocabulary hints to Cloudflare. Local mode does not automatically switch to cloud.