YouTube caption exports
Download the captions for a video as WebVTT and give it a new voice, or a voice in another language, without touching the edit.
Upload the WebVTT file you already have, from YouTube, Zoom, Teams or your video player, and download a voiceover where each cue is spoken in its own time slot. No conversion to SRT. No timeline work.
10 free minutes · No credit card required
Upload the .vtt file
Drop in the file exactly as it was exported. Recastr reads WebVTT directly.
Review the cue text
Header, notes, styling and positioning are already gone. Fix wording or delete cues you do not want spoken.
Choose a voice
Pick a voice, or translate the text first and pick a voice for the target language.
Download MP3 or WAV
Every cue is synthesised into its own start-to-end window and assembled into one file.
WebVTT carries more than words and times. This is what happens to each part. To see it for yourself, download a sample .vtt that uses all of them.
Download the captions for a video as WebVTT and give it a new voice, or a voice in another language, without touching the edit.
Both export meeting transcripts as .vtt. Upload it as is, tidy the text, and generate a clean narration of what was said.
Vimeo, Wistia, Kaltura and most LMS players store captions as WebVTT for the HTML5 track element. Reuse that file instead of re-transcribing.
Captions written for compliance are already timed and proofread. That makes them the cleanest possible script for an audio version.
Converters rewrite the timestamps into SRT and hand the file back to you, and the voiceover is still your problem. Recastr reads WebVTT natively, so there is no intermediate SRT, nothing to re-upload, and the cue times you got from YouTube or Zoom are the ones the audio is generated against.
Teams marks speakers with <v> tags, which are removed automatically. Zoom writes the speaker name as ordinary text at the start of every cue. Recastr shows you those cues in the editor before generating, so you can remove the names before anything is spoken.
Remove a cue in the editor and nothing is generated for its time window. A caption for on-screen text or a [MUSIC] marker never turns into narration.
10 free minutes included · No credit card required