WebVTT · .vtt · no conversion step

Turn a VTT file into audio.
Keep every cue.

Upload the WebVTT file you already have, from YouTube, Zoom, Teams or your video player, and download a voiceover where each cue is spoken in its own time slot. No conversion to SRT. No timeline work.

10 free minutes · No credit card required

How it works

  1. 1

    Upload the .vtt file

    Drop in the file exactly as it was exported. Recastr reads WebVTT directly.

  2. 2

    Review the cue text

    Header, notes, styling and positioning are already gone. Fix wording or delete cues you do not want spoken.

  3. 3

    Choose a voice

    Pick a voice, or translate the text first and pick a voice for the target language.

  4. 4

    Download MP3 or WAV

    Every cue is synthesised into its own start-to-end window and assembled into one file.

What Recastr reads from a WebVTT file

WebVTT carries more than words and times. This is what happens to each part. To see it for yourself, download a sample .vtt that uses all of them.

WEBVTT header and NOTE blocks
Skipped. The signature line and any comments never reach the voice.
STYLE and REGION blocks
Skipped. Audio has no layout, so CSS-style rules and named regions are ignored.
Cue identifiers
The optional label line before each timestamp is ignored. Cues are placed by time, not by name.
Timestamp format
HH:MM:SS.mmm or MM:SS.mmm with a dot before the milliseconds. The hours part is optional and both forms are read to the millisecond.
Cue settings
position, line, align and size after the arrow are dropped. They place text on a screen; the audio does not need them.
Inline tags
Voice <v>, class <c>, bold, italic, underline, ruby and karaoke timestamp tags are removed. The words inside them are kept.
Multi-line cues
Line breaks inside a cue become spaces, so a two-line caption is spoken as one sentence.

Where VTT files come from

YouTube caption exports

Download the captions for a video as WebVTT and give it a new voice, or a voice in another language, without touching the edit.

Zoom and Teams recordings

Both export meeting transcripts as .vtt. Upload it as is, tidy the text, and generate a clean narration of what was said.

Course and video platforms

Vimeo, Wistia, Kaltura and most LMS players store captions as WebVTT for the HTML5 track element. Reuse that file instead of re-transcribing.

Accessibility teams

Captions written for compliance are already timed and proofread. That makes them the cleanest possible script for an audio version.

This is not a VTT-to-SRT converter

Converters rewrite the timestamps into SRT and hand the file back to you, and the voiceover is still your problem. Recastr reads WebVTT natively, so there is no intermediate SRT, nothing to re-upload, and the cue times you got from YouTube or Zoom are the ones the audio is generated against.

Zoom and Teams transcripts, as exported

Teams marks speakers with <v> tags, which are removed automatically. Zoom writes the speaker name as ordinary text at the start of every cue. Recastr shows you those cues in the editor before generating, so you can remove the names before anything is spoken.

Cues you delete stay silent

Remove a cue in the editor and nothing is generated for its time window. A caption for on-screen text or a [MUSIC] marker never turns into narration.

Watch it work

Frequently asked questions

Do I need to convert my VTT to SRT first?
No. Upload the .vtt directly. Recastr parses WebVTT itself, including files with cue identifier lines, NOTE comments and STYLE blocks.
What happens to speaker <v> tags?
The tag is removed and the words are kept. A subtitle import is treated as a single speaker. To give different people different voices, run the original recording through transcription instead, where speakers are detected from the audio.
Can I upload a Zoom or Microsoft Teams transcript?
Yes, both export WebVTT. Teams marks speakers with <v> tags, which are stripped automatically. Zoom writes the name as plain text at the start of each cue, so delete those prefixes in the editor before generating or they will be read out.
Are positioning and styling kept?
No. Cue settings such as position and align, and STYLE blocks, describe where text sits on a screen. The output is audio, so they are dropped and never affect timing.
How large can the VTT file be?
Subtitle uploads are capped at 1 MB, which is far more than a feature-length caption file needs. Generating the voiceover draws on your minutes: 10 free, then by plan.
What if I have an SRT or ASS file instead?
SRT files go through the SRT-to-audio workflow and ASS files through ASS-to-audio. All three formats land in the same editor and produce the same kind of timed audio.

Ready to hear your captions?

10 free minutes included · No credit card required