WebVttParser

Reads WebVTT (.vtt) subtitles, the caption format the web standardised out of SubRip.

The same philosophy as SubRipParser: accept what real files contain. The differences that matter here, and only these, are handled:

  • The WEBVTT signature line, optionally after a byte order mark, optionally with a trailing description. A file without it is still read, because files without it exist.

  • NOTE and REGION blocks are skipped whole. STYLE blocks are read for the ::cue rules a cue here can carry, and a rule with anything more is ignored whole, never half applied (#498).

  • The millisecond separator is a full stop and hours are optional, which the shared timestamp grammar already accepts.

  • A cue's identifier (the line before its timing line) is kept only for ::cue(#id) rules.

  • Cue settings after the timing (position:, line:, align:) are read for the one thing the text path draws today, the horizontal alignment; the rest is recorded nowhere rather than misdrawn.

  • Inline <b>, <i>, <u>, <c> classes, <v Speaker> voices, <lang> and <ruby>: bold, italic and underline map to styles, the standard's colour classes such as <c.yellow> and <c.bg_blue> colour the text and its background, and the file's rules style classes and voices. Ruby keeps its reading after the base text in parentheses. Timestamp tags (<00:00:01.000>, karaoke) are stripped, because painting karaoke honestly is libass's job.

Functions

Link copied to clipboard

Parses text into cues, sorted by start time. Never throws on malformed input. A file with more than 100,000 cues keeps the first 100,000 in file order.

Link copied to clipboard

One cue's body from a container track, timing already on the packet, with no stylesheet.

Link copied to clipboard

A parser for the cues of one container track whose header is the start of a WebVTT file, up to its first cue, as Matroska keeps it in the track's private data. Its STYLE blocks style every cue of the track. A header with none, or none that reads, styles nothing.