SubRipParser

Reads SubRip (.srt) subtitles.

SubRip has no specification. What exists is twenty years of files written by dozens of tools, so this parser is written to accept what real files contain rather than what a grammar would allow. Every tolerance below corresponds to files that exist in the wild.

Accepted deviations from the common shape:

  • A missing or non-numeric index line. The index is ignored entirely, because it is wrong in many files and nothing depends on it.

  • Comma or full stop as the millisecond separator. Both appear.

  • One or two digit hours, and a missing hour field.

  • Windows, Unix or old Mac line endings, mixed within one file.

  • A byte order mark at the start.

  • Blank lines inside a cue's text, when the following line is not a timing line.

  • Cues out of chronological order. The result is sorted by start time.

  • A final cue with no trailing blank line.

The inline markup SubRip files carry in practice is HTML-like: bold, italic, underline, strike and a font colour. Those are parsed into StyledSpans. An unknown HTML-like tag is passed through as literal text rather than dropped, because dropping text loses meaning and showing a stray tag only looks untidy.

Many files also carry ASS override tags in braces, above all {\an8} to lift a line to the top when burned-in text covers the bottom. Those are tags, as FFmpeg reads them: the first \an1 to \an9 places the cue, \b, \i, \u and \s set the style, and every other {\...} run is dropped.

Properties

Link copied to clipboard

A cue whose end does not follow its start would otherwise never display: the selector requires the time to sit strictly before the end. The parser resolves it against the next cue's start, or holds it for this documented default when no cue follows.

Functions

Link copied to clipboard

Parses text into cues, sorted by start time. Never throws on malformed input. A file with more than 100,000 cues keeps the first 100,000 in file order.

Link copied to clipboard
fun parseCue(body: String, startMicros: Long, endMicros: Long): SubtitleCue.Text?

One cue from its body and its timing, placed where a {\anN} tag in the body asks. This is what a Matroska SubRip packet needs, because parseCueBody answers the text alone. Null when the body holds no text.

Link copied to clipboard

Parses ONE cue's body, the shape a Matroska SubRip track's packets carry: the text alone, timing already on the packet. Same markup rules as whole-file parsing.