# Animated captions API, TikTok styles in one call

Canonical: https://rendobar.com/blog/animated-captions-api/
Author: Abdelrahman Essawy
Published: 2026-09-27
Updated: 2026-09-27

---

## Key takeaways

- Animated captions are two problems, not one. You need a timestamp for every word, then a renderer that highlights each word on cue, and FFmpeg only solves the second.
- ASS karaoke tags do the rendering. \kf sweeps a colour across a word over a duration in centiseconds, and FFmpeg's ass filter burns the result into the picture.
- Raw word timestamps are the weak part. Speech recognisers report some words as a few hundredths of a second long and stretch others across a pause, and a naive script turns both straight into visible glitches.
- Grouping words into on-screen phrases is its own job. Splitting on a fixed word count strands single words on screen, so break at pauses and punctuation.
- captions.animate does transcription, phrasing and styling in one job, but only for English speech today. For RTL or other source languages, the ASS route is the one that works now.

An animated captions API takes a video with speech and returns it with the
words burned in and highlighted one at a time as they are spoken, the style
TikTok and Reels made standard. Rendobar's `captions.animate` does it in one
call: it transcribes the audio, splits the words into short phrases, and
renders them in a preset such as `hormozi`, `karaoke` or `pill`. Below is the
same clip in three presets, the FFmpeg route you would build yourself, and the
places that route breaks.

_One clip, one frame (13.0 s), three presets. Top to bottom: hormozi (job_25590f0ee8fe4858), karaoke (job_4d62ca0a850c40d0) and pill (job_034a09aeda314fd6)._

The source is 27.5 seconds of President Kennedy's 1962 Rice University speech,
starting at "not because they are easy". The recording is NASA's and in
the public domain, and I took it from the
[Wikimedia Commons copy](https://commons.wikimedia.org/wiki/File:President_Kennedy%27s_Speech_at_Rice_University.ogv).
It is 640x480, 1962 film audio, with a crowd in the room, so it is harder than
a podcast mic. To cut the same clip:

```bash
ffmpeg -ss 557.5 -i "President_Kennedy's_Speech_at_Rice_University.ogv" -t 27.5 \
  -c:v libx264 -crf 20 -pix_fmt yuv420p -c:a aac -b:a 128k \
  -movflags +faststart rice-moon.mp4
```

## What an animated caption needs

A static subtitle needs a line of text and a start and end time. An animated
caption needs a start and end time for every word, because the highlight moves
on each one. That splits the work in two:

1. Word timestamps, from a speech recogniser with word-level output, usually
   Whisper or one of its ports.
2. A renderer that can highlight a word on cue. In FFmpeg that means an ASS
   subtitle file and libass, through the `ass` or `subtitles` filter.

Static burning, where FFmpeg draws an SRT or ASS file onto the picture as-is,
has its own guide: [burning subtitles into a video with FFmpeg](/blog/ffmpeg-burn-subtitles/).
It covers the filter escaping and `force_style`. This post is about the moving
highlight.

## The DIY route, Whisper plus ASS karaoke

ASS has karaoke tags built in. `{\k50}` holds the next syllable for 50
centiseconds, then switches it from the style's `SecondaryColour` to its
`PrimaryColour`. `{\kf50}` sweeps the colour across the word from left to right
over the same half second, which is the look most people mean by karaoke
captions. `{\ko50}` does it with the outline instead.

Getting word timestamps from the reference Whisper looks like this. I did not
run this step for the post, because Whisper is not installed on the machine I
wrote it on:

```bash
pip install -U openai-whisper
whisper rice-moon.mp4 --model small.en --word_timestamps True --output_format json
```

The JSON it writes carries `segments[].words[]`, each with `word`, `start` and
`end` in seconds.

To show the rendering step, I wrote the first two lines of the clip by hand.
The timings below are typed in, rounded to centiseconds from the word
timestamps the Rendobar job returned, not produced by Whisper on my machine:

```text
[Script Info]
ScriptType: v4.00+
PlayResX: 640
PlayResY: 480

[V4+ Styles]
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
Style: Karaoke,Arial,34,&H0000FFFF,&H00FFFFFF,&H00000000,&H80000000,-1,0,0,0,100,100,0,0,1,3,0,2,20,20,150,1

[Events]
Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
Dialogue: 0,0:00:00.58,0:00:02.40,Karaoke,,0,0,0,,{\kf50}not {\kf33}because {\kf13}they {\kf46}are {\kf30}easy,
Dialogue: 0,0:00:02.42,0:00:04.80,Karaoke,,0,0,0,,{\kf42}but {\kf25}because {\kf12}they {\kf38}are {\k8}{\kf113}hard.
```

ASS colours are `&HAABBGGRR`, so `&H0000FFFF` is yellow and `&H00FFFFFF` is
white. `PlayResX` and `PlayResY` set the coordinate space the font size is
measured in. Match them to the video, or a size of 34 means something else at
render time. The `{\k8}` before "hard." is an 8-centisecond silent gap, because
`\k` durations run back to back and a pause has to be spelled out.

Then burn it:

```bash
ffmpeg -i rice-moon.mp4 -vf "ass=karaoke.ass" -c:v libx264 -crf 20 -c:a copy out.mp4
```

That ran on FFmpeg 8.0 with libass 0.17.4 and produced the sweep:

_The hand-timed \kf line at 1.3 s. The yellow fill is partway through because, the second word._

To go from Whisper's JSON to a full file, you group the words into lines and
write one `Dialogue` per line. This is the shortest version I would trust to
run, and I tested it by feeding it the job's word timestamps in Whisper's
shape:

```python
import json, sys

HEADER = open("header.ass", encoding="utf-8").read()  # the [Script Info] and [V4+ Styles] above

def ts(t):
    cs = round(t * 100)
    return f"{cs // 360000}:{cs // 6000 % 60:02}:{cs // 100 % 60:02}.{cs % 100:02}"

def lines(words, max_words=5, max_gap=0.3):
    line = []
    for w in words:
        if line and (len(line) == max_words or w["start"] - line[-1]["end"] > max_gap):
            yield line
            line = []
        line.append(w)
    if line:
        yield line

words = [w for seg in json.load(open(sys.argv[1]))["segments"] for w in seg["words"]]
out = [HEADER, "\n[Events]\nFormat: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text\n"]
for line in lines(words):
    text, cursor = "", line[0]["start"]
    for w in line:
        gap = round((w["start"] - cursor) * 100)
        if gap > 0:
            text += f"{{\\k{gap}}}"
        text += f"{{\\kf{max(round((w['end'] - w['start']) * 100), 1)}}}{w['word'].strip()} "
        cursor = w["end"]
    out.append(f"Dialogue: 0,{ts(line[0]['start'])},{ts(line[-1]['end'])},Karaoke,,0,0,0,,{text.strip()}\n")
open("captions.ass", "w", encoding="utf-8").write("".join(out))
```

It renders. It also shows every problem in the next section.

## Where the DIY route breaks

### Word timing

Raw word timestamps are not what a viewer hears. On this clip the recogniser
reported "easy," as 42 milliseconds long and "skills." the same, which a `\kf`
sweep turns into a flash. It went the other way on "hard.", which it gave 1.13
seconds. FFmpeg's `silencedetect` at -30 dB puts a 0.85 s pause right where
that word starts, so most of the sweep crawls through silence. Each of these
sat at the end of a clause, just before punctuation.

In my hand-written file I stretched "easy," to 30 centiseconds. The converter
passes the raw values through, which is safe to write and wrong to watch. Fixing it means rules about minimum durations and pauses, and checking
the audio levels to decide where a word really ends. Whisper's own CLI
labels `--word_timestamps` experimental, and this clip shows why.

### Line breaking

The converter splits a line at a 0.3 s gap or after five words. On this clip
that stranded "postpone," alone on screen for 0.12 seconds, because the
five-word limit fell one word before the end of the clause:

_17.87 s. The naive five-word split leaves one word on screen (top). The job keeps the clause together (bottom, job_25590f0ee8fe4858)._

The job's caption data split the same 27.5 seconds into 13 phrases of two to
six words, mostly ending at punctuation and pauses: "not because they are easy," then "but
because they are hard." Getting there takes punctuation, pause length and a
width budget in pixels, not a word count, because "unwilling" and "to" do not
take the same room. ASS will also wrap a line that is too wide by itself
(`WrapStyle` 0 balances the lines), but a wrapped karaoke line jumps mid-sweep,
so it is better to never hand libass a line that needs wrapping.

### Fonts

The styles people want use display faces such as Bebas Neue or Montserrat
Black, which are not on a render server by default. libass does not fail when
a font is missing. I changed `Fontname` to `Bebas Neue` on a machine without
it, and FFmpeg rendered the file anyway, in Arial:

```text
[Parsed_ass_0 @ 000001a30dde5980] fontselect: (Bebas Neue, 700, 0) -> Arial-BoldMT, 0, Arial-BoldMT
```

The fix is to ship the font file and point the filter at it with
`ass=karaoke.ass:fontsdir=fonts`, and to use the family name stored inside the
file, which is not always the filename.
[Why custom fonts in a video API fail silently](/blog/custom-fonts-video-api/)
goes through that failure in detail.

### RTL and other scripts

Arabic, Hebrew and Persian need text shaping and bidirectional layout. The
`ass` filter has a `shaping` option, and complex shaping (HarfBuzz) has to be
in effect for joined scripts to render correctly. You also need a font that
covers the script. The karaoke tags on top of that are the part I would test
by eye before trusting on right-to-left text. I did not test RTL for this
post.

## The one-call version

Three runs of `captions.animate` on the same clip, one per preset in the
figure above:

| Preset | Look | Job | Cost |
|---|---|---|---:|
| `hormozi` | Bold capitals, spoken word in yellow | job_25590f0ee8fe4858 | $0.0106 |
| `karaoke` | Two to three words at a time, spoken word in green | job_4d62ca0a850c40d0 | $0.0098 |
| `pill` | Sentence case, a box behind the spoken word | job_034a09aeda314fd6 | $0.0097 |

Each was one request with the video as the only input. The hormozi run also set
`captionData: true`, which returns the word timestamps, an SRT and a VTT
alongside the video. That is where the timings in the DIY section came from.
Other presets include `mrbeast`, `tiktok`, `word-pop`, `neon`, `gold` and
`reveal`, and a `style` object overrides the font, size, colours, highlight and
position on top of any of them. Fonts take any Google Fonts family by name, or
your own file as the `font` input.

const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });

const job = await rb.jobs.create({
    type: "captions.animate",
    inputs: { source: "https://example.com/rice-moon.mp4" },
    params: {
      preset: "hormozi",
      // Also return the word timestamps, an SRT and a VTT on output.data.
      captionData: true,
    },
});

console.log(job.id);`}
  curl={`curl -X POST https://api.rendobar.com/jobs \\
    -H "Authorization: Bearer $RENDOBAR_API_KEY" \\
    -H "Content-Type: application/json" \\
    -d '{
      "type": "captions.animate",
      "inputs": { "source": "https://example.com/rice-moon.mp4" },
      "params": { "preset": "hormozi", "captionData": true }
    }'`}
/>

**It is English-only.** The `language` parameter accepts `en` and nothing
else today, and the default and fast modes run English-only speech models. `translateTo`
can turn English speech into Spanish, French, German, Portuguese, Italian or
Dutch captions, all Latin script. There is no way to caption Arabic, Hebrew,
Hindi or Chinese speech with this job right now, and no RTL output. For those,
the Whisper and ASS route above is the one that works today. You can burn
that ASS file with your own FFmpeg, or pass it as an inline input to
Rendobar's [FFmpeg API](/ffmpeg/) and run the same `ass` filter there.

The job does not make word timing perfect. It uses the same class of
recogniser, and its timestamps for "easy," and "hard." are the ones quoted
above. What it saves you is the rest: installing and running a model, the
phrase splitting, the font files and the preset styling, which is most of the
code in the DIY route. On this 640x480 source the presets also rendered
smaller than I would ship, and `style.fontSize` is the override for that.

If your speech is English, I would use the job and spend the saved time
reviewing the output. If it is not, write the ASS file yourself, with
phrase breaks at punctuation, a minimum duration on every word, and the font
file shipped next to it.
