# How we benchmark FFmpeg

Canonical: https://rendobar.com/blog/how-we-benchmark/
Author: Abdelrahman Essawy
Published: 2026-09-26
Updated: 2026-09-26

---

## Key takeaways

- Every benchmark post runs against the same public 5-second clip, so two numbers from different posts were measured on the same frames.
- Trust the size column to the byte. The same command wrote the same file every time it ran.
- Treat the time column as one draw from a wide distribution. Unless a post says it repeated the run with pinned threads, a gap under 2x is not a result.
- Cost is what the API billed for the whole job, so it carries the timing noise and reads best as a range.

Every benchmark on this blog is a sweep: one FFmpeg command run several times
through the Rendobar API, with one flag changed per run, against the same
5-second clip. The output size of each run is exact and repeatable. The encode
time and the cost are single samples unless a post says otherwise, and I do not
treat a gap under 2x between two single samples as a finding.

This page holds the setup so the individual posts do not have to repeat it.

## The source clip

All video, audio and still-image sweeps read
[`sample.mp4`](https://cdn.rendobar.com/assets/examples/sample.mp4), a public
file you can download and probe yourself.

| Property | Value |
|---|---|
| Container | MP4, 352,121 bytes |
| Video | H.264 High, 1280x720, 24 fps, 120 frames, yuv420p 8-bit, about 459 kbps |
| Audio | AAC-LC, stereo, 48 kHz, about 96 kbps |
| Duration | 5.0 s of video, 5.013 s in the container |
| Content | Animation that fades in from black, then a slow move across clouds and trees |

Four properties of this clip shape what the sweeps can show, and the posts that
depend on one of them say so.

- It is short. Fixed costs such as the first keyframe and the container
  header are a larger share of a 5-second file than of a 10-minute one, so
  percentages that depend on them come out larger here.
- It is already compressed. Every encode is a re-encode of H.264. The video
  is 4:2:0, so a 4:4:4 output holds interpolated chroma, not new colour detail.
  The audio is already 96 kbps AAC, so a higher-bitrate audio encode cannot
  restore what that encode discarded.
- It runs at 24 fps. Flags counted in frames, such as `-g 60`, cover
  2.5 seconds here, not the 2 seconds they would at 30 fps.
- It is animation. Smooth gradients and a slow camera compress more easily
  than grainy camera footage, so absolute sizes on your content will be larger.
  The ordering between settings is the part that tends to travel.

## What stays fixed

Each sweep changes exactly one thing. A video sweep runs a command of this shape
and varies a single flag:

```bash
ffmpeg -i https://cdn.rendobar.com/assets/examples/sample.mp4 \
  -c:v libx264 -preset medium -crf 23 -an -t 5 out.mp4
```

`libx264 -preset medium -crf 23` is the control, because those are x264's
defaults. `-an` strips the audio so the size
column is pure video, and audio sweeps use `-vn` the other way round. Every
benchmark table prints its own "Held constant" line under the title, so the
fixed settings for that sweep are always next to its numbers.

## Sizes are exact, timings are not

The same control command appears in ten sweeps. It wrote 443,351 bytes all ten
times, while its encode time ranged from 898 ms to 1,912 ms. The full record,
and the published claim it forced me to retract, are in
[FFmpeg benchmark variance](/blog/measurement-noise-ffmpeg-benchmarks/).

**A size needs one run. A timing needs repeats.** The encode time on every table
is the `execute-ffmpeg` step of the job alone, separated from the download and
the upload, but it still runs on a shared host whose spare CPU moves from run to
run. When I repeated a sweep five times with `-threads` pinned, each variant
landed within 1.4% to 8.9% of itself. With x264's default automatic threading,
the same work swung 85%. The output bytes stayed identical in every case, so
the swing is contention for the CPU, not a change in the work.

So the rule on every post is:

- A size difference from one run is a finding, quoted to the byte.
- A timing difference under 2x from one run is not a finding. Posts either say
  so or leave the claim out.
- A sweep whose conclusion depends on timing is repeated (`--repeat 3` or more)
  and reports a median with its range. The
  [SVT-AV1 preset sweep](/blog/svt-av1-presets-measured/) and the
  [thread-count sweep](/blog/ffmpeg-threads-measured/) are the two that do.

## How cost is measured

The cost column is what the API billed for the job, read from the job's `cost`
field. It is metered on the whole run, including the download and the upload,
not on the encode step alone. It therefore moves with time and inherits the
same noise: identical work billed anywhere from 0.23 to 0.39 US cents. Read one
job's cost as a point inside that kind of range, and compare costs only when
they differ by more than it.

## FFmpeg and encoder versions

The runner installs FFmpeg from
[BtbN's linux64-gpl build](https://github.com/BtbN/FFmpeg-Builds), the
`ffmpeg-master-latest` release. That is a build of FFmpeg's development branch,
not a tagged release, and it bundles the encoders the sweeps use: libx264,
libx265, libsvtav1, libaom, libvpx, libopus, libmp3lame and libwebp, plus
FFmpeg's own `aac`, `flac`, `mjpeg` and `png` encoders.

The sweeps did not record `ffmpeg -version`, so I cannot give you exact version
strings for them, and I will not guess. The nearest evidence is an output file
from the same build pipeline five days before the August sweeps. It carries the
muxer tag `Lavf63.5.101`, which places that build after FFmpeg 8.0 (whose
libavformat is 62.x). The runner picks up a newer build on each deploy, so a
rerun can use newer encoders than the published numbers did. Every table shows
the date it was measured for that reason.

You can read the build from any output file:

```bash
# The muxer version FFmpeg stamped into the file
ffprobe -v error -show_entries format_tags=encoder -of default=nw=1:nk=1 out.mp4

# libx264 also writes its own version string into the stream
strings out.mp4 | grep -m1 "x264 - core"
```

## Rerun a sweep

Each row of a benchmark table is one API call, so a sweep is a loop. The script
that produced the published numbers submits every variant as an `ffmpeg` job,
four at a time, and records the encode step's duration, the output size and the
billed cost. Four is not arbitrary: at eight in flight, polling hit the API rate
limit and silently lost results, so the script trades speed for complete data.

Here is the same loop for a three-point CRF sweep, run one job at a time:

const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });
const source = "https://cdn.rendobar.com/assets/examples/sample.mp4";

for (const crf of [18, 23, 28]) {
  const job = await rb.jobs.run({
    type: "ffmpeg",
    params: {
      command:
        "ffmpeg -i " + source +
        " -c:v libx264 -preset medium -crf " + crf + " -an -t 5 out.mp4",
    },
  });

  // The FFmpeg step alone, without download and upload.
  const encode = job.steps.find((s) => s.id === "execute-ffmpeg");
  const encodeMs =
    encode && encode.startedAt && encode.completedAt
      ? encode.completedAt - encode.startedAt
      : null;

  // cost can still be null for a moment after completion. Re-read the job
  // with rb.jobs.get(job.id) if you need it.
  console.log(crf, job.output?.file?.size, encodeMs, job.cost?.formatted);
}`}
  curl={`curl -X POST https://api.rendobar.com/jobs \\
  -H "Authorization: Bearer $RENDOBAR_API_KEY" \\
  -H "Content-Type: application/json" \\
  -d '{
    "type": "ffmpeg",
    "params": { "command": "ffmpeg -i https://cdn.rendobar.com/assets/examples/sample.mp4 -c:v libx264 -preset medium -crf 23 -an -t 5 out.mp4" }
  }'`}
/>

Swap the source URL for one of your own files and the same loop answers the
question for your content, which is the only answer that matters for your
library. If your question is about speed, add `-threads 4` to the command and run
each variant at least three times before comparing.

The part I would not skip is the size column. It costs one run per variant, it
does not move, and it is the number most encoding advice gets wrong.
