How we benchmark FFmpeg
Learn how the FFmpeg benchmarks on this blog are run: the source clip, what stays fixed, why sizes are exact and timings are not, and how to rerun a sweep.
Every benchmark on this blog is a sweep: one FFmpeg command run several times through the Rendobar API, with one flag changed per run, against the same 5-second clip. The output size of each run is exact and repeatable. The encode time and the cost are single samples unless a post says otherwise, and I do not treat a gap under 2x between two single samples as a finding.
This page holds the setup so the individual posts do not have to repeat it.
The source clip
All video, audio and still-image sweeps read sample.mp4, a public file you can download and probe yourself.
| Property | Value |
|---|---|
| Container | MP4, 352,121 bytes |
| Video | H.264 High, 1280x720, 24 fps, 120 frames, yuv420p 8-bit, about 459 kbps |
| Audio | AAC-LC, stereo, 48 kHz, about 96 kbps |
| Duration | 5.0 s of video, 5.013 s in the container |
| Content | Animation that fades in from black, then a slow move across clouds and trees |
Four properties of this clip shape what the sweeps can show, and the posts that depend on one of them say so.
- It is short. Fixed costs such as the first keyframe and the container header are a larger share of a 5-second file than of a 10-minute one, so percentages that depend on them come out larger here.
- It is already compressed. Every encode is a re-encode of H.264. The video is 4:2:0, so a 4:4:4 output holds interpolated chroma, not new colour detail. The audio is already 96 kbps AAC, so a higher-bitrate audio encode cannot restore what that encode discarded.
- It runs at 24 fps. Flags counted in frames, such as
-g 60, cover 2.5 seconds here, not the 2 seconds they would at 30 fps. - It is animation. Smooth gradients and a slow camera compress more easily than grainy camera footage, so absolute sizes on your content will be larger. The ordering between settings is the part that tends to travel.
What stays fixed
Each sweep changes exactly one thing. A video sweep runs a command of this shape and varies a single flag:
ffmpeg -i https://cdn.rendobar.com/assets/examples/sample.mp4 \ -c:v libx264 -preset medium -crf 23 -an -t 5 out.mp4libx264 -preset medium -crf 23 is the control, because those are x264’s defaults. -an strips the audio so the size column is pure video, and audio sweeps use -vn the other way round. Every benchmark table prints its own “Held constant” line under the title, so the fixed settings for that sweep are always next to its numbers.
Sizes are exact, timings are not
The same control command appears in ten sweeps. It wrote 443,351 bytes all ten times, while its encode time ranged from 898 ms to 1,912 ms. The full record, and the published claim it forced me to retract, are in FFmpeg benchmark variance.
A size needs one run. A timing needs repeats. The encode time on every table is the execute-ffmpeg step of the job alone, separated from the download and the upload, but it still runs on a shared host whose spare CPU moves from run to run. When I repeated a sweep five times with -threads pinned, each variant landed within 1.4% to 8.9% of itself. With x264’s default automatic threading, the same work swung 85%. The output bytes stayed identical in every case, so the swing is contention for the CPU, not a change in the work.
So the rule on every post is:
- A size difference from one run is a finding, quoted to the byte.
- A timing difference under 2x from one run is not a finding. Posts either say so or leave the claim out.
- A sweep whose conclusion depends on timing is repeated (
--repeat 3or more) and reports a median with its range. The SVT-AV1 preset sweep and the thread-count sweep are the two that do.
How cost is measured
The cost column is what the API billed for the job, read from the job’s cost field. It is metered on the whole run, including the download and the upload, not on the encode step alone. It therefore moves with time and inherits the same noise: identical work billed anywhere from 0.23 to 0.39 US cents. Read one job’s cost as a point inside that kind of range, and compare costs only when they differ by more than it.
FFmpeg and encoder versions
The runner installs FFmpeg from BtbN’s linux64-gpl build, the ffmpeg-master-latest release. That is a build of FFmpeg’s development branch, not a tagged release, and it bundles the encoders the sweeps use: libx264, libx265, libsvtav1, libaom, libvpx, libopus, libmp3lame and libwebp, plus FFmpeg’s own aac, flac, mjpeg and png encoders.
The sweeps did not record ffmpeg -version, so I cannot give you exact version strings for them, and I will not guess. The nearest evidence is an output file from the same build pipeline five days before the August sweeps. It carries the muxer tag Lavf63.5.101, which places that build after FFmpeg 8.0 (whose libavformat is 62.x). The runner picks up a newer build on each deploy, so a rerun can use newer encoders than the published numbers did. Every table shows the date it was measured for that reason.
You can read the build from any output file:
# The muxer version FFmpeg stamped into the fileffprobe -v error -show_entries format_tags=encoder -of default=nw=1:nk=1 out.mp4
# libx264 also writes its own version string into the streamstrings out.mp4 | grep -m1 "x264 - core"Rerun a sweep
Each row of a benchmark table is one API call, so a sweep is a loop. The script that produced the published numbers submits every variant as an ffmpeg job, four at a time, and records the encode step’s duration, the output size and the billed cost. Four is not arbitrary: at eight in flight, polling hit the API rate limit and silently lost results, so the script trades speed for complete data.
Here is the same loop for a three-point CRF sweep, run one job at a time:
Run a small CRF sweep and print size, encode time and cost per variant
import { createClient } from "@rendobar/sdk";
const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });const source = "https://cdn.rendobar.com/assets/examples/sample.mp4";
for (const crf of [18, 23, 28]) {const job = await rb.jobs.run({ type: "ffmpeg", params: { command: "ffmpeg -i " + source + " -c:v libx264 -preset medium -crf " + crf + " -an -t 5 out.mp4", },});
// The FFmpeg step alone, without download and upload.const encode = job.steps.find((s) => s.id === "execute-ffmpeg");const encodeMs = encode && encode.startedAt && encode.completedAt ? encode.completedAt - encode.startedAt : null;
// cost can still be null for a moment after completion. Re-read the job// with rb.jobs.get(job.id) if you need it.console.log(crf, job.output?.file?.size, encodeMs, job.cost?.formatted);}Install with npm i @rendobar/sdk. jobs.run() submits and waits, so it returns the finished job in one call.
curl -X POST https://api.rendobar.com/jobs \-H "Authorization: Bearer $RENDOBAR_API_KEY" \-H "Content-Type: application/json" \-d '{ "type": "ffmpeg", "params": { "command": "ffmpeg -i https://cdn.rendobar.com/assets/examples/sample.mp4 -c:v libx264 -preset medium -crf 23 -an -t 5 out.mp4" }}'Returns immediately with a job id. Poll GET /jobs/{id} or register a webhook rather than blocking on the request.
Swap the source URL for one of your own files and the same loop answers the question for your content, which is the only answer that matters for your library. If your question is about speed, add -threads 4 to the command and run each variant at least three times before comparing.
The part I would not skip is the size column. It costs one run per variant, it does not move, and it is the number most encoding advice gets wrong.
