Why FFmpeg benchmark times vary run to run

Learn why one FFmpeg command gave byte-identical output ten times while its encode time swung 2.13x, and how pinning -threads makes the timings repeatable.

Share

FFmpeg benchmark times vary run to run because x264’s default threading competes for CPU with whatever else shares the host, not because the encode changes. I ran one libx264 command ten times and got the same 443,351 bytes every time, with encode times anywhere from 898 ms to 1,912 ms. Pin -threads and repeat the run: pinned runs landed within 1.4% to 8.9% of each other, while the default auto setting swung 85% on identical work.

This post also retracts two things I published: a speed claim about NVENC, and my first explanation of where the noise comes from. The shared test setup is on how we benchmark.

One command, ten runs, one output size

Terminal window
ffmpeg -i sample.mp4 -c:v libx264 -crf 23 -preset medium -an -t 5 out.mp4

This is the control in ten of my sweeps. Six rows run it exactly. Four add a flag that equals the default on this source (-vf scale=1280:-2 on a 1280x720 input, -vf format=yuv420p, -g 250 and -threads 0), so all ten reduce to the same encode.

SweepEncodeCostOutput
presets / medium1,912 ms$0.0039443,351 B
ratecontrol / CRF 231,596 ms$0.0031443,351 B
threads / auto1,033 ms$0.0024443,351 B
crf / CRF 23982 ms$0.0025443,351 B
gop / -g 250945 ms$0.0024443,351 B
ladder / 720p944 ms$0.0023443,351 B
gpu / libx264 medium913 ms$0.0030443,351 B
tune / no tune912 ms$0.0025443,351 B
pixfmt / yuv420p899 ms$0.0024443,351 B
codecs / H.264898 ms$0.0024443,351 B

The output never moved. The clock moved by 2.13x.

The threads row is the median of that sweep’s five runs, which is what the committed measurement file records. The other nine rows are single runs. An earlier version of this table showed 924 ms and $0.0023 for the threads row, taken from a first single-run pass of that sweep, so it disagreed with the file the page is built from. It now matches the file. The band is the same either way.

Why the bytes match and the clock does not

x264 with fixed settings and a fixed thread count is deterministic. Same input, same flags, same thread count, same bytes. That is why the size column on every benchmark here can be trusted from one run.

The thread count matters to that sentence. In the FFmpeg threads benchmark, one thread wrote 397,121 bytes and the default wrote 443,351. So ten identical outputs on the default setting also tell you something else: x264 started the same number of threads every time.

Cost is noisier than encode time. Billing covers the whole run (download, encode, upload and job overhead), not the FFmpeg step alone. The table shows it: 913 ms of encoding cost $0.0030 because that job’s upload was slow, while 982 ms cost $0.0025. Identical work cost $0.0023 to $0.0039.

Pinned threads make the clock repeatable

Repeating the runs was the obvious next step, so I repeated a timing sweep five times per variant. The result pointed at the cause.

ThreadsMedianRangeSpread
13,644 ms3,590 to 3,6571.9%
22,122 ms2,108 to 2,1813.5%
41,510 ms1,490 to 1,5111.4%
81,220 ms1,183 to 1,2888.9%
auto1,033 ms918 to 1,69985%

With the thread count pinned, five runs of the same encode land within a few percent. The wide variance lives almost entirely in auto.

On auto, x264 sizes its thread pool from the CPUs it can see, about 1.5 threads per core. On a shared host that is likely more threads than the container’s CPU share. The work is identical every run, as the bytes prove. What changes is how quickly those threads get scheduled, and that depends on what else the host is running at that moment. A pinned, smaller thread count asks for less than the share, so the neighbours matter much less.

Correction: my first explanation was wrong

The first version of this post said: “x264’s default threading adapts to the machine: it asks how many cores are available and sizes itself accordingly. On shared infrastructure that answer changes between runs, so the encoder is doing a different amount of parallel work each time.”

That does not fit the data. A different thread count writes different bytes, and all ten outputs were identical, so the thread count did not change. The workload was the same every run. The contention was not. The fix (pin -threads) was right, and the reason I gave for it was wrong.

What the rule retracted

I had published, in NVENC vs libx264, that hardware encoding was not faster than libx264 on a short clip. The evidence was NVENC at 1,225 ms against libx264 at 913 ms.

913 ms is the bottom of a band that runs to 1,912 ms, and NVENC’s 1,225 ms sits inside it. These samples cannot separate the two, and that page now says so.

The cost claim on the same page survived in a weaker form. I had said hevc_nvenc was 3.8x cheaper. Against a CPU cost band of $0.0023 to $0.0039, its $0.0008 is 2.9x to 4.9x cheaper, so the direction held and the decimal did not. The like-for-like h264_nvenc comparison is narrower still, and the NVENC post now leads with it.

The rule I use now

If a conclusion depends on timing, pin -threads and repeat the run. If it depends on size, one run is enough, because size is deterministic. Applied to what is already published:

  • ultrafast at 238 ms against veryslow at 4,635 ms is a 19x gap. It stands.
  • Adjacent x264 presets a few hundred milliseconds apart cannot be separated from one sample. That page reads the size column only.
  • The scaling-filter sweep, where five algorithms landed within 17 ms of each other, already said “no measurable difference in time”. The noise band makes that conclusion stronger.
  • Every size comparison on the site stands unchanged.

Running these sweeps through the API made the check cheap: every variant is one job, and each job records the encode step separately from download and upload, which is what exposed the noise in the first place.

I would read any benchmark that does not state its run count as one sample, including the ones on this site from before this post. A page that shows its own correction is worth more to a reader than one that never checked.

Frequently asked questions

Does the thread count change FFmpeg's output?

Yes, for libx264. The same CRF 23 encode wrote 397,121 bytes with one thread and 443,351 bytes on the default auto setting. Within one thread count the output never changed across repeated runs.

How many times should I repeat an encoding benchmark?

Once for size, because the encoder is deterministic. For timing, pin -threads and repeat it, five runs in my case, then report the median and the range. Without pinned threads, repeating mostly measures the neighbours.

Sources

Tags #ffmpeg#benchmarks#methodology#x264
All posts
Share
  1. How we benchmark FFmpeg Engineering blog
  2. Custom fonts in a video API fail silently Engineering blog
  3. Opus vs AAC vs MP3, requested vs delivered bitrate Engineering blog