FFmpeg encoding settings, ranked by what they cost

Compare the FFmpeg settings that change output size, measured on one file: preset, keyframe interval, CRF, tune, scaler and pixel format, ranked by effect.

Share

-tune and the keyframe interval are the two x264 settings most likely to multiply your file size by accident, and CRF and the scaler flag are the cheapest decisions you can make. I ran every setting that changes output size through the same 720p source, so the results compare across sections. Sizes here are exact. Quality was not scored except where a section says so.

Codec choice matters more than any of these, but it has to be measured as quality at a fixed file size, not as bytes. Held to the same size, AV1 scored 20.5 VMAF points above H.264, and the video codec comparison has that test. The shared setup is on how we benchmark.

SettingLargest swing measured on this clipWhat it trades
Presetultrafast wrote 6.6x the bytes of veryslowencode time, and the middle is not ordered
Keyframe intervalall-keyframes wrote 3.4x the bytes of -g 12seek and segment granularity
CRF3.6x across CRF 18 to 32quality
-tunezerolatency added 121%latency, and nothing else
Scaler flag66% between the largest and smallestsharpness
Pixel formatyuv444p added 5% over yuv420pplayback compatibility

The settings that cost the most

Keyframe interval

A keyframe is a complete picture, and every other frame stores only what changed. The interval decides how often the encoder discards its compression history and starts over, so the cost compounds.

Keyframe interval
What does a shorter keyframe interval cost, and what does it buy?
Held constant: libx264 CRF 23 preset medium, same source, audio stripped
Bar chart. Keyframe interval. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
-g 12 (0.4s)836 ms778.7 KB$0.0024
-g 30 (1s)939 ms584.6 KB$0.0024
-g 60 (2s)960 ms504.6 KB$0.0024
-g 250 (default)945 ms433.0 KB$0.0024
all keyframes446 ms fastest2661.4 KB$0.0039
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 5 of 5 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

The labels assume 30 fps. This clip runs at 24 fps and is 120 frames long, so -g 30 and -g 60 put a keyframe every 1.25 and 2.5 seconds, and -g 250 never reaches a second interval (one keyframe for the clip, apart from any scene-cut keyframes). That row is a baseline no real video has, so compare the intervals to each other. Going from -g 60 to -g 30 added 16% to the file, -g 30 to -g 12 added 33%, and making every frame a keyframe wrote 3.4x the bytes of -g 12. More keyframes did not make the encode slower (the all-keyframes run was the quickest), so this is a bandwidth decision. For streaming I would use 2 seconds. For a download, keep the default interval. Detail in the keyframe interval sweep.

-tune is the most expensive single word

x264 tune presets
What does -tune actually change, and does it cost anything?
Held constant: libx264 CRF 23 preset medium, same source, audio stripped
Bar chart. x264 tune presets. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
no tune912 ms fastest433.0 KB$0.0025
film1502 ms446.6 KB$0.0030
animation2045 ms404.9 KB$0.0036
grain1133 ms450.0 KB$0.0026
stillimage959 ms536.6 KB$0.0023
fastdecode982 ms725.7 KB$0.0024
zerolatency1444 ms955.5 KB$0.0029
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 7 of 7 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

-tune zerolatency disables lookahead and B-frames and swaps frame threading for sliced threads, so the encoder never holds a frame back. Each of those features exists to improve compression. -tune fastdecode cost 68% for a different reason. It turns off CABAC entropy coding and the deblocking filter to make decoding cheaper.

Neither flag is misnamed. The trap is that “tune for low latency” reads like a scheduling hint and behaves like a bitrate multiplier. If your pipeline already buffers seconds of HLS segments, zerolatency buys you nothing. Detail in FFmpeg tune presets compared.

CRF is the cheapest decision available

x264 CRF
What does each step of CRF actually cost in bytes?
Held constant: libx264, preset medium, 5 seconds of the same source
Bar chart. x264 CRF. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
CRF 182216 ms665.8 KB$0.0038
CRF 201938 ms551.9 KB$0.0036
CRF 23982 ms433.0 KB$0.0025
CRF 262384 ms338.1 KB$0.0038
CRF 28967 ms282.5 KB$0.0025
CRF 301200 ms233.8 KB$0.0026
CRF 32854 ms fastest186.8 KB$0.0023
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 7 of 7 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

Each 3 CRF points removed about 20% of the remaining bytes, and moving from the default 23 to 26 removed 22%. Encode time showed no pattern across CRF values, so raising it is a pure size saving. What it costs in quality depends on the content, and this sweep measured bytes only. Detail in FFmpeg CRF explained and measured.

The settings people get wrong

Preset ordering is not what you think

x264 presets
Is a slower x264 preset worth the extra encode time?
Held constant: libx264, CRF 23, 5 seconds of the same 1280x720 source
Bar chart. x264 presets. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
ultrafast238 ms fastest2185.1 KB$0.0029
superfast388 ms732.2 KB$0.0024
veryfast542 ms388.4 KB$0.0021
faster737 ms440.0 KB$0.0022
fast985 ms463.1 KB$0.0025
medium1912 ms433.0 KB$0.0039
slow2450 ms422.8 KB$0.0040
slower2685 ms358.6 KB$0.0041
veryslow4635 ms329.5 KB$0.0062
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 9 of 9 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

The ends behave as advertised. The middle does not: veryfast wrote a smaller file than faster, fast, medium and slow.

Nothing is broken. CRF targets constant quality, so a slower preset can spend the bytes it saves on detail a faster one discarded. At fixed CRF, size is a side effect, not a dial. The same inversion shows up in SVT-AV1 presets compared, where preset 12 wrote a smaller file than preset 10.

Pick a preset by measuring your own content, and do not assume slower means smaller. The x264 preset sweep has the full curve.

The pixel format that breaks Safari

Pixel formats
What does chroma subsampling cost in bytes and compatibility?
Held constant: libx264 CRF 23 preset medium, same source, audio stripped
Bar chart. Pixel formats. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
yuv420p899 ms fastest433.0 KB$0.0024
yuv422p1152 ms475.0 KB$0.0026
yuv444p1487 ms456.7 KB$0.0031
yuv420p10le1257 ms467.9 KB$0.0029
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 4 of 4 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

yuv420p wrote the smallest file, and it is the format most hardware decoders expect. If a video plays in Chrome and not in Safari or QuickTime, check this first. A filter somewhere in the chain probably output a higher-chroma or RGB format and nothing converted it back. Ending a delivery chain with format=yuv420p does nothing when the format is already right. Detail in FFmpeg yuv420p vs yuv444p.

One oddity from that sweep: yuv422p wrote more bytes than yuv444p, which is not what the subsampling arithmetic predicts. Do not estimate output size from chroma volume.

The scaler flag is free, so use a good one

Scaling algorithms
Does the scaling flag matter, and what does the good one cost?
Held constant: 720p down to 480p, libx264 CRF 23 preset medium, audio stripped
Bar chart. Scaling algorithms. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
neighbor653 ms394.0 KB$0.0022
bilinear636 ms fastest236.9 KB$0.0035
bicubic (default)649 ms265.3 KB$0.0022
lanczos645 ms272.5 KB$0.0024
spline641 ms268.8 KB$0.0020
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 5 of 5 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

All five algorithms took the same time within 17 ms while output size varied 66%. Scaling is a per-pixel filter pass and encoding is a search, so the encoder dominates the clock.

Nearest-neighbor wrote 49% more bytes than bicubic, because its stair-stepped edges are high-frequency detail the encoder then pays to describe. That generalises: any filter that adds artificial sharpness costs encoder bytes.

Smaller is not better here. Bilinear wrote the smallest file because it is the blurriest. Bicubic, spline and lanczos landed within 3% of each other, and that is the cluster to choose from. Detail in FFmpeg lanczos vs bicubic scaling.

Rate control: quality, bitrate, or both

CRF against target bitrate
Should you ask for a quality or ask for a bitrate?
Held constant: libx264 preset medium, same source, audio stripped
Bar chart. CRF against target bitrate. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
CRF 23 (quality)1596 ms433.0 KB$0.0031
700k ABR948 ms fastest483.0 KB$0.0023
700k capped VBV1074 ms428.1 KB$0.0025
CRF 23 capped VBV1602 ms417.0 KB$0.0030
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 4 of 4 runs succeeded. Sizes are exact and repeatable. Timings are a single sample and vary up to 2x run to run, so treat small differences as noise. Why.

Single-pass ABR at 700k wrote 494,611 bytes where 700 kbit/s for five seconds predicts 437,500, overshooting its own target by 13%. It aims without knowing what is coming, so it corrects as it goes.

CRF with a VBV cap is the mode I would use for streaming. Quality drives the encode, and the cap bounds the peaks a player has to buffer:

Terminal window
-crf 23 -maxrate 900k -bufsize 1400k

Set both flags. x264 ignores -maxrate without -bufsize and logs a warning saying so, which makes it one of the most commonly half-configured things in FFmpeg. Detail in FFmpeg CRF vs bitrate.

Threads, and why your benchmarks disagree

Thread count
Where does adding encoder threads stop helping?
Held constant: libx264 CRF 23 preset medium, same source, audio stripped
Bar chart. Thread count. Encode time and output size for each variant, each metric scaled to its own maximum.
VariantEncodeSizeCost
1 thread3644 ms387.8 KB$0.0053
2 threads2122 ms387.9 KB$0.0036
4 threads1510 ms387.7 KB$0.0029
8 threads1220 ms387.2 KB$0.0026
auto (0)1033 ms fastest433.0 KB$0.0024
Measured 2026-08-20 on the Rendobar API. Encode time is the FFmpeg step alone, separated from download and upload. 5 of 5 runs succeeded. Sizes are exact and repeatable. Timings are the median of 5 runs, and cost is from the first run. Why.

Each pinned thread count ran five times. Eight threads was 2.99x faster than one, not 8x, because encoding has serial parts. Four threads reached 2.41x, which is most of the benefit.

The larger finding is about measurement. With -threads pinned, repeated runs landed within 1.4% to 8.9%. On auto, the identical command ranged from 918 to 1,699 ms, because x264 sizes itself to whatever the machine has free. That spread invalidated a timing claim I had already published, and FFmpeg benchmark variance documents the retraction. If your timings wander between runs, pin the threads before concluding anything.

What to set

For web delivery, this avoids every costly setting above:

Terminal window
ffmpeg -i input.mp4 \
-c:v libx264 -crf 23 -preset medium -g 60 \
-vf "scale=1280:-2:flags=lanczos,format=yuv420p" \
-c:a aac -b:a 128k \
-movflags +faststart \
output.mp4

CRF 23 is the quality anchor. -g 60 is a 2-second GOP at 30 fps. lanczos costs no measurable time. format=yuv420p prevents the Safari failure, and +faststart lets a browser start playing before the download finishes. Before any of it runs, validating the upload with ffprobe stops you paying for an encode that was always going to fail.

For archival, drop -g 60, use a slower preset, and look at AV1 at a matched size. To hit a size and a quality target at once, see how to compress video to a target size.

For live, -tune zerolatency is correct, and the extra bytes are the price of the latency you are buying.

Method and limits

Every row is one job run through Rendobar’s FFmpeg API against the same 5-second 1280x720 H.264 source, audio stripped, with the command recorded. Each variant was one call, which is what made a sweep of this size practical to repeat. The runner uses BtbN’s floating FFmpeg master static build, and exact x264 and FFmpeg versions were not recorded for this sweep. Timings are single samples except the threads sweep, so no timing difference under 2x on this page is a finding.

One clip is one content type. A static screencast or high-motion sports footage will move several of these numbers. The ordering and the mechanisms should travel. Keyframes are expensive, zerolatency disables compression features by design, and filters that add artificial detail cost bytes downstream. I would start any pipeline review by grepping for -tune.

Frequently asked questions

What are the best FFmpeg settings for web video?

libx264 at CRF 23, preset medium, a 2-second keyframe interval, format=yuv420p and -movflags +faststart. It plays everywhere, starts playing before the download finishes, and avoids the costly settings measured on this page.

Does a slower FFmpeg preset always produce a smaller file?

No. At a fixed CRF the encoder targets quality, so file size is a side effect. veryfast wrote a smaller file than faster, fast, medium and slow on this source.

Why does FFmpeg ignore -maxrate?

x264 ignores -maxrate unless -bufsize is also set, and it logs a warning saying so. Set both, for example -maxrate 900k -bufsize 1400k.

Sources

Tags #ffmpeg#x264#encoding#benchmarks#guide
All posts
Share
  1. Animated captions API, TikTok styles in one call Guides for the video API
  2. Compress video for Discord under 20 MB, every time Guides for the video API
  3. Compress video for WhatsApp under 16 MB Guides for the video API