How to compress video to a target size with FFmpeg
Convert a file size target into FFmpeg settings with the two-pass formula or a CRF search, and see how little already-encoded video shrinks.
To compress a video to a target size with FFmpeg, divide the target size in kilobits by the duration in seconds, subtract the audio bitrate, and run a two-pass encode at that video bitrate. That hits the size and ignores how the result looks. If you want a size and a quality floor together, it takes several measured encodes, and the data below shows why and what they found.
# 10 MB target, 60 second video, 128k audio# (10 * 8192) / 60 - 128 = 1237 kbit/s video# Use the same -preset in both passes. On Windows, replace /dev/null with NUL.ffmpeg -y -i input.mp4 -c:v libx264 -preset medium -b:v 1237k -pass 1 -an -f null /dev/nullffmpeg -i input.mp4 -c:v libx264 -preset medium -b:v 1237k -pass 2 -c:a aac -b:a 128k output.mp4If you want it small and still good, and the exact size is flexible, use CRF and stop specifying a size:
ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k output.mp4The formula, and what it does not know
Two-pass distributes bits across the timeline well, but the total is fixed by you, not by the content. A static screencast and a handheld shot of falling leaves get the same budget for the same duration. One of them wastes most of it and the other cannot look acceptable.
A bitrate target answers “how big” and refuses to answer “how does it look”. CRF answers “how does it look” and refuses to answer “how big”. Neither one answers both, which is why hitting a size with predictable quality takes more than one encode.
Searching CRF against a size yourself
Size at a given CRF depends on the content, so the practical DIY method is to search. This bisects CRF on a 10 second sample, scales each sample’s size up to the full duration, and returns the lowest CRF (the best quality) that fits:
in=input.mp4target_kb=3000 # budget for the video stream, in KBdur=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$in")lo=17; hi=52 # 52 means "no CRF fits"while [ $((hi - lo)) -gt 1 ]; do crf=$(( (lo + hi) / 2 )) # Encode a 10 s sample from 20 s in, scale its size to the full duration. ffmpeg -v error -y -ss 20 -t 10 -i "$in" -c:v libx264 -preset medium -crf "$crf" -an sample.mp4 est_kb=$(awk -v b="$(wc -c < sample.mp4)" -v d="$dur" 'BEGIN { printf "%d", b / 1024 * d / 10 }') echo "crf $crf -> about $est_kb KB" if [ "$est_kb" -gt "$target_kb" ]; then lo=$crf; else hi=$crf; fidone[ "$hi" -eq 52 ] && echo "No CRF fits. Lower the resolution." || echo "Use -crf $hi"On a synthetic 40 second test pattern it settled on CRF 35 after five sample encodes, and the full encode at CRF 35 came out 3% over the estimate, so leave a margin. One sample misses scenes that are harder than the one you picked. The script also measures size only. It cannot tell you whether CRF 35 looks acceptable, which is the question a quality metric answers.
What a measured search found
Rendobar’s compress.target job runs that search with a quality metric on every candidate. Each probe is a real encode of a sample at a candidate quality level, scored against the original with VMAF, and the search moves based on what it measured. Every job reports its own probe count, metric, score and verdict, so the figures below come from reading finished jobs.
These are the video jobs in the archive as of 2026-08-20 (outputs in H.264, HEVC or AV1), about half of all compression jobs. Many are repeat runs against the same short test clip, described in how we benchmark, so read the shape more than the exact values.
| On video jobs | Median | 90th percentile | Highest |
|---|---|---|---|
| Probe encodes per job | 6 | 7 | |
| Compression ratio (input size / output size) | 1.39x | 3.08x | 6.48x |
The median quality landed at VMAF 88.5, with a top score of 96.3. The median job cost $0.0160, since six encodes plus scoring is more work than one encode.
At the 90th percentile a job cost $0.0466.
Most video compresses far less than you expect
A 1.39x ratio means the output was about 28% smaller than the input. Not 5x, not 10x. This is what happens when the input is already an H.264 file that another encoder optimised. The tail comes from sources that were encoded wastefully to begin with.
If your mental model is “compression makes videos much smaller”, it was built on unoptimised sources. On video that is already sensibly encoded, expect tens of percent.
Sometimes the right answer is not to encode
About one video job in five returned passthrough. The search evaluated candidates, found that none beat the original at acceptable quality, and handed back the source untouched.
Re-encoding an already-efficient file can make it larger, because you pay a fresh generation of lossy coding on top of the old one. A compressor that always produces output will hand you a bigger file and call it success. So whether or not you use a service: compare the output to the input, and keep the smaller one. It is three lines, and the compressor that refuses to compress covers the behaviour in detail.
What a hard byte ceiling costs
The lowest scores, down to VMAF 30.7, all came from jobs given a byte ceiling of 50 to 150 KB for a clip that needed more. When the size is non-negotiable, quality pays for it, and no setting avoids the trade. Measuring tells you before you ship.
Here is that trade on the same five seconds, at an 8x difference in bytes:
Play both at full size. To my eye the gap is smaller in motion than on a paused frame, which is why picking a target from a still goes wrong in both directions.
The codec follows the content
On video outputs the search picked H.264 for about 82% of jobs, AV1 for about 12% and HEVC for about 5%. AV1 compresses best of the three, and it was chosen rarely because it is slow and most of these short clips did not need it. With three HEVC picks the per-codec split is an anecdote, not a ranking. Video codec comparison compares codecs properly on matched inputs.
The same job also compresses images and audio, which the figures above leave out. Across every media type, the pick shifted to AVIF for about a third of jobs and Opus for about 8%, and about one job in six returned passthrough. The largest ratio in that wider set, 825.7x, was an audio job. An earlier version of this post called it an image, which was wrong. Image jobs are scored with SSIMULACRA2, which is on a different scale from VMAF, so the two are not averaged together here.
Doing it over HTTP
This runs the size-target version of the search on the sample clip, with a 150 KB ceiling:
import { createClient } from "@rendobar/sdk";
const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });
const job = await rb.jobs.run({type: "compress.target",inputs: { source: "https://cdn.rendobar.com/assets/examples/sample.mp4" },// { maxBytes } when the size is fixed. A named posture or a 1-100 number// targets quality instead.params: { target: { maxBytes: 153600 }, for: "web" },});
// output.data is typed unknown in the SDK. The job reports its own search.const d = job.output.data as { verdict: string; codec: string; probes: number; achievedScore: number };
// "passthrough" means nothing beat the original.console.log(d.verdict, d.codec, d.probes, d.achievedScore);Install with npm i @rendobar/sdk. jobs.run() submits and waits, so it returns the finished job in one call.
curl -X POST https://api.rendobar.com/jobs \-H "Authorization: Bearer $RENDOBAR_API_KEY" \-H "Content-Type: application/json" \-d '{ "type": "compress.target", "inputs": { "source": "https://cdn.rendobar.com/assets/examples/sample.mp4" }, "params": { "target": { "maxBytes": 153600 }, "for": "web" }}'Returns immediately with a job id. Poll GET /jobs/{id} or register a webhook rather than blocking on the request.
target also takes a named posture such as "balanced" or a quality number from 1 to 100. Setting dryRun: true returns the predicted size, quality and cost without encoding. About one job in ten in this sample was a dry run, so its figures are predictions.
Every quality figure here is a metric score, not a panel of viewers, and metrics disagree with eyes often enough that this blog has published a case where they ranked the visibly worse option higher. My advice holds either way: when the size is flexible, use CRF, and when it is fixed, measure quality on every candidate and keep the original if nothing beats it.
Frequently asked questions
What is a good VMAF score for compressed video?
VMAF runs from 0 to 100. A common rule of thumb treats scores above about 93 as visually indistinguishable from the source and above 80 as good for streaming, but the threshold depends on screen size and viewing distance, so treat it as a guide, not a guarantee.
Why is my compressed video larger than the original?
Re-encoding an already-efficient file adds a fresh generation of lossy coding on top of the existing one, and the new encoder does not know the old one already threw that detail away. Keep whichever file is smaller.
