# FFmpeg vs video API, one edit rendered three ways

Canonical: https://rendobar.com/blog/ffmpeg-vs-video-api/
Author: Abdelrahman Essawy
Published: 2026-09-27
Updated: 2026-09-27

---

## Key takeaways

- The choice is not FFmpeg or a template API. An FFmpeg API that also runs JSON timelines gives you the command you already wrote and a layout format for the edits that need one, on one account.
- On one machine, one render, local FFmpeg wins on speed. What a hosted API sells is everything around the encode: queueing, long jobs, retries, fonts and a URL at the end.
- Describe the edit as a timeline when positioning and layering are the hard part, and send the command when you already know it. The timeline cost several times more for the same 4-second edit.
- Template APIs price by the minute of output on top of a base fee or a monthly bucket. Measure your own clip lengths before comparing, because short renders are where fixed costs dominate.

There are three ways to render video from code, not two: run FFmpeg yourself,
call a template or timeline API, or call an FFmpeg API that also runs timelines.
I rendered one edit (trim, upscale to 1080p, logo in the corner, a title bar)
locally, as an FFmpeg command on an API and as a JSON timeline on the same API,
three times each. The hosted FFmpeg command cost under a cent a
render and the same edit as a timeline cost several times that, while local
FFmpeg beat both on speed and left me owning the queue, the retries and the
server.

## Three routes, not two

The usual framing, including Shotstack's own
[FFmpeg vs video API](https://shotstack.io/learn/ffmpeg-vs-video-api/) guide,
puts raw FFmpeg on one side and a managed template API on the other. You get
full control of codecs and filters on your own infrastructure, or you get
rendering infrastructure and a JSON template format with no FFmpeg in sight. It
treats those as a package deal.

They are two separate decisions. Who runs the servers is one. Whether you
describe the edit as an FFmpeg command or as a timeline is the other. A third
kind of service answers the first with "we do" and the second with "either":

| | Who runs it | How you describe the edit |
|---|---|---|
| Self-hosted FFmpeg | You | FFmpeg command |
| Template or timeline API (Shotstack, Creatomate) | The vendor | Their JSON timeline or template |
| FFmpeg API with timelines (Rendobar) | The vendor | FFmpeg command, or a JSON timeline |

[Video processing APIs fall into four categories](/blog/video-processing-api/)
maps the wider market. This post measures the three routes above on one job.

## The edit as an FFmpeg command

The job is a realistic social cut: take 4 seconds of a clip starting at 1
second, scale it to 1920x1080, put a logo in the top right and a title on a
translucent bar near the bottom. The source is
[`sample.mp4`](https://cdn.rendobar.com/assets/examples/sample.mp4), a public
5-second 1280x720 H.264 clip at 24 fps, and the logo is a 6000x6000 PNG of the
Rendobar mark.

```bash
ffmpeg -ss 1 -t 4 -i sample.mp4 -i logo.png -filter_complex \
  "[0:v]scale=1920:1080:flags=lanczos,drawtext=fontfile=/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf:text='Launch week':fontsize=72:fontcolor=white:box=1:boxcolor=black@0.5:boxborderw=20:x=(w-text_w)/2:y=h-th-80[bg];[1:v]scale=160:-1[lg];[bg][lg]overlay=W-w-48:48[v]" \
  -map "[v]" -map 0:a -c:v libx264 -preset veryfast -crf 23 \
  -c:a aac -b:a 128k -movflags +faststart out.mp4
```

`-ss` before `-i` seeks the input, `-t 4` caps it at 4 seconds, `drawtext` burns
the title with its box, and the logo is scaled to 160 px wide before `overlay`
places it 48 px in from the top right corner.

I ran it three times on my desktop, an AMD Ryzen 5 7600 (6 cores, 12 threads) on
Windows 11 with the gyan.dev FFmpeg 8.0 essentials build, inputs already on
local disk. Each run took 823 to 866 ms. Windows has no DejaVu at that path, so
the local runs pointed `fontfile` at Arial, which is why their output size is
not comparable with the API runs below.

Under a second is the number self-hosting is chosen for, and it is real. It is
also the cost of the encode alone, with the file already on the machine and
nobody else's job in the way.

## The same edit on the API, as a command and as a timeline

Then I ran the edit on Rendobar six times on 2026-09-27: the command above
unchanged as an `ffmpeg` job, and the same layout written as a `compose`
timeline, three runs each. Both read the source and logo from their public URLs.
Time here is the API's own measure, from submit to a finished, downloadable
output, so it includes fetching inputs, starting the render and storing the file.

| Route | Runs | Time to finished file | Cost per render | Output bytes |
|---|---:|---:|---:|---:|
| Local FFmpeg, Ryzen 5 7600 | 3 | 0.82 to 0.87 s | your own hardware | not comparable (Arial) |
| Rendobar, `ffmpeg` command | 3 | 5.4 to 15.9 s | $0.0046 to $0.0077 | 998,390 every run |
| Rendobar, `compose` timeline | 3 | 30.9 to 67.5 s | $0.030 to $0.034 | 966,683 every run |

Two things in that table are solid and one is noise. The byte counts repeated
exactly on every run of each route, because FFmpeg is deterministic. The cost
gap holds too. Across three runs each, the timeline averaged about 5.4 times
the command's cost, because Rendobar bills the seconds of compute a job uses and
every timeline render ran longer. The timings swung nearly 3x within the command
route alone, so read them as a range, not a ranking.

_The same edit from the ffmpeg command (left, job_1d851195ceb6483c) and the compose timeline (right, job_a2c21cd7e7754f9e), 0.5 s in. Logo and title land in the same places._

The frames line up. The logo sits in the same corner at the same size and the
title bar lands in the same spot. The typeface differs, and that is on me: I
named DejaVu Sans in the timeline, the compose engine loads fonts by Google Fonts
family name, DejaVu is not one, and the render drew a fallback without an error.
The FFmpeg route has the same trap in a different form, as
[custom fonts in a video API](/blog/custom-fonts-video-api/) showed. A frame
catches it and an exit code does not.

What the timeline bought was the authoring. The logo is "160 px box, 94% across,
12% down" and the title is "centered at 88% down". There is no filter graph to
get wrong, no `W-w-48`, no label wiring, and the same JSON renders at any output
size. For a one-off edit you already know as a command, that premium buys
nothing. For a template that product code fills in a thousand ways, it pays for
itself in code you never write.

## What self-hosting costs besides the render

The sub-second local run leaves out everything a production pipeline wraps around
it.

The binary comes first. A static FFmpeg build is tens of megabytes. The
ffmpeg-static binary is 76 MB, which fits Vercel but sets the shape of every
function you deploy it in, as [FFmpeg on Vercel](/blog/ffmpeg-on-vercel-size-limit/)
measured. Fonts, codecs and filters you need have to be in that build.

Then scaling and queueing. One machine renders one job fast. Fifty jobs at once
need a queue, workers, a concurrency limit per customer and something that
notices a worker died mid-encode. None of that is FFmpeg, and all of it is yours.

Long jobs hit a wall on serverless. Functions cap wall time. AWS Lambda stops a standard
invocation at 900 seconds, so a long encode needs a container service or a VM,
with its own scaling story.

Failure handling is the last piece. A crashed encode needs a retry that does not double-bill
or double-deliver, a temp directory that gets cleaned up, and an input URL that
has not expired by the time the retry runs.

Self-hosting is still the right answer for a team that already runs a job
platform, renders steadily at high volume, and wants the encode cost to be the
whole cost. For everyone else those four lines are the project.

## What template APIs charge

Template and timeline APIs price by the minute of rendered output, with a base
fee or a monthly bucket. These figures come from each vendor's pricing page, read on 2026-09-27
for the [Rendobar vs Shotstack comparison](/vs/shotstack/).

| Vendor and plan | Per minute of output | Base |
|---|---|---|
| Shotstack Subscription | $0.20 | $39 a month |
| Shotstack pay as you go | $0.30 | $75 one-time, credits valid one year |
| Creatomate Essential | about $0.84 effective at 1080p25 | $54 a month for 2,000 credits, reset monthly |

Creatomate bills in credits by pixels. At 1080p25 a minute uses about 31
credits, so the 2,000-credit plan buys about 64 minutes.

A 4-second render sits awkwardly against per-minute pricing. Prorated, Shotstack's
Subscription rate would be about $0.013 for this edit, below the timeline run
above and above the command run. I did not test whether Shotstack bills
fractions of a minute, so treat that as arithmetic, not a measurement. The larger
difference is the base: $39 a month or $75 up front before the first render.

## Running the edit on Rendobar

Rendobar is the default I would give a developer choosing today. The command you
tested locally runs unchanged on [Rendobar's FFmpeg API](/ffmpeg/) for less than
a cent a render here, the timeline format is on the same account for the edits
that are easier to describe as layout, and there is no base fee. The Free plan
starts with a $5 credit grant, and Pro is $9 a month with $5 of credit included
and jobs up to 9 hours.

The command, with the same inputs mapped by name:

const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });

const job = await rb.jobs.run({
  type: "ffmpeg",
  inputs: {
    source: "https://cdn.rendobar.com/assets/examples/sample.mp4",
    "logo.png": "https://cdn.rendobar.com/assets/brand/logo-mark.png",
  },
  params: {
    command:
      "ffmpeg -ss 1 -t 4 -i source -i logo.png -filter_complex " +
      "\\"[0:v]scale=1920:1080:flags=lanczos," +
      "drawtext=fontfile=/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf:text='Launch week':fontsize=72:fontcolor=white:box=1:boxcolor=black@0.5:boxborderw=20:x=(w-text_w)/2:y=h-th-80[bg];" +
      "[1:v]scale=160:-1[lg];[bg][lg]overlay=W-w-48:48[v]\\" " +
      "-map \\"[v]\\" -map 0:a -c:v libx264 -preset veryfast -crf 23 -c:a aac -b:a 128k -movflags +faststart out.mp4",
  },
});

console.log(job.output?.file?.url, job.cost?.formatted);`}
  curl={`curl -X POST https://api.rendobar.com/jobs \\
  -H "Authorization: Bearer $RENDOBAR_API_KEY" \\
  -H "Content-Type: application/json" \\
  -d @- <<'EOF'
{
  "type": "ffmpeg",
  "inputs": {
    "source": "https://cdn.rendobar.com/assets/examples/sample.mp4",
    "logo.png": "https://cdn.rendobar.com/assets/brand/logo-mark.png"
  },
  "params": {
    "command": "ffmpeg -ss 1 -t 4 -i source -i logo.png -filter_complex \\"[0:v]scale=1920:1080:flags=lanczos,drawtext=fontfile=/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf:text='Launch week':fontsize=72:fontcolor=white:box=1:boxcolor=black@0.5:boxborderw=20:x=(w-text_w)/2:y=h-th-80[bg];[1:v]scale=160:-1[lg];[bg][lg]overlay=W-w-48:48[v]\\" -map \\"[v]\\" -map 0:a -c:v libx264 -preset veryfast -crf 23 -c:a aac -b:a 128k -movflags +faststart out.mp4"
  }
}
EOF`}
/>

The same edit as a [compose timeline](/compose/). Three tracks stack bottom to
top: the trimmed clip, the logo, the title. This version names Inter, a bundled
family, where my measured runs named DejaVu Sans and got a fallback.

const rb = createClient({ apiKey: process.env.RENDOBAR_API_KEY });

const job = await rb.jobs.run({
  type: "compose",
  params: {
    schemaVersion: 1,
    output: {
      format: "mp4",
      resolution: { width: 1920, height: 1080 },
      fps: 24,
      crf: 23,
      encoderPreset: "veryfast",
    },
    timeline: {
      tracks: [
        { name: "video", clips: [{ asset: {
          type: "video",
          src: "https://cdn.rendobar.com/assets/examples/sample.mp4",
          trim: { from: 1, to: 5 },
        } }] },
        { name: "logo", clips: [{
          asset: { type: "image", src: "https://cdn.rendobar.com/assets/brand/logo-mark.png" },
          transform: { position: { x: "94%", y: "12%" }, size: { width: 160, height: 160 }, fit: "contain" },
        }] },
        { name: "title", clips: [{
          asset: {
            type: "text",
            text: "Launch week",
            style: { font: "Inter", size: 72, color: "#FFFFFF", background: "#00000080", align: "center" },
          },
          transform: { position: { x: "50%", y: "88%" } },
        }] },
      ],
    },
  },
});

console.log(job.output?.file?.url, job.cost?.formatted);`}
  curl={`curl -X POST https://api.rendobar.com/jobs \\
  -H "Authorization: Bearer $RENDOBAR_API_KEY" \\
  -H "Content-Type: application/json" \\
  -d @- <<'EOF'
{
  "type": "compose",
  "params": {
    "schemaVersion": 1,
    "output": { "format": "mp4", "resolution": { "width": 1920, "height": 1080 }, "fps": 24, "crf": 23, "encoderPreset": "veryfast" },
    "timeline": { "tracks": [
      { "name": "video", "clips": [{ "asset": { "type": "video", "src": "https://cdn.rendobar.com/assets/examples/sample.mp4", "trim": { "from": 1, "to": 5 } } }] },
      { "name": "logo", "clips": [{ "asset": { "type": "image", "src": "https://cdn.rendobar.com/assets/brand/logo-mark.png" }, "transform": { "position": { "x": "94%", "y": "12%" }, "size": { "width": 160, "height": 160 }, "fit": "contain" } }] },
      { "name": "title", "clips": [{ "asset": { "type": "text", "text": "Launch week", "style": { "font": "Inter", "size": 72, "color": "#FFFFFF", "background": "#00000080", "align": "center" } }, "transform": { "position": { "x": "50%", "y": "88%" } } }] }
    ] }
  }
}
EOF`}
/>

## Which one to pick

If you have a working FFmpeg command and a server you do not want to run, send
the command to an API and keep the timeline for edits where layout is the hard
part. That is where I would start, and it is why the middle category exists: it
removes the choice between your own command and someone else's infrastructure.

Self-host when renders are steady, high in volume and already inside a job
platform your team runs. Pick a template API such as Shotstack or Creatomate when
the people designing the video are marketers working in a visual editor, or when
you need a white-label editor inside your product, which neither FFmpeg nor
Rendobar ships. For a wider ranking by what each service lets a job run for,
[the best FFmpeg API by runtime cap](/blog/best-ffmpeg-api/) lines them up.
