Running a paid screenshot API on one small server

What a 1-core, 2 GB VPS actually does with headless Chrome - measured slot-seconds per render, a paid-first queue, the free-tier budget that protects it, and the honest limits of a single box.

The hosted version of Screenshotline runs on one server. One core, 2 GB of RAM, a KVM slice that costs a few dollars a month, in a single region. There is no autoscaling group, no load balancer, no Kubernetes.

That is a deliberate choice with real consequences, and this post is both halves: what a box that size genuinely does with headless Chrome, and what it cannot do.

What one render actually costs

Measurements from the production image, POOL_SIZE=1, an isolated container on loopback, 46 pages from the benchmark with default options:

  • 9.44 slot-seconds per render, mean. That's wall-clock occupancy of the one rendering slot, not CPU.
  • 7.3 CPU-seconds per render.
  • 77% of the core busy while a single render is in flight.
  • Peak memory 1,228 MB of the ~1,400 MB available to the container.

Three things fall out of those numbers.

Chrome costs real CPU, but not much money. At full utilisation, the compute behind a thousand renders costs a couple of cents. The fee my payment processor takes on a single small subscription is worth more than the compute for a great many renders. When people assume a screenshot API is expensive to run, they are usually thinking of egress or GPU time, neither of which is in this path.

One render uses most of one core. So POOL_SIZE is not a throughput dial you turn up for free: a second browser on a single-core box does not double throughput, it splits the core and makes both renders slower while doubling the memory.

Memory is the wall, not CPU. Peak was 1,228 MB against a 1,400 MB ceiling. That is the number that decides whether this box can run two browsers, and it says no.

A queue, not a scale-up

With one slot, the interesting question is who gets it next. The answer is explicit: paid requests jump the queue.

Every render acquires the slot through a small pool that sorts waiters by priority, and free and demo traffic sits behind paying customers. On a fixed price box, that ordering is what protects the thing that pays for the box.

Two more fences sit behind it, and both are about capacity, not money:

freeSlotShare:   0.75,  // share of pool time free + demo traffic may use, rolling hour
freeCallerShare: 0.2,   // share of THAT budget any one free caller may take

freeSlotShare at 0.75 on a one-slot box means roughly 45 minutes of rendering per hour is available to people who are not paying. That is deliberately generous. The box costs the same whether it is busy or idle, so unused capacity is not saved money, it is wasted capacity. The paid-first queue already protects customers minute to minute; this is a backstop against a sustained flood, and an emergency lever: set FREE_SLOT_SHARE=0.3 and restart, or 0 to turn free rendering off entirely.

freeCallerShare exists because of a specific hole. A free key sending pages that run to the 30-second timeout is never billed and never touches its quota, because failures aren't charged. Without a per-caller cap, one key could spend the entire free tier's hour and lock every other free user out, at no cost to itself.

There's a small implementation detail in that config worth stealing:

const share = (v, d) => {
  const n = Number.parseFloat(v ?? '');
  return Number.isFinite(n) && n >= 0 ? n : d;
};

Number('') is 0, so an empty FREE_SLOT_SHARE= line in an env file would otherwise switch the whole free tier off silently. Parsing it by hand and falling back to the default turns a silent outage into a no-op.

Set limits by what they cost on the day, not by what sounds prudent

An earlier version of the free budget was a 50% rolling share. It sounded responsible. Played forward against a launch-day traffic spike, it would have switched the public demo off mid-surge to protect paying customers who did not exist yet. The demo, with no signup and no key, is the thing that convinces people the product works, and it would have gone dark exactly when it mattered.

The per-domain limit taught the same lesson from the other direction. I set free renders of any one domain to 10 a minute, which felt reasonable, then noticed the landing page demo is pre-filled with example.com. Every visitor who clicks "Try it" without editing the box renders the same domain, so they all share one bucket. Ten a minute would have rate-limited the demo against itself on the busiest day of the year. It is 30 now.

Before setting any shared limit, check what the most common input actually is.

Knowing when the box is unwell

Two health endpoints, because they answer different questions:

  • /healthz says the process is answering. It stays green while Chrome is broken, so alone it will cheerfully report that all is well during an outage.
  • /healthz/render renders a real page and returns 503 when that fails. It is fenced so it can't become a free renderer or a load lever: at most one real render a minute, everyone else gets the cached last result, queued at low priority, never metered.

The uptime monitor points at the second one, and it runs on someone else's infrastructure, not this box. A status page on the same server dies with the server and tells you everything is fine.

Backups go off the box too, for the same reason. A VPS is one machine. The SQLite database with accounts, key hashes and usage counters is the only state that cannot be regenerated, and a snapshot that lives only on the machine it is protecting is not a backup.

What this setup cannot do

I would rather write these down than have someone discover them.

  • One region. Latency from the other side of the world includes the trip. The hosted p50 is about 7.1s, against 5.7s on my benchmark laptop, and part of that gap is simply where the box is.
  • A deploy is a blip. Rebuilding and restarting the container takes about 30 to 45 seconds, and the status page shows it. No rolling deploy, because there is nothing to roll.
  • The box is a single point of failure. If it dies, everything is down until I move it. There is no SLA, and I don't pretend otherwise.
  • Throughput has a ceiling and it is not large. Sustained heavy traffic means a bigger box, then a second one, and the queue is what keeps paying customers first while that happens.

None of this is advice to run your production on one small server. It is what an honest version of "we're small" looks like: measure what the machine does, decide explicitly who waits behind whom, and publish the limits rather than discovering them in public.

If you'd rather own the box yourself, the whole thing self-hosts with Docker, and the sizing numbers above are the ones to start from.

Get a free key — 500 renders a month Try it without signing up

More from the blog