How do you test a screenshot API when every failure returns 200 OK?

The benchmark behind Screenshotline - 202 real sites, every page tagged ok or known-hard in advance, blank frames failing the run, and why a captcha that scores as a pass is the reason someone still has to open the image.

Most test suites answer the question "did it work?". A screenshot renderer makes that question almost useless, because the interesting failures all return 200 OK with an image attached.

A cookie wall is a valid PNG. So is a page captured mid-paint, a blank white frame from a bot wall, and a captcha. Every one of those passes a status-code check, passes a "is the response an image?" check, and passes a "is the file bigger than zero bytes?" check. Then it sits in someone's pipeline for a month.

So the benchmark behind Screenshotline is built around a different question: which pages surprised me this time? This post is how that works, because the shape of it applies to anything that scrapes, renders or crawls the real web.

Every page carries a claim, written before the run

The suite is 202 real sites across 15 categories: news, shops, SaaS, docs, government, education, finance, WebGL and visualisation, TLS edge cases, HTTP semantics, and 13 languages in 10 scripts. Each entry looks like this:

['nyt', 'https://www.nytimes.com/', 'Fides consent banner injected after networkidle'],

and each page is tagged in advance with what I claim my own product should do:

  • ok: this must work. A failure here is a bug.
  • known-hard: I expect a bot wall or a block. A failure here is information, not a regression.

That tag converts a run from a score into a set of predictions that can be wrong, which is the version that teaches you something. The runner prints the two kinds of broken prediction under a heading I actually read:

  BUGS - expected to work, failed (0):
  Expected hard, but worked (14) - verify the image is real content, not a bot wall

A page that was supposed to work and didn't is a bug. A page that was supposed to be hard and passed means something changed. Maybe my renderer got better, maybe the site dropped a defence, maybe my own test URL has rotted.

A pass on a hard page means nothing until you look

This is the rule that cost me the most to learn. There are three completely different things behind a 200 on a page I marked known-hard, and only one is a success:

  • A login wall is what any logged-out visitor sees. Capturing it is correct. facebook, instagram, quora.
  • A captcha is a challenge shown only to suspected bots. Capturing it is a failure dressed as a pass. walmart, wsj.
  • A 404 means the URL in my own suite died. amazon, where an ASIN went away and I was quietly benchmarking a "product not found" page.

No automatic check tells those apart. I tried. A text heuristic scored wsj as "looks real" on zero extracted characters, because the page around the challenge was full of navigation and footer text.

So when the runner says 14 known-hard pages returned something, I open all 14. The last time: 11 were genuinely the real page, 2 were captchas, and 1 was that dead URL. That is a result I can publish. "14 of 20 hard pages passed" is not.

It is also why I will never quote a single headline percentage that mixes the two groups together. The honest version has two numbers: of the 182 pages marked ok, all 182 capture correctly; of the 20 marked known-hard, 14 returned something and each one was checked by eye.

Blank frames fail the run

The suite's sharpest rule is about blank captures, because that is the worst thing this product can do: return a 200 that a customer's pipeline files as a success while the image shows nothing.

Some pages really are blank for everyone: a bot wall that serves an empty document, or an IP block. Those are marked blankOk in the suite file, one page at a time, after I have looked. A blank frame on any other page fails the whole run:

if (unexpectedBlank.length) {
  console.log(`  FAILED: ${unexpectedBlank.length} unexpected blank capture(s).`);
  process.exitCode = 1;
}

The non-zero exit is deliberate. Without it, a script or a CI job that only checks the status code would let a silent blank regression through, which is precisely the class of failure the suite exists to catch.

The report also separates "blank, and known to be" from "blank, and new". The first list is context. The second is a stop-everything.

File size is not quality, in either direction

The tempting shortcut is to compare byte counts: bigger image, better capture. It is wrong both ways, and I have been fooled in both directions.

Removing a consent modal makes a capture smaller. A full-screen dark overlay compresses into a lot of bytes. The better picture is the lighter one.

Collapsing blocked ad slots makes it bigger. On france24.com, collapsing the empty slots took a capture from 122 KB to 234 KB of actual content, because real content moved up into the frame.

A larger competitor capture can still be worse. One competitor's france24 image was bigger than mine and blurry, because it was taken mid-paint.

And some pages simply swing. nytimes.com moves between 80 KB and 650 KB depending on how much of its skeleton has hydrated. That is page timing, not a renderer bug, and delay is the parameter for it.

Which leaves the same conclusion as everywhere else in this project: open the image.

To make that bearable across 202 pages and several providers, the runner writes every capture to disk and a second script builds one HTML page with every provider's version of every URL side by side. Ten minutes of scrolling tells you more than any metric I have found.

What the suite is not

It is not a scoreboard, and the pass rate is the least useful line it prints. Any suite can be made to look perfect by lowering its expectations, and a benchmark whose numbers only ever improve is measuring the author's optimism.

It is not stable across machines. Several of these sites block datacenter IP ranges, so a page that passes from my laptop may fail from a cloud host. I'd rather know that before a customer discovers it: it is why the hosted API's p50 (about 7.1s, single region) is measured separately from the benchmark machine's 5.7s rather than reusing the nicer number.

It is not a competitor takedown. I run the same 202 URLs against other providers because side-by-side images are how I find the pages where I am worse. The honest result of the last head-to-head, on 18 news front pages: we captured 18/18, ScreenshotOne 17/18, and their median was faster than ours, 5.6s against 6.0s. Publishing the part that flatters you and quietly dropping the rest is how benchmarks became a genre people don't trust.

If you are building something similar

The parts worth stealing, in order of how much they saved me:

  1. Write the expectation next to the URL, before you run anything. A test that can only pass or fail tells you less than one that can surprise you.
  2. Make the silent failure loud. Decide what "wrong but successful" looks like in your domain, detect it, and let it fail the run with a non-zero exit.
  3. Keep a category that you expect to fail. It stops you from tuning the suite until it is green and keeps the hard cases in view.
  4. Look at the output with your own eyes, on a schedule. Not because automation is bad, but because the failures that matter are the ones your checks were not built to see.

The suite lives in bench/ in the repo: the runner, the 202 URLs with their tags and notes, and the comparison page builder. It runs against a local instance, so if you self-host you can point it at your own and see what your hardware does with the same pages.

Get a free key — 500 renders a month Try it without signing up

More from the blog