← The Bench Journal
Tool of the WeekBy Hamed Arab·6 September 2026·4 min read

Why a dry-run flag catches what your tests miss

The flag that found three bugs the tests didn’t

One of my scheduled automations picks and publishes a daily social post for Silux London, my jewellery brand, drawing from a rotating bank of stories and lore behind each piece. It had two passing test suites behind it. Green across the board, every time I checked.

I added a --dry-run flag anyway: a way to run the whole pick-and-publish logic and print out exactly what it would post, without actually posting it. The idea was just a safety habit before letting something run completely unattended. Instead, the first time I ran it, it turned up three real bugs in one pass.

One was a seasonal story scheduled for the wrong month, sitting there ready to publish something autumn-themed in the middle of summer. One was a duplicate pick: two different data files referred to the same story, but one wrote “and” where the other wrote ”&”, so a string match meant to catch duplicates silently failed to notice they were the same thing. And one was a slugify regression, a bug that had crept in through an earlier fix to how the tool turns titles into URL-safe text, breaking links that should have worked.

Two test suites, both passing, had missed all three.

Tests answer a narrower question than you think

That isn’t a knock on testing. Passing tests genuinely prove something: that your code does what you told it to do, under the conditions you thought to write a test for. What they can’t prove is that what you told it to do is still correct, today, against the real data your automation is actually going to touch when it runs.

A test suite for that publishing script checked things like “does the picker return a story” and “does the scheduler respect the configured slot”. Both true. Neither test knew that one particular story’s month field had been typed wrong, because nobody wrote a test that checked every story against every month, and nobody would think to. That’s not a gap in diligence. It’s the nature of unit tests: they check the shape of the logic, not the correctness of the specific, messy, real-world data sitting behind it on any given day.

A dry run asks a completely different question. Not “does this code work correctly”, but “given what’s actually in my data right now, what is this thing about to do”. That’s the question that catches a typo in a date field, a mismatched string, a broken URL, because it runs the real logic against the real data and just refuses to take the last, irreversible step.

Build the preview before you trust the automation

If you’re automating any part of your business with AI, and letting it run on a schedule without you watching, this is the one habit I’d put ahead of almost anything else: give it a preview mode before you ever let it run unattended.

It doesn’t need to be complicated. A single flag that walks through the exact same logic your automation will use on the day, and then prints or logs what it would have done instead of doing it, is enough. Run it once before the schedule goes live, and ideally run it again any time you change the data the automation reads from, not just when you change the code.

The value isn’t really in catching bugs at the code level. It’s in catching the moment where correct code meets incorrect or unexpected data, which is precisely the failure your tests are structurally unable to see. Mine found a wrong month, a silent duplicate, and a broken link, on the first run, in data that had already sailed through two green test suites. That’s not a coincidence. It’s what a dry run is for.

Want to go deeper?

My book The CAD/CAM Jeweller covers these topics in production-ready detail. For one-to-one teaching, book a free 30-minute call to talk about where you are with your CAD or your jewellery brand.