All Work

Creative Operations · Tool Evaluation · Saint-Gobain North America

The Machine Does the Rollout

Six tools, one test brief, one recommendation.

  • Tool Evaluation
  • Creative Operations
  • Technical Validation
  • Stakeholder Recommendation

Role

Web Designer; sole evaluator, tester, and author of the recommendation

Timeline

Assigned July 31, presented August 11, 2026

Scope

Whether AI ad-production tooling could take one master concept to a full Google and Meta size list

Status

Recommendation accepted and cleared security review; in procurement as of September 2026, with the pilot to follow

The Problem

Twenty-Three Sizes, Built by Hand

A single campaign concept has to ship as twenty-three different ad units — the full Google display and Meta size list, static and animated. Every one of those units is built by hand: a designer opens the master, resizes the canvas, and rebuilds the composition to fit a shape it wasn’t designed for. A 300×1050 skyscraper and a 1200×628 link ad have nothing in common except the artwork inside them.

A full twenty-three-size rollout runs about twelve hours, measured from master to delivered set: the per-size design work, the export, and the revision rounds. That’s an average rather than a stopwatch reading — my own build times checked against what other designers reported — and the spread around it is wide. A straightforward master with one round of edits finishes well under. Unusual size requests and three rounds of changes run well over. Twelve is the number I’d defend as typical, not the number any single rollout costs.

The work is spread across several designers, and it is the least interesting thing any of them do. It is also the thing that eats the hours that would otherwise go to concept work. The studio manager asked me to find out whether that was still necessary — specifically, whether one tool a colleague had used previously could do it, and whether anything else could do it better. One week to research, a call the following week to present.

The craft is in the master. Everything after it is repetition, and the repetition is what was eating the week.

What I Did

One Test Brief, Run Identically

I built one standardized test brief and ran it identically through every candidate that could plausibly clear the bar.

The brief: a real, already-launched master as a layered source file. A fixed list of twenty-three sizes across Google display and Meta. Static and animated output. Then a second brand kit swapped in, to see whether multi-brand was a feature or a marketing line.

I scored against source-file compatibility, reflow intelligence, per-size editability, HTML5 and GIF export validity, multi-brand architecture, font fidelity, learning curve, and security posture. Setup cost was tracked separately from rollout cost, so that a one-time expense couldn’t hide inside a recurring saving.

I also labeled every claim in the findings by how I’d obtained it — tested, analyzed, vendor-claimed, or scoped and ruled out. That labeling turned out to matter more than any single finding, and I’ll come back to it.

Tool Source files Reflow Per-size edits HTML5 export Multi-brand Evidence
The Brief Layered PSD; no INDD Recomposes per size Yes, every unit Native — 23 of 23 passed Brand kits; two tested Tested
Adobe Express PSD and INDD — the only tool that opened InDesign; import degraded fidelity Repositions only Yes None — no clickTag Brand libraries; limited Tested
Bannerflow PSD and INDD (claimed) None yet — manual, pushed across sizes Yes Native Cross-brand layout copying Demoed
Figma + plugin Rebuild in-tool None — auto-layout is container logic Manual, per size Third-party plugin, unvalidated Libraries and tokens, manual Analyzed
Celtra Same enterprise tier as Bannerflow, which I evaluated as the representative of that tier rather than running two enterprise processes in a week. Scoped
Smartly Creative production bolted to media buying — an adjacent category, and a media-team tool rather than a design-team one. Scoped
Six tools against the same brief. The evidence column is the part that mattered: there is no composite score, because a single number would imply a precision this evaluation didn’t have.

What I ruled out

Adobe Express got two structured tests, because we already license it and “can’t we just use what we have” is the first question anyone sensible asks. It was also the only platform in the set that would open an InDesign file at all — everything else started from Photoshop — which is a real advantage and the reason it got the second test. But the import degraded fidelity. Its resize repositions elements rather than recomposing them — building a template natively inside the tool improved background handling, but text and elements stayed where they were put. And it has no HTML5 ad export with clickTag support, which is a hard requirement rather than a preference. Express is good at what it’s built for: fast social and image work, which the team should keep using it for. It isn’t built for layered ad recomposition.

Adobe Express · InDesign source
An ad master placed on a much wider canvas: the artwork sits unchanged at its original size in the middle, with empty white canvas either side of it
A resize that repositions. The artwork keeps its original dimensions and the canvas grows around it — recomposition is the thing that isn’t happening. There is no second panel here because there couldn’t be one: this master is an InDesign file, and Express was the only tool that would take it.

Figma with a third-party plugin I assessed without a full hands-on build, because the assessment answered itself. There is no AI reflow in that lane. Auto-layout is container logic, not recomposition, so the work stays manual — and the plugin carries its own cost on top of a rebuild. Figma stays in the toolkit for interoperability with agencies who deliver in it. It wasn’t a candidate for this pipeline.

Celtra and Smartly I scoped and ruled out without testing. Celtra sits at the same enterprise tier as Bannerflow, so I evaluated Bannerflow as the representative of that tier rather than running two enterprise processes in a week. Smartly is an adjacent category — creative production bolted to media buying, which makes it a media-team tool, not a design-team one.

That left one tool I could test end to end and one enterprise platform I could only see demonstrated.

The evaluation cost nothing but a week, and four of the six answers came from ruling things out properly.

What I Found

The Recommendation, and Its Limits

I recommended The Brief (formerly Creatopy).

Raw generation of all twenty-three sizes took about twenty minutes. That figure is measured. From there I projected roughly three hours per rollout once a designer finishes each unit to production quality — same scope as the baseline, revision rounds included — for a modeled reduction of about 75%. The three-hour figure is an extrapolation from a tested generation time, not an observed result, and I presented it that way. The point of asking for a pilot is to replace the estimate with a real number.

How the HTML5 was validated

I exported the full set as HTML5 and ran every unit through Google’s hosted h5validator. All twenty-three passed its structural checks — markup, clickTag, asset references. Weight I treated as a separate pass, checking initial load against the 150KB class two ways: the size the validator reported and the exported file weights on disk. There the set wasn’t clean on the first run. The 300×1050 came in at 207KB, and adjusting export settings brought it under.

Google h5validator · 23 units, full size list

Markup well-formed, no blocking errors 23 / 23
clickTag present and correctly bound 23 / 23
Assets every reference resolves 23 / 23
Initial load under the 150KB class, first export 22 / 23
Initial load after export settings adjusted 23 / 23

300×1050 — 207KB on the first export. The only unit over the class, and the reason weight was checked separately from the structural pass.

Rebuilt from the run rather than screenshotted. Structural checks cleared on the first pass; weight was a second, separate check, and one unit failed it before its export settings were changed.

Worth noting that the endpoint I used now carries a no-longer-maintained notice; the requirements behind it are Google’s published ad specs, which haven’t moved. This is the part of the recommendation I’d defend hardest, because it’s the part where a tool either meets a published spec or it doesn’t.

The multi-brand question, answered partway

I swapped a second brand kit into the same master and regenerated. It held — type, color, and logo placement carried across without a rebuild, and I exported the two versions side by side as proof of concept rather than as a claim about the feature’s depth. What I did not test is sub-brand inheritance or how the kits behave across every campaign-level variation, and I said so: that’s a pilot question, not a one-week question. It was enough to answer the thing I actually needed to know, which was whether multi-brand was a real architecture or a line on a pricing page.

Norton
CertainTeed
One master, regenerated against the Norton and CertainTeed brand kits. Proof of concept, not a feature claim — type, color and logo placement carried without a rebuild, but sub-brand inheritance was left to the pilot.

What I said plainly, before anyone asked

The outputs are not pixel-perfect. Designers finish every unit. That’s the intended division of labor, not a defect I was hoping nobody would notice — the machine does the rollout, the designer does the craft. The fastest way to lose a room of designers is to imply the tool replaces their judgment.

The enterprise option, followed through

I’d presented Bannerflow as vendor-claimed rather than tested, with a demo already on the calendar. I attended it the following week and wrote up the comparison. It’s a capable platform and it would solve the problem. Three things decided against it: its design studio behaves more like Figma than Adobe, which is the less familiar lane for this team; there’s no AI recomposition yet, so a designer still makes each change and pushes it across sizes; and it requires an up-front commitment with no pilot available, which means signing before any evidence exists. Its strongest differentiator, live updates to already-served ads, solves a problem we rarely have.

Cost closed it, and the shape of the cost mattered more than the figure: an enterprise-wide commitment against a subscription one person could trial, for a team that needed evidence before it could justify either. Tracking setup separately from rollout was what made that comparison legible — without it, a large one-time expense disappears into a per-rollout saving and the numbers stop meaning anything. I ended the vendor conversation in writing and left the door open if scale changes the math.

Where It Stands

Approved, and yet still short of a result.

I presented on August 11 to the creative director and the studio manager, and recommended a thirty-day pilot of The Brief on the next real campaign, run by me, with four success metrics: hours per rollout against the twelve-hour baseline, sizes shipped, revision rounds, and direct feedback from the designers using it.

The recommendation went into internal review and came out the other side: it cleared security review in September 2026 and moved to procurement, which should put the platform in the team’s hands and the pilot on the calendar.

What that doesn’t give me yet is a number. The pilot is the thing that produces one, and it hasn’t run. So the seventy-five percent is still modeled, the three hours still extrapolated, and the only measured figures in this study remain twenty minutes of generation time and twenty-three units through a validator. That’s the state of it, and I’d rather say so than dress an approval up as an outcome.

In Hindsight

What I’d Do Differently

I didn’t ask the follow-up question.

The enterprise quote came to me verbally, at the end of a call, without a timeframe attached. I wrote it down and carried it into the deck on the wrong terms, because the other tool I was evaluating priced on different ones and that’s the shape my head was already in. I caught it going back through my notes after the demo and corrected it in writing to both stakeholders the same week.

What bothers me isn’t the number. It’s that a one-second question would have prevented it, and I didn’t ask because the meeting was ending and it felt like a detail. Every other figure in that deck was labeled with how I’d obtained it, precisely so nobody had to guess how much weight to give it. The one figure I took on trust from a conversation is the one I got wrong. The method was right; I just stopped applying it when I stepped out of the document and into a room.

I estimated the baseline when I could have measured it.

Twelve hours is my own timing plus what other designers told me, and the entire business case rests on it — every saving I projected is a percentage of that one number. The better version was available and I didn’t do it: build a real twenty-three-size rollout by hand, clock it start to finish, log each revision round separately, and come in with a measured figure and a documented range instead of a defensible average. That’s a day of work. It would have made the comparison unarguable, and it would have given the pilot a like-for-like number to be measured against rather than an estimate to be reconciled with. I had a week for the whole evaluation and spent it on breadth. Given the same week I’d now spend a day of it on depth, because the baseline is the number every other number in the deck depends on.

The evidence labels did more work than the findings.

Marking every claim as tested, analyzed, vendor-claimed, or scoped meant nobody had to guess how much weight to put on any line, and the one enterprise tool I hadn’t touched couldn’t quietly borrow credibility from the one I had. I was uneasy while building it that the weaker labels would read as thin diligence. The opposite happened — being explicit about the limits of what I knew is what made the rest of it trustworthy.

The Takeaway

A Recommendation Is a Claim About Evidence

Six tools, a week, and one standardized brief. Four of them were answered by ruling them out properly, which is the cheapest work in the whole evaluation and the part most likely to be skipped. The two that survived got the time that freed up.

What I’d carry into the next one is the labeling. Naming how I knew each thing — tested, analyzed, vendor-claimed, scoped — is what let a one-week evaluation say something useful without overclaiming, and it is also what made the one unlabeled figure stand out as the mistake it was.

The recommendation was accepted. The pilot hasn’t run, so the number that matters doesn’t exist yet.

Created as an employee of Saint-Gobain North America. Shown here as a demonstration of capability, not as a service offered for hire. Company-owned work; presented with role clearly attributed. Product names are public; no commercial terms from any vendor conversation appear here.