Beyond A/B Testing: Why Incremental Wins Have a Ceiling
"We've done A/B testing for decades," the head of marketing said. Nearly doubling revenue was still on the table.
"We do A/B testing since decades," a head of marketing told me. Said with quiet pride, the way you'd mention a long marriage.
And still — nearly doubling the revenue of their direct mailing turned out to be realistic.
How can both be true? How can a team that has optimized relentlessly for years still have that much left on the table?
Because of what A/B testing is. And, more importantly, what it isn't.
The Infinite-Tweaks Problem
In an A/B test, you compare one "tweak" against the original. Headline B against headline A. Green button against blue. One thing changes; you measure which wins.
That's the strength. It's also the trap.
Because there are infinite ways to optimize. Infinite headlines, offers, layouts, sequences, tones. Every test you run picks exactly one of them and ignores the rest. You are drawing one card from an endless deck and asking whether it beats the card in your hand.
And where do those cards come from? Mostly out of the blue. A competitor's mailing you admired. A hunch in a meeting. Your own genius on a good day. The hypotheses aren't wrong — but they're inspired, not derived. Nobody in the room knows why the customer responds, so the ideas are guesses dressed as experiments.
Incremental by Design
Here's the part that matters for your budget.
A/B testing is, almost by definition, incremental. Each test measures a small change against the current best. So you climb — slowly, honestly, one percentage point at a time. And as the obvious ideas get used up, the wins get smaller and the tests get more expensive to justify.
It's all good. I'm not telling you to stop testing.
But understand the ceiling. Reaching a genuinely new, higher level of performance by tweaking is prohibitively expensive — because tweaking never questions the underlying idea driving the response. You can perfect the wording of an appeal that was aimed at the wrong motive, and you'll never test your way out of that. You'll optimize a hill and never learn there's a mountain next to it.
This is the same failure that haunts most stated-preference research: it refines the answer to a question the customer's subconscious never actually asked. I've written about that at length in why market research is broken.
The Alternative: Read the Mind, Don't Guess It
The way past the ceiling isn't a better test. It's a better hypothesis.
And you get a better hypothesis by understanding the subconscious mechanics of the customer's mind when it's exposed to your stimuli. What actually fires. What association forms. What need lights up before a single conscious thought arrives.
That's what we do at SUPRA with Deep Implicit Research. Instead of pitting variant against variant and waiting weeks for a verdict, we measure what drives the response beneath awareness — and then use Causal AI to identify which levers actually move behavior, not just which ones correlate with it.
The distinction is everything. A/B testing tells you that B beat A. It never tells you why. And without the why, your next test is another guess. Understand the mechanism and you don't stumble onto the next level — you design it. Most of what customers can't tell you is the exact thing that decides whether they buy; that's the say-do gap, and it's where the real gains hide.
A/B testing vs. reading the subconscious
- A/B testing evaluates one guess at a time; implicit research reveals the driver behind all of them.
- A/B testing tells you which won; implicit research tells you why — so the next idea isn't a guess.
- A/B testing gains shrink over time; understanding the mechanism resets the ceiling.
- A/B testing needs traffic and patience; the insight often pays off immediately.
It Pays Off Immediately
The objection I hear next is always the same: sounds deep, sounds slow, sounds academic.
It isn't. Our work with clients shows it works — and it pays off immediately, because the improvements are structural, not cosmetic. When you change the driver instead of the wording, you're not shaving a point off conversion. You're moving the whole curve.
A recent example. An insurance client set the bar plainly: "It would be a success if you can raise conversion already by 10%." We cleared it with ease.
Not by finding a cleverer button. By understanding what the customer's mind was actually reaching for — and speaking to that instead.
That's the difference between optimizing the answer and changing the question.
This is how you 10× the ceiling on what your marketing can do.
A/B testing and Deep Implicit Research: frequently asked questions
What is the main limitation of A/B testing?
It compares one tweak against the original, and there are infinite possible tweaks. Each test evaluates a single hypothesis, and those hypotheses usually come out of the blue — inspired by a competitor or someone's own genius. A/B testing is excellent at climbing the hill you're on. It rarely tells you a taller hill exists.
Is A/B testing incremental or transformational?
Incremental by design. Each round measures a small change against the current version, so gains compound slowly and shrink as the obvious ideas run out. It's often worth doing — but it's prohibitively expensive to reach a genuinely new, higher level of performance by tweaking, because tweaking never questions the underlying idea driving the customer's response.
What is Deep Implicit Research?
It's SUPRA's approach to understanding the subconscious mechanics of the customer's mind when exposed to your stimuli. Instead of testing one variant against another and waiting, it measures what actually drives response beneath conscious awareness, then uses Causal AI to identify which levers move behavior. It tells you why something works, so you can design the next level rather than stumble onto it.
Can understanding the subconscious really lift conversion quickly?
Yes, and often immediately. Because the method reveals the mechanism behind response rather than a single winning variant, improvements tend to be structural rather than cosmetic. In one recent insurance engagement the client framed a 10% conversion lift as the bar for success — and that bar was cleared with ease.
Dr. Frank Buckler is the founder of SUPRA and a pioneer in Causal AI for marketing. For over 25 years he has helped brands understand what actually drives customer behavior — beneath what customers say, and beyond what a single test can reveal.
Hit the ceiling on your testing?
If your optimization has gone flat and every test wins less than the last, the problem usually isn't the test — it's the hypothesis. That's exactly the conversation we have on a Growth Diagnostic.
Get my AI Diagnostic →