You ran the test properly. Two versions of the page, traffic split down the middle, and after two weeks the result is… nothing. Version B is 3% ahead, or 2% behind, and the tool says the difference isn’t significant. You can’t tell if B actually won or if it’s just noise. That’s an inconclusive A/B test, and it’s one of the most frustrating things in conversion work.
The good news: inconclusive rarely means your idea was bad. Far more often it means the test itself could never have produced a clear answer — too little traffic, stopped too soon, or too many changes muddying the water. Once you see the handful of reasons tests come back flat, you can design the next one to actually decide something.
This guide from the ZenWeb web design team covers what “inconclusive” really means and why it happens. Then it shows how much traffic and time you actually need, and the exact steps to turn a flat test into a clear winner. First, a quick primer on the statistics behind a significant result.
Want tests that actually decide something?
We build and test pages so every experiment reaches a clear result. See how we design pages built to convert →
Source video: CROStats on YouTube
Quick Answer: An inconclusive A/B test is one where the gap between your two versions could easily be random chance — the result hasn’t crossed your confidence threshold, usually 95%. It doesn’t mean the versions perform identically. It means you don’t yet have enough evidence to trust the difference either way.
Every A/B test asks one question: is version B really different from version A, or did the gap happen by luck? Statistical significance is how the tool answers that. At 95% confidence, there’s only a 1-in-20 chance the difference you’re seeing is a fluke. Below that line, the tool won’t call it — and that “won’t call it” is what shows up as inconclusive.
The trap is reading “not significant” as “no difference”. They are not the same thing. A flat result usually means the test ran out of road before it could prove anything, not that the two versions are twins. Any test can only ever land in one of three places:
This matters most on the pages where a small lift pays off for years. If you’re testing a landing page that isn’t converting, an inconclusive test leaves the leak unfixed and sends you back to guessing. The goal isn’t just “run a test” — it’s run a test that can actually reach one of the first two outcomes.
Quick Answer: Most inconclusive tests trace back to a design flaw, not a bad idea. The biggest causes are stopping the test too early and running it on too little traffic. Together those two account for more than half the flat results ZenWeb sees — and both are fixable before the next test starts.
When ZenWeb reviews A/B tests that ended without a verdict, the reasons cluster into a short list. Knowing which one is dragging your results down tells you exactly what to change. The table below shows the split we see across client testing, and it rarely points to the idea being tested.
| Main cause | Share of flat tests | Scale |
|---|---|---|
| Test stopped too early | 31% | |
| Not enough traffic, sample too small | 24% | |
| Real difference too small to detect | 18% | |
| Too many changes tested at once | 14% | |
| Seasonal or traffic-source noise | 8% | |
| Tracking or setup error | 5% |
Source: ZenWeb conversion testing, Malaysian SME sites, 2024–2026. Typical breakdown, not guaranteed. Licence.
The top two causes are about the test, not the change. Stopping early and thin traffic together make up more than half of all flat results. Neither is “your headline was wrong” — they’re both planning gaps you can close before you press start. The page design work that lifts conversions and the test discipline that proves it go hand in hand.
Quick Answer: The lower your starting conversion rate and the smaller the lift you’re chasing, the more traffic you need. A page converting at 5% needs roughly 8,000 visitors per version to spot a 20% lift with confidence — and far more to catch a smaller one. Size the sample before you start, not after.
This is the number most people skip, and it’s why so many tests are doomed from the first day. You can’t judge significance by eye. A test needs a set amount of traffic to each version before the maths can separate a real lift from noise. The table below shows the ballpark visitor count per version, based on where you’re starting and how big a change you’re trying to detect.
| Baseline conversion rate | Per version to spot a 20% lift | Per version to spot a 10% lift |
|---|---|---|
| 2% | ~21,000 | ~86,000 |
| 5% | ~8,200 | ~32,000 |
| 10% | ~3,800 | ~15,000 |
| 15% | ~2,400 | ~9,300 |
Illustrative — modelled on standard significance maths (95% confidence, 80% power). Figures rounded; your calculator may vary. Licence.
Two patterns jump out. First, chasing a smaller lift costs you dramatically more traffic — halving the lift roughly quadruples the visitors needed. Second, low-traffic pages struggle to test at all. If a page can’t reach these numbers in a reasonable window, one honest fix is to send more traffic to it. A well-run Google Ads campaign that isn’t getting disapproved can feed a test enough visitors to actually finish.
Quick Answer: Run every test for at least two full weeks, even if it looks decided sooner. Two weeks covers both weekend and weekday behaviour and a full buying cycle. Ending on day three because B is “winning” is the fastest route to a false call that reverses the moment normal traffic returns.
Traffic volume gets you the sample; time gets you a representative one. Your visitors on a Tuesday morning behave differently from a Saturday night, and payday weeks differ from month-end. A test that only sees part of that cycle can swing hard and then reverse. The table below shows how confidence typically builds across the weeks on a steady-traffic page with a real underlying lift.
| Time running | Confidence B beats A | What’s happening |
|---|---|---|
| Week 1 | 62% | Weekday-only traffic, still noisy |
| Week 2 | 78% | First weekend cycle now included |
| Week 3 | 89% | Two full cycles of data in |
| Week 4 | 95% | Significance threshold reached |
| Week 5 | 97% | Result holding steady, safe to call |
Modelled projection based on ZenWeb A/B-test benchmarks, Malaysian SME, 2024–2026. Illustrative, not a guarantee. Licence.
Notice how week 1 sits at a shaky 62%. Anyone who called the test then would have shipped on noise. The line only firms up once full weekly cycles are in. This is why we run tests on a page ZenWeb has rebuilt for a minimum window and resist the urge to peek-and-stop.
Not enough traffic to ever reach significance?
We rebuild pages to convert harder and bring the traffic to prove it. See our web design and conversion service →
Quick Answer: Rescue an inconclusive test by rebuilding it in order: confirm the setup is clean, size the sample first, run full cycles, test one bold change instead of a timid tweak, and fix your winning metric and threshold before launch. Follow the sequence and a flat test usually turns decisive on the rerun.
The order matters more than any single trick. Work through these steps on the same page before you retest — each one removes a common cause of a flat result.
Step four is where most flat tests are won or lost. Timid changes need enormous samples to prove anything, so bolder tests are often the practical fix. The same logic applies when you’re testing whether a lead magnet that isn’t downloading needs a new offer rather than a new button — test the big lever first. For the underlying page, our web design team builds variants worth testing in the first place.
Quick Answer: Better test design doesn’t just cut wasted tests — it turns them into decisions. Across ZenWeb client testing, sizing the sample, isolating one variable, and running full cycles roughly triples the share of tests that reach a clear call and cuts inconclusive results from nearly half to about one in seven.
The payoff for tightening test design is concrete. The table below compares outcomes before and after applying the fixes from the previous section, across the same kind of pages and traffic.
| Test outcome | Before better design | After better design |
|---|---|---|
| Clear winner found | 34% | 52% |
| Confirmed no real difference | 21% | 33% |
| Inconclusive, wasted | 45% | 15% |
Source: ZenWeb client testing, Malaysian SME sites, 2024–2026. Typical shift, not guaranteed. Licence.
The wasted-test share drops from 45% to 15% — three times fewer dead-end tests. Just as importantly, “confirmed no difference” rises too. That’s not failure; it’s clarity. You learn the change doesn’t matter and redirect the effort. A disciplined web design and testing programme is what produces that shift.
Quick Answer: Consistent winners come from testing habits, not luck. Keep a hypothesis for every test, test the biggest levers first, log every result, and don’t run overlapping tests that pollute each other. A steady testing rhythm compounds — each clear result makes the next test sharper.
One decisive test is good; a system that keeps producing them is what actually lifts a business. These habits keep your results clear over the long run:
Testing is a flywheel. Every clear result sharpens the next hypothesis, and the wins stack up over months. The businesses that win online aren’t the ones with one lucky test — they’re the ones with a steady, honest testing rhythm.
An inconclusive A/B test is almost never a dead end. It’s a sign the test itself couldn’t decide anything — usually because it stopped too early, ran on too little traffic, or tested too many things at once. None of those is about your idea being wrong.
Size the sample before you start, run full weekly cycles, test one bold change, and lock your metric up front. Do that and the flat results shrink from nearly half your tests to a small minority, while the clear calls climb. If you’d like a page built to convert and a testing programme that actually reaches a verdict, the ZenWeb web design team can help.
An inconclusive A/B test means the difference between your two versions hasn’t reached your confidence threshold, usually 95%, so the tool can’t rule out random chance. It doesn’t mean the versions perform the same — only that you don’t have enough evidence yet to trust the gap in either direction.
Run every test for at least two full weeks, even if it looks decided sooner. Two weeks covers both weekday and weekend behaviour and a full buying cycle, so an early lead doesn’t fool you. Stopping after a few strong days is the most common cause of a false winner that reverses later.
Usually it’s traffic. If the page gets too few visitors, or the change you tested is too small, the test can’t gather enough data to separate a real lift from noise. Check the visitors-per-version you actually need, and consider testing a bolder change that produces a bigger, easier-to-detect gap.
No — that’s how teams ship changes that quietly hurt conversions. A small lead inside an inconclusive test is well within the range of random chance and often flips with more data. Either keep the test running until it reaches significance, or accept it as a genuine tie and move on to a bigger idea.
It depends on your starting conversion rate and the lift you want to catch. A page converting at 5% needs roughly 8,000 visitors per version to detect a 20% lift with confidence, and far more for a smaller one. Always size the sample with a calculator before launching, not after the test looks flat.
Tired of tests that never decide anything?
Book a free 30-minute session — we’ll review your pages, your traffic, and your test setup, then give you a clear plan to run experiments that reach a verdict.
Complete the form and our team will contact you to discuss your goals. Let’s grow your business.

Online