Randomised test of versions A–H, 1 June – 9 August 2026 (70 days, visitors). Success metric: purchases per visitor.
Visitors tested
~/day, split 8 ways
Current page (A)
purchases per 1,000 visitors
Best page (F)
purchases per 1,000 visitors
F vs current page
95% CI , p
Worst page (H)
vs current page, 95% CI
What we recommend
Ship version F to 100% of traffic. It is the only redesign that sells more: more purchases per visitor than the current page (95% CI , p ). At today's traffic that is roughly extra purchases a day (about a week) from this page.
Retire version H. It is reliably worse than what we have today (, 95% CI ). Whatever it changed, it should not be repeated in future designs.
Drop B, D, E and G. None of them moved purchases (every one is within of the current page). The test was big enough that if they helped at all, the help is small — not worth the cost and risk of a redesign.
Do not be fooled by C. It is the best page in the test at getting clicks and add-to-carts, and the joint-worst at turning them into purchases. Net effect on sales: nothing. See "the click trap" below.
Next test: take what F changed and push it further, and re-check F on revenue per visitor (this test only counted purchases, not basket value).
Results at a glance
Each visitor saw one page for the whole test. The funnel is strictly sequential: visitors → clicked the call-to-action → added to cart → purchased. Rates below are conditional on the previous step, so multiplying the three of them gives the bottom-line number in the "purchases per 1,000 visitors" column.
Page
Visitors
Clicked CTA (% of visitors)
Added to cart (% of clickers)
Purchased (% of carts)
Carts per 1,000 visitors
Purchases per 1,000 visitors
Lift vs A
95% CI
p-value
Verdict
Verdict uses a two-proportion test on purchases per visitor against page A, with a stricter threshold to account for making seven comparisons at once (p < 0.007). "No difference" means we could not detect one, not that the pages are identical.
The bottom line: purchases per 1,000 visitors
Bars are the ten-week totals; the whiskers are 95% confidence intervals. F stands clearly apart from the pack; H stands clearly below it. The other six pages overlap each other.
Where the differences come from
Each step of the funnel, by page
Three separate rates, each measured on the people who reached that step. Notice that a page can win big on one step and give it all back on the next.
F wins a little at every step — and that compounds
F is not dramatic anywhere: more clicks per visitor, higher add-to-cart rate among clickers, higher purchase rate among carts. Multiply the three small gains and you get the gain in purchases. This is the healthy pattern: the page attracts more interest without diluting the quality of that interest.
H loses a little at every step — and that compounds too
H is slightly behind the current page on clicks, on add-to-cart and on purchase completion. None of the three gaps looks alarming on its own; together they cost of purchases. It is the mirror image of F.
The click trap: version C
C is by far the best page in the test at the top of the funnel — more clicks per visitor than the current page and more add-to-carts per visitor. Both differences are enormous and unambiguous.
And yet it sells nothing extra: purchases per visitor (95% CI ) — a tie with today's page. The reason is visible in the last step: only of C's carts turn into a purchase, versus on the current page. C is very good at persuading people to click and cart who were never going to buy.
The lesson for future tests: if we had run this test on click-through rate, C would have been declared a spectacular winner (+40%) and shipped, and the business would have gained nothing. Judge landing pages on purchases per visitor.
The mirror image: version E
E does the opposite of C: it drives the fewest clicks of any page ( vs today), but the people who do click are the most committed — of its carts convert, the highest in the test. The two effects cancel out almost exactly ( on purchases). Useful to know: E's copy/CTA is a strong filter, not a strong seller.
Did the result hold up over the ten weeks?
Yes. Traffic and conversion swing noticeably from day to day and week to week (weekends, promo cycles) — but all eight pages swing together, which is exactly what a randomised split should look like. F is above the current page in of the 10 weeks, and H below it in of 10.
Weekly purchases per 1,000 visitors
Click a legend entry to hide/show a page. A (current), F (winner) and H (worst) are drawn boldest.
Running result: cumulative lift vs the current page
Each line is "purchases per visitor so far, compared with page A so far". After the first couple of weeks the picture stops changing: F settles around +20%, H around −13%, everything else hugs the zero line. There is nothing to gain from running the test longer.
How much is F worth?
The page received about visitors a day across all eight versions during the test. If every one of those visitors had seen F instead of the current page A, we would expect roughly:
Extra purchases / day
vs the current page
Extra purchases / week
~ per year at this traffic
Relative increase
purchases from this page
Plausible range
95% confidence interval
Multiply by average order value to get revenue — but see the caveats: we counted purchases, not basket size.
Method, in plain terms
What we compared. For each page: purchases ÷ visitors over the whole ten weeks. Daily rates are noisy; pooling the whole period is the right unit for the decision.
Was the split fair? Yes — the eight pages received between and visitors (a spread of ), consistent with an even random allocation.
How we decided "real" vs "noise". A standard two-proportion z-test per page against A, plus a 95% confidence interval on the lift. Because we made seven comparisons, we required p < 0.007 (0.05 ÷ 7) before calling a winner or loser. F (p ) and H (p ) pass that bar by a wide margin; nothing else comes close.
How sensitive was the test? Very. For the five "no difference" pages, the confidence intervals rule out anything better than about . So those pages are, at best, marginal.
Data shape. Counts are aggregated per day and per page, and each funnel step is counted only among people who completed the previous one (clicks among visitors, carts among clickers, purchases among carts).
Caveats and things we could not see in this data
Purchases, not money. We have no revenue or basket-size data. If F wins by pushing a cheaper product, the revenue gain will be smaller than the purchase gain. Check revenue per visitor in the first two weeks after rollout.
No segments. The data are daily totals, so we cannot break the result down by device, channel, new vs returning, or geography. F is the best page on average; a follow-up with segmented logging would tell us whether it is best for everyone.
One page, one metric. We measured what happened after landing. Effects further downstream (returns, cancellations, repeat purchase) are invisible here; C's low-intent carts, for example, could also carry extra support or abandonment cost.
No novelty effect visible, but worth noting: the lift is flat across all ten weeks, so this is not a short-lived curiosity bump.
Roll out with a guardrail. Ship F to 100%, and keep a small holdback (say 5% on A) for two weeks to confirm the lift in production and to have a clean baseline if anything else on the site changes at the same time.