Landing page experiment: what happened, and what to do next
Eight versions (A = current page, B–H = redesigns) · 1 June – 9 August 2026 (70 days) · visitors, randomly split
The short version
- Ship version F. It turned of visitors into buyers, compared with on today's page. That is about more purchases from the same traffic. The plausible range is . F beat the current page in every week of the test.
- Don't be fooled by version C. It got by far the most clicks () and even more add-to-carts, but no more purchases (). The extra activity came from people who never bought.
- Drop version H. It cost us sales: purchases versus the current page.
- B, D, E and G performed the same as the current page. Any differences are small enough to be chance.
Version F: purchases per visitor
vs current page (A). Clear, repeatable win.
Version C: clicks vs purchases
Click lift → purchase lift. Clicks didn't turn into sales.
Version H: purchases per visitor
vs current page (A). Clearly worse.
Extra orders if F had run on all traffic
Over the 10 weeks, at the traffic level seen in the test.
1. Which page sells more?
The number that matters is purchases per visitor: out of everyone who lands on the page, how many end up buying. The chart shows each redesign's change against the current page. The dot is our best estimate. The bar is the range the true effect plausibly falls in (a 95% interval). If a bar crosses the zero line, we cannot tell that version apart from the current page.
Chart library could not be loaded; see the tables below for all numbers.
Change in purchases per visitor vs version A, with 95% range. Green = clearly better, red = clearly worse, grey = no clear difference. The verdicts account for the fact that we tested seven redesigns at once (see "How sure are we?").
Only two versions clear the bar. F is better, and by a wide margin: its whole range sits well above zero. H is worse. Everything else is statistically indistinguishable from what we have today.
2. Clicks are not sales: the version C trap
If we had only tracked clicks on the call-to-action, version C would look like a runaway success. The chart below shows three measures side by side for each redesign, all compared with the current page: clicks, add-to-carts and purchases, each counted per visitor.
Chart library could not be loaded; see the tables below for all numbers.
Change vs version A in clicks per visitor, add-to-carts per visitor and purchases per visitor.
What happened with C: The most likely reading is that C's call-to-action pulls in people who are curious but not ready to buy. They click and even drop something in the cart, then leave. Optimising for clicks here would have shipped a page that does nothing for sales.
Version E is the mirror image: This is another reminder that click rate is a poor proxy for revenue on this page.
Where in the journey each version wins or loses
The purchase journey has three steps: land → click the call-to-action → add to cart → buy. The table shows the conversion at each step, with the change against version A underneath.
3. Is F's win real, or a lucky streak?
A result that shows up in one burst and then fades is a warning sign, for example one that comes from the novelty of a new design. F's advantage does not behave like that. The chart shows each version's weekly purchase rate. F (green) sits on top almost every week.
Chart library could not be loaded; see the tables below for all numbers.
Purchases per visitor by week (weeks run Monday–Sunday). Highlighted: A (current), F, C, H. Other versions in grey. Click legend items to show or hide lines.
4. Full results
"Adjusted p" is the probability of seeing a gap at least this large by chance if the version truly made no difference. It is corrected for testing seven redesigns at once (Holm method). Values below 0.05 are treated as a real effect. "Day-by-day check" repeats the comparison using the 70 daily results. It does not assume every visitor is independent, so it is a stricter test, and it agrees with the main result.
5. How sure are we?
The split was fair
We guarded against "lucky winners"
When you test seven redesigns, one of them can look good purely by chance. We raised the bar for calling something a win to account for this. F and H pass that stricter bar comfortably. None of the others come close.
The data are internally consistent
What this experiment can't tell us
- Order value and revenue. We counted purchases, not money. If F's layout nudges people toward cheaper items, revenue could rise by less than 21%. Check average order value for F before and after launch.
- Returns and repeat buying. We saw only the first purchase after landing, not what happened later.
- Seasonality. The test ran through summer. The ranking is very unlikely to flip, but the exact size of the lift may differ in other seasons.
- Visitor counting. Figures are daily per-version totals. Some people may have visited on several days and been counted more than once. This affects all versions equally, so it does not change the comparison.
6. Recommendation
- Roll out version F to all visitors. Expected impact: roughly more purchases from landing-page traffic, or about extra orders per week at test-period traffic levels.
- Keep a small holdout (e.g. 5% on version A) for a few weeks after launch to confirm the gain holds and to check average order value, which this test did not measure.
- Retire H, and don't ship C, however good its click numbers look.
- Mine C and E for lessons. C shows that a louder call-to-action can attract low-intent clicks. E shows that fewer, better-qualified clicks can sell just as well. A natural next test is F combined with E's more selective call-to-action.
- Going forward, judge page tests on purchases per visitor (and ideally revenue per visitor), not click-through rate.
Method notes: rates are totals over the full 70 days. Differences vs A use a two-proportion comparison, with 95% intervals from the normal approximation. The relative-lift intervals divide the absolute interval by A's rate. Multiple-comparison correction uses Holm's method across the seven redesigns. The day-by-day check is a paired t-test on daily purchase-rate differences (70 days). All numbers on this page are computed from the raw data embedded in this file.