Bottom line
Ship variant F. It produced more purchases per visitor than the current page — vs purchases per 1,000 visitors — worth roughly extra purchases over the 10-week window at observed traffic. The result is decisive () and comes from small improvements at every step of the funnel, not from one gimmick. The lead held — and slightly widened — across all ten weeks, so it does not look like a novelty effect.
Retire H (, convincingly worse) and do not ship C: C’s redesign makes the call-to-action much more clickable (click-through up ) yet produces no extra purchases. The remaining variants — B, D, E and G — are statistically indistinguishable from the current page (all within ±).
Scoreboard
Totals and conversion rates over the full 70 days. “Purchases / 1,000” is the end-to-end rate (purchases ÷ visitors) — the metric that matters for revenue. Lifts and p-values compare each variant against A (two-proportion z-test).
| Variant | Visitors | CTR | Click→Cart | Cart→Buy | Purchases / 1,000 (95% CI) | Lift vs A | p-value | Verdict |
|---|
End-to-end purchases: only F beats the current page
Purchases per 1,000 visitors by variant, with 95% confidence intervals (error bars). The dashed line marks the current page’s rate.
Clicks ≠ purchases: the lesson from variant C
Lift versus A for two different metrics: click-through rate (grey bars) and end-to-end purchase rate (variant-coloured bars). C is the eye-catcher — and the trap.
C moves the CTA click rate from to (), and add-to-carts rise too — yet purchases end at . The extra carts rarely check out: the cart→purchase close rate collapses from on A to on C. That pattern reads as curiosity clicks or low-intent engagement rather than real demand. Shipping C would inflate engagement dashboards without adding revenue — we recommend against it, but its design is worth studying (see recommendations).
Why F wins: small gains at every step, compounded
Funnel conversion at each stage for the current page (A), the click-bait case (C), the winner (F) and the loser (H).
F is not dramatic on any single metric: click-through , click→cart , cart→purchase . Multiplied together, those modest gains compound to more purchases. Improvements spread across the whole funnel are harder to attribute to a single quirk (misclicks, miscaptured events) and more likely to persist after launch. H, meanwhile, loses ground at the top of the funnel and never recovers.
Stable across all ten weeks — no novelty effect
Weekly purchase rate (purchases per 1,000 visitors) for each variant. B, D, E and G are shown in light grey; A, C, F and H are emphasised.
F (green) sits above A (slate) in essentially every week, and the gap does not decay — if the lift were a novelty effect we would expect it to fade as the design gets familiar. H (red) sits below A throughout, with no sign of recovery. The ten-week window also covers ten full Monday–Sunday cycles, so weekday/weekend shopping patterns average out evenly for every variant.
Recommendations
- Roll variant F out to 100% of landing-page traffic. Expected impact: roughly more purchases from this page (≈ extra purchases per 10 weeks at current traffic levels).
- Verify after launch. Keep instrumentation on and watch purchase rate and revenue per visitor for two weeks post-rollout to confirm the lift holds outside the experiment.
- Retire H permanently — it costs the business of purchases with no offsetting benefit.
- Don’t ship C as-is, but investigate it. Its layout/copy makes people click and add to cart far more often, yet those carts rarely convert (close rate vs ). Session recordings, heatmaps and cart-abandonment analysis could reveal why — and feed a stronger next iteration.
- Close out B, D, E, G. None differs detectably from the current page; further testing of these exact designs would not pay off.
How we analysed this (and the caveats)
- Design. Parallel A–H test; each visitor randomly assigned to one variant with equal probability; consecutive days; ~ visitors total.
- Primary metric. End-to-end purchases per visitor, because that is what drives revenue. Click-through, click→cart and cart→purchase are secondary diagnostics.
- Statistics. Each variant was compared to A with a two-proportion z-test (two-sided); error bars are 95% confidence intervals. In plain language: F’s p-value of means that if F and A were truly equal, a gap this large would appear by chance less than once in ten thousand tries.
- Multiple comparisons. We compared seven variants to A. Even under a strict Bonferroni correction (significance threshold 0.05/7 ≈ 0.007), only F (positive) and H (negative) stand out — exactly the two we act on. All “no difference” verdicts also survive this correction.
- Randomisation check. Visitor totals per variant ran from to (χ² = on 7 degrees of freedom), .
- Limitations. The data is daily aggregates, so we cannot separate new vs returning visitors or inspect individual sessions. We counted purchases, not revenue — verify average order value during rollout in case F changes basket size. Nothing here measures longer-term retention effects.