F turns more landings into purchases.
C turns more landings into clicks.
Prioritize the end of the funnel, not the loudest signal at the top. F is the strongest rollout candidate; C’s striking click-through rate does not translate into a demonstrated purchase gain.
Recommendation: move F toward a staged rollout
Calculating the full-period purchase results…
Before committing: confirm whether the daily visitor counts include repeat visitors. Assignment was at visitor level, so repeat observations need visitor-level or clustered inference. Also check revenue, margin, performance and customer-experience guardrails; these data contain none of them.
Only a purchase gain earns a rollout
Each redesign is compared with A over the entire experiment. The intervals below allow for testing seven candidates against the same control—not just selecting the best-looking result.
Dots show the observed difference in purchases per 10,000 logged visitors. Lines are approximate, Bonferroni-adjusted intervals with 95% family-wise coverage across the seven comparisons. Crossing zero means the data do not clearly establish improvement or harm. These intervals assume independent visitor observations; see limitations below.
| Version | Visitors* | Purchases | Purchase rate | Lift vs A | Adjusted interval, pp | Evidence vs A† |
|---|
pp = percentage points, not percent change. *Visitors are summed daily counts, not verified unique people. †“No clear difference” does not mean equivalence. Comparisons here are against A, not all candidate-to-candidate pairs.
More clicks are not necessarily more customers
The same 10,000-visitor starting point makes the trade-offs visible. Stage counts below are cumulative; transition percentages describe the subset that reached the previous stage.
| Version | Clicks / 10k | Carts / 10k | Purchases / 10k | Visitor → click | Click → cart | Cart → purchase |
|---|
C: a top-of-funnel win, not a purchase win
F: more people make it through
These are diagnostic descriptions, not separate causal effects at each step. Redesigns can change who clicks or adds to cart. A lower cart-to-purchase rate does not, by itself, prove that the checkout experience became worse. Layout, copy and CTA changed together, so this test does not isolate which element caused the result.
Check the ten weeks, not just the final total
Weekly purchase rates use the same Monday–Sunday windows for every version. Use the controls to compare any candidate with the current page.
Does F’s lead persist?
H: keep the time pattern in context
Weekly and half-period views are exploratory consistency checks, not additional winner-selection tests. Shared changes in traffic mix or calendar conditions can move all versions together. Neither a late rebound nor a single good week overturns the full-period comparison.
Show weekly purchase rates for all versions
Act on the purchase signal; validate the business impact
Resolve the unit of analysis
Audit visitor IDs, repeat visits, event attribution and conversion windows. Re-estimate the F–A effect at the randomized visitor level, including any delayed purchases. Keep the observed aggregate lift separate from its statistical certainty.
Retain an A holdout
If the visitor-level result holds, stage the rollout with predefined stop rules. Monitor purchases, revenue and contribution margin per visitor, plus page performance, cancellations and returns. The projected extra purchases are not a revenue forecast.
Use the funnel to design the next test
Do not ship C for its click rate or H based on a late-period improvement. Treat E’s lower-click, higher-completion pattern as a research lead. Isolate promising components of F in a subsequent test rather than inferring which bundled change worked.
What these data support—and what they do not
How the analysis was done
- Decision endpoint: purchases divided by visitors, the furthest downstream measured outcome. This report chooses that endpoint for the product decision; no prespecified analysis plan was supplied.
- Aggregation: all 70 days, weighted implicitly by visitor counts. Every version was running concurrently. Weekly results are seven-day aggregates; halves are the first and last 35 days.
- Uncertainty: for candidate v, the standard error of the purchase-rate difference is √[pv(1−pv)/nv + pA(1−pA)/nA]. Intervals use a normal critical value of 2.6901, equivalent to a two-sided 0.05/7 threshold per comparison.
- Multiple candidates: Bonferroni controls the chance of any false positive among these seven comparisons under the model assumptions. It does not correct for unreported interim stopping, metric searching or earlier experiments.
Important boundaries
- Repeat visitors: daily aggregates cannot establish whether observations are independent. Sticky assignment prevents switching arms, not repeated counting. Person-level conversion and valid person-level uncertainty require visitor IDs; day-level grouping alone cannot resolve this.
- Follow-up: timing and attribution of purchases are not specified. Unequal or incomplete follow-up near the end of the test could affect the measured endpoint.
- Selection: the leading observed lift may overstate the future effect. F’s lead versus A is not a formal claim that it beats every other redesign, nor a guarantee that the effect will persist.
- Business value: no order values, costs, device/channel segments or operational guardrails were supplied. Purchase lift is not necessarily profit lift. No minimum commercially worthwhile effect was specified.