Week one said 'no effect.' Week four said +41%.
25 days, 63,173 randomized visitors. The first few days of this test read flat — then the full window settled at +41% checkout conversion and +40% revenue per visitor at z = 6.5. A case study in why we don't call tests early.
Two weeks into this test the honest read was 'roughly zero' — early purchase counts were small and the arms were trading places day to day. We reported it that way internally, because a split you'd only publish when it's winning isn't a split.
By the full window the picture was unambiguous: 2.62% checkout conversion for escaped visitors vs 1.86% for the webview control (+41%, z = 6.5), and $1.90 per visitor vs $1.36 (+40%). No outlier orders on either arm; 1,399 total purchases.
The lesson cuts both ways. Small early samples can hide a real effect just as easily as they can invent a fake one — which is why every number on this site is a completed window with its sample size and z-score attached, not a screenshot from a good afternoon.
After the test: The test is still running with a live control arm as the brand scales its Instagram spend.
| Arm | Visitors | Orders | CVR | Rev / visitor |
|---|---|---|---|---|
| A · escaped | 29,660 | 776 | 2.62% | $1.90 |
| B · control | 33,513 | 623 | 1.86% | $1.36 |
Brand anonymized. Window still accruing at time of publication — figures are the full randomized window through Aug 4, 2026.
Run this exact test on your traffic.
One script tag, 60-second install. Randomized 50/50 from the first visitor — in 7–14 days this page is your data.