What Is Statistical Significance?
Statistical significance is the threshold at which a difference between two test variants (an A vs B creative, landing page, audience) is unlikely to be explained by random chance. The industry standard is 95 percent confidence (p < 0.05), meaning there is less than a 5 percent probability the observed lift is a coincidence.

Formula
Where p is your baseline conversion rate and MDE is the minimum detectable effect you care about (as a decimal). A 2 percent baseline CVR and 20 percent relative lift target needs roughly 8,000 users per variant to hit 95 percent confidence.
Why 'A beat B by 12 percent' is usually noise
A brand tests two Meta creatives for 3 days. Ad A: 41 purchases from 1,200 clicks. Ad B: 36 purchases from 1,180 clicks. That looks like a 14 percent lift for A. Plug into a significance calculator: p-value 0.58. The test needs about 4x more traffic before you can trust the winner. Calling A the winner and scaling it is exactly how brands scale losing creatives.
Benchmarks
- Confidence threshold: 95 percent minimum for scaling decisions; 90 percent for directional signal.
- Minimum conversions per variant: 50 as a floor, 100 to be comfortable.
- Test duration: at least one full purchase cycle (usually 7 to 14 days), regardless of significance.
- Meta A/B tools bake in significance; ad-set-level manual splits do not.
Why it matters
Undersized creative and audience tests are the single biggest source of self-inflicted CPA damage. Every 'winner' scaled off 30 conversions has roughly a coin flip's chance of actually being better than the loser. Treating significance as a gate (not a nice-to-have) is what separates operators from tinkerers.
Common mistakes
- 1.Calling winners at low conversion counts. Under 50 conversions per variant the noise dominates.
- 2.Peeking daily and stopping the test the moment one variant leads. That inflates false-positive rate to 30+ percent.
- 3.Ignoring baseline CVR when planning sample size. Low-CVR pages need much more traffic to prove a lift.
- 4.Confusing statistical significance with practical significance. A 2 percent lift can be statistically real and commercially meaningless.
Put Statistical Significance to work
Free calculators
Related services
FAQs about Statistical Significance
What confidence level should I use for ad tests?
95 percent for scaling decisions. 90 percent is fine as a directional read to inform which tests to expand, but not to reallocate real budget.
How many conversions do I need per variant?
At least 50; ideally 100 or more. Below that the confidence interval is so wide that most 'winners' are noise.
Is Meta's A/B Test tool statistically valid?
Yes, when the test runs to completion and Meta declares a winner. Stopping it early or running below-recommended budget invalidates it, same as any other test.
Can I test creatives at the ad level instead of using A/B tests?
Only with big caveats. Meta's delivery algorithm will spend unevenly across creatives to hit the objective, so what you observe is a mix of creative effect and delivery bias. A/B tests randomise properly; ad-level splits do not.
Related terms
Controlled test isolating one variable to prove which version wins.
Repeatable process for launching, judging, and killing ad variants.
% of clicks or sessions that complete the target action.
Measures the lift ads caused vs what would have happened anyway.
Randomised users kept ad-free so you can measure true lift.