Strategy

    What Is Statistical Significance?

    Statistical significance is the threshold at which a difference between two test variants (an A vs B creative, landing page, audience) is unlikely to be explained by random chance. The industry standard is 95 percent confidence (p < 0.05), meaning there is less than a 5 percent probability the observed lift is a coincidence.

    Md Morshed Parvej Patwary
    By
    Updated · Next review

    Formula

    Required per-variant sample size approx 16 * p * (1 - p) / MDE^2

    Where p is your baseline conversion rate and MDE is the minimum detectable effect you care about (as a decimal). A 2 percent baseline CVR and 20 percent relative lift target needs roughly 8,000 users per variant to hit 95 percent confidence.

    Why 'A beat B by 12 percent' is usually noise

    A brand tests two Meta creatives for 3 days. Ad A: 41 purchases from 1,200 clicks. Ad B: 36 purchases from 1,180 clicks. That looks like a 14 percent lift for A. Plug into a significance calculator: p-value 0.58. The test needs about 4x more traffic before you can trust the winner. Calling A the winner and scaling it is exactly how brands scale losing creatives.

    Benchmarks

    • Confidence threshold: 95 percent minimum for scaling decisions; 90 percent for directional signal.
    • Minimum conversions per variant: 50 as a floor, 100 to be comfortable.
    • Test duration: at least one full purchase cycle (usually 7 to 14 days), regardless of significance.
    • Meta A/B tools bake in significance; ad-set-level manual splits do not.

    Why it matters

    Undersized creative and audience tests are the single biggest source of self-inflicted CPA damage. Every 'winner' scaled off 30 conversions has roughly a coin flip's chance of actually being better than the loser. Treating significance as a gate (not a nice-to-have) is what separates operators from tinkerers.

    Common mistakes

    • 1.Calling winners at low conversion counts. Under 50 conversions per variant the noise dominates.
    • 2.Peeking daily and stopping the test the moment one variant leads. That inflates false-positive rate to 30+ percent.
    • 3.Ignoring baseline CVR when planning sample size. Low-CVR pages need much more traffic to prove a lift.
    • 4.Confusing statistical significance with practical significance. A 2 percent lift can be statistically real and commercially meaningless.

    Put Statistical Significance to work

    FAQs about Statistical Significance

    What confidence level should I use for ad tests?

    95 percent for scaling decisions. 90 percent is fine as a directional read to inform which tests to expand, but not to reallocate real budget.

    How many conversions do I need per variant?

    At least 50; ideally 100 or more. Below that the confidence interval is so wide that most 'winners' are noise.

    Is Meta's A/B Test tool statistically valid?

    Yes, when the test runs to completion and Meta declares a winner. Stopping it early or running below-recommended budget invalidates it, same as any other test.

    Can I test creatives at the ad level instead of using A/B tests?

    Only with big caveats. Meta's delivery algorithm will spend unevenly across creatives to hit the objective, so what you observe is a mix of creative effect and delivery bias. A/B tests randomise properly; ad-level splits do not.