Significance vs. Probability in Ad Testing

Statistical significance answers a narrower question than most media buyers think it does. A p-value tells you how surprising your results would look if the two ads were, in truth, identical performers. It does not tell you the probability that ad B beats ad A, and it says nothing about how much better B might be or what it costs you to be wrong. Those are the three things an advertiser actually needs, and none of them come out of a significance test.
What a p-value actually says
Run an A/B test and get a p-value of 0.03. The honest reading is: if A and B truly convert at the same rate, a gap this large or larger would show up about 3% of the time by chance alone. That is a statement about the data assuming no difference exists. It is not a statement about how likely it is that a difference exists, and it is not the “5% chance this is a fluke” shorthand that gets repeated in most marketing explainers.
This distinction matters because of how it gets used. A buyer who treats p < 0.05 as “B is proven better” will act with more confidence than the number supports, especially with small samples where a real 3% edge and a real 30% edge can produce similar p-values. See statistical significance in ads for the mechanics of computing this correctly, including the traps of peeking early and stopping a test the moment it crosses the line.
The question a Bayesian framing actually answers
A Bayesian approach starts from your prior belief about the likely range of outcomes, updates it with the data you collected, and outputs something closer to what you want: the probability that B beats A, given everything you have seen, plus a distribution of how much better it might be. That last part matters more than advertisers give it credit for. A 60% chance of a 2% lift is a different decision than a 60% chance of a 40% lift, and a frequentist significance test collapses that distinction into a single pass or fail threshold.
Neither framework is free of assumptions. Bayesian results depend on the prior you choose, and a bad prior can bias the answer as badly as a small sample biases a p-value. The honest case for the Bayesian frame in advertising is not that it is more rigorous. It is that it answers the question you are actually asking, which is “what should I do with my budget tomorrow,” not “would this gap be surprising under a null hypothesis.”
The sample size problem nobody budgets for
Detecting a real but modest lift requires more data than most ad sets ever accumulate. As a rough illustration: to reliably detect a 10% relative improvement in a 2% baseline conversion rate (moving it to 2.2%) at conventional confidence levels (95% confidence, 80% power), a standard two-proportion sample size calculation puts you in the range of 70,000 to 90,000 visitors per variant, translating to roughly 1,500 to 1,800 conversions per side. Most ad sets, even ones spending a few thousand dollars a week, do not clear that bar before creative fatigue or seasonality changes the underlying rate anyway.
This is not a reason to ignore the math. It is a reason to stop expecting textbook significance from ad tests and to build a decision process that works honestly with the sample sizes you actually get.
A decision rule that does not require textbook significance
Given that most tests never reach formal significance, treat the decision as an expected-value problem instead of a pass or fail gate. A workable threshold: if a variant is showing at least 70% probability of being the true winner (frequentist or Bayesian, either can approximate this) after accumulating at least 100 conversions per side, and the observed lift is large enough to matter to the account’s margin, act on it. Below 100 conversions per side, treat any lead as directional at best, and do not kill the trailing variant outright. Above 70% probability with a trivial lift, the decision does not matter enough to make either way, so save the budget for a bigger test.
The threshold is not magic. It is a stated line that stops the two failure modes that actually cost money: waiting forever for a “significant” result that will never arrive at typical ad-set volumes, and calling a winner off ten conversions because the graph moved.
Where this does not apply
None of this applies to compliance-critical or safety-critical decisions where a false positive is expensive regardless of ad spend, and it does not apply to tests running across wildly different audiences or placements, where the two arms are not actually comparable. In those cases, use the full rigor of a proper A/B test with pre-registered stopping rules, not a probability threshold borrowed from a marketing blog.
How YieldBI helps
YieldBI does not run its own significance calculator. What it does is surface which ad sets have accumulated enough conversions to make a call worth trusting, and flags the ones sitting in the no man’s land of low volume and shaky signal, so a buyer is not staring at a dashboard trying to guess whether ten days of data means anything.
The uncomfortable truth is that most “the data is inconclusive” moments in ad accounts are not actually inconclusive. They are decisions the buyer does not want to make without more certainty than the channel can supply. Waiting for that certainty is itself a decision, and it is usually the expensive one.