The test finished and the confidence interval still straddles zero
In short
An interval that contains zero is not evidence a channel does nothing. Multiply the half-width of the 95% interval by about 1.4 and you have the smallest effect the test could reliably have detected: if that number is larger than any effect the channel could plausibly produce, the test was uninformative by construction and the result says nothing about the channel at all.
Key takeaways
- Half-width of the 95% interval times 1.4 is the effect the test had a fair chance of detecting. Compute it before arguing.
- Inconclusive means the detectable floor sits above the plausible effect. Null means the floor sits below it and nothing was found.
- Halving the detectable effect needs roughly 4 times the units, which is why a rerun at the same size changes nothing.
- Extra weeks are worth less than extra markets once weekly observations share seasonality.
- An accepted bound is a publishable result: the effect is smaller than X, and X is a number you can put in a budget argument.
A geo test ends, the estimate is 1% with an interval running from minus 9% to plus 11%, and somebody in the room says the channel does not work. That conclusion is not available from this data. The interval is telling you the experiment could not resolve an effect of the size you were looking for, which is a statement about the design, not about the channel.
Before anything is decided, work out what this test could ever have detected. The shortcut takes 30 seconds and needs nothing but the readout you already have: take the half-width of the 95% interval and multiply by roughly 1.4. That is the smallest true effect the test had a reasonable chance of returning as significant. In the example above, half-width 10% gives a detectable floor near 14%.
The number that settles the argument in 30 seconds
The 1.4 comes from the 2 quantities every power calculation contains. A conventional 95% two-sided test needs an estimate about 2 standard errors from zero to be significant, and having an 80% chance of getting there needs the true effect to sit about 2.8 standard errors out. The reported interval is already 1.96 standard errors either side of the estimate, so the ratio between the detectable effect and the half-width is 2.8 divided by 1.96.
- Take the half-width of the reported 95% interval, in the same units as the estimate. For an interval of minus 9% to plus 11%, the centre is 1% and the half-width is 10%.
- Multiply by 1.4. That is the minimum detectable effect the test actually had, given the variance it actually encountered.
- Work out the largest effect the channel could plausibly produce. A channel whose spend equals 2% of revenue, returning 1.5 units of revenue per unit of spend, moves total sales by about 3%.
- Compare the 2. A detectable floor of 14% against a plausible ceiling of 3% means the test was incapable before it started, and no analysis rescues it.
- Only if the floor sits below the plausible effect does the result carry information about the channel.
Inconclusive and null are different findings, and the difference is arithmetic
Both produce an interval containing zero. Only one of them licenses a decision about the channel. Reading them the same way is how a working channel gets cut and a dead one gets kept.
| Interval on sales | Detectable floor | Reading | Next move |
|---|---|---|---|
| Minus 9% to plus 11% | About 14% | Inconclusive. The design could not see an effect of any plausible size | Rerun with more units, or accept and say so |
| Minus 1.2% to plus 1.6% | About 2% | Informative null. The effect is smaller than 1.6% at the top of the interval | Act on it: reallocate, and re-test at a larger spend step |
| Plus 0.4% to plus 9% | About 6% | Positive but imprecise. Direction established, magnitude not | Size any reallocation on the lower bound, not the point estimate |
A test that could not have detected the effect you were looking for has not measured the channel. It has measured your sample size.
Five faults behind a wide interval, cheapest to rule out first
Work through these in order. The first 2 are settled from the test design document, the next 2 from the data you already have, and only the last one is a statement about the world.
- Too few units. In a matched-pair geo design the standard error scales with the pair-difference variability divided by the square root of the number of pairs. With 10 pairs and pair differences that vary by about 6% of baseline week to week, the detectable floor lands near 5%. Halving that floor needs about 40 pairs, not 15.
- Duration shorter than the conversion lag. If a meaningful share of conversions land more than a week after the interaction, the final weeks of a 4-week test are counting a truncated version of the effect and diluting the estimate toward zero.
- Spillover between arms. Check whether the measured effect shrinks in control markets adjacent to test markets. Contamination compresses the difference and widens nothing, so it produces a small estimate with a normal-looking interval — the most misleading failure of the 5.
- Baseline divergence that predates the test. Plot the treated-minus-control series for 8 to 12 weeks before the start. A trend or a step there means the pre-period was not parallel and the estimate carries that gap.
- The effect is genuinely small. Real, positive, and below anything this design could resolve. That is a legitimate finding, and it belongs in the readout as a bound rather than as a failure.
One fault sits outside that list because it is not statistical: the outcome definition changing mid-test. A window edit, a tag migration or a switch of attribution model between the pre-period and the test period restates the outcome column underneath the experiment, and the restatement is indistinguishable from lift. Anything of that kind belongs behind the dual-reporting discipline set out in switching attribution models mid-year. The same care applies to the market-level numbers feeding the test, which usually arrive through a rollup that has its own reconciliation problems, described in when the multi-market rollup does not add up.
Why another 4 weeks buys less than another 10 markets
Extending a test feels cheaper than expanding it, and it is, which is why it is the usual answer. It is also the weaker lever. Adding weeks reduces the standard error only to the extent that each new week is independent information, and weekly observations from the same markets share seasonality, promotional calendars and weather. The effective sample grows more slowly than the calendar does.
Adding markets adds genuinely new units, provided the new pairs are matched as carefully as the originals. When the footprint is too small to supply them, the design has to change rather than stretch, which is the situation handled in pairing markets for a holdout in a small footprint. Extending is still the right move in one specific case: when the conversion lag, not the sample size, is what truncated the effect.
Four endings, and the condition that selects each
| Condition established | Ending | What it requires |
|---|---|---|
| Detectable floor far above the plausible effect, units available | Rerun larger | Roughly 4 times the units to halve the floor, and a bigger spend step to lift the signal |
| Lag truncated the effect, sample was adequate | Extend the measurement window | Run the read past the flight by at least the lag, on the same arms |
| Spillover or unmatched baselines confirmed | Redesign | A different unit of assignment, a synthetic control, or an audience split instead of a geo split |
| Floor already below the plausible effect and nothing found | Accept the bound | Publish the upper limit and stop paying for repeat tests of the same question |
The branch teams skip is the last one. An accepted bound is a real result: it says the effect is smaller than a stated number, which is often exactly what a budget conversation needs. It is also the honest ending for questions that experiments answer badly, such as brand search, where an inconclusive test is routinely read as proof of value in whichever direction the reader already preferred — the trap examined in when branded search absorbs credit for everything.
Writing up a test that did not resolve
State the detectable floor in the first paragraph of the readout, beside the estimate and the interval. A result reported as no significant effect will be quoted for a year as the channel does nothing; the same result reported as the test could not detect effects below 14%, and the plausible effect is 3% stops that in the room. Keep the pre-period parallel-trends plot in the document, because the next test will be argued against this one.
Then feed the bound forward rather than discarding it. An experiment's interval is the most defensible constraint available to a mix model, and the place it earns its keep is exactly where the model produces something implausible, as traced in the model that hands a channel a negative coefficient. Keeping test design, power assumption, arms and result in one queryable place, rather than in slide decks, is ordinary internal tooling and it is what stops the same underpowered test being commissioned twice. Both sit inside the attribution, incrementality and mix work we do for marketing and advertising teams.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
Does a confidence interval containing zero mean the channel has no effect?
No. It means the data cannot distinguish the effect from zero at the chosen confidence level, which depends as much on the test's precision as on the channel. Compute the detectable floor by multiplying the interval's half-width by about 1.4: if that floor sits above any effect the channel could plausibly produce, the test never had a chance of a positive result and says nothing about performance.
How long should a geo lift test run?
Long enough to cover the conversion lag after the flight ends, and only as long after that as adds independent information. Duration fixes truncation, not sample size: weeks from the same markets share seasonality, so the effective sample grows more slowly than the calendar. If the interval is wide because there are too few markets, extending the test will not fix it and adding matched markets will.
How much bigger does an underpowered test need to be?
Roughly 4 times the units to halve the detectable effect, because precision improves with the square root of sample size. That scaling is why a rerun at slightly larger size is usually wasted: going from 10 matched pairs to 15 moves the floor by about a fifth. The other lever is the size of the spend change being tested, since a larger intervention produces a larger true effect against the same noise.
Can an inconclusive test still be used to make a budget decision?
Yes, if it is reported as a bound rather than as a verdict. A test that could detect a 6% effect and found nothing supports the statement that the effect is probably below 6%, which is enough to cap what the channel is worth in a planning model. What it does not support is cutting the channel on the grounds that the estimate was near zero, since an underpowered test produces near-zero estimates whatever the truth.
- incrementality
- experiments
- statistics
- measurement
The work behind this page
Builds from our portfolio that this page draws on.
Read next
- The model hands a channel a negative coefficient nobody believesAn implausible coefficient is usually a statement about your spend history, not about the channel. Five causes, the checks that separate them, and where each branch ends.diagnostic
- Adstock: the assumption that last week's spend is still selling this weekAdstock carries a share of this week's media into next week's model input. The half-life is chosen, rarely identified from the data, and it moves the answer.definition
- The lookback window: the assumption inside every conversion figureA lookback window has two dimensions set independently, so two platforms reporting different totals for the same sales are not disagreeing. They are counting different event classes.definition
- A slice of events lands before the visitor has answered the bannerEvents arriving with an absent consent field are not a compliance abstraction. They are a race between two scripts, and the race has a rate you can measure this afternoon.diagnostic
- Click identifiers: the URL parameters that let a server-sent conversion find the ad that caused itA campaign tag describes where traffic came from. A click identifier is the key that joins a sale back to a specific click — and only one of the two is load-bearing.definition
- Consent state: a typed field on each event, not a switch on the pageThe pageview before the banner answer and the purchase after it are both correct, and they carry different consent values. That only works if consent travels per event.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation