Skip to content
Experiment paths connected to a confidence gauge and balanced decision scale, representing evidence under uncertainty
Rich Tank27 Jul 265 min read

Why 90% confidence can be enough for marketing experiments

For reversible, low-risk marketing decisions, a pre-set 90% confidence threshold can help teams learn faster. It must still be paired with effect size, guardrails and a clear decision rule.

Quick answer

Use a 90% confidence level only for reversible, low-cost marketing decisions with a pre-set decision rule. It is not a shortcut for weak data, and 99% is not medicine's universal standard.

Marketing teams often inherit a simple rule: do not act until a test reaches 95% confidence. The more useful question is how much evidence is enough for this specific decision.

A 90% threshold can be defensible for small, reversible marketing bets. The threshold still needs to be chosen before the data is inspected, with a practical-effect and guardrail plan.

Start with the decision, not the percentage

A significance threshold is a tolerance for false-positive risk, not a universal badge of truth.

In a conventional frequentist test, alpha is set before the experiment. A two-sided alpha of 0.10 corresponds to a 90% confidence interval; alpha of 0.05 corresponds to a 95% interval. A p-value is compared with that pre-set threshold. It is not the probability that the hypothesis is true, and it does not tell you whether the lift is commercially worthwhile.

For a low-cost landing-page message test, acting at 90% may be reasonable if the alternative can be rolled back quickly and the expected upside is material. For pricing, brand claims, a major media reallocation or a change that can harm users, choose a stricter rule.

  • Set alpha, the primary metric, minimum worthwhile effect and stopping rule before launch.
  • Record what a positive, negative or inconclusive result will change.
  • Keep a reversal path where the downside of a wrong decision is meaningful.

Why 90% can work in marketing

A 90% threshold can reduce time to learn when the cost of a false positive is lower than the cost of waiting.

Marketing is often a sequence of small bets: an ad angle, landing-page hierarchy, CTA description or on-site prompt. Waiting for 95% or 99% on every small decision can mean fewer learning cycles and less useful evidence overall.

That argument only works when the test protects against self-deception. A 90% result with a tiny effect, a noisy metric, repeated peeking or a biased audience is not a reason to declare success. Treat it as a decision signal with stated uncertainty, then monitor the outcome after rollout.

Use expected value and reversibility

Ask what a false positive costs and how easily it can be undone. If a change is easy to reverse and has bounded downside, accepting more uncertainty in exchange for speed can be rational. If it affects trust, compliance, accessibility or a large budget, raise the evidence bar.

Pair confidence with practical significance

Set a minimum effect that would justify the work. A statistically detectable improvement that will not change revenue, qualified demand or visitor experience is not automatically a win. Review the estimated effect and its interval, not only a pass/fail label.

Psychology and medicine do not map neatly to 95% and 99%

The common 95% and 99% shorthand is too simple: thresholds vary by design, consequences, multiplicity and the decision being made.

Psychology has commonly used p < .05 and sometimes p < .01; the American Psychological Association defines alpha as a choice that depends on the consequence of a Type I error. That is a risk decision, not a rule that makes psychology intrinsically a 95% discipline.

UK medical guidance does not make 99% the universal rule. MHRA guidance for clinical investigations describes 5% as the conventional level for two-sided tests, matching two-sided 95% confidence intervals; a typical one-sided non-inferiority test uses 2.5%. NICE likewise describes p < 0.05 as a convention and separates statistical significance from clinical significance. Stricter control can be appropriate for multiple endpoints, safety-critical claims or a particular pre-specified design, but 99% is not the default shorthand for UK medical evidence.

The case for 90% is not that evidence can be relaxed casually. It is that the threshold should match the harm of being wrong, while the analysis still reports uncertainty and protects people from avoidable downside.

A practical operating rule for marketing teams

Choose 90%, 95% or a stricter bar before launch, then make the rollout proportionate to the risk.

Classify the decision as a reversible optimisation, material commercial change or high-risk customer/trust impact. Pre-set alpha, minimum worthwhile effect, primary metric, guardrails, sample rule and the date or size at which the test will be read.

For a reversible optimisation, consider 90% confidence with a measured rollout and post-launch monitoring. For material or hard-to-reverse decisions, use a stricter threshold and seek replication, holdouts or additional evidence. Report the estimated effect, interval, decision, limitation and next checkpoint.

  1. Classify the downside of a false positive.
  2. Set the threshold and guardrails before looking at results.
  3. Check practical effect, not only statistical significance.
  4. Roll out in proportion to the risk and monitor the result.

Further reading

Sources

  1. MHRA: Statistical considerations for clinical investigations — Medicines and Healthcare products Regulatory Agency
  2. NICE glossary: p value — National Institute for Health and Care Excellence
  3. GOV.UK: Analyse your data for digital health products — Department of Health and Social Care
  4. APA Dictionary: significance level — American Psychological Association
  5. ICH E9: Statistical Principles for Clinical Trials — International Council for Harmonisation

Frequently asked questions

Does 90% confidence mean there is a 90% chance the result is true?

No. In the usual frequentist interpretation, it describes the long-run performance of the interval procedure, not the probability that this specific hypothesis is true.

When should marketers use 95% or 99% instead?

Use a stricter pre-set bar when a false positive could create substantial spend, customer harm, accessibility or trust risk, or when the change is hard to reverse.

avatar
Rich Tank
Rich Tank is a founder, consultant and product-minded marketer with experience across growth, CRM, digital strategy and user experience. He has spent his career helping businesses improve how they attract, convert and support customers online. He writes about digital experience, website journeys, marketing technology and the broader challenge of creating websites that are both effective and easy to use. Based in London, Rich is particularly interested in the intersection of user behaviour, conversion and product thinking.

RELATED ARTICLES