In short: Choose the business decision, outcome, and smallest change worth detecting before you set a test budget. Use a baseline conversion rate and a sample-size method to estimate how many eligible people each version needs; then translate that volume into time and spend. If the required sample is outside the team’s reach, change the question or use another form of evidence instead of calling an underpowered result conclusive.
A small test budget may be affordable but too small to answer the question. A larger one may generate plenty of clicks while still missing the outcome that matters. The budget becomes useful only after the team defines what evidence it needs and how much relevant traffic it can realistically reach.
This process is for controlled conversion tests, such as comparing two landing-page or form versions with a valid split between them. A channel launch or broad campaign test has different sources of uncertainty and should not be treated as if it were a simple A/B test.
1. Write the decision and primary outcome
Start with the decision the test will support: keep the current page, adopt a proposed change, or gather more evidence. Define one primary outcome that connects to that decision, such as a completed qualified form—not a list of clicks, scrolls, and secondary events that can produce a preferred answer after the test.
Also define the unit of observation. It may be an eligible user, account, or session, depending on the test and how the experiment assigns versions. Count each unit consistently, prevent repeat visits from being counted as independent people when that would distort the result, and check that both versions receive comparable eligible traffic.
For a B2B site, a form completion may arrive sooner than a qualified opportunity or signed engagement. If the latter outcomes are too sparse for a short test, use a nearer-term primary outcome only when it is a defensible signal. Keep lead quality and later pipeline outcomes as guardrails or follow-up measures; do not describe a conversion lift as a revenue lift without evidence.
2. Choose a baseline and the smallest useful effect
Estimate the current conversion rate from comparable traffic: the same offer, audience, page intent, and tracking definition. Use a time period that reflects the normal traffic pattern and note any recent campaign, site, or qualification changes. If the baseline is noisy or unreliable, fix the measurement first or state the uncertainty in the calculation.
Then define the smallest change that would matter to the decision. For example, ask what improvement would justify the implementation or what decline would be unacceptable. Express the change consistently as an absolute percentage-point difference or as a relative percentage change; the two are not interchangeable.
A smaller difference is harder to distinguish from ordinary variation and generally requires more observations. NIST’s sample-size guidance for proportion tests makes clear that the required sample depends on the assumed baseline and detectable difference as well as the selected error and power criteria. There is no universal traffic requirement that works for every conversion rate and test question.
3. Estimate the required sample before choosing spend
Use a sample-size or power calculator that supports your outcome and experimental design, or ask an analyst to calculate it. Record the baseline rate, smallest useful effect, confidence or significance settings, desired power, and traffic split. Check whether the output means the total across both versions or the number needed in each version; do not mistake one for the other.
The result is a planning estimate, not a guarantee that a test will identify the right business decision. It depends on assumptions, clean event tracking, consistent assignment, and an outcome that is measured the same way across versions. A calculator cannot repair a changing offer, a broken form event, or a test in which the audience composition differs between versions.
Keep the test simple enough for the available volume. Testing several variants or many page elements at once divides attention and can make the result difficult to interpret. If you need to learn which of two changes matters, isolate the most important difference first.
4. Translate eligible traffic into duration and cost
Once you have an estimated sample, use the traffic that can actually enter the test—not all site sessions. Account for the fraction of visitors who match the target audience, can be assigned to a version, and can reach the outcome event. In paid media, use the expected cost of an eligible visit or click, not the cheapest reported click if it rarely reaches the test page.
A simple planning sequence is:
- Eligible sample required: total valid observations across the test, with the split accounted for.
- Expected eligible volume per week: based on the traffic source and the share that meets the test rules.
- Estimated test duration: required eligible sample divided by expected eligible weekly volume.
- Estimated media spend: required paid eligible visits multiplied by the expected cost per eligible visit, plus any other direct test costs.
Keep the assumptions beside the result. If the source mix, weekly volume, cost per visit, test split, or conversion delay changes, update the duration and spend estimate. Keep the test long enough to include the normal weekly pattern and the time needed for the chosen event to register; do not set a convenient end date that prevents the required sample from arriving.
5. Check whether the learning is affordable
Compare the estimated time and spend with the cash the business can risk, the campaign’s practical life, and the value of the decision. If the required sample takes longer or costs more than the business can support, a smaller budget does not make the test more conclusive.
Instead, consider whether you can narrow the decision, reduce the number of variants, use a higher-volume but still relevant event with a lead-quality guardrail, or collect evidence for longer. If a randomized test remains infeasible, use interviews, usability reviews, a small pilot, or other directional evidence and label it accordingly. These approaches may inform a choice without proving that one version caused a specific lift.
6. Set the stop rule and interpret an inconclusive result
Before launch, record the planned sample or test window, the primary metric, what counts as a technical failure, and who can stop the test for a safety or business reason. Avoid ending the experiment as soon as a favorable number appears unless the method was designed for that kind of sequential monitoring. Repeatedly checking results and changing the end point can make chance movement look more convincing than it is.
An inconclusive result does not prove that the versions are equal. It can mean the true difference is small, the test did not reach its planned sample, or the measurement and audience were too variable. Google’s Ads experiment guidance lists insufficient traffic or an experiment that has not run long enough among reasons a clear winner may not appear. That is a reminder to diagnose the setup and volume—not to keep increasing spend without a new decision and approved risk limit.
Conversion-test sizing worksheet
- Decision the test will support: ______
- Primary outcome and unit of observation: ______
- Baseline rate, source, and date range: ______
- Smallest useful absolute or relative change: ______
- Sample-size method, settings, and sample per version: ______
- Eligible traffic available per week: ______
- Estimated duration and direct spend: ______
- Technical checks and lead-quality guardrails: ______
- Stop rule, decision owner, and review date: ______
- If the sample is not feasible, alternate evidence and its limits: ______
A well-sized test makes its limits visible before the money is committed. If the planned traffic cannot answer the question, change the question or the method; do not turn a low-volume result into proof.
If your acquisition tests lack a clear outcome, sample plan, or decision rule, request a marketing diagnostic to define what the next experiment can reliably tell you.
Sources and scope
- NIST/SEMATECH e-Handbook: Sample Sizes Required for Proportion Tests — describes how assumptions about the baseline proportion, detectable difference, and test error criteria affect required sample size; accessed October 8, 2026.
- Google Ads Help: Monitor Your Experiments — documents that a Google Ads experiment may not show a clear winner when traffic or run time is insufficient; accessed October 8, 2026.
This is an operating framework, not statistical advice or a promise that a sample-size estimate will produce a commercially decisive result. Match the calculation to the experiment design, verify measurement before launch, and ask a qualified analyst to review complex or high-impact tests.
How did this article land?
Choose one reaction. You can change it anytime.