How Long Should a Landing Page A/B Test Run?

How long should a landing page A/B test run?

How Long Should a Landing Page A/B Test Run? Long enough to collect the evidence your experiment needs, account for relevant patterns in visitor behaviour, and allow the conversions you are measuring to mature. That might take days for a high-traffic page testing a substantial change, or several weeks or longer for a low-traffic page trying to detect a small improvement. There is no single duration that makes every test reliable.

The practical approach is to calculate the required sample size before launch, estimate how quickly your page can collect it, and define when you will evaluate the results. You should also consider whether the experiment covers a representative period of visitor behaviour and whether the metric reflects the business outcome you care about.

This guide explains how to estimate testing duration, recognise unreliable results, handle inconclusive experiments and decide when the evidence is sufficient to act.

How long should a landing page A/B test run?

A landing page A/B test should run until it meets its pre-established statistical and operational requirements. Those requirements depend on the experiment, so a fixed duration such as seven days or two weeks cannot guarantee a reliable result.

A test needs enough eligible visitors to distinguish the effect you care about from ordinary random variation. It also needs a suitable design, accurate tracking, and a consistent experience for visitors in each group.

A practical starting point is to plan for three things:

  • Enough observations: Calculate the sample size needed to detect an improvement that would matter to the business.
  • Representative conditions: Account for weekly traffic patterns, promotions, seasonal changes and other conditions that influence visitor behaviour.
  • Mature conversions: Allow enough time for visitors to complete the action being measured, particularly when purchases or qualified leads take time to materialise.

The first requirement is statistical. The other two help establish whether the statistical result answers the business question you intended to investigate.

A complete weekly cycle can be a sensible operational consideration for many landing pages. It is not, however, a substitute for a sample-size calculation. Equally, collecting the required sample does not make a poorly instrumented or badly designed test trustworthy.

What determines how long a landing page A/B test should run?

Several factors influence how quickly an experiment can produce useful evidence. Understanding them helps you estimate duration before committing traffic and budget.

1. Eligible traffic volume

More eligible visitors generally allow an experiment to collect its required sample sooner. The important word is eligible.

If your website receives 30,000 visits a month but only 12,000 belong to the audience included in the experiment, you should calculate duration using the relevant 12,000 visits, not the site's total traffic.

For a two-variant experiment with an even split, half the eligible visitors will usually see the control and half the variation. If you need 8,000 visitors per variant, you need approximately 16,000 eligible visitors in total.

Check how your testing platform counts traffic. Unique users, sessions, and page views are not interchangeable. A returning visitor may generate several sessions, while a user-based experiment may keep that visitor assigned to the same variant across visits.

Your sample-size calculation and duration estimate should use a denominator consistent with the experiment's randomisation unit and conversion-rate definition.

2. Baseline conversion rate

The baseline conversion rate is the proportion of eligible visitors who complete the target action under the existing experience.

If 500 out of 10,000 eligible visitors submit a form, the baseline conversion rate is

The baseline helps establish how much data you may need to detect a change.

Pages with low conversion rates often produce relatively few conversion events, which can make comparisons difficult. However, a high baseline conversion rate does not automatically mean a test will finish sooner. The effect you want to detect and the statistical method also matter.

For example, detecting a change from 5% to 6% is a different task from detecting a change from 5% to 9%. The second difference is larger and generally easier to detect, assuming the other design choices remain the same.

Use historical data from a comparable audience, offer, page, and traffic source. A baseline from a different campaign or a period with unusual demand may not represent the experiment you are about to run.

3. Minimum detectable effect

The minimum detectable effect (MDE) is the smallest difference your experiment is designed to detect with the chosen statistical assumptions.

You should choose an MDE that reflects the decision you need to make, rather than selecting a number simply because it produces a convenient testing duration.

There are two ways to express an improvement:

  • Absolute change: The difference between two conversion rates, expressed in percentage points.
  • Relative change: The difference expressed as a proportion of the original rate.

Suppose a page converts 5% of visitors and a variation is expected to convert 6%.

The absolute increase is one percentage point. The relative increase is: 20%

These figures describe the same hypothetical change in different ways. They are not interchangeable.

A business might decide that an improvement of one percentage point would justify implementing a change. But if available traffic is insufficient to detect an effect that small within a reasonable period, the team may need to test a more substantial hypothesis, gather more traffic, or use a different decision method.

Do not simply increase the MDE in a calculator to shorten the test if you would still make the same decision about a smaller improvement. The MDE should reflect the effect you genuinely need to be able to detect.

4. Statistical significance and power

Sample-size planning also depends on two statistical choices.

The significance threshold controls the tolerated false-positive risk under the specified test and its assumptions. A common threshold is 5%, often described as a 0.05 significance level. It is a design choice, not proof that a result is correct.

Statistical power is the probability that a test will detect an effect of a specified size when that effect is genuinely present, under the assumptions used to plan the experiment. A commonly used planning value is 80%, or 0.80.

Higher power generally requires more observations. So does choosing a stricter significance threshold. The smaller the effect you want to detect, the more observations the experiment will generally need.

These relationships are described in Optimizely's explanation of Type II errors and statistical power and the statsmodels documentation for two-sample power calculations.

Do not interpret a statistically significant result as a 95% probability that your variation is better. A conventional significance test answers a narrower question: whether the observed evidence would be sufficiently unusual under a specified null hypothesis and statistical model.

5. Traffic allocation

Traffic allocation determines how eligible visitors are distributed between the control and the variation. A balanced 50/50 split is common for a two-variant test because it generally makes efficient use of the available sample when the cost and value of observations are comparable.

Suppose your calculation requires 8,000 visitors per variant. If 1,000 eligible visitors enter the experiment each day and allocation is evenly split, approximately 500 visitors will reach each variant daily.

You would therefore need about 16 days to collect 8,000 visitors per variant, before accounting for conversion delays, operational checks, or a decision to cover complete weekly cycles.

An uneven allocation changes the calculation. If the control receives 80% of traffic and the variation receives 20%, the smaller group will take longer to reach its required sample.

Unequal allocation can be appropriate for some experiments, particularly when limiting exposure to a risky change. However, it should be deliberate and reflected in the sample-size calculation.

Also investigate unexpected allocation imbalances. If a 50/50 experiment sends substantially more eligible users to one variant, check the assignment logic, exclusions, and instrumentation before interpreting the conversion rates.

6. Conversion delays and sales cycles

The time required to collect visitors is not always the same as the time required to measure their outcomes.

For a newsletter signup, conversion may occur during the same session. For an ecommerce purchase, a visitor may return later. For a B2B product, a form submission might precede qualification, a sales conversation, and an eventual contract by weeks or months.

Consider a test that finishes collecting its required visitors on Friday. If many visitors exposed to the variation have not yet had a reasonable opportunity to purchase, comparing their incomplete outcomes with more mature outcomes from earlier visitors can mislead you.

Define a conversion window that reflects normal customer behaviour. Then allow the relevant observations to mature before drawing the final conclusion.

For B2B landing pages, distinguish between:

  • Form submissions.
  • Qualified leads.
  • Demo attendance.
  • Sales opportunities.
  • Closed revenue.

A test may produce more form submissions without producing more qualified opportunities. If sales quality matters, track downstream outcomes and consider whether the sales cycle is long enough to affect when you can make a final decision.

7. Weekly patterns and seasonality

Visitor behaviour can vary by day, time, campaign and season.

A B2B page may receive more relevant traffic during working hours. An ecommerce page may behave differently at weekends or during a major sale. A payday promotion can change both the number of visitors and their likelihood of buying.

Because a conventional A/B test shows variants to visitors during the same period, randomisation helps protect the comparison from many time-related differences. However, ending a test after an unusual traffic spike can still limit how well its results represent the audience and conditions you care about.

Covering a complete weekly cycle can help when weekday and weekend behaviour differ. Longer periods may be needed if monthly purchasing patterns, holidays or seasonal demand matter to the decision.

Do not extend every test automatically to two weeks. Decide which cycles are relevant to your business, then combine that operational judgement with the statistical plan.

8. Test design and implementation quality

A longer test does not repair a faulty experiment.

Before launch, verify that:

  • Visitors are assigned to variants as intended.
  • Both versions work on the devices and browsers in scope.
  • The primary conversion event fires correctly.
  • Duplicate conversions are not being counted.
  • The audience and exclusion rules are consistent.
  • The variants remain stable during the test.
  • Other experiments or page changes are not interfering with the comparison.

If the tracking event fails for one variant, collecting more data may make the wrong result look more convincing. If the control and variation receive different audiences because of an implementation error, the comparison may no longer isolate the effect of the page change.

For pages that can be indexed by search engines, follow Google Search Central's guidance on website testing. It advises against showing Googlebot different content from that shown to users, recommends canonical links for alternative test URLs, and recommends removing test elements when an experiment ends.

How to calculate how long a landing page A/B test should run

A useful duration estimate begins with the number of observations required. You then work out how quickly the experiment can collect those observations.

Step 1: Define the experiment and its primary metric

Write down the hypothesis, the control, the variation, the eligible audience, and the event that counts as a conversion.

For example:

Hypothesis: Reducing the number of fields in a demo-request form will increase completed submissions without materially reducing lead quality.

The primary metric might be completed demo-request submissions per eligible visitor. A guardrail metric could be the proportion of submitted leads that meet the team's qualification criteria.

Choose the primary metric before launch. If you decide which metric matters only after seeing the results, you increase the risk of selecting the measure that makes your preferred variant look best.

Step 2: Establish the baseline

Suppose the current landing page converts 5% of eligible visitors.

Use relevant historical data to estimate this rate. Check that the period reflects the audience, offer, traffic mix, and conversion definition used in the planned experiment.

If the baseline is unstable or tracking has recently changed, resolve those issues before relying on it for sample-size planning.

Step 3: Choose the minimum detectable effect

Assume, for illustration, that an increase from 5% to 6% would justify implementing the change.

That is a one-percentage-point absolute increase and a 20% relative increase.

These are hypothetical assumptions, not a benchmark for landing page performance. Your business may require a smaller or larger improvement to justify the cost and risk of the change.

Step 4: Set the statistical assumptions

For the example below, assume:

  • Baseline conversion rate: 5%.
  • Expected variant conversion rate: 6%.
  • Significance level: 5%, using a two-sided test.
  • Statistical power: 80%.
  • Traffic allocation: 50/50.
  • Outcome: One binary conversion per eligible experimental unit.
  • Statistical method: A conventional two-sample comparison of proportions, with the required sample estimated using a normal approximation.

The sample-size estimate depends on these assumptions. A different statistical method, allocation, baseline, or effect size can produce a different answer.

For an experiment with multiple variants, repeated measures, clustering, or a different outcome type, use a calculator or statistical method designed for that particular setup.

Step 5: Estimate the required sample size

Under the assumptions above, a conventional normal-approximation calculation gives an estimated requirement of approximately 8,143 eligible observations per variant, or about 16,286 in total. This is a hypothetical planning estimate, not a universal minimum.

The calculation uses the two conversion rates, a 5% significance level, and 80% power. The method is consistent with the kind of two-sample power calculation documented by statsmodels.

The estimate illustrates why small improvements can require substantial traffic. If you want to detect a smaller change, the sample requirement will generally increase. If the true effect is larger, a test may have a better chance of detecting it with the same sample.

The number is not a promise that the experiment will finish with a definitive result. It is a planning estimate based on assumptions that must be appropriate for the real test.

Step 6: Convert sample size into duration

Now suppose the experiment receives 1,000 eligible visitors per day and uses a 50/50 allocation.

Each variant receives approximately 500 visitors per day.

The estimated collection time is 16 days

You would therefore expect to collect the planned sample in about 17 days, assuming traffic remains steady and all eligible visitors are included as expected.

This is the sample-collection estimate. It does not automatically include a conversion-maturation window or extra time needed to investigate tracking problems.

If covering complete weekly cycles is important, you may decide to round the planned endpoint up to a suitable full-cycle boundary. That is an operational choice, not a statistical correction to the sample-size calculation.

If traffic is lower than expected, the test will take longer. If the page receives only 400 eligible visitors daily, with approximately 200 assigned to each variant, collecting 8,143 per variant would take roughly 41 days.

The estimate assumes a stable 50/50 split. Recalculate if actual allocation or eligible traffic differs materially.

Step 7: Plan for conversion maturity and quality checks

Before launch, decide how long visitors normally take to convert and how you will check that the data are valid.

For example, an experiment measuring completed purchases may need a defined period for visitors to return and buy. A B2B experiment may need a separate follow-up period to evaluate lead quality.

Do not add an arbitrary buffer percentage to make the estimate feel safer. Instead, identify the actual uncertainties: traffic variability, conversion delay, data exclusions, relevant business cycles, and operational risks.

The result should be a planned analysis point or stopping rule that you can explain and follow.

How long should you run an A/B test in different situations?

The same number of days can represent very different amounts of evidence. Use the following table to understand how circumstances change the plan.

Situation What affects duration Recommended approach
High-traffic page testing a substantial change A large effect may be easier to detect, so the required sample can accumulate quickly. Calculate the sample size and avoid stopping just because the early result looks strong.
High-traffic page testing a small improvement Small effects generally require more observations than large ones. Check whether the expected business value justifies the sample and time required.
Low-traffic landing page The experiment may take months to collect enough observations. Prioritise stronger hypotheses, qualitative research or a different decision method if the wait is not worthwhile.
Low baseline conversion rate Conversion events may be rare, making comparisons less precise. Check the sample-size estimate and verify that the chosen metric is appropriate.
Page with delayed conversions Visitors may need time to complete the action after exposure. Separate sample collection from the conversion-maturation window.
B2B page with a long sales cycle Form submissions can be observed sooner than qualified opportunities or revenue. Use an appropriate early metric for monitoring, but assess downstream quality before making a commercial decision.
Seasonal or promotion-driven campaign The audience and purchase intent may differ from normal conditions. Decide whether the test is intended to answer a seasonal question or a general one. Interpret results accordingly.
Uneven traffic allocation The smaller group may take longer to reach its required sample. Check assignment and exclusions, then recalculate using the actual allocation plan.

These situations do not have fixed durations because the answer depends on the experiment's assumptions and traffic. A duration estimate should be calculated for the specific test rather than copied from a general benchmark.

Statistical significance versus practical significance

Statistical significance and practical significance answer different questions.

Statistical significance concerns whether the observed data provide sufficient evidence against a specified null hypothesis under the chosen statistical method and assumptions.

Practical significance concerns whether the size of the difference is important enough to justify a business decision.

Suppose a high-traffic page produces a small but statistically significant increase in form submissions. If those additional submissions are mostly unqualified leads, the change may have little commercial value.

Now suppose a variation shows a promising increase in qualified leads, but the sample is too small to distinguish the result confidently from random variation. The potential business value may be high, but the evidence remains uncertain.

Neither situation should be decided by the significance label alone.

Consider the likely business impact, implementation costs, potential harm, lead or customer quality, and the cost of waiting for more evidence. A small improvement may justify implementation when it is easy to reverse and has little downside. A change to pricing, checkout, or a high-value acquisition funnel may warrant more caution.

The decision should reflect both what the data support and what the business stands to gain or lose.

When should you stop a landing page A/B test?

You should stop at the analysis point or under the stopping rule defined for the experiment, provided the data are valid and the relevant conversions have matured.

Before concluding, check that the experiment has met its planned requirements:

  1. The required sample size or the chosen method's analysis criteria have been reached.
  2. The primary conversion metric was defined in advance and tracked correctly.
  3. Visitors were allocated to the variants as intended.
  4. The variants remained consistent during the test.
  5. The conversion window has elapsed for the observations being evaluated.
  6. Relevant traffic patterns, promotions, and external changes have been considered.
  7. Guardrail metrics and downstream outcomes have been reviewed.
  8. The result is interpreted in light of uncertainty and commercial value.

Reaching a target date does not guarantee that the test has enough information. Reaching a target sample does not guarantee that the experiment was implemented correctly.

If tracking is broken or allocation is compromised, investigate the problem before treating the result as evidence. Depending on the severity, you may need to correct the implementation and restart the experiment.

If the experiment reaches its planned endpoint but remains inconclusive, do not keep extending it automatically. Decide whether additional observations are likely to change the business decision and whether the value of that information justifies the wait.

Why you should not stop as soon as the results look significant

Repeatedly checking a conventional fixed-horizon test and stopping as soon as it crosses a significance threshold can inflate the risk of false-positive conclusions.

This practice is often called peeking. Early results can fluctuate considerably because a small number of conversions can move the apparent conversion rate sharply. If you repeatedly check until the result looks favourable, you give random variation many opportunities to produce a misleading signal.

Research on online experiments has examined this problem and methods for handling it. See Peeking at A/B Tests: Why It Matters, and What to Do About It and the broader review of statistical challenges in online controlled experiments.

There are three approaches to distinguish:

  • Fixed-horizon testing: Set the required sample and planned analysis point before launch, then evaluate the result at that point.
  • Sequential testing: Use a method designed for planned interim analysis, with statistical error controlled under its assumptions.
  • Informal peeking: Repeatedly check ordinary fixed-horizon results and stop whenever the preferred outcome appears.

The third approach is the problem. Sequential methods can support ongoing monitoring, but they must be used as designed. A dashboard that refreshes continuously does not, by itself, make continuous significance checking statistically valid.

Also avoid repeatedly testing multiple variants, metrics, or audience segments and then reporting only the most favourable result. Multiple comparisons create additional opportunities to find an apparently positive result by chance.

Can an A/B test run for too long?

Yes, if the experiment continues without a clear purpose, if the underlying conditions change substantially, or if the cost of waiting outweighs the value of more information.

Longer tests are not inherently invalid. A low-traffic page may genuinely need more time. But a test should not continue indefinitely simply because it has not produced a winner.

As time passes, promotions, campaigns, product changes, and shifts in traffic can complicate interpretation. Meanwhile, the team may be delaying a decision or missing the opportunity to test a more important hypothesis.

Follow the stopping rule, review the value of further data, and document why you decided to continue or conclude.

What if your landing page does not receive enough traffic?

A conventional A/B test may not be the best use of time when a page attracts too few eligible visitors to detect a meaningful effect within a reasonable period.

Start by estimating the traffic requirement. If the experiment would need several months to collect enough data, ask whether the expected improvement is valuable enough to justify that delay.

You have several alternatives.

  • Prioritise larger, evidence-based changes. A substantial change to an unclear offer or a high-friction form may be more useful to investigate than a subtle wording adjustment. A larger effect is generally easier to detect, although that does not guarantee a result.
  • Use qualitative research to identify problems. Session recordings, heatmaps, usability tests, visitor interviews and form analytics can reveal where people struggle. These methods help generate and prioritise hypotheses, but they do not prove that one variant will convert better.
  • Fix clear technical problems directly. If the form is broken on mobile or the conversion event is not firing, fix it. You do not need an A/B test to establish that a known defect should be corrected.
  • Improve traffic quality and measurement. More visitors will not help if they are the wrong audience or if conversions are being recorded incorrectly. Check campaign targeting, message match, form functionality, and tracking before launching another experiment.
  • Consider a different decision method. A Bayesian approach or a properly designed sequential method may fit some situations, but neither automatically solves a lack of information. The assumptions and evidence still matter.

You can also compare performance before and after a change when a controlled test is impractical. Treat that comparison cautiously: changes in seasonality, traffic sources, promotions, and other conditions may explain some or all of the difference.

The responsible outcome may be to make a decision using the best available evidence while documenting the uncertainty, rather than claiming that an underpowered experiment has identified a winner.

How to handle inconclusive A/B test results

An inconclusive result means the experiment has not established a sufficiently clear difference under the chosen method. It does not prove that the variants are identical.

Start by examining the uncertainty around the estimated difference.

If the range of plausible effects includes both a meaningful improvement and a meaningful decline, the result may not support a confident decision. If the data rule out an improvement large enough to justify the change, you may have a useful answer even without a statistically significant winner.

For example, imagine that a new landing page headline appears to increase conversions, but the estimate is imprecise. The result could be consistent with a substantial improvement, no real change, or a small decline.

There are two separate questions to investigate:

  • Was the experiment adequately designed and executed? Check sample size, allocation, tracking, audience consistency, and the planned analysis method.
  • What effects remain plausible? Review the estimated difference and its uncertainty, then consider whether the remaining possibilities would change the business decision.

If the test was underpowered, you can decide whether additional traffic is worth collecting. If the test was compromised, more data may not help; you may need to correct the implementation and start again.

If the result is inconclusive but the plausible effects are too small to matter commercially, retaining the control and moving to a stronger hypothesis may be more useful than continuing indefinitely.

Document what the experiment established, what it could not establish, and what you plan to do next. An inconclusive result can still narrow the range of plausible effects and improve future decisions.

How seasonality, weekly cycles and external changes affect duration

Testing duration should reflect the environment in which the experiment runs. The aim is to gather enough evidence under conditions relevant to the decision, without allowing unrelated changes to undermine the comparison.

Consider these situations:

Change or pattern Why it matters How to handle it
Weekday versus weekend behaviour Visitor intent and conversion patterns may differ by day. Consider covering a full weekly cycle when the difference is relevant.
Holidays and seasonal demand Visitors may behave differently during unusual periods. Decide whether the experiment is intended to answer a seasonal question or a general one.
Promotions or price changes An offer can change conversion behaviour independently of the page variation. Keep the offer consistent across variants or define the experiment around the promotion.
Paid campaign changes Different targeting, budgets, or creative can change who reaches the page. Record campaign changes and investigate material shifts in traffic composition.
Product or checkout changes The conversion path may no longer be the same throughout the test. Avoid unrelated changes where possible; investigate whether a major change invalidates the comparison.
Shifts in traffic source or device Different audiences may respond differently to the same page. Check source and device patterns, while avoiding unplanned subgroup analysis to find a convenient winner.

Random assignment helps make the control and variation comparable during the same period. It does not make every result universally applicable to every season, campaign, or audience.

If you are testing a Black Friday offer, for example, the relevant question may be how the page performs during that promotion. You should not assume that the same result will hold under ordinary demand.

Testing duration for different landing page goals

The conversion you choose affects both the sample-size calculation and the time required to evaluate the result.

A page designed to collect newsletter signups may generate many conversion events quickly. A page designed to generate enterprise sales opportunities may produce far fewer, with substantial delays between the first visit and the eventual outcome.

Landing page goal Potential primary metric Additional consideration
Newsletter signup Completed subscriptions per eligible visitor Confirm that the signup event fires once and that signups are valid.
Lead-generation form Completed submissions per eligible visitor Review lead quality and whether the form attracts the intended audience.
Demo booking Completed bookings per eligible visitor Consider cancellations, attendance, and the quality of the resulting opportunities.
Free-trial registration Completed registrations per eligible visitor Trial activation and later conversion may provide important context.
Ecommerce purchase Completed purchases per eligible visitor Allow for purchase delays, refunds where relevant, and changes in order value.
Download Completed downloads per eligible visitor Confirm that the download is the outcome that matters, rather than a button click.
Consultation request Completed requests per eligible visitor Consider whether requests become attended consultations or paying customers.
Qualified B2B opportunity Qualified opportunities per eligible visitor Allow for qualification and sales-cycle delays before evaluating the final outcome.

You can monitor an earlier event when the ultimate outcome takes too long to mature, but be clear about its limitations.

For instance, trial registrations may be a useful early signal for a SaaS product. They are not necessarily a reliable substitute for paid conversions if the tested page attracts a different mix of users.

Choose the metric that best represents the decision you want to make, then use secondary measures to understand whether the result creates downstream problems.

How to choose an A/B testing tool or calculator

Different tools solve different parts of the testing problem. Choosing the right one starts with understanding which job you need it to perform.

Sample-size calculators

A sample-size calculator estimates how many observations you need under specified assumptions. Depending on the tool, it may ask for the baseline conversion rate, expected variant rate, significance threshold, power, and allocation.

Check which statistical method it uses and whether its inputs match your experiment. Calculators may produce different estimates because they use different assumptions, tests, or approximations.

The statsmodels power-calculation documentation describes supported power and sample-size methods.

Duration estimators

A duration estimator converts a sample-size requirement into an estimated collection period using traffic volume and allocation.

It is only as good as its inputs. If it assumes 1,000 eligible visitors daily but your experiment receives 600, the estimate will be too short. It may also omit conversion-maturation time or the operational need to cover a particular business cycle.

A/B testing platforms

An experimentation platform manages the experiment itself. Depending on the product, it may assign visitors to variants, record conversion events, calculate results, and support stopping rules.

Verify the specific capabilities of the platform you intend to use. Do not assume that every tool offers valid sequential testing, the same statistical method, or automatic sample-size planning.

Analytics platforms

Analytics tools help you understand traffic sources, conversion funnels, device differences, and visitor behaviour. They can help you diagnose a problem or verify tracking, but standard analytics reports do not necessarily provide the statistical design and analysis needed for a controlled experiment.

Statistical analysis tools

Statistical software can support more specialised designs, custom metrics and analysis. It may be appropriate when your experiment involves multiple variants, non-standard outcomes, clustered observations or a more advanced testing method.

Whatever you choose, remember that a calculator estimates requirements; the experiment still needs a sound design and trustworthy data.

Fixed-horizon, sequential and Bayesian testing

These approaches differ in how they use evidence and support decisions. They should not be treated as interchangeable labels for the same process.

Fixed-horizon testing

In a fixed-horizon test, you define the planned sample size and analysis point before launch. You then evaluate the experiment according to that plan.

This approach is straightforward to manage when the expected traffic and conversion behaviour are reasonably predictable. Its main discipline is to avoid changing the stopping rule simply because the interim result looks attractive.

Sequential testing

Sequential methods can allow planned interim analysis while accounting for repeated looks at the data. They are useful when teams need to monitor experiments during their run, but the method must control the relevant statistical error under its assumptions.

Do not use ordinary fixed-horizon significance calculations as though they automatically support continuous monitoring.

Bayesian approaches

Bayesian methods use a different framework for reasoning about uncertainty, combining a model with prior information to produce a posterior distribution.

They can support questions such as how plausible different effect sizes are under the specified model. The interpretation depends on the model, prior assumptions, data, and decision rule.

Bayesian analysis does not make a small or biased sample automatically reliable. Nor does it eliminate the need to define a meaningful outcome and assess whether the available evidence is useful.

Choose a method that fits your decision, and follow the analysis and stopping rules designed for it. If the experiment has high financial or operational stakes, involve someone with appropriate statistical expertise.

Common mistakes when estimating A/B test duration

Many testing problems begin before the first visitor enters the experiment. The following mistakes can lead to unreliable results or waste time.

Mistake Why it matters Better approach
Using the same number of days for every test Different traffic levels and effect sizes produce different sample requirements. Calculate duration for the specific experiment.
Treating a calculator output as a guarantee The estimate relies on assumptions that may not hold in practice. Verify the inputs and check actual traffic and data quality.
Stopping when one variant pulls ahead Early differences can reflect random variation. Follow the pre-established stopping rule or a valid sequential method.
Ignoring statistical power A small sample may fail to detect an effect that matters. Plan the sample size around the effect you need to detect.
Confusing percentage points with relative lift The size of the proposed improvement can be misunderstood. State both measures when they help clarify the business case.
Changing the success metric after launch Selecting a metric after seeing the results can bias the decision. Define the primary metric before the experiment starts.
Ignoring conversion delays Some visitors have not had enough time to complete the target action. Define and respect the conversion window.
Overlooking traffic allocation problems The variants may not be receiving visitors as intended. Check assignment and investigate unexplained imbalances.
Pooling unrelated audiences Different audiences may have different conversion behaviour. Define the target audience and investigate meaningful differences.
Running overlapping tests without a plan Experiments may interact or make results harder to interpret. Coordinate experiments and document interactions.
Ignoring tracking errors The observed conversion rates may not represent real behaviour. Test the full conversion path and validate event collection.
Treating an inconclusive result as proof of no effect The experiment may simply lack enough information. Examine uncertainty and the range of effects still plausible.
Declaring a winner based only on the higher observed rate A higher rate may not represent a reliable or commercially useful difference. Consider the planned statistical analysis and business impact.
Assuming more time always produces better evidence Conditions can change, and the extra wait may not change the decision. Continue only when additional data have a clear purpose.

A checklist before ending your landing page A/B test

Use this checklist as a final review, not as a replacement for the statistical design.

  • The hypothesis and primary metric were defined before launch.
  • The required sample size or planned analysis criteria have been met.
  • The statistical method and stopping rule have been followed.
  • The relevant conversion window has elapsed.
  • Tracking and data quality have been checked.
  • Traffic allocation is behaving as expected.
  • Audience, campaign and device changes have been investigated where relevant.
  • The test has covered the relevant business or weekly cycle, if required.
  • Guardrail metrics and downstream outcomes have been reviewed.
  • The observed effect is commercially meaningful enough to consider implementation.
  • The conclusion reflects the uncertainty in the data.
  • The decision and its limitations have been documented.

If a key condition fails, do not declare a winner merely because the test has reached a date you selected earlier. Investigate the issue and decide whether the experiment remains valid.

A practical decision framework

The right next step depends on the evidence you have collected and the condition of the experiment.

Situation What it means Recommended action What to avoid
Planned sample reached, and the analysis supports a decision The experiment may be ready for evaluation. Verify data quality, conversion maturity, and business impact. Assuming significance alone proves business value.
Target date reached, but sample size is insufficient The calendar deadline has arrived, but evidence remains limited. Reassess whether continuing is worthwhile and document the uncertainty. Declaring a winner because the deadline passed.
Interim result looks favourable, but the planned endpoint has not been reached The apparent lead may reflect random variation. Follow the predefined stopping rule or a valid sequential method. Stopping informally because the result looks promising.
Traffic is too low for a useful test The experiment may take too long to answer the question. Consider qualitative evidence, stronger hypotheses, or another decision method. Claiming certainty from an underpowered test.
Results are inconclusive The experiment has not established a clear difference. Review uncertainty, data quality, and the size of the effect that matters. Claiming the variants are identical.
Tracking or allocation is compromised The observed comparison may be invalid. Investigate the implementation and decide whether the test must restart. Treating a larger sample as a repair for bad data.
The result is statistically clear but commercially trivial The effect may not justify the implementation cost or risk. Compare the estimated benefit with business costs and consequences. Shipping a change solely because it is significant.
Conversions have not matured Some visitors have not had a fair opportunity to complete the action. Wait for the defined conversion window before final evaluation. Comparing mature outcomes with incomplete observations.

How Episode can help with landing page testing

Episode's published material describes a workflow for building campaign landing pages, creating page variations, and using analytics to understand traffic and conversion behaviour. Its guide to improving landing page conversion rates discusses page variants and experiments alongside conversion measurement, traffic-source reporting, forms, and downstream lead quality.

That makes Episode relevant to the page-production and measurement parts of the workflow. Teams can use a landing page builder to create and manage variants, then use appropriate analytics and experimentation functionality to assess them.

The distinction matters: creating a page variant is not the same as proving that it performs better. Before relying on a platform for an experiment, verify that its current functionality supports the required traffic allocation, conversion measurement, statistical analysis, and stopping method.

Episode's published landing page builder comparison identifies A/B testing as part of Episode's offering, but this article does not independently verify every current product setting or statistical capability.

For any tool, including Episode, confirm the current feature documentation and plan limits before assuming it calculates sample sizes, controls false-positive risk during sequential monitoring, or evaluates downstream sales outcomes. A separate analytics or statistical tool may still be needed for aspects the builder does not support.

Frequently asked questions

Is seven days enough for an A/B test?

Sometimes, but seven days is not a universal minimum or guarantee. A high-traffic test with a substantial effect may collect enough data in less than a week. A low-traffic test may need much longer. Consider the required sample size, conversion delay, and relevant traffic patterns.

Is two weeks enough for an A/B test?

Two weeks can be enough if the experiment reaches its planned sample size and meets its other requirements within that period. It may be too short for a page with limited traffic, rare conversions, or a small minimum detectable effect. Use the calendar as an operational consideration, not the sole stopping rule.

How many visitors do you need for an A/B test?

There is no universal visitor count. The required sample depends on the baseline conversion rate, the effect you want to detect, the significance threshold, statistical power, traffic allocation, and the selected method. Calculate the requirement before launch.

How many conversions are needed for statistical significance?

There is no fixed number of conversions that guarantees significance. The answer depends on the conversion rates, sample sizes in each group, the difference between variants, and the statistical method. A platform's minimum page-view setting should not be confused with a universal requirement.

When should you stop an A/B test?

Stop according to the experiment's pre-established analysis or stopping rule, once the required evidence has been collected and the data are valid. Allow the relevant conversions to mature, check tracking and allocation, and assess whether the result matters commercially.

Why should you not stop an A/B test early?

Stopping when an interim result looks favourable can select a random fluctuation rather than a real effect. Repeatedly checking an ordinary fixed-horizon test and stopping when it crosses a significance threshold can inflate the false-positive risk. Use the planned endpoint or a method designed for sequential monitoring.

Can an A/B test run for too long?

Yes. Continuing without a clear purpose can delay decisions and tie up traffic that could be used for other experiments. Long tests can also span changing campaigns, seasons, or website conditions. Continue only when additional observations are likely to improve the decision.

What happens if an A/B test is inconclusive?

An inconclusive result means the experiment has not established a sufficiently clear difference. Review the uncertainty, sample size, tracking, and original MDE. You may continue if more data are valuable, retain the control, investigate qualitatively, or move to a more promising hypothesis.

Does low traffic make A/B testing impossible?

No, but it can make conventional tests slow or impractical. Focus on higher-impact hypotheses, improve measurement and traffic quality, and use qualitative methods to find problems. If you use a different statistical approach, remember that it still depends on appropriate assumptions and adequate information.

What is the difference between statistical significance and practical significance?

Statistical significance concerns evidence against a specified null hypothesis under a chosen method. Practical significance concerns whether the size of the effect is worth acting on. A small effect can be statistically detectable without being commercially useful.

How do you calculate A/B test duration?

First estimate the required sample size from the baseline conversion rate, minimum detectable effect, significance threshold, power, and allocation. Then divide the sample needed per variant by the expected eligible traffic per variant per day. Add time for relevant conversions to mature and consider business cycles or operational constraints.

Conclusion

How Long Should a Landing Page A/B Test Run? Until the experiment has collected enough valid evidence to support the decision you need to make, under the statistical method and stopping rule you chose in advance.

Start by defining the primary conversion, selecting a meaningful minimum detectable effect, and calculating the required sample size. Use eligible traffic—not total website visits—to estimate how long collection will take. Then account for conversion delays, relevant weekly or business cycles, and the quality of the data.

Avoid treating seven days, two weeks, or a particular significance threshold as a universal answer. A test is ready to conclude when its planned requirements have been met, its results are valid, and the remaining uncertainty is understood well enough to make a business decision.

Before launching your next experiment, calculate the required sample size and estimated duration. If the wait is not worthwhile, revisit the hypothesis or use another source of evidence rather than forcing a weak test to produce a winner.