What Should I Test First on a Low Converting Landing Page?

What should I test first on a low-converting landing page? The answer starts with the cause of the problem and not with a page element. The first test belongs on the most plausible high-impact reason the page is underperforming, based on what your funnel data and the page itself show. Sometimes that turns out to be the headline, sometimes the offer, and sometimes the right move is not a test at all, because the tracking is broken or the wrong people are arriving.
This article discusses how to find where the page is losing people, how to separate problems you should simply fix from hypotheses worth testing, how to rank candidate tests by evidence, and how to run the first one so the result means something.
What should I test first on a low-converting landing page? The short answer
The short answer has three parts: diagnose, rank by evidence, and then test the biggest uncertainty. Skipping the first step is the usual reason a test program burns months on button colors while the real problem sits somewhere else. This section gives you the principle, shows how the published guides disagree, and explains why no single element deserves the top spot every time.
a. The principle: Test the most plausible high-impact cause, not the easiest element.
A landing page can underperform for at least eight different reasons: the wrong traffic, a mismatch between ad and page, a weak offer, a confusing page, friction in the form, missing trust, technical faults, or broken measurement. Each one calls for a different change. The element that is easiest to edit is rarely the one that explains the problem, but it feels like the obvious move because the page is the part you can see and change. So the first question is where the evidence says visitors are lost, and the second is what change would most plausibly fix that.
b. What the references say: Every guide ranks differently, and none shows data for its ranking.
Leadpages puts offer first, then message match, headline and first screen, call to action, form friction, proof, mobile, technical basics, and visual polish, and it presents this as its own judgment. KlientBoost also leads with the offer, arguing that it offers the biggest upside for the least effort. Knak starts with the headline, call to action, and form because they sit closest to the conversion. Contentful emphasizes headlines and gives no ranking. Global iTech Systems leads with matching the click. WSI says to begin with analytics and look for drop-off points. Webflow and Equinet Academy give no priority order at all.
Those lists overlap more than they conflict, since most start with something that shapes what the visitor expects (offer, message match, headline) before moving to mechanics (forms, speed). But they are practitioner opinions, and a ranking that fits one page can mislead on another. A page whose form throws an error on phones needs a fix before any headline test, whatever order a blog post prints.
c. Why no single element wins: The cause of the problem decides.
It's tempting to say always test the headline, or always test the call to action, or always test the form. Each of those is right for some pages and wasteful for others. A strong headline over a weak offer still fails. A short form on a page with the wrong audience produces many low-quality leads. A new hero section on a page that takes eight seconds to load on a phone changes nothing for the people who leave before it appears. Which element deserves a test depends on where the funnel is leaking and why.
d. Fixing and testing: Not every conversion problem is an A/B testing problem.
An A/B test answers a question you can't answer by looking, such as whether this headline persuades better than that one. It doesn't help with problems you can see directly, such as a form that fails on mobile, a conversion tag that never fires, or an ad that promises something the page doesn't mention. Those get repaired. Treating a broken page as a testing opportunity slows you down, and the section on when not to test below covers the cases in detail.
Diagnose before you test: Where is the page losing people?
Testing should follow diagnosis, because the right first test depends on where in the path visitors drop out. This section walks through the path from traffic to business outcome, and at each stage it describes what a problem looks like and what it points to. The signals below are clues, and any of them can have several causes, so treat each as a hypothesis to check. Our guide on how to diagnose a landing page that gets traffic but no conversions covers the same ground with a ten-step checklist, and this section focuses on what each stage means for the question "what should I test first on a low-converting landing page?".
1. Traffic: Are the right people arriving?
Break visitors down by source, campaign, search term, device, and location. A page can convert well for one audience and badly for another, and an overall average hides that. If most of your visitors come from one broad campaign and almost none of your conversions do, the page may be fine and the targeting may be the problem. That's a fix in the ad platform, and no page test will address it.
2. Click: Does the ad or link set up the right expectation?
A low click-through rate on the ad suggests an ad or targeting issue before the visitor ever sees your page. A healthy click-through rate paired with a very low page response suggests the ad promised something the page doesn't deliver. Comparing the ad wording, the search terms, and the page headline side by side takes a few minutes and often settles the question.
3. Landing page: Can a visitor tell what this is within a few seconds?
Show the first screen to someone who doesn't know your business and ask what is offered and who it's for. If they can't answer, the page has a clarity problem, and clarity problems sit in the value proposition and headline. This is also the stage where mismatch between the promise that earned the click and the page's opening shows up as a quick exit.
4. Engagement: Are visitors reading, scrolling, and interacting?
Scroll depth, clicks on elements, and exits show whether visitors find anything worth engaging with. Little engagement with decent traffic points to relevance or opening clarity. Strong engagement with no action points to an offer or call-to-action problem, because people are interested but unconvinced. Engagement numbers are indirect signals, so read them alongside recordings and comparisons between segments.
5. Call to action and form: Do interested visitors take the next step?
Separate the people who click the button, the people who start the form, and the people who finish it. A big drop between starting and finishing points to friction, trust, or technical errors, and few button clicks at all point to the offer or how the action is presented.
6. Conversion: Is the result being recorded correctly?
Check that the conversion event fires once per real conversion on every browser and device. Google's help page on key events in Google Analytics describes a key event as an event that measures an action that's particularly important to the success of your business, and it notes that you can turn a key event into a Google Ads conversion. A conversion that happens but isn't recorded looks the same as one that never happened.
7. Lead quality: Are the people who convert the right people?
Follow leads into your sales process and look at the share that is qualified, accepted by sales, and turned into opportunities. A page with a healthy conversion rate and poor lead quality has an offer, targeting, or qualification problem, and it needs a different test from a page with low conversion and good leads.
8. Business outcome: Did the page produce revenue or pipeline at an acceptable cost?
Connect the page to opportunities, revenue, and the cost of acquiring a customer. This is the stage that stops you from winning the wrong test. A change that cuts cost per lead but raises cost per customer has made the business worse, however good the page report looks.
Reading the pattern: Symptoms point to different first moves.
The table links common patterns to a likely cause and to whether the first move is a fix or a test. It is a set of starting hypotheses, since a single pattern can have more than one cause and most pages have more than one problem.
| What the data shows | Likely cause | First move |
|---|---|---|
| Low click-through on the ad | Ad or targeting | Fix or test in the ad platform |
| Plenty of clicks, quick exits, little scrolling | Message match or opening clarity | Compare ad and headline; test the headline or first screen |
| Visitors read to the call to action but don't click | Offer or call to action | Test the offer or the call to action |
| Many form starts, few submissions | Form friction, errors, or trust | Check for errors first; then test the form |
| Desktop converts, mobile doesn't | Mobile layout or speed | Fix obvious mobile faults; then test |
| Conversions look low, sales look fine | Tracking | Fix the measurement |
| Plenty of conversions, few good customers | Offer or qualification | Test the offer or form qualification |
| Weak engagement from every source | Offer or audience | Test the offer; review targeting |
Fix first, test second: When an A/B test is the wrong move
An experiment is the right tool when you have a real uncertainty and enough traffic to resolve it. When the problem is visible, or when the data can't support a test, the better move is to repair the page or gather other evidence. This section covers the situations where the answer to what should I test first on a low-converting landing page is nothing yet, and it separates fixing an obvious problem from testing an uncertain hypothesis.
a. Broken analytics or conversion tracking: You can't test what you can't measure.
If the conversion event misses some browsers, counts duplicates, or doesn't separate traffic by source, every test result inherits the error. Landing Doctor lists unreliable analytics as one of four preconditions for testing, and Leadpages tells readers to rule out technical problems because they distort every other test. Complete a real conversion on a phone and a computer and confirm it shows up in analytics, the ad platform, and your CRM before anything else.
b. Poor-quality or mismatched traffic: The page can't convert people who never wanted the offer.
If search terms are broad, audiences are loose, or placements are poor, the weak result comes from who clicked. Tightening targeting is a cheaper and larger move than any page change. The Google Ads help page on creating a good landing page says to choose a landing page that closely matches your ad and keywords, which makes the match between the click and the page a basic requirement and not an optimization.
c. Severe message mismatch: Repair it, then see what is left.
If the ad promises a free audit and the page headline sells a platform, you don't need an experiment to know the visitor will be confused. Bring the headline and offer in line with the promise, and then test refinements.
d. Obvious technical or mobile faults: Ship the repair.
Forms that fail on certain browsers, buttons that don't respond on phones, layouts that break on small screens, and very slow loads are defects. A test would only measure how much those defects cost you, which you rarely need to know before fixing them.
e. A fundamentally weak offer: Cosmetic tests won't rescue it.
If the evidence says people understand the page and still don't want what it offers, tweaking the layout wastes effort. This is where the offer itself becomes the test, which the section on offers below covers. KlientBoost makes the same argument, saying strong copy or design can't offset a weak offer.
f. Too little traffic or too few conversions: A test can't reach a trustworthy answer.
A page with a few hundred visits a month and a handful of conversions can run an A/B test for a year without a clear result. WSI's own example of a test with only 20 visitors is one it calls too small to trust. In that situation, use qualitative research, bigger changes, longer windows, and sequential fixes. The traffic section below explains how to judge whether you have enough.
g. A page that needs a redesign: Small variations won't answer a big question.
If the page's structure, message, and offer are all wrong, a headline test can't tell you much. Rebuild the page around a clear hypothesis, and test the rebuilt page against the old one as a whole, as the section on test size describes.
h. Problems you can diagnose without an experiment: Don't test what you can decide.
Not every change needs an A/B test. If a five-second review, a recording, or a customer complaint shows exactly what is wrong, fix it and watch the numbers afterward. Save experiments for choices where two reasonable options exist, and the evidence doesn't say which is better.
How to decide what to test first
Once you have ruled out the fixes, you will usually be left with several plausible hypotheses and room to run only one or two at a time. This section gives you the factors to weigh and a way to use them without inventing a formula. Several published scoring models exist, including ICE, PIE, and PXL, and a comparison from Mida notes that the inputs to ICE are all subjective and that PXL can penalize unusual tests that might have produced large lifts. Treat any scoring model as a way to organize your judgment, and not as an authority.
a. Potential impact: How much could this move the outcome?
A change to something every visitor meets, such as the offer, the headline, or the form, can move more than a change to something few visitors reach, such as a footer link. Impact also depends on where the funnel leaks. If most visitors leave before the first scroll, improving content further down touches only the few who got there.
b. Evidence and confidence: What supports the hypothesis?
A hypothesis backed by recordings, survey answers, and funnel data deserves more confidence than one that a team member suggested in a meeting. Leadpages states the test it wants to see as a problem you observed, a change that should fix it, and a metric that shows whether you were right, and it adds that if you can't write the hypothesis, you shouldn't run the test. The more independent evidence points to the same cause, the earlier that test should run.
c. Ease and speed: What does it take to build and launch?
A new offer or a rewritten headline can launch in a day, and a rebuilt page can take weeks. Ease matters, but it ranks below impact and evidence. Low-effort tests on low-impact elements keep a team busy and teach it little.
d. Traffic available: Can this test reach a result?
A test aimed at a small segment, such as visitors on one browser, may never collect enough data. Prefer tests that affect a large share of your traffic, and aim them at the pages with the most visitors and conversions, which both Knak and WSI recommend.
e. Measurability: Can you tell whether it worked?
Choose tests whose effect appears in a metric you can track reliably. A test of trust signals whose only visible effect is a change in lead quality needs lead outcomes wired back to the page, which many teams haven't done yet.
f. Risk and business importance: What if the variation does harm?
A pricing page or a checkout step carries more risk than a blog sidebar, so test changes there with care and send a smaller share of traffic to the variation. Business importance also matters, since a campaign that drives a large budget deserves attention before one that drives little.
g. Conversion quality: Will the result be worth having?
Some tests can raise the number of conversions while lowering their value. Before you commit, decide how you will check lead quality and downstream results for each variation.
h. A practical sequence: Write the hypothesis, rate the evidence, and pick the largest uncertainty.
List the two or three most plausible causes, write a one-line hypothesis for each in the form "Because we saw this, we believe changing that will improve this metric for this audience", and rate how strong the evidence for each is. Run the test that combines the highest potential impact with the strongest evidence, and let ease and speed break ties. A good first test usually targets a meaningful conversion hypothesis, such as the offer, the message, or a friction point in the form, and a cosmetic variable such as button color rarely qualifies unless contrast is demonstrably the problem.
Use qualitative research to find what deserves a test
Numbers tell you where visitors leave, and they rarely tell you why, which is why the question "what should I test first on a low-converting landing page?" needs research behind it. Qualitative evidence fills that gap, and it is the best source of strong hypotheses. This section separates research that generates hypotheses from experiments that validate them, and it describes where each type of evidence helps and where it misleads.
a. Two jobs: Research generates hypotheses, and experiments validate them.
Recordings, interviews, and survey answers suggest what might be wrong, based on a handful of people, and they can be misread. An A/B test measures whether a specific change shifts behavior across many visitors, but it can't tell you what to try. Using research to choose the test and the test to confirm the choice gets more from both. Equinet Academy frames this as moving from an observation to a problem to a hypothesis, then to a change and a metric.
b. Session recordings and heatmaps: See what visitors do.
Microsoft Clarity describes click maps as showing which areas attract the most clicks, scroll maps as showing how far visitors engage down a page, and session recordings as giving detail on individual behavior. These help you spot rage clicks, ignored buttons, and pages where most people never reach the offer. They show behavior and not motive, so use them to form a question and then check it another way.
c. On-page surveys and post-conversion questions: Ask what nearly stopped them.
A one-question survey on the page or after the conversion, such as asking what almost kept someone from signing up, surfaces objections you hadn't thought of. Global iTech suggests a post-conversion question about what influenced the decision. Answers come only from the people who stayed or converted, so they reveal objections that were overcome and miss the ones that caused exits.
d. Customer support and sales conversations: Collect the objections people actually raise.
The questions that support and sales hear repeatedly are a ready list of what the page fails to answer. If prospects keep asking about pricing, security, or setup time and the page says nothing, that is a hypothesis with strong evidence behind it.
e. Search-query data and ad comments: Learn what people expected.
Search terms and the comments under social ads show the language and the expectations visitors bring. They help you judge message match before the visitor lands, and they point to audience segments that may need their own page.
f. Form analytics and funnel data: Locate the exact step.
Field-level drop-off, error rates, and funnel steps show where a form loses people. Pair them with a recording of a failed submission to separate friction from a technical bug.
g. User interviews: Deepen the picture.
A few conversations with recent customers or lost prospects can show why the page persuaded or failed to. They are slow and small in scale, so weigh them as strong hypothesis fuel and not as proof.
What to test, area by area
After diagnosis, "what should I test first on a low-converting landing page?" becomes a question about which part of the page goes into the experiment. This section covers eleven areas, and for each one it says when it makes a sensible first test, what a good hypothesis looks like, and what to measure. Use it as a reference once the evidence has narrowed your choice, and not as a menu to work through in order.
a. Value proposition: Is it clear what the visitor gets and why it matters?
- Test this first when: First-time visitors can't say what the page offers, or recordings show people leaving in the first seconds.
- Example hypothesis: Because new visitors in our survey describe the product in three different ways, a first screen that states one concrete outcome for one audience will raise CTA clicks.
- Measure: CTA click-through rate first, then conversion rate, with scroll depth as a check that people stay.
b. Headline and supporting copy: Does it say the specific thing this visitor came for?
- Test this first when: The ad or link promised something specific and the headline is generic, or the page copy describes features when visitors ask about outcomes.
- Example hypothesis: Because search terms are mostly about a specific problem, a headline that names that problem will lift conversion among paid-search visitors.
- Measure: Conversion rate by source, with bounce or early exits as a companion. Test one idea per variation, such as benefit-led against feature-led, so you know what moved.
c. Message match: Does the page continue the promise that earned the click?
The chain runs from audience or search intent, to the ad or message, to the landing page headline, to the offer, to the call to action. A page can contain good copy and still fail because one link in that chain breaks. If the ad says "free audit" and the headline says "Grow your business with our platform," the visitor has to work out whether they are in the right place, and many won't.
- Test this first when: Clicks are healthy, engagement is weak, and the page opens with something different from what the source promised.
- Example hypothesis: Because the ad offers a free audit, a headline and form that repeat "free audit" will raise form starts.
- Measure: Form starts and conversion rate by campaign.
d. Offer: Is this what the visitor wants to trade their details or money for?
Changing the offer, such as a trial length, a consultation, a discount, an assessment, or a different asset to download, can have a larger effect than changing how the same offer looks. KlientBoost argues that the offer carries the biggest upside for the least effort, and that claim is its own judgment, but the logic holds: if the proposition is weak, nothing else rescues it.
- Test this first when: Traffic matches, the message is clear, friction is low, and visitors still don't act, or when lead quality is poor.
- Example hypothesis: Because most visitors reach the pricing section and leave, a free assessment as the entry step will convert more than a request for a demo.
- Measure: Conversion rate, and then qualified-lead rate and sales acceptance, because an easier offer can lower quality.
e. Call to action: Is the next step clear, relevant, and in the right place?
Call-to-action tests cover far more than color. They include the wording, how prominent the button is, where it sits, how many competing actions the page offers, and whether the action fits the visitor's readiness. A cold visitor asked to "Book a demo" may need a lighter first step.
- Test this first when: Visitors read to the end and few click, or the page offers several equally weighted actions.
- Example hypothesis: Because visitors who scroll deep still don't click, a single primary action that states what happens next will increase clicks.
- Measure: CTA click-through rate, then completed conversions.
f. Forms: How much do you ask, and what does it cost the visitor?
You can change the number of fields, which fields are required, whether the form sits on the page or opens after a click, whether it runs as one step or several, the reassurance text near it, and how errors appear. Fewer fields often raise completions, but they can also raise the share of poor leads, so the decision depends on how much qualification your sales process needs.
- Test this first when: Many people start the form and few finish, or lead quality is the problem.
- Example hypothesis: Because analytics show most abandonment at the phone field, making it optional will raise completions without hurting qualification.
- Measure: Form completion rate, qualified-lead rate, and field-level drop-off.
g. Trust and proof: Does the page answer the visitor's doubts?
Testimonials, reviews, customer logos, case studies, security and privacy reassurance, certifications, and guarantees all help when they address the objection that is actually stopping the visitor. A security badge does little for a visitor who doubts the product works. Specific proof, such as a named customer describing a measurable result, usually carries more weight than generic praise.
- Test this first when: Surveys, sales calls, or support questions show repeated doubts, especially on higher-priced or higher-risk offers.
- Example hypothesis: Because prospects keep asking whether the product works for companies our size, a case study from a similar company near the form will raise conversion.
- Measure: Conversion rate and the quality of resulting leads.
h. Page structure and hierarchy: Does the order of the content match the order of the visitor's questions?
This covers the hero section, the sequence of benefits, features, proof, and objections, what appears above the fold, and how dense the page is. Structure tests usually follow evidence that visitors are missing important content, such as scroll maps showing most people never reach the proof section.
- Test this first when: Scroll data shows visitors stop before the content that would persuade them.
- Example hypothesis: Because few visitors reach the testimonials, moving proof above the first call to action will raise clicks.
- Measure: Scroll depth, CTA clicks, and conversion rate.
i. Design and UX: Does anything get in the way?
Layout, visual hierarchy, distracting navigation, readability, accessibility, and interaction friction all belong here. Design tests should have a reason behind them: a recording that shows people missing the button, or a navigation bar that leads away from the conversion. Equinet Academy and Leadpages both advise tying changes to an observation, and Leadpages says button color matters only when contrast or visibility is the actual problem.
- Test this first when: Recordings show repeated confusion with a specific element.
- Example hypothesis: Because recordings show visitors tapping the headline thinking it is a link, removing the underline style will cut dead taps.
- Measure: CTA clicks and conversion rate, with segments by device.
j. Speed and technical performance: Fix it before you test.
A slow or broken page distorts every other test, since the visitors who leave before it loads never see your variation. Open the page on a real phone, run a speed report, and repair what you find. Google's Ads help page on landing pages names ease of navigation and quick mobile experience among the basics. A test belongs here only when a repair involves a real trade-off, such as a heavy video that helps persuasion but slows the page.
Measure: Load and interaction metrics alongside conversion rate.
k. Personalization and audience-specific variants: When does one page stop being enough?
If two audiences arrive with different needs, one generic page serves neither well. Separate pages or sections for each audience can outperform a compromise, and the test is whether the added work pays off. Our guide on how many landing pages a PPC campaign should have covers when splitting by audience or ad group is worth it.
- Test this first when: Segment data shows one audience converts well and another poorly on the same page.
- Example hypothesis: Because visitors from the pricing keywords convert far better than those from the educational keywords, a separate page for the educational group will lift its conversion.
- Measure: Conversion rate for each segment, plus lead quality.
Should you test one thing, several, or a whole new page?
Whatever you decide about "what should I test first on a low-converting landing page?" the size of the test decides how clearly you learn and how large an effect you can find. A small change gives a clean answer and a small effect, and a large change can move the result a lot while leaving you unsure which part did the work. This section lays out the options and the trade-off between them.
a. One variable: The cleanest read, best for isolating a specific cause.
Changing a single element lets you attribute the result to it. Most references, including Knak, WSI, and Equinet Academy, recommend this. It works well when you have a precise hypothesis and enough traffic to detect a modest effect. The cost is speed, because testing one thing at a time takes many rounds.
b. Several related changes: A larger effect with a less certain cause.
Sometimes a hypothesis involves a bundle, such as a new offer with a matching headline and call to action. Testing the bundle as one variation gives you a bigger potential effect, and you learn whether the direction works while not learning which piece mattered. This is reasonable when the changes serve one idea and would look strange apart.
c. A complete page variation: A split-URL test for major redesigns.
Webflow describes split-URL testing as comparing entirely different pages, which suits a major redesign. This is the right choice when the page's structure and message are in question. If the URLs are public, Google's guidance on website testing and search tells you to use a rel=canonical link on alternate URLs, use 302 redirects and not 301s, avoid showing Googlebot different content from visitors, and remove test elements promptly once the test ends.
d. Sequential testing: One change after another with the winner as the new baseline.
The winner of one test becomes the control for the next. This works when traffic is limited and decisions need to be made one at a time, and it needs discipline so that seasonal swings or campaign changes aren't mistaken for the effect of your edits.
e. Multivariate testing: Useful when traffic is high, and not automatically better.
A multivariate test varies several elements at once and estimates the effect of each combination. WSI's example of two headlines, three images, and two buttons makes twelve combinations, and each needs enough visitors to be measured. WSI itself notes that this isn't practical for every site. For most pages with modest traffic, a simple A/B test is the better choice.
f. The trade-off: Learning clarity against the size of the change.
| Approach | What you learn | Traffic needed | Best when |
|---|---|---|---|
| One variable | Exactly what moved | Moderate | You have a precise hypothesis |
| Related changes | Whether a direction works | Moderate | Several changes serve one idea |
| Whole-page variation | Whether a new approach beats the old | Moderate to high | The page needs a rethink |
| Multivariate | Which combination works | High | You have a lot of traffic and several candidate elements |
How to run a landing page test that teaches you something
Once you know the answer to "what should I test first on a low-converting landing page?", a well-run test starts before you build the variation and ends after you have written down the result. This section covers the steps, and the next two sections cover the metrics and the traffic questions that go with them. Our guide on the build, measure, and improve loop describes a routine for repeating them.
1. Define the hypothesis: Name the problem, the change, the audience, and the reason.
Write it as a sentence: because we saw this, we believe that changing this will improve this metric for this audience. A hypothesis that can't be written down in that form isn't ready.
2. Set the primary goal: One conversion that the test is judged on.
Pick the primary conversion before the test starts and don't change it after you see the results. Knak's advice to set one conversion goal per page points the same way, and choosing the metric after seeing the results is one of the easiest ways to fool yourself.
3. Choose diagnostic metrics: Secondary numbers that explain the result.
Track things like CTA clicks, form starts, and scroll depth so that you can see how a result came about, and use them as explanations, not as winners.
4. Define the control and the variation: Document every change.
Record exactly what differs. If the variation changes several things, say so in the log, since you will want to know later what the result does and doesn't prove.
5. Check tracking before launch: Test the conversion on both variations.
Complete a real conversion on each version, on phone and computer, and confirm it records. A test with broken tracking in one arm produces a confident wrong answer.
6. Split traffic randomly and run both versions at the same time: Avoid before-and-after comparisons.
Comparing this month's variation with last month's control mixes the page change with every other change in traffic, season, and campaign. A simultaneous split avoids that.
7. Don't change the control mid-test: Leave the test alone.
Edits to the control, to the ads sending traffic, or to the tracking during a test make the result hard to interpret. WSI warns that changes made mid-test can corrupt the results, and Global iTech warns against stacking simultaneous changes.
8. Decide the stopping rule in advance: Don't stop on a good-looking early number.
Evan Miller's article on how not to run an A/B test explains that checking results repeatedly and stopping when they look significant raises the rate of false positives. In one worst-case example he describes, a change with no real effect was flagged as significant 26.1% of the time and not the nominal 5%. His recommendation is to decide on a sample size in advance and wait until the experiment is over, and he points to sequential and Bayesian designs for teams that need to stop early with valid results.
9. Document and decide: Implement, iterate, or reject.
Write down the hypothesis, the result, the segments you checked, and what you will do next. A test that doesn't beat the control still tells you that one idea is unlikely to be the cause. The log is what turns a series of tests into a strategy.
What to measure: Conversion rate is not the only number
A test should be judged on the outcome you care about, and the easiest metric to move is often not that outcome. The answer to "what should I test first on a low-converting landing page?" also depends on which result you plan to judge it by. This section describes primary, diagnostic, and business metrics and shows with a hypothetical example how a winning conversion rate can hide a worse result. Our guide on what conversion rate is and how to measure it covers the basics.
a. Primary metric: The conversion the page exists to earn.
This is usually a completed form, a purchase, a booking, or a trial start. Google's help page on key events describes the setup in Google Analytics, and the definition is yours to choose. Whatever you pick, keep it fixed for the whole test.
b. Diagnostic metrics: Form completion rate, CTA click-through rate, and engagement.
These show where the change had its effect. A headline test that raises CTA clicks but not completions tells you the headline is working, and something later is not. Bounce or exit behavior and scroll depth are useful as supporting evidence, though Knak calls clicks and time on page directional at best.
c. Quality metrics: Qualified lead rate and sales acceptance.
Count how many conversions are valid, how many sales accepts, and how many become opportunities. Without these, you can't tell a better page from an easier one.
d. Business metrics: Pipeline, revenue, customer acquisition cost, and lead-to-customer rate.
These connect the page to money. They arrive slowly, since a sale can take weeks, so decide up front how long you'll wait and what interim quality signal you'll use.
e. A hypothetical example: The variation that wins on conversions and loses on customers.
Suppose a control and a variation each receive 10,000 visitors from $12,000 of ad spend, and 20% of qualified leads become customers.
| Control | Variation (shorter form) | |
|---|---|---|
| Leads | 300 (3%) | 500 (5%) |
| Qualified leads | 90 (30% of leads) | 70 (14% of leads) |
| Customers | 18 | 14 |
| Cost per lead | $40 | $24 |
| Cost per customer | about $667 | about $857 |
The variation wins on conversion rate and cost per lead and loses on customers and cost per customer. A team that stopped at the conversion rate would ship the worse page. The numbers are hypothetical, and the pattern is common enough that you should plan the quality check before the test begins.
How much traffic do you need, and how long should a test run?
Before you commit to a first test, check that the page can support one. The honest answer is that it depends on your baseline conversion rate, the size of the lift you want to detect, and how certain you want to be. This section explains why the published rules of thumb conflict, shows how sample size works with one sourced example, and offers options when your traffic is low.
a. The references disagree: Durations from days to weeks, with no derivation.
WSI's process step says to run a test for 24 to 48 hours, while its mistakes section says a couple of weeks. Knak suggests a week or two at minimum for most B2B pages, and Contentful and Global iTech both default to two weeks. KlientBoost says sample sizes are usually in the thousands per variant and aims for 90 to 95% confidence. Leadpages gives no figure at all. None of them show how they arrived at their numbers, and they can't all be right for every page.
b. Sample size depends on your baseline and the lift you want to detect.
Evan Miller's sample size calculator shows the idea in a worked example: with a 10.2% baseline and a target of 13.2%, it calls for 2,545 visitors per variation. The calculator also asks for a significance level and statistical power. A lower baseline, or a smaller lift you want to detect, needs more visitors, so a page with a 1% conversion rate may need far more than that example. Run the calculator with your own baseline and the smallest improvement that would matter to the business.
c. Cover full cycles: Run long enough to include normal variation.
Leadpages advises running tests long enough to cover normal traffic patterns, and warns against calling a winner on very few conversions or after one unusually good or bad day. Weekday and weekend traffic often differ, and campaigns have their own rhythms, so a test that ends mid-cycle can mislead. That is a reason to avoid stopping early, and it does not set a fixed number of days.
d. Confidence is not a probability that you're right: Read significance carefully.
One reference describes 95% confidence as a 95% chance that the result is reliable. A significance level doesn't measure that. It describes how often a no-effect change would look this good by chance under the test's assumptions. Statistical significance is also not the only consideration, since a significant result can be too small to matter or can come from a segment that doesn't match your target customer.
e. When traffic is low: Change the approach, not the standard of evidence.
If the sample size you need is out of reach, you have four options. Test bigger changes that could plausibly produce bigger effects, such as a new offer in place of a new button label. Extend the measurement window. Combine similar pages into one comparison. Lean on qualitative evidence and fix clear problems directly. Leadpages and WSI both point to qualitative data and larger changes for low-traffic pages. Don't declare a winner on three conversions.
How the first test changes by traffic source
What should I test first on a low-converting landing page can have a different answer depending on where visitors come from, because intent, familiarity, and buying stage differ by channel. The table below describes the visitor's likely mindset and a sensible first thing to check, and it reflects reasoning from how each channel works and not measured data. Check your own segment numbers before acting on any row.
| Channel | Visitor's likely state | First thing to check or test |
|---|---|---|
| Google Search ads | Looking for a specific answer now | Match between search terms, ad, and headline |
| Microsoft Ads | Similar search intent, often a different audience mix | The same match checks, with segments reviewed separately |
| Meta ads | Interrupted while browsing, not searching | The promise, the offer, and a low-commitment first step |
| LinkedIn ads | Professional context, B2B decision-makers | Offer fit for the role and form qualification |
| Display | Low intent, early awareness | Whether the offer suits a cold visitor |
| Retargeting | Already knows you | Objection handling and proof, since the basics are covered |
| Organic search | Intent varies with the query | Match to the query that ranks, then clarity |
| Warm, chose to hear from you | Continuity from the email and a clear single action | |
| Referral | Depends on the referrer's context | Whether the page continues the referrer's promise |
a. Search ads: Intent is high, so mismatch hurts most.
Google's help page on landing pages says to choose a landing page that closely matches your ad and keywords, and that the page should mirror the ad's call to action. With search traffic, the first test is often the headline or the offer's wording against the actual queries. Our guide on ads that get clicks but few conversions walks through this case in more detail.
b. Social and display: Interruption means the first test is often the offer or the first step.
Visitors who weren't looking for you need a reason to care and a small first step. A lighter offer or a clearer statement of value often matters more than form tweaks.
c. Retargeting and email: Familiar audiences need different questions answered.
People who already know you are weighing objections, not learning what you do, so proof, pricing clarity, and risk reduction move up the list. Email visitors expect continuity with the message they just read.
d. Test in the ad platform when the problem is in the ad: Use experiments where the change lives.
When the weak point is the ad copy, bid strategy, or targeting, test it where it is configured. Google Ads' help page on custom experiments describes a test copy of an existing campaign that runs alongside the original, shares its traffic and budget, and lets you choose the share allocated to the experiment, with 50% recommended. For Search, it recommends a cookie-based split so each user stays in one version; the original campaign has to stay active, and edits to the original during the test can make results harder to read. Landing page changes themselves are usually tested in your page builder or testing tool.
Hypothetical examples: Choosing the first test from the evidence
The examples below are hypothetical and show how to answer "what should I test first on a low-converting landing page?" by reasoning from the evidence to the first test. They are not case studies, and the numbers are made up for illustration.
a. A Google Ads page with clicks and few conversions.
- What the evidence shows: 400 clicks a week, 1.2% conversion, search terms that are broad, and a generic headline.
- Likely problem: Message match and traffic quality.
- First test: Tighten the search terms first, then test a headline that repeats the main query.
- Why this comes first: The visitors may not want the offer, and a headline test on the wrong audience would mislead.
- What to measure: Conversion rate by search-term group, then cost per conversion.
b. A SaaS demo page with strong traffic and weak form completion.
- What the evidence shows: 60 form starts and 10 submissions a week, most abandonment at a required phone field, mobile worse than desktop.
- Likely problem: Form friction, and possibly a mobile fault.
- First test: Check the form on a phone for errors and fix any. Then test making the phone field optional.
- Why this comes first: Funnel data points at one field, so the hypothesis has strong evidence and low effort.
- What to measure: Form completion rate and the qualified-lead rate.
c. A B2B page with many leads and few qualified opportunities.
What the evidence shows: 200 leads a month from a free template, and sales accepts 6.
Likely problem: The offer attracts the wrong audience.
First test: Offer a more specific asset aimed at the buyer role.
Why this comes first: Changing the form wouldn't change who wants a generic template.
What to measure: Sales-accepted lead rate and pipeline from the page.
d. An ecommerce page with high engagement and low purchases.
What the evidence shows: Visitors scroll through the product page and view the delivery section, and few add to cart.
- Likely problem: A risk or cost surprise, such as delivery costs or returns terms shown late.
- First test: Show delivery cost and return terms near the price.
- Why this comes first: Engagement is strong, so the objection is probably at the point of decision.
- What to measure: Add-to-cart rate and purchases.
e. A retargeting page for visitors who already know the product.
- What the evidence shows: The audience has seen the pricing page and the page repeats the introduction.
- Likely problem: The page answers questions they no longer have.
- First test: Replace the introduction with answers to pricing and risk objections, plus a case study.
- Why this comes first: The audience has context, so the page should skip the basics.
- What to measure: Conversion rate and the cost per conversion.
f. A high-ticket service page where trust matters more than the button.
- What the evidence shows: Consultations cost the visitor an hour, and the service costs $15,000; recordings show visitors reading testimonials and leaving.
- Likely problem: Trust and perceived risk.
- First test: Add a case study from a similar client and a clear description of what happens on the call.
- Why this comes first: The button color is not the barrier at this price.
- What to measure: Booking rate and the share of bookings that become clients.
g. A mobile-heavy campaign with form friction.
- What the evidence shows: 80% of visits come from phones, and form completion is far lower on mobile.
- Likely problem: Mobile form usability.
- First test: Fix visible mobile faults, then test a shorter or multi-step form.
- Why this comes first: The largest segment has the weakest result.
- What to measure: Mobile form completion and lead quality.
h. A page with a strong conversion rate and poor lead quality.
- What the evidence shows: A 14% conversion rate, with 3% of leads accepted by sales.
- Likely problem: The form or offer doesn't qualify.
- First test: Add a qualifying question, such as team size or budget range, and state who the offer is for.
- Why this comes first: The rate is not the problem. Quality is.
- What to measure: Qualified-lead rate, cost per qualified lead, and pipeline.
Common testing mistakes
Most bad answers to "what should I test first on a low-converting landing page?" fall into a few groups. Spotting which group a mistake belongs to is half the fix.
a. Starting in the wrong place: Testing before diagnosing.
- Testing button color first without evidence: Unless contrast is the problem, it rarely addresses the cause.
- Copying competitor tests: Their traffic, offer, and audience differ from yours.
- Testing random elements: Without a hypothesis, you can't learn from the result.
- Ignoring traffic quality, message match, and the offer: These cause more underperformance than page elements do.
b. Running the test badly: Design and discipline errors.
- Testing too many things at once: You won't know what worked.
- Changing the control during the test: The result becomes uninterpretable.
- Running tests with too little traffic: You'll get noise that looks like a finding.
- Ending tests too early: Early leads often reverse, as Equinet Academy warns.
- Treating significance as the only consideration: Size of effect and quality of conversions matter too.
c. Reading the result badly: Metrics and meaning.
- Optimizing only for conversion rate: The earlier example shows how it can mislead.
- Assuming a winning variation is universally better: It won for this audience and period.
- Failing to document learnings: The same idea gets tested twice.
d. Running tests without a strategy: Process errors.
- Endless tests with no backlog or priority: Activity without learning.
- Creating AI-generated variants without a hypothesis: You get volume and no insight.
Using AI to speed up landing page testing
AI can shorten the time between evidence and a test, and it can't tell you the answer to "what should I test first on a low-converting landing page?" for your situation. This section separates the tasks where it helps from the judgments that stay with you. Our guide on creating a landing page with AI covers the building side.
a. Where AI helps: Reading, sorting, and drafting.
An assistant can summarize survey answers and sales-call notes, group recurring objections, draft headline and copy variations from a stated hypothesis, write messaging for specific audiences, organize a testing backlog, and produce test documentation. Each of those saves time when you give it real inputs and check what comes back.
b. Where AI doesn't decide: Evidence, design, and judgment.
AI is not a substitute for real customer evidence, analytics, experiment design, statistical judgment, business context, or human review. An assistant can generate fifty headline variants in a minute, and none of them is a good test unless it addresses a cause you have evidence for. It can also state a plausible benchmark or finding with nothing behind it, so verify any figure it gives you.
c. Faster testing versus knowing what to test: Keep the two separate.
Speed in producing variations makes it easy to run many weak tests. The scarce skill is picking the right problem, and that comes from the diagnosis and research described above.
How Episode helps
Episode is a campaign landing page and growth platform, and it helps with the workflow around testing and not with deciding the answer to "what should I test first on a low-converting landing page?" It can't decide what your evidence means, repair a weak offer, or replace your analytics and CRM, and it won't make a test meaningful when traffic is too low. Where it helps is in building and changing pages quickly, capturing leads, showing where visitors drop out, and running page variants.
a. Build and change pages without a development queue.
Episode creates a campaign page from a brief and your company URL, applies a brand kit, and lets you edit it in a visual studio. That shortens the time from hypothesis to variation, a point we cover in our guide to launching a campaign landing page in under a day. Our piece on how a campaign page differs from a normal website page explains why that speed matters for campaigns.
b. Capture leads and send them to your CRM.
Forms and lead capture on the page send leads to your CRM. Connecting the page to the place where leads are qualified is what lets you check lead quality by page and source, which is what the earlier example depends on.
c. See where visitors drop out.
Episode's analytics include funnel views, heatmaps, and UTM and traffic-source reporting. These support the diagnostic steps above. They don't replace your main analytics setup, and you still need to define the conversion and check tracking.
d. Run page and section variants.
Episode supports page and section experiments, so you can create a variant and compare it without rebuilding the page. You still need to write the hypothesis, decide the stopping rule, and check quality. Recommendations from the platform are suggestions that you review and apply or reject.
e. Where it doesn't fit.
The benefit is smaller if you rarely launch campaign pages, and Episode doesn't necessarily replace your main website or content system. If your main constraint is traffic or the offer, the tool won't change that.
Conclusion: What should I test first on a low-converting landing page?
So, what should I test first on a low-converting landing page? Test the most plausible high-impact cause of the problem, as shown by your funnel data, your qualitative research, and the page itself. Fix tracking, traffic, message match, and technical faults before you experiment, because tests can't rescue those. Then rank the remaining hypotheses by impact and evidence, write each one down, and run the best one with a sample size and stopping rule decided in advance.
Judge the result by the outcome that matters to the business, which is usually qualified leads or revenue and not the conversion rate alone. If you do that consistently, each test teaches you something, and your next test starts from stronger evidence than the last.