Get My Free Acquisition Audit
Blog / CRO / LPO

Landing Page A/B Testing: How to Design Tests That Actually Move Results

Most A/B tests produce no statistically significant result, not because testing is broken but because the tests are designed incorrectly, run too briefly, or applied to the wrong elements. This guide walks through the correct testing hierarchy, minimum sample size, common design errors, and how to read results honestly.

Key Takeaways
  • Test in impact order: headline, CTA, proof placement, form length, visual, not button colour first.
  • Run tests simultaneously, not sequentially; sequential testing confounds time effects with variant effects.
  • Calculate required sample size before launch, not after, and stop at that number regardless of early results.
  • A non-significant result is still information: the change was undetectable or genuinely had no effect.

Why most tests fail to produce useful results

A/B testing is structurally sound as a method. The common failures are not with the method but with how tests are designed and stopped. The three most frequent problems are testing elements that cannot produce large enough effects to be detectable, stopping tests when they look good rather than when they have reached the required sample size, and testing multiple things simultaneously with traffic volumes that make interaction effects undetectable.

The result is a testing programme that produces mostly inconclusive results, or worse, confidently wrong ones. Correctly designed tests, starting with high-impact elements, run to adequate sample sizes, and read at a fixed significance threshold, give you something useful: knowledge of what actually moves results on your specific page with your specific traffic.

What to test and in what order

The hierarchy is driven by potential impact. Test the elements that, if changed substantially, would most affect the visitor’s decision.

PriorityElementWhy Here
1HeadlineSeen first, determines if the visitor continues reading; even small wording changes can produce large effects
2CTA copy and placementDirectly attached to the action; label and position both affect completion rate
3Proof type and locationMoving testimonials from footer to adjacent to the form reliably produces measurable changes on high-consideration pages
4Form lengthEvery field removed reduces friction; the SaaS demo funnel reached 64% completions by cutting from eleven fields to four
5Hero visualIllustrative vs decorative visuals can produce significant differences but require clean image production to test fairly

Button colour, font size, and similar cosmetic changes belong at the end of the list, not the beginning. They can produce results but rarely as large as the structural elements above, and testing them first spends traffic budget on low-yield questions.

Statistical validity basics

Four numbers define a valid test design. Work out all four before launching.

  • 01

    Baseline conversion rate

    Your current rate for the specific action on this specific page for paid traffic. Do not use blended site rates.

  • 02

    Minimum detectable effect

    The smallest improvement worth detecting. If you are converting at 3% and a 0.1% gain is not commercially meaningful, set the MDE at 0.5% or 1%. Smaller MDEs need larger samples.

  • 03

    Required sample size per variant

    Use any free sample-size calculator with 95% confidence and 80% statistical power. This gives you the number of visits each variant needs before you can trust the result.

  • 04

    Expected test duration

    Divide required sample size by your average daily paid traffic to the page. If the answer is more than eight weeks, either accept a larger MDE or wait until traffic volume supports testing.

Running the test cleanly

  • Split traffic 50/50. Unequal splits extend the time needed to reach significance without reducing risk materially.
  • Run both variants simultaneously. Sequential testing, variant A for two weeks then variant B for two weeks, confounds time effects with variant effects. Promotions, seasonality, and algorithm changes all move results.
  • Do not change anything on the page mid-test. Any edit to either variant invalidates the data collected to that point.
  • Set a stop date before you start. Decide the sample size in advance and stop when you hit it, not when the result looks the way you hoped. Early stopping with a leading variant frequently reverses as data accumulates.
  • Track the downstream metric, not just the page action. A form that converts at a higher rate but produces worse-quality leads is a net negative. Connect the test to CRM or order data where possible. This is covered in the offline conversion tracking guide.

Reading results honestly

A result that does not reach statistical significance is still information. It means either the change was too small to detect with available traffic, or there genuinely was no difference. Neither conclusion is a failure; both tell you something about where to focus next.

A statistically significant result tells you the effect is real, not how large it will be permanently. Winning variants often show smaller effects over time as novelty wears off or traffic mix changes. Record the result, implement the winner, and move to the next test. The full optimisation context is in the LPO framework, which sequences testing inside the wider conversion programme. Our CRO service and landing page optimisation run exactly this process on client accounts.

Frequently Asked Questions

Statistical significance is the probability that an observed difference between variants is not due to random chance. Most testing platforms use 95% confidence as the standard, meaning the result would occur by chance less than 5% of the time. Running a test until you hit 95% confidence, regardless of how long that takes, is more reliable than running for a fixed time.

It depends on your current conversion rate and the size of difference you want to detect. A page converting at 3% that you hope to improve to 4% needs roughly 10,000 visits per variant to reach significance. Lower starting conversion rates or smaller target uplifts require more traffic. Most free sample-size calculators will give you this number in 30 seconds.

Call the test early because one variant is leading. Early leaders frequently flip as more data comes in, which is why tests with inadequate sample sizes produce unreliable conclusions. Also avoid changing the page mid-test, running overlapping tests on the same traffic pool, or adjusting traffic splits after launch.

One thing at a time for most accounts. Multivariate tests require traffic volumes that most landing pages do not have: to detect interactions between multiple changing elements with statistical validity, the required sample size multiplies. Test the headline, then the CTA, then proof placement, each separately.

Ready to fix what’s costing you conversions?

We’ll review your acquisition funnel, show you exactly what’s underperforming, and hand you a clear, prioritised plan, whether or not you choose to work with us.

Get a clear diagnosis of your acquisition performance.

  • 30-minute review
  • Free, with no obligation
  • Clear findings and priorities
  • You keep the findings
Get My Free Acquisition Audit

30 minutes · Free · No obligation · You keep the findings

30 min · Free · Keep the report Get My Free Acquisition Audit