The short version: you judge a landing page design by what it does with real traffic over a proper sample size, not by what your team thinks looks nicer in a meeting. Track the whole funnel, not just the click, because the version with the higher conversion rate sometimes brings in worse leads. Give any test at least two to four weeks and roughly 1,000 visitors per variant before you trust the number.
Stop judging with your eyes first
Here’s the bit that annoys clients when I say it out loud: your opinion of a landing page design is worth almost nothing. Mine too. I’ve sat in rooms with marketing directors who hated a page because the colour clashed with the brand guidelines, and that page went on to convert at nearly double the rate of the “on-brand” version. Taste and conversion are not the same thing, and treating them as the same thing is how businesses waste months arguing about button shades instead of running a test that would settle it in three weeks.
If you want a working principle, use this one: a landing page design is only “good” in relation to a specific audience, a specific offer, and a specific traffic source. A page that converts brilliantly for cold Facebook traffic can flop for warm email traffic sent to the same URL. So the first honest answer to “which design converts best” is: best for what, and best for whom.
The story of the two landing pages
A training company I worked with in Kent last year had two versions of a lead page for a webinar. Version A was clean, minimal, one headline, one form, lots of white space, the kind of design that wins awards. Version B was longer, had three testimonials, a short video, and a bullet list of what attendees would learn.
Over four weeks with paid traffic split 50/50, Version A converted at 3.4 percent. Version B converted at 2.1 percent. On raw numbers, A won by a mile, and the founder wanted to kill B immediately.
I asked him to wait for the sales data instead of the form data. Leads from Version A closed at 18 percent. Leads from Version B closed at 40 percent, because the longer page had pre-sold and pre-qualified people before they ever filled in the form. Fewer, better leads. When you multiply it out, Version B produced more paying customers per pound spent even though it “lost” the conversion rate test. That’s the uncomfortable truth most people judging landing pages never get to, because they stop measuring at the form submit.
What you’re measuring matters more than the design
Before you compare two designs, decide what a win looks like, and pick a number that connects to revenue, not just activity. For a lead gen page that’s usually cost per qualified lead or lead-to-sale rate, not conversion rate on its own. For an ecommerce page it’s usually average order value alongside conversion rate, since a page that converts more people into smaller orders can leave you worse off. I’ve seen this trip up businesses that had strong copywriting on the page but were still judging success on the wrong metric entirely, so the winning version kept getting picked based on a number that didn’t matter.
Write your success metric down before you launch the test. If you decide afterwards, based on which version happens to be winning, you’ll unconsciously pick the metric that flatters your favourite design. I’ve done this myself, more than once, and caught myself doing it.
A simple step by step process for judging design
This is the process I use with clients when they can’t agree on which version to run with.
- Step 1: Define the single primary goal of the page (booked call, form fill, purchase, email signup) and the one metric that proves it worked.
- Step 2: Build two versions that differ in more than one small element, at least at first. Testing button colour when you have low traffic wastes weeks you don’t have.
- Step 3: Send an even split of the same traffic source to both, at the same time, for the same length of campaign. Never compare a page that ran in January against one that ran in July.
- Step 4: Wait until each version has had roughly 1,000 visitors or 100 conversions, whichever comes first, before you look at the result seriously. Below that, you’re reading noise, not a signal.
- Step 5: Check the downstream numbers, not just the form or button click, ideally two to four weeks after the lead or sale happened, so you can see close rate or refund rate too.
- Step 6: Roll out the winner, then test the next single variable against it. One test never ends a design conversation, it starts the next one.
The sample size problem nobody likes to mention
I’ll say the thing most landing page articles skip: the vast majority of small business “tests” I see are decided after two or three days and a few dozen visitors. That’s not a test, that’s a coin flip you’ve dressed up in a chart. If your page gets 40 visitors a day, you don’t have enough traffic to run a meaningful A/B test in under a month, full stop. In that situation, judging design by split test is the wrong tool entirely, and you’re better off judging by user recordings, heatmaps, and a handful of real phone calls where you ask people what nearly stopped them buying.
I made this mistake early on with a webinar funnel of my own. I called a test after 260 total visitors because the numbers looked clear, switched the whole funnel over, and watched conversion drop the following month. The first result was luck, not a genuine pattern. Now I set a minimum visitor count before I’ll even open the results dashboard, and I stick to it even when I’m impatient, which is often.
What good landing page design does, regardless of style
Strip away the visual trends and every landing page that consistently converts well shares the same underlying structure: one offer, one call to action repeated at sensible intervals, proof placed close to the point of doubt, and no navigation menu pulling people away before they’ve decided. This is also why a landing page nearly always outperforms sending traffic to a homepage; a homepage has to serve ten different visitor intentions at once, and a landing page only has to serve one. I’ve written before about why landing pages beat homepages for conversions, and the short reason is focus. Judging design in isolation from that structure is how businesses end up praising a beautiful page that quietly underperforms an ugly one built around a single clear ask.
Design elements that consistently move the needle across the tests I’ve run for clients: a headline that states the specific outcome rather than the category (so “Get 20 qualified leads a month without cold calling” rather than “Marketing Services”), a form with fewer fields than feels comfortable, a single primary button colour used nowhere else on the page, and testimonials that include a number or a name rather than a vague quote. None of that is about taste. It’s about removing friction and doubt in that order.
When AI tools help and when they get in the way
There’s a growing pile of AI landing page builders and heatmap-with-AI-summary tools that will happily tell you which version “performed better” after a few dozen visits. Be careful here. These tools are built to give you an answer quickly, and quickly is not the same as correctly. If you’re weighing up which of the many AI tools out there to bring into your testing process, it’s worth reading through how to choose between the flood of AI tools available now before you hand over the decision to software that has no idea what a qualified lead looks like for your business.
What AI is good for is generating the variants faster, drafting five headline options in the time it used to take to write one, or spotting patterns in session recordings across hundreds of visits that a human would take a day to review. It’s a poor judge of which variant made you money, because it usually only sees the click, not the close.
Judging design across traffic sources
The same landing page can behave differently depending on where the visitor came from, and this is where a lot of people get their judgement wrong. A page that wins with LinkedIn traffic, where visitors have already seen your face and your posts and arrive with some trust built in, can lose badly with cold search traffic that needs far more convincing on the page itself. I see this constantly with founders who build their whole authority through an active LinkedIn profile and then wonder why the same page underperforms when they run paid ads to strangers who’ve never heard of them. Judge each traffic source against its own version if you can afford to build one, and never assume a design win in one channel will repeat in another.
The same logic applies if you’re pulling traffic from somewhere visual and slow-burn like Pinterest, where people are browsing rather than buying in the moment; a page built for someone learning Pinterest marketing from scratch and sending that traffic to a landing page needs more warmth and story on the page, because the visitor hasn’t shown buying intent the way a Google search visitor has.
The checklist I use before I trust a result
- Did both versions run for the same calendar period, including weekends?
- Did I hit at least 1,000 visitors or 100 conversions per version?
- Did I check revenue or close rate, not just the top-of-funnel conversion rate?
- Was the traffic source identical for both versions?
- Did anything external change mid-test, a price rise, a press mention, a seasonal spike?
- Would I be comfortable explaining this result to someone who paid for the traffic?
If I can’t answer yes to most of those, I don’t call the test. I let it run longer instead of picking a winner to make a Monday meeting go smoothly.
Frequently asked questions
How long should I run a landing page A/B test before judging the winner?
Run it for at least two to four weeks and until each version has roughly 1,000 visitors or 100 conversions, whichever arrives first. Calling a winner earlier than that means you’re mostly reading random noise rather than a real pattern in behaviour.
What’s the biggest mistake people make when judging landing page design?
Judging on conversion rate alone rather than on what happens after the click, such as close rate, refund rate, or average order value. A page can win the click and still lose the sale, so track the number that connects to actual revenue.
Can I trust design opinions from my team over test data?
No, and this is where most disagreements go wrong. Team opinions about which design “looks better” are shaped by taste and internal habit, not by how a stranger with no context reacts to the page in the first eight seconds. Test it with real traffic instead of settling it by vote.
Do I need enough traffic to run a proper split test?
If you’re getting fewer than a few hundred visitors a week, a formal split test will take too long to be useful. In that case, judge design through session recordings, heatmaps, and direct conversations with buyers instead of waiting months for statistical significance.
Prefer to hand this over? Start here: become a web design contributor.