EN
Webmail

Email A/B Testing: What to Test and How to Read the Results

Email A/B Testing: What to Test and How to Read the Results

Most email platforms make A/B testing look effortless: write two subject lines, pick a percentage of the list, and let the software send the “winner” to everyone else. The trouble is that many of these tests prove nothing. The sample is too small, the difference is random noise, or the metric is open rate, which modern privacy features have made unreliable. Teams then build their email strategy on conclusions that were never real.

This guide explains how to run email A/B testing properly: what is worth testing, how many recipients you need, which metrics to trust, and how to turn individual tests into knowledge that improves every future campaign. It is written for small marketing teams and business owners who send newsletters, promotions or automated flows and want better results without a data science department.

What Email A/B Testing Is and What It Is Not

An A/B test sends two versions of an email that differ in one element to two randomly chosen groups from the same list, then compares how each group behaves. Because the groups are random and receive the emails at the same time, a clear difference in results can be attributed to the element you changed.

It is not the same as comparing this week’s newsletter with last week’s. Different weeks have different content, different news, different weather and different moods. Those comparisons are interesting, but they cannot tell you whether a specific change caused a result.

It is also not a way to rescue a weak list or a poor offer. Testing optimises what you already send. If your list is full of inactive addresses or your emails land in spam, fix deliverability and list quality first, as described in our guide to segmenting your email list.

What to Test: The Elements That Move Results

Every element of an email can be tested, but some have far more influence than others. Start where the potential gain is largest.

Subject line and preview text

These decide whether an email is opened at all. Useful contrasts include a specific benefit versus curiosity, a question versus a statement, including a number or not, and short versus descriptive. Preview text, the snippet shown after the subject in most inboxes, is often left to chance and is an easy win.

Sender name

People decide whether to trust an email partly by who it is from. “Anna from Company” and “Company” can perform very differently, and the result depends on your audience and relationship with them.

Offer and call to action

The offer usually has more impact than any design detail: free shipping versus a percentage discount, a free guide versus a webinar, “Book a call” versus “See pricing”. One clear primary call to action generally beats several competing ones, but test it with your own audience.

Content length and layout

Some audiences respond to short, plain emails that read like a personal message; others prefer a designed newsletter with images and sections. Test the format, not just the words.

Send day and time

Timing tests are popular but tricky, because the two versions are necessarily sent at different moments. Run them over several sends before drawing conclusions, and remember that automated flows are triggered by behaviour, so timing matters less there.

Choosing the Right Metric

The metric decides what “winning” means. Choosing the wrong one is the most common reason tests mislead.

MetricGood for testingReliability todayWatch out for
Open rateSubject line, sender name, preview textLow to mediumPrivacy features that preload images inflate opens, especially on Apple devices
Click rate (clicks per delivered)Subject line, content, call to actionHighBot clicks from security scanners on some corporate addresses
Click-to-open rateContent and design, given an openMediumInherits the open-rate problem in its denominator
Conversion rate or revenue per recipientOffer, call to action, whole-email changesHigh, if tracking worksNeeds enough conversions; tracking must be tagged consistently
Unsubscribe and complaint rateFrequency, tone, aggressive offersHighLow numbers; watch as a guardrail rather than a goal

Since the introduction of mail privacy features that load images automatically, open rates count many emails that nobody actually looked at. That does not make open rate useless, but it means a subject line test judged only on opens can pick the wrong winner. Wherever possible, judge tests on clicks or conversions, and use unsubscribes and spam complaints as guardrails. Mailbox providers pay close attention to complaints; Google’s email sender guidelines set clear expectations for bulk senders.

How Many Recipients Do You Need?

This is where most small-list tests fail. The smaller the difference you want to detect, and the lower your baseline rate, the more recipients each version needs.

A rough intuition: if your click rate is around 2 percent and you want to detect a lift to 2.5 percent, you need many thousands of recipients in each group. If you only want to detect a doubling, from 2 to 4 percent, a few hundred per group may be enough. A free sample size calculator gives you the number for your own baseline and target in seconds. Run it before the test, not after.

What to do with a small list

  • Test bigger differences. Two subject lines that differ by one word will never show a reliable result on a list of 1,500. Two completely different approaches might.
  • Send the test to the whole list. Split the entire list 50/50 instead of testing on 20 percent and sending a “winner” to the rest. You lose the automatic winner feature but gain enough data to learn something.
  • Repeat the same test. Run the same contrast across several campaigns and combine the results.
  • Test in automated flows. A welcome email sent to every new subscriber accumulates recipients over months, which makes it an excellent place for tests that need volume. Our guide to email automation flows covers the welcome, abandoned cart and re-engagement sequences where this works best.

Running a Test Step by Step

  1. Write a hypothesis. “A subject line with a concrete saving will get more clicks than a curiosity subject line, because our readers are price-sensitive.” A hypothesis makes the result useful even when it loses.
  2. Change one element. If you change the subject, the image and the button at once, you will not know which one made the difference. Whole-email tests are fine, but then you are testing concepts, not elements.
  3. Choose the metric and sample size in advance. Decide what counts as a win before you see any numbers.
  4. Randomise properly. Let the platform split the list randomly. Do not send version A to one segment and version B to another.
  5. Wait long enough. Most clicks happen in the first day or two, but purchases can take longer. For conversion metrics, wait several days before deciding.
  6. Check significance. Many platforms show a confidence figure. If yours does not, use a simple significance calculator. If the result is not clear, record it as “no difference detected”, which is also a finding.
  7. Record the result. Write down the hypothesis, versions, sample size, metric, result and decision in a shared log.

Reading the Results Without Fooling Yourself

Beware of early peeking

Checking results every hour and stopping as soon as one version looks ahead greatly increases the chance of declaring a false winner. Early differences often shrink as more data arrives. Decide the stopping point in advance and stick to it.

Look at absolute numbers

“Version B increased clicks by 40 percent” sounds impressive until you see it went from 10 clicks to 14. Always look at the counts behind percentages.

Check for side effects

An aggressive subject line can win on opens while raising unsubscribes and complaints. A discount can win on conversion while lowering revenue per order. Look at guardrail metrics before rolling out a winner.

Segment results carefully

It is tempting to slice results by device, country or customer type until something looks significant. With enough slices, something always will, by chance. Treat segment differences as ideas for the next test, not as conclusions.

Test Ideas by Type of Email

Different emails have different jobs, so the most useful tests differ too. These ideas are starting points, not rules; your own audience decides.

Newsletters

The goal is usually clicks to content and long-term engagement. Test the order of stories (lead with the most useful item or the newest one), a single featured article versus a digest of several, and a personal introduction from a named person versus going straight to the links. Measure clicks per delivered email and unsubscribes over several issues, because newsletter habits form slowly.

Promotional campaigns

Here the goal is revenue. Test the framing of the same offer (“Save 20 percent” versus “Save 30 euros on orders over 150”), a deadline versus no deadline, and one hero product versus a selection. Judge on revenue per recipient, not on click rate, because a curious click is not a sale.

Welcome and onboarding emails

New subscribers are at their most attentive. Test whether the first email should deliver the promised incentive only or also introduce your best content, and whether a short plain-text message from the founder performs better than a designed welcome. Because every new subscriber receives these emails, results accumulate steadily without extra work.

Re-engagement emails

For inactive subscribers, test a direct question (“Do you still want to hear from us?”) against a strong incentive. Measure clicks and the number of people who confirm they want to stay, and remove those who do not respond. A smaller, active list protects deliverability for everyone else.

Building a Testing Programme, Not Just Tests

A single test answers a single question. A programme builds knowledge about your audience that compounds over time.

  • Keep a test log. A simple spreadsheet with date, hypothesis, versions, sample, metric, result and decision is enough. After a year it becomes your most valuable marketing document.
  • Prioritise by impact and effort. Offer and call-to-action tests usually matter more than button colour.
  • Retest important findings. Audiences change. A result from two years ago may no longer hold.
  • Share results beyond email. A subject line style that wins in email often works in ad headlines and social posts too, and vice versa.

If you want help setting up tests, tracking and reporting across campaigns and automated flows, that is part of our e-mail marketing service.

Frequently Asked Questions

What should I A/B test first in email?

Start with the elements that have the most influence: the offer and call to action, then subject line and preview text. Design details such as button colour rarely produce measurable differences on small lists.

Is open rate still a valid metric for email A/B testing?

Only partly. Privacy features that load images automatically inflate opens, so open rate is less reliable than before. Judge tests on clicks or conversions whenever you have enough volume.

How big should my test groups be?

It depends on your baseline rate and the difference you want to detect. Use a sample size calculator before sending. Small lists should test big differences or split the whole list in half.

How long should I wait before choosing a winner?

For clicks, at least a full day and preferably two. For purchases or sign-ups, wait several days. Decide the waiting period before the test starts and do not stop early because one version looks ahead.

Can I test more than two versions at once?

Yes, but each extra version needs more recipients to reach a reliable result. For most small and medium lists, two versions are the practical limit.

What if my test shows no difference?

Record it as a result. It tells you that element does not matter much for your audience, so you can focus testing effort elsewhere and choose the version that is easier to produce.

The Bottom Line

Email A/B testing works when it is treated as a small experiment, not a button in the send screen. Test elements that matter, change one thing at a time, size the sample before you send, judge results on clicks or conversions rather than inflated opens, and write every result down. Small lists can still learn a lot by testing bold differences and using automated flows to build volume. Over a year, a disciplined testing log will teach you more about your customers than any benchmark report.