Ready to send better email?
Brew is the AI-native ESP. Define your brand once, ship on-brand email in minutes.
Brew is the AI-native ESP. Define your brand once, ship on-brand email in minutes.
Master email A/B testing with proven strategies. Learn what to test, how to run valid tests, and how to interpret results for continuous improvement.
Email A/B testing transforms opinions into data. Instead of guessing what works, you let your audience tell you through their actions.
This guide covers everything from basic split testing to advanced multivariate strategies. By the end, you'll know how to run tests that drive real improvements.
A/B testing (split testing) sends two versions of an email to see which performs better. You change one element, measure results, and apply the winner.
Simple example:
Subject Lines
Send Time
From Name
Email Content
CTA (Call to Action)
Images
Change only one element per test. If you test subject line AND send time, you won't know which caused the difference.
Good test: Subject A vs. Subject B (same everything else) Bad test: Subject A at 9am vs. Subject B at 2pm
Statistical significance requires adequate recipients:
| Total List Size | Per Variant | Confidence Level |
|---|---|---|
| 1,000 | 300+ | Moderate |
| 5,000 | 500+ | Good |
| 10,000+ | 1,000+ | High |
Rule of thumb: 1,000+ per variant for reliable results.
What defines "winning"?
| Test Type | Primary Metric |
|---|---|
| Subject line | Open rate |
| Content | Click rate |
| CTA | Click rate or conversion |
| Send time | Open rate + click rate |
Treat subject-line opens as noisy after Apple Mail Privacy Protection prefetches images. Confirm with clicks when the change is in the body.
Is the difference statistically significant?
Quick rule: If winner is 5%+ better with 1,000+ recipients per variant, it's likely significant.
For precise analysis, use a statistical significance calculator. Pair the result with metrics that still matter after prefetch clients inflate opens.
1. Personalization
2. Curiosity vs. Clarity
3. Numbers
4. Emoji
5. Length
6. Question vs. Statement
| Version A | Version B |
|---|---|
| "Buy Now" | "Get Yours" |
| "Learn More" | "See How It Works" |
| "Start Free Trial" | "Try Free for 14 Days" |
| "Download" | "Get Your Free Copy" |
One change at a time. Multiple changes = meaningless results.
At least 2-4 hours for opens, 24 hours for clicks. B2B may need 48 hours.
Statistical significance requires sample size. Small lists mean less certainty.
Create a testing log. Patterns emerge over time.
Focus on high-impact elements first. Button shade differences rarely matter.
Some testing beats no testing. Start somewhere.
Track improvements over time:
| Metric | Baseline | After 3 Months | Improvement |
|---|---|---|---|
| Open rate | 20% | 25% | +25% |
| Click rate | 2% | 3% | +50% |
| Conversion | 1% | 1.5% | +50% |
Small improvements compound. 10% better opens × 10% better clicks = 21% more conversions.
Test multiple variables simultaneously:
Requirements: Large lists and statistical software.
Keep a percentage that never gets optimization. Compare long-term performance to measure cumulative impact.
Build improvements incrementally:
Brew makes testing easier:
Instead of spending hours writing test variants, describe what you want and AI creates options.
Ready to optimize your emails? Try Brew free and create test variants in seconds with AI.
A winning variant still has to reach the inbox. Email deliverability and Gmail's sender requirements do not pause for a test. Commercial tests still need a working unsubscribe under CAN-SPAM.
To contact us about this article, write Brew support.
Email A/B testing is sending two versions to a sample and keeping the one that wins on a named metric.
A valid test is one change, a large enough sample, and a metric that is not a prefetch pixel.
A winner still has to reach the inbox.
For example, a subject-line test should not also change the hero image.
| Move | Why it matters |
|---|---|
| Change one thing | Mixed tests cannot tell you what worked |
| Confirm with clicks | Opens moved after Mail Privacy Protection |
US commercial mail still follows the FTC CAN-SPAM guide and 16 CFR 316.3. Bulk sending still has to meet Gmail's sender requirements. To contact us about this article, use the support link above.
A/B testing sends two versions of an email to see which performs better. You change one element, measure the result, and keep the winner. A simple case is two subject lines to equal halves of the list, then you count opens. The same method works for send time, from name, body, and CTA once you pick the matching metric.
Start with subject lines, then send time, then from name. Those change opens and trust with the least production work. After that, test length, tone, CTA text and placement, and images. Preview text and footer tweaks are lower impact. One change per test, or you cannot name the cause.
Test one variable. Send both variants at the same time. Size the groups so each variant has enough recipients. Define the win metric before you send. Wait at least a few hours for opens and a day for clicks. Do not call a winner from a peek. Document the result and plan the next test.
Apple Mail Privacy Protection prefetches images, including tracking pixels, when the message arrives. That can count as an open whether or not anyone read the mail. Use opens as a directional signal. Confirm with clicks when the test is about the body, and read email metrics that matter.
Testing two things at once, stopping after a couple of hours, declaring a winner on a tiny list, skipping the log, and arguing over button shade are the usual failures. Some testing still beats none. Keep the calendar simple: subjects weekly, CTAs monthly, design quarterly.