A/B Testing Cold Email: What to Test First
Subject lines are the weakest first test in cold email. The handful of variables worth testing before them, and how to run a test that means something.
, 3 min read, Cold email
Key takeaways
- A subject line only decides whether a message gets opened, not whether it gets a reply, which makes it the variable with the least leverage over the metric that actually matters in cold email.
- Sender identity, personalization approach, call-to-action size, sequence length and send time all carry more weight on reply rate than the subject line and deserve to be tested first.
- A valid test at cold-email volume changes one variable at a time, holds everything else constant, and waits for at least a few hundred sends and the full close of the response window before concluding anything.
Most teams start A/B testing cold email with the subject line, because it is the easiest thing to change and the easiest to imagine mattering. It is also usually the wrong place to start, because a subject line only decides whether a message gets opened, and open is not the outcome anyone is actually paid on.
Why subject lines are the wrong first test
A subject line test measures one narrow event in a chain that has to succeed several more times before it produces a reply. Even a large, statistically real lift in open rate can produce no measurable lift in reply rate if the body of the email is what is actually holding performance back. Testing subject lines first optimizes the variable with the least leverage over the metric that pays the bills, and it consumes list volume that could have gone toward testing something with a bigger ceiling.
The variables worth testing before subject lines
| Variable | Why it has more leverage | What a real test looks like |
|---|---|---|
| Sender identity | A rep's name versus a generic company alias changes trust before the message is even read | Same content, two different sending personas |
| First-line personalization approach | Trigger-based versus generic opener changes whether the email reads as researched | Same offer, two different opening sentences |
| Call-to-action size | A low-commitment ask versus a meeting request changes who is willing to reply at all | Same body, two different closing asks |
| Sequence length and spacing | Determines how many total opportunities a prospect has to engage | Two cohorts on different cadences, measured on cumulative reply rate |
| Send time and day | Affects whether the email is seen during a decision-ready moment | Same content sent at two different times |
Each of these sits closer to the actual decision a prospect makes than the subject line does, which is why they move reply rate more reliably once tested properly.
How to actually run a valid test at cold-email volume
Cold email lists rarely have the volume that consumer marketing tests assume. Testing five variables at once across a list of 2,000 splits the sample so thin that no result will be statistically meaningful. Pick one variable, hold everything else constant, including the list segment itself, and run each variant against at least a few hundred sends before drawing a conclusion. Measure the primary outcome as reply rate across the full sequence, not open rate on a single touch, since open tracking is unreliable and the sequence, not the individual message, is the unit that actually produces a reply.
Common testing mistakes
The most common mistake is changing more than one variable between the two versions being compared, which makes it impossible to know which change caused any difference in outcome. The second is calling a test at too small a sample size, often after only a few dozen sends, which produces a result driven by noise rather than a real effect. The third is testing on a mixed list that spans multiple segments or personas, which can hide a real effect in one segment behind a flat result in another. A fourth common mistake is running a test for too short a window and missing effects that only show up once the initial send-day spike settles and later-arriving replies come in over the following one to two weeks; call results only after the sequence's full response window has closed.
Where subject lines fit once the bigger variables are settled
Subject line testing still has a place, just later in the process, once sender identity, personalization approach and call-to-action size are already working well. At that point, a subject line test refines an already-working system rather than trying to compensate for a body and offer that are not converting, and the lift it produces is real rather than a rounding error mistaken for a signal.
Frequently asked questions
- Why shouldn't I start A/B testing with the subject line?
- Because the subject line only decides whether a message gets opened, not whether it gets a reply, which makes it the variable with the least leverage over the outcome that actually matters in outbound.
- What variables should I test before the subject line in a cold email A/B test?
- Sender identity, the first-line personalization approach, the size of the call to action, sequence length and spacing, and send time all carry more weight on reply rate than the subject line.
- How many sends do I need for a reliable cold email A/B test?
- At least a few hundred sends per variant, changing only one variable at a time, and you should wait until the sequence's full response window has closed before drawing any conclusions.