GrowthStack Advisory / Free tools / Cold email A/B test calculator
Free tool
Cold email A/B test calculator
The short answer
Variant A got 15 replies from 500 sends. Variant B got 22 from 500. Do you switch everyone to B? Almost every sales team says yes, and almost every sales team is wrong, because at the reply rates typical of cold outbound the eye cannot separate a real improvement from ordinary random variation. This calculator answers two questions from your own counts: whether the difference you are looking at is real yet, and how many more sends it would take to know. It uses a pooled two-proportion z test for the verdict and Wilson score intervals for context. No reply rate is assumed anywhere, and nothing is stored.
- Low base rates need large samples. The smaller the reply rate, the more volume it takes to detect a difference
- Decide the sample size first, then read the result once
- Stopping early when one variant leads is the single most common way teams fool themselves
- Overlapping confidence intervals do not mean the difference is not real, which is why the verdict here does not use them
Your outbound readiness scorecard asks whether messaging is tested with a defined sample size before it scales across the list. This is the instrument for computing that number.
Interactive tool
Is the difference real, or is it noise?
Enter what each variant actually sent and received. Counts, not percentages. Nothing is sent anywhere and nothing is stored.
The starting values are counts from a hypothetical test, not a claim about typical reply rates. Nothing on this page asserts what your reply rate should be. Replace all four counts with your own.
Why raw reply rates mislead at outbound volumes
The default figures above show the trap clearly. Fifteen replies from five hundred is 3.0%. Twenty two from five hundred is 4.4%. That reads as a 47% improvement, and it is the kind of number that gets announced in a Monday meeting. Run it through the test and the difference is comfortably inside what chance produces at those volumes.
The reason is the base rate. When something happens 3% of the time, the count in any given sample bounces around a lot in relative terms. Seven extra replies out of five hundred is not a signal, it is a Tuesday. The same seven-reply gap across twenty thousand sends per side would be conclusive. Volume is what converts a difference into evidence.
This is also why testing works better on the parts of the funnel with higher base rates. A reply-to-meeting rate of 30% reaches significance far faster than a 3% reply rate does, because the underlying number is larger. If you want faster answers, test the stages where more people convert.
The mistake that costs more than the maths
Most teams do not get this wrong through arithmetic. They get it wrong through timing. The sequence runs, someone checks the dashboard each morning, and the test is declared over the first day one variant is comfortably ahead.
That practice has a name, peeking, and it quietly destroys the guarantee the test is supposed to give you. A 5% false positive rate assumes you look once, at a point decided in advance. If you look every day and stop whenever you like what you see, you will eventually find a winner in two variants that are identical, because random variation crosses the line sooner or later.
The fix costs nothing. Compute the required sends per variant before the test starts, run to that number, then read the result one time. If the answer is "not yet", the honest options are to keep running or to accept that the difference is too small to matter, not to keep refreshing until it looks better.
The method, stated plainly
The verdict uses a pooled two-proportion z test. Both variants are combined to estimate a shared reply rate, that estimate produces a standard error for the difference, and the observed difference is expressed as a number of standard errors. The p value is the probability of seeing a gap at least this large if the two variants were genuinely identical. Below 0.05 the tool reports a real difference.
The required sample size uses the standard formula for comparing two proportions at 95% confidence and 80% power. In plain terms, 95% confidence means a one in twenty chance of calling a difference real when it is not, and 80% power means that if a difference of the size you specified genuinely exists, the test will detect it four times out of five.
Each variant also gets a Wilson score interval, which is the plausible range for its true reply rate given the sample. These are displayed for context only. The verdict deliberately does not come from whether the two intervals overlap. That comparison is a well known error and is too strict: two 95% intervals can overlap while the difference between them is significant at p below 0.05.
What this calculator does not tell you
- Whether the reply is any good. Replies are counted, not qualified. A subject line that lifts replies while attracting the wrong people is a worse variant with a better number. Test on meetings held once you have the volume to do it.
- Whether the two groups were comparable. The test assumes each variant went to a random split of the same list. If A went to enterprise accounts and B went to mid-market, the tool will still return a p value and it will be meaningless.
- Whether anything else changed. Deliverability drift, a new sending domain, a holiday week or a competitor's announcement all move reply rates. Running the two variants in the same window is the only way to keep the comparison honest.
- How many variants you are testing at once. Comparing five variants pairwise multiplies the chance of a false positive. This tool compares exactly two.
Related reading
This tool exists because of one line in the outbound readiness scorecard, which asks whether messaging is tested with a defined sample size before it scales. For how the messages being tested should be built in the first place, see ICP, persona and messaging and outbound sequence structure.
If you need to size the programme rather than the test, the pipeline capacity calculator works backwards from a revenue target to the sends and headcount required. For how engagements are structured, see pricing.
