GrowthStackAdvisory
Book 30 min

GrowthStack Advisory / Free tools / Cold email A/B test calculator

Free tool

Cold email A/B test calculator

The short answer

Variant A got 15 replies from 500 sends. Variant B got 22 from 500. Do you switch everyone to B? Almost every sales team says yes, and almost every sales team is wrong, because at the reply rates typical of cold outbound the eye cannot separate a real improvement from ordinary random variation. This calculator answers two questions from your own counts: whether the difference you are looking at is real yet, and how many more sends it would take to know. It uses a pooled two-proportion z test for the verdict and Wilson score intervals for context. No reply rate is assumed anywhere, and nothing is stored.

Your outbound readiness scorecard asks whether messaging is tested with a defined sample size before it scales across the list. This is the instrument for computing that number.

Shobhit Gupta, founder of GrowthStack Advisory

By Shobhit Gupta

Founder, GrowthStack Advisory. 10+ years building SDR and GTM systems at Locus, GoComet, and Landmark Group.

· Free, no signup

Interactive tool

Is the difference real, or is it noise?

Enter what each variant actually sent and received. Counts, not percentages. Nothing is sent anywhere and nothing is stored.

The starting values are counts from a hypothetical test, not a claim about typical reply rates. Nothing on this page asserts what your reply rate should be. Replace all four counts with your own.

Why raw reply rates mislead at outbound volumes

The default figures above show the trap clearly. Fifteen replies from five hundred is 3.0%. Twenty two from five hundred is 4.4%. That reads as a 47% improvement, and it is the kind of number that gets announced in a Monday meeting. Run it through the test and the difference is comfortably inside what chance produces at those volumes.

The reason is the base rate. When something happens 3% of the time, the count in any given sample bounces around a lot in relative terms. Seven extra replies out of five hundred is not a signal, it is a Tuesday. The same seven-reply gap across twenty thousand sends per side would be conclusive. Volume is what converts a difference into evidence.

This is also why testing works better on the parts of the funnel with higher base rates. A reply-to-meeting rate of 30% reaches significance far faster than a 3% reply rate does, because the underlying number is larger. If you want faster answers, test the stages where more people convert.

The mistake that costs more than the maths

Most teams do not get this wrong through arithmetic. They get it wrong through timing. The sequence runs, someone checks the dashboard each morning, and the test is declared over the first day one variant is comfortably ahead.

That practice has a name, peeking, and it quietly destroys the guarantee the test is supposed to give you. A 5% false positive rate assumes you look once, at a point decided in advance. If you look every day and stop whenever you like what you see, you will eventually find a winner in two variants that are identical, because random variation crosses the line sooner or later.

The fix costs nothing. Compute the required sends per variant before the test starts, run to that number, then read the result one time. If the answer is "not yet", the honest options are to keep running or to accept that the difference is too small to matter, not to keep refreshing until it looks better.

The method, stated plainly

The verdict uses a pooled two-proportion z test. Both variants are combined to estimate a shared reply rate, that estimate produces a standard error for the difference, and the observed difference is expressed as a number of standard errors. The p value is the probability of seeing a gap at least this large if the two variants were genuinely identical. Below 0.05 the tool reports a real difference.

The required sample size uses the standard formula for comparing two proportions at 95% confidence and 80% power. In plain terms, 95% confidence means a one in twenty chance of calling a difference real when it is not, and 80% power means that if a difference of the size you specified genuinely exists, the test will detect it four times out of five.

Each variant also gets a Wilson score interval, which is the plausible range for its true reply rate given the sample. These are displayed for context only. The verdict deliberately does not come from whether the two intervals overlap. That comparison is a well known error and is too strict: two 95% intervals can overlap while the difference between them is significant at p below 0.05.

What this calculator does not tell you

This tool exists because of one line in the outbound readiness scorecard, which asks whether messaging is tested with a defined sample size before it scales. For how the messages being tested should be built in the first place, see ICP, persona and messaging and outbound sequence structure.

If you need to size the programme rather than the test, the pipeline capacity calculator works backwards from a revenue target to the sends and headcount required. For how engagements are structured, see pricing.