Are you measuring impact, or just hoping?
Every company invests in initiatives to sell more: a new training program, a process change, a new tool. And almost every company evaluates those initiatives the same way: it looks at sales afterwards, sees they went up, and concludes it worked.
The problem is that "it went up after" is not "it went up because". Sales could have risen for a thousand reasons: a hot season, a change in quota, a competitor that stumbled. Without a way to isolate the initiative's effect, you are not measuring impact — you are hoping the coincidence is on your side.
Experiment design is a common discipline in product and marketing, and strangely rare in sales training and performance initiatives. This article shows how to apply it, without heavy math.
1. The problem: correlation is not causation
Two things moving together does not mean one caused the other. It is the most common mistake, and the most convincing. Ice cream sales and drownings both rise in the summer, but ice cream drowns nobody: heat drives both. In your context, the reps who used the new tool most may already have been the best reps before they touched it. The tool may have caused nothing at all.
If all you look at is two numbers rising side by side, it is easy to "prove" almost anything. To claim causation you need a point of comparison: what would have happened without the initiative?
2. The solution: the controlled experiment
The answer is the good old A/B test, or controlled experiment. The idea is simple: one group receives the initiative (treatment), a comparable group does not (control), over the same period. If the treatment group rises more than the control group, you have isolated the initiative's effect from whatever is merely "market tide" — because the market hit both groups equally.
For the control group to work, both groups need to be comparable and assigned at random. That way, on average, the only systematic difference between them is the initiative. And they need to run in parallel, over the same months, so that any market movement shows up on both sides and cancels out.
3. Signal or noise: when the difference is real
Say the group with the initiative hit 85% of quota and the control hit 80%. Are those 5 points a real effect, or are they within the normal variation between any two groups? The answer depends on the noise — how much results naturally vary from person to person and from month to month.
This is where statistical significance comes in. No formula needed: it answers "what are the odds I would see a gap this large even if the initiative did nothing?". If those odds are low (the convention is below 5%), we call the result significant: hard to explain by chance alone. The more people in the test and the lower the noise, the easier it is to separate signal from noise.
An honest way to report the result is the confidence interval: instead of a bare number, a range — "a gain between 3% and 11%". If the range is narrow and entirely positive, you have solid evidence. If it crosses zero, the result is inconclusive.
4. The three decisions behind an experiment
Before you run it, you define three things:
- The metric. What exactly are you measuring? Prefer a performance metric on the same scale for everyone, such as percentage of quota attainment, rather than raw values that mix territories of very different sizes.
- The test group and the control group. Who receives the initiative and who does not. Both groups need to be comparable and assigned at random, so that the only systematic difference between them is the initiative, and they must run in parallel, over the same months, so the market cancels out.
- The duration. How long to measure before comparing. Longer windows reduce noise and stabilise the result.
5. How many people does the experiment need?
Sample size is not a guess: it falls out of three factors. Put the three together and you arrive at the number of people per group.
- Variability (the standard deviation within the groups). First, measure how much the result naturally varies, between people and from month to month. The larger the standard deviation, the more noise, and the more people you need to see the effect through it.
- The minimum effect that matters (MDE). The smallest gain that would still justify adopting the initiative. It is the most expensive lever: detecting half the effect costs roughly four times as many people.
- The rigour (confidence and power). How little you are willing to be wrong — both about declaring an effect that does not exist and about missing one that does. More rigour, more people.
You do not have to do this by hand: any sample size calculator will work it out from those three factors. The real work is defining the three honestly, before you run the pilot.
6. What if a clean control group is impossible?
Splitting the team into two groups is not always feasible. There are two fallbacks, in order of reliability.
Before/after compares the same group before and after the initiative. It is simple, but dangerous: it credits the initiative with anything that changed during the period, seasonality and market included. Only use it when there is no alternative and you can argue that nothing major changed outside.
Better than that is difference-in-differences: you compare the treatment group's change with the control group's change. Whatever the market did shows up in both and cancels out; what is left is the initiative's effect. It is the robust middle ground when you have a control group and some history.
7. A step-by-step you can apply tomorrow
- Define the performance metric, on the same scale for everyone.
- Define the test and control groups, assigned at random.
- Define the duration of the experiment.
- Estimate the noise (standard deviation) from your operation's history.
- Define the minimum effect and the rigour, and compute the sample size.
- Run in parallel, measuring both groups over the same months.
- Read the result through the confidence interval, not a standalone number.
The goal is not to turn sales into a laboratory. It is to stop making million-dollar decisions on the strength of a chart that went up. A well-designed pilot costs a few extra weeks of planning and returns something no before-and-after spreadsheet can give you: confidence that what you are about to scale actually works.
