Share article

Kolik odpovědí stačí (zdroj: chat GPT)
Kolik odpovědí stačí (zdroj: chat GPT)

Company X launches a quarterly satisfaction survey. NPS (Net Promoter Score, a loyalty metric calculated as the difference between the share of promoters and detractors) falls from 42 to 37. Someone at the meeting says the word “drop”. The head of customer care is tasked with finding out what went wrong. An internal investigation kicks off, maybe even a process change.

The problem is that there may have been no drop at all.

If only 60 people filled in the survey, a five-point difference on the NPS scale is very likely just statistical noise. A random fluctuation. Companies routinely change processes, evaluate teams and present results to leadership based on numbers like these. The data is consistent on this point: the smaller the sample, the bigger the spread of results you can expect purely by chance. And this is exactly where statistical significance comes in.

What statistical significance actually means

Statistical significance answers one specific question: is the observed difference between two measurements large enough that it can’t simply be explained by who happened to respond?

If a company asked every single customer, it would get an exact figure. In practice, though, it only asks a fraction, a sample. And every sample, even if actual customer satisfaction hasn’t changed at all, will produce a slightly different number. Simply because different people answered. Statistics describe how big this natural spread is. Significance then tells you whether the difference you’re seeing is bigger than that spread could produce on its own.

The standard threshold across both market research and academia is a 95% confidence level. This means that if the same survey were repeated a hundred times under identical conditions, the result would fall within the calculated range in 95 of those hundred cases. But there’s a key condition that often gets overlooked: this certainty is tied to the sample size, not to how much the result matters to the company.

Confidence intervals and margin of error

In practice, significance is calculated using the margin of error (MOE) – the range within which the true value lies, at a given confidence level. For a simple proportion, say the percentage of satisfied customers, at 95% confidence the rough relationship is MOE ≈ 1/√n, where n is the sample size.

In numbers, that looks like this. With a sample of 100 respondents, the margin of error is roughly ±10 percentage points. With 400 respondents, it drops to around ±5 points. With 1,000 respondents, around ±3 points.

Notice the pattern. Quadruple the sample, and the error only halves. Accuracy grows with the square root of the number of responses, not in a straight line with the number itself. The law of diminishing returns applies here too.

For NPS, the situation is even more sensitive than for a simple percentage. NPS is calculated as the difference between two proportions, promoters minus detractors, and the difference between two random variables always has more variance than either one on its own. In practice, this means the confidence interval around an NPS score tends to be wider than intuition would suggest. For smaller samples, it can easily be ±8 to ±10 points even with a hundred responses.

Fred Reichheld, who first published the NPS methodology in Harvard Business Review in 2003 (“The One Number You Need to Grow”), designed it as a simple, actionable loyalty indicator, not a lab-grade precise measurement. That simplicity comes at a price, and the price is wider uncertainty ranges.

When a difference is real, and when it’s just noise

The practical rule is simple: if the confidence intervals of two measurements overlap, you can’t say with certainty that anything has changed. The difference might be real. But statistically, it can’t be distinguished from chance.

Here’s a worked example. A company measures NPS on a sample of 80 respondents and gets a score of 45. A quarter later, on another 80 respondents, it measures 38. A seven-point difference looks like a clear signal. But with a sample of 80 people, the NPS margin of error typically sits around ±10 to ±12 points. So the confidence intervals of both measurements very likely overlap. A seven-point difference falls squarely within the range that can be explained simply by a different group of people answering, with a different random mix of experiences.

So the real question isn’t whether the score dropped. It’s whether the drop is bigger than what the sample size alone could account for. Without knowing the sample size and its margin of error, that question can’t be answered.

So how many responses are actually enough

This is where a finding comes up that surprises a lot of managers: above a certain threshold, the sample size you need doesn’t depend on the size of your total customer base. A company with 5,000 customers and a company with 500,000 customers need practically the same sample size for the same level of accuracy.

This follows from the central limit theorem, one of the fundamental principles of statistics. As a population grows beyond a certain point, each additional respondent adds less and less to the accuracy of the estimate, quickly approaching zero.

In market research practice, a sample of around 384 respondents is commonly cited as the standard reference point, for 95% confidence and a margin of error of ±5 percentage points. This holds true across both large and smaller populations. For rougher, faster tracking, with a margin of error of around ±7 to ±8 points, a sample approaching 150 to 200 respondents may be enough. Below that, with samples in the tens of responses, you need to reckon with a margin of error exceeding ten points. At that scale, comparing small month-to-month or quarter-to-quarter score differences stops making sense.

But size isn’t everything. A sample of 400 responses, of which 350 come from a single customer segment, represents the whole customer base no better than a sample ten times smaller. Sample size deals with random error. Lack of representativeness creates systematic error (bias), and no amount of extra responses will fix that on its own.

Practical rules for not misreading data that means nothing

Before deciding on a strategy change at the next meeting based on a quarterly shift in a CX metric, it’s worth running through a few checks:

How many responses make up the scores being compared, tens, or hundreds? For samples under a hundred respondents, any month-to-month difference needs to be treated with serious caution.

Do the confidence intervals of the two measurements overlap? If so, the difference can’t be called proven, however compelling the story around it sounds.

Is the rise or fall consistent across several periods, or is it a one-off blip between two points? A trend built from three or more consecutive measurements carries far more information than a comparison of just two points.

Is the sample representative of the customer base structure, or is it skewed by who happened to have the time or reason to respond?

In CX, data is often treated as objective fact simply because it comes in numerical form. But a number on its own says nothing about how much uncertainty is hiding behind it. So the real question for any organisation working with CX metrics isn’t what number it measured. It’s whether it has enough data to actually trust that number.

Full magazine experience. Zero desk required.

xpulse_app_store
Dan Bauer
Dan je náš investigativní AI novinář, využívající všemožné zdroje a AI k tomu, aby Vám články o CX poskytl v co možná nejvyšší kvalitě. Nikdy ho ještě nikdo neviděl, i když by každý chtěl.

Full magazine experience. Zero desk required.

xpulse_app_store
Dan Bauer
Dan je náš investigativní AI novinář, využívající všemožné zdroje a AI k tomu, aby Vám články o CX poskytl v co možná nejvyšší kvalitě. Nikdy ho ještě nikdo neviděl, i když by každý chtěl.