[{"Value":"","Discard":false,"Expires":9999999999}]
Both camps have a problem not with the maths, but with the premise. On its own, 45 says nothing at all. It’s a good score for an insurance company, but a weak one for a telecoms provider. It’s an excellent result for B2B software (businesses selling to other businesses), but below average for consumer electronics. Without a frame of reference, 45 is just a number between zero and fifty that can be read however suits whoever’s reading it.
The evidence here is consistent: benchmarking systematically comparing your own performance against a reference point is one of the most widely used, and most widely misused, disciplines in CX (customer experience). Yet one crucial condition keeps getting overlooked: a comparison only has value if it compares like with like.
Before a company goes looking for a benchmark, it should first be clear about which question it’s actually trying to answer. There are, in fact, two fundamentally different disciplines that in practice tend to get merged into one.
The first is internal benchmarking comparing the company with itself over time. 41 last year, 45 this year. The question isn’t “are we good in absolute terms,” but “are we improving or declining, and why.” This discipline has one major advantage: it eliminates most of the variables that would otherwise undermine a comparison. Data-collection methodology, question wording, customer segment, seasonality all of that stays (when set up correctly) constant. A change in score then genuinely reflects a change in customer experience, not a change in how the company measures it.
The second discipline is external (industry) benchmarking comparing yourself against competitors or the sector as a whole. The question here is “how do we stack up against other players in the market.” This discipline is more appealing to leadership and marketing, because it provides context on a bigger scale. At the same time, it’s far more methodologically fragile, because it pulls in a whole range of factors that have nothing to do with the quality of the customer experience.
Most companies instinctively gravitate towards external benchmarking, because it lands better in a meeting. “We’re ahead of the competition” sounds better than “we improved by four points since last year.” Yet from a CX management perspective, it’s exactly the other way round: the internal trend is the more reliable and useful indicator, while the external benchmark only has value as a supplement, not as the headline metric.
This is illustrated well by recent trends in data from Forrester, a research and consulting firm that publishes an annual CX Index rating the quality of customer experience across hundreds of brands, industries and countries. According to Forrester’s 2025 rankings, scores fell for 21% of the brands rated globally, improved for just 6%, and stayed unchanged for 73%. In the US it was even more pronounced: for the second year running, ratings worsened for a quarter of brands, while only 7% improved, and among those that declined, the average drop was four points. This continues a four-year downward trend that began after scores peaked in 2021, when the overall CX Index score reached 72 points. Only the 2026 US data showed the first year-on-year improvement since 2021.
What does this mean for benchmarking? If a company benchmarked itself against the industry average in 2024, and that average has been falling ever since, then “being at the industry average” in 2026 means being worse off than the company itself was two years earlier even if its own score has, on paper, stayed the same. An external benchmark without an internal trend can neatly hide this decline. The company reassures itself that it’s “in line with the industry,” while the whole industry slides downward and takes it along.
The temptation to compare yourself with whoever happens to have a similar number to hand is strong. A manager reads that Apple, or some other iconic brand, has an NPS above 60, and starts asking why their own company isn’t hitting similar figures. The question is flawed on two levels.
The first problem is the industry itself. Bain & Company, the consultancy that co-created the NPS metric together with Fred Reichheld, runs a benchmarking platform called NPS Prism and states openly on its own site that scores vary substantially between industries, and that a leading score in one sector can be below average in another. Grocery retail follows a different logic of customer relationship than telecoms or banking. According to NPS Prism data, in 2026 grocery chains (based on responses from over 50,000 consumers across more than 40 brands) achieved an average relational NPS of 34. Comparing that figure with a software company or an insurer makes no sense, because a customer has a completely different type of emotional relationship with their weekly grocery shop than with a life insurance provider.
The second problem is geography and culture. A seemingly identical question (“on a scale of 0 to 10, how likely are you to recommend us”) doesn’t produce the same response behaviour in every country. Academic research published in the Journal of International Marketing (Van Herk, Poortinga and Verhallen, 2004, replicated by Harzing in 2006) found that respondents in Southern Europe (Italy, Spain, Greece) show a markedly stronger tendency towards extreme and agreeable answers than respondents in Northern Europe (the UK, Germany, France) even on entirely identical scales. In research methodology, this phenomenon is known as cultural response bias, and it includes at least two components: the tendency to agree regardless of the question’s content (so-called acquiescence), and the tendency to pick extreme points on the scale. A company comparing the NPS of its Prague branch with its Milan branch is therefore partly measuring not a difference in customer experience quality, but a difference in how people in that culture generally respond to surveys.
This doesn’t mean international comparisons are pointless. It means they require correction, or at least an awareness that raw figures between countries aren’t directly comparable without further context.
A third, less discussed problem lies in the methodology of the benchmarking sources themselves. In preparing this article, data from several commercial providers of industry NPS benchmarks were compared, and the spread of figures for the same industry varied by tens of points between sources. The reason is simple: each provider collects data using a different methodology, over a different period, on a different sample of companies, and often with different question wording. This isn’t proof that benchmarks are useless it’s proof that a benchmark is only ever as good as its methodology, and that blindly adopting a single figure from a single source without checking how it was derived is risky.
So the real question isn’t “what’s the industry score.” The question is: who is my relevant reference group, and where can I get data I can trust.
A relevant reference group usually needs to meet four conditions at once. The same industry, ideally at sub-category level, not just a broad sector (retail banking behaves differently from investment banking). The same business model B2B (selling to other businesses) shouldn’t be compared with B2C (selling to end consumers), because the dynamics of the relationship, purchase frequency and drivers of loyalty all differ. A comparable geographic and cultural context, ideally the same market, or at least one with a known, corrected difference in response behaviour. And finally, a comparable data-collection methodology the same or at least similar question wording, the same type of scale, a similar point in the customer journey at which the survey is sent.
As for the sources themselves, there are several types worth considering. Industry associations and chambers of commerce often collect aggregated data from their members and can be a surprisingly good, if less well-known, source, because the methodology is consistent across the whole sector. Commercial benchmarking platforms such as Bain’s NPS Prism or Forrester’s CX Index offer sophisticated methodology and large samples, but they come at a cost, and their data primarily reflects large, often US or Western European markets something smaller, local companies need to bear in mind. Running your own competitive benchmark directly comparing yourself against two or three specific competitors using the same questionnaire sent to a comparable sample of customers is methodologically the cleanest option, but it takes time and often requires working with a research agency. And finally, mystery shopping or auditing a competitor’s customer journey provides qualitative context that numbers alone can’t capture.
The practical recommendation is simple: the smaller and more specific the reference group, the higher its informative value but the lower the availability of data. Most companies therefore need to combine several sources and be transparent with leadership about the degree of uncertainty they’re working with. Presenting an external benchmark as a precise figure, when it’s actually a rough estimate from a heterogeneous sample, is more dangerous from a management standpoint than not having a benchmark at all.
The last, and perhaps most important, question isn’t about methodology, but about how the benchmark is used. A benchmark can work in two opposing ways: either as a tool that shows where you need to move to, or as an alibi that justifies standing still.
The phrase “we’re at the industry level” is a dangerous one from a strategic management perspective, because it quietly assumes that the industry level is an acceptable target. As the Forrester CX Index data shows, the industry level has been falling in most sectors in recent years. Being at the industry average therefore often means sharing in a trend of worsening experience along with the rest of the market, rather than maintaining the status quo.
A functional approach uses the benchmark the other way round: the external comparison serves to identify the gap and set an ambition, while the internal trend serves to measure genuine progress towards that ambition. Forrester’s own commentary on the data explicitly states that even a small improvement in customer experience quality can reduce customer churn and increase share of spend. This is the key point: the value of a benchmark isn’t in finding out where you rank in a league table, but in helping you decide how many points of improvement are realistically achievable, and how much of that improvement will actually show up in the business.
In practice, this means running two parallel metrics. Your own trend, tracked with a consistent methodology on a quarterly or half-yearly basis, is the primary indicator of whether CX initiatives are succeeding or failing. The external benchmark, refreshed once a year from a credible source, acts as a corrective lens, telling you whether your pace of improvement is sufficient given what the rest of the market is doing.
The number 45 hasn’t moved since the start of this article. It’s still 45. But the question of whether that’s good or bad now has a clearer answer: it depends on whether that 45 sits in an industry where scores typically hover around 30, or one where 60 is the standard. It depends on whether the company scored 38 or 52 last year. And it depends on whether the whole industry is currently declining or growing.
Benchmarking in CX isn’t about obtaining a single magic number that can be wheeled out in a meeting as proof of success. It’s about building the frame of reference within which a number is even worth interpreting. Companies that ignore this framework and settle for an isolated score ultimately aren’t answering the question of whether they’re good. They’re only answering the question of whether they like the number.
]]>The problem is that there may have been no drop at all.
If only 60 people filled in the survey, a five-point difference on the NPS scale is very likely just statistical noise. A random fluctuation. Companies routinely change processes, evaluate teams and present results to leadership based on numbers like these. The data is consistent on this point: the smaller the sample, the bigger the spread of results you can expect purely by chance. And this is exactly where statistical significance comes in.
Statistical significance answers one specific question: is the observed difference between two measurements large enough that it can’t simply be explained by who happened to respond?
If a company asked every single customer, it would get an exact figure. In practice, though, it only asks a fraction, a sample. And every sample, even if actual customer satisfaction hasn’t changed at all, will produce a slightly different number. Simply because different people answered. Statistics describe how big this natural spread is. Significance then tells you whether the difference you’re seeing is bigger than that spread could produce on its own.
The standard threshold across both market research and academia is a 95% confidence level. This means that if the same survey were repeated a hundred times under identical conditions, the result would fall within the calculated range in 95 of those hundred cases. But there’s a key condition that often gets overlooked: this certainty is tied to the sample size, not to how much the result matters to the company.
In practice, significance is calculated using the margin of error (MOE) – the range within which the true value lies, at a given confidence level. For a simple proportion, say the percentage of satisfied customers, at 95% confidence the rough relationship is MOE ≈ 1/√n, where n is the sample size.
In numbers, that looks like this. With a sample of 100 respondents, the margin of error is roughly ±10 percentage points. With 400 respondents, it drops to around ±5 points. With 1,000 respondents, around ±3 points.
Notice the pattern. Quadruple the sample, and the error only halves. Accuracy grows with the square root of the number of responses, not in a straight line with the number itself. The law of diminishing returns applies here too.
For NPS, the situation is even more sensitive than for a simple percentage. NPS is calculated as the difference between two proportions, promoters minus detractors, and the difference between two random variables always has more variance than either one on its own. In practice, this means the confidence interval around an NPS score tends to be wider than intuition would suggest. For smaller samples, it can easily be ±8 to ±10 points even with a hundred responses.
Fred Reichheld, who first published the NPS methodology in Harvard Business Review in 2003 (“The One Number You Need to Grow”), designed it as a simple, actionable loyalty indicator, not a lab-grade precise measurement. That simplicity comes at a price, and the price is wider uncertainty ranges.
The practical rule is simple: if the confidence intervals of two measurements overlap, you can’t say with certainty that anything has changed. The difference might be real. But statistically, it can’t be distinguished from chance.
Here’s a worked example. A company measures NPS on a sample of 80 respondents and gets a score of 45. A quarter later, on another 80 respondents, it measures 38. A seven-point difference looks like a clear signal. But with a sample of 80 people, the NPS margin of error typically sits around ±10 to ±12 points. So the confidence intervals of both measurements very likely overlap. A seven-point difference falls squarely within the range that can be explained simply by a different group of people answering, with a different random mix of experiences.
So the real question isn’t whether the score dropped. It’s whether the drop is bigger than what the sample size alone could account for. Without knowing the sample size and its margin of error, that question can’t be answered.
This is where a finding comes up that surprises a lot of managers: above a certain threshold, the sample size you need doesn’t depend on the size of your total customer base. A company with 5,000 customers and a company with 500,000 customers need practically the same sample size for the same level of accuracy.
This follows from the central limit theorem, one of the fundamental principles of statistics. As a population grows beyond a certain point, each additional respondent adds less and less to the accuracy of the estimate, quickly approaching zero.
In market research practice, a sample of around 384 respondents is commonly cited as the standard reference point, for 95% confidence and a margin of error of ±5 percentage points. This holds true across both large and smaller populations. For rougher, faster tracking, with a margin of error of around ±7 to ±8 points, a sample approaching 150 to 200 respondents may be enough. Below that, with samples in the tens of responses, you need to reckon with a margin of error exceeding ten points. At that scale, comparing small month-to-month or quarter-to-quarter score differences stops making sense.
But size isn’t everything. A sample of 400 responses, of which 350 come from a single customer segment, represents the whole customer base no better than a sample ten times smaller. Sample size deals with random error. Lack of representativeness creates systematic error (bias), and no amount of extra responses will fix that on its own.
Before deciding on a strategy change at the next meeting based on a quarterly shift in a CX metric, it’s worth running through a few checks:
How many responses make up the scores being compared, tens, or hundreds? For samples under a hundred respondents, any month-to-month difference needs to be treated with serious caution.
Do the confidence intervals of the two measurements overlap? If so, the difference can’t be called proven, however compelling the story around it sounds.
Is the rise or fall consistent across several periods, or is it a one-off blip between two points? A trend built from three or more consecutive measurements carries far more information than a comparison of just two points.
Is the sample representative of the customer base structure, or is it skewed by who happened to have the time or reason to respond?
In CX, data is often treated as objective fact simply because it comes in numerical form. But a number on its own says nothing about how much uncertainty is hiding behind it. So the real question for any organisation working with CX metrics isn’t what number it measured. It’s whether it has enough data to actually trust that number.
]]>In the two decades since it was created, Net Promoter Score (NPS) has earned its place as the crown metric of CX (customer experience). It’s simple, comparable across companies, and investors love it. But precisely because it’s so universal, it’s also a fairly blunt instrument. It measures how a customer feels about the brand as a whole, not what just happened to them. This is where CSAT (Customer Satisfaction Score) comes in: a metric that’s less impressive on a slide for the board, but in many situations, far more precise.
CSAT asks one specific thing, at one specific moment: “How satisfied were you with [the purchase / the support call / the delivery]?” Usually on a scale of 1–5 or 1–10, sometimes using emojis or star ratings. It’s what’s known as a transactional metric, measuring satisfaction with a single interaction rather than the overall relationship with the company.
This is a crucial difference from NPS, which is a relational metric, asking how favourably a customer feels towards the company as a whole, regardless of what happened yesterday on a support chat. So if you want to know whether your new onboarding process is working, or whether the latest change to your complaints form has annoyed customers more than before, NPS won’t help you. CSAT will, because it’s sensitive to change and can be measured immediately after an interaction, while the experience is still fresh.
I’ll admit I long underestimated how much this distinction matters. CSAT is most useful exactly where you need fast feedback on a specific step of the customer journey: after a purchase, after a support ticket is resolved, after onboarding is completed, after a product installation. Anywhere you’re asking “did this work?” rather than “do you like us?”
There’s a third metric worth mentioning: CES (Customer Effort Score), which asks how easy or difficult it was for the customer to get something resolved. A typical question reads: “The company made it easy for me to handle my issue,” with an agreement-scale response.
CES was born in 2010, when Matthew Dixon, Karen Freeman and Nicholas Toman of the research firm CEB (now part of Gartner) published a Harvard Business Review article with the provocative title “Stop Trying to Delight Your Customers.” Their analysis of almost 97,000 customers revealed something that turned conventional wisdom on its head at the time: there was practically no difference in loyalty between customers whose expectations were exceeded and those whose expectations were merely met. What actually predicted loyalty was effort: 96% of customers who had a high-effort interaction became disloyal, compared with just 9% of those with a low-effort one. Later Gartner research found that CES is roughly 1.8 times better at predicting loyalty than CSAT, and about twice as good as NPS, specifically in the context of service and support interactions.
So when should you reach for which?
NPS makes sense as a periodic temperature check on the relationship, quarterly or twice a year, when you want to know how overall loyalty is developing and whether your CX initiatives are working in the long run. CSAT belongs with specific interactions, where you want fast, sensitive feedback on quality, and ideally where satisfaction can be influenced in the short term (a new process, retrained agents, a product change). CES belongs wherever ease is the topic, primarily in customer support and service processes, where you want to know how much effort the whole thing cost the customer.
The important thing is not to try to track all three metrics with equal intensity everywhere. Realistically, most CX teams don’t have the capacity for that, and three half-heartedly tracked metrics are worse than one you actually do properly. A proven approach is to treat CSAT and CES as operational metrics at individual touchpoints, with NPS as the strategic compass sitting above it all.
This is where I want to be a bit pedantic, because this is exactly where most of the confusion around CSAT comes from. The standard formula is:
CSAT (%) = (number of satisfied responses ÷ total number of responses) × 100
A “satisfied response” doesn’t mean the average of the scale, though, it means what’s called a top-box (or top-two-box) approach: on a 1–5 scale, you count responses of 4 and 5; on a 1–10 scale, usually 7–10. So if 78 out of 100 respondents answer with a 4 or 5, your CSAT is 78%.
This has one important implication: CSAT isn’t an average mood, it’s the proportion of people who cleared a certain satisfaction bar. Two companies could theoretically have the same “average” on the scale, yet very different CSAT scores, depending on how the responses are distributed.
As for the numbers themselves, available benchmarks put the average CSAT across industries at around 78%. A score above 80% is generally considered strong; below 70% is a signal that something needs addressing immediately. Industries vary significantly, though: consulting services sit at around 84% according to the data, while e-commerce and retail sit around 82%. So before judging your own number as “good” or “bad,” you need to know your industry’s benchmark, not some generic figure from the internet.
This is the part I’d want anyone working with CSAT to read twice, because it’s surprisingly easy to do a lot of damage here even with the best of intentions.
The first problem is non-response bias. Satisfied customers respond to surveys more readily than dissatisfied ones, because frustrated customers often leave (or stop being customers) before the survey even reaches them, or they simply ignore it, since they don’t have the energy to fill in another form after something has annoyed them. The result is a systematically inflated CSAT that doesn’t reflect the reality of the whole customer base, just the portion that still bothered to respond.
The second problem is central tendency: people avoid extremes on a scale and gravitate towards the middle. On a 1–5 scale, most responses fall into the 3–5 range, because the bottom two options are psychologically tied to a strongly negative rating and require the respondent to actively choose something unpleasant.
The third mistake relates to what happens to the data after it’s collected: many teams only look at the extremes, “very satisfied” versus “very dissatisfied”, and ignore neutral responses, even though the neutral group is often the largest one. A neutral customer isn’t the same as a satisfied customer who’s simply not complaining, it’s a segment still waiting to be swayed one way or the other.
The fourth mistake is timing. Asking about satisfaction too long after an interaction means the customer is responding based on what they remember or how they feel in general, rather than what actually happened. A survey sent a week after a ticket is resolved measures something different from one sent five minutes after the call.
And the last, perhaps most insidious mistake: treating CSAT as a universal success metric that can be compared across completely different touchpoints without segmentation. CSAT after a complaint and CSAT after a routine purchase measure fundamentally different situations with different emotional weight, and merging them into a single dashboard number leads to decisions with no grounding in reality.
NPS is a great compass for overall direction, but it’s blind to detail. CES is the best predictor of loyalty where ease of resolving a problem is what matters. And CSAT, when used honestly and with the right context, is a tool that can pinpoint exactly where in the customer journey something is breaking down or, just as often, going right. It’s not as flashy a number as NPS, and it doesn’t sound as clever as CES in a board presentation. But for the everyday improvement of specific processes, it’s often exactly the tool you need.
]]>In CX, this law plays out with remarkable precision and with consequences that can be fatal to a business, even when they’re completely invisible at first glance.
Imagine a company whose NPS climbs every quarter. Leadership is pleased. Bonuses are paid. Then the annual data arrives, and it turns out that churn has stayed the same or got worse.
How is that possible?
The answer has nothing to do with customers. It lies in a reward system that has quietly learned to optimise the metric rather than the experience. Contact centre employees whose bonuses depend on NPS will naturally start doing things that lift the score without necessarily improving the experience. They call customers before sending a survey. They ask for a high rating. They time the survey for a moment when the customer is in a good mood. Or they skip it altogether if the interaction went badly.
Ipsos has named this behaviour survey gaming the manipulation of surveys by employees or internal stakeholders to artificially inflate scores. Their analysis shows it can range from directly asking customers for a specific rating to tapping satisfaction terminals in stores with no controls over who presses the button or how many times (Ipsos, Gaming the Score: When Feedback Becomes Fiction, 2026).
The data is consistent. A 2024 study published in the International Journal of Market Research confirmed that the NPS calculation methodology itself introduces statistical noise that makes it difficult to detect genuine shifts in customer sentiment and that’s before anyone even tries to game the results (CMSWire, Why NPS Is Lying to You About Customer Experience, 2026).
To be clear that this isn’t a CX-specific problem, two classic examples are worth revisiting.
In the 1980s, British hospitals were given a target: no patient should wait more than 18 weeks for treatment. The target was met. Waiting times fell. But a closer look revealed that doctors were deliberately delaying formal referrals to avoid starting the clock. Complex cases were avoided because they dragged down the statistics. Simple cases multiplied at the expense of those that genuinely needed urgent attention. The metric was hit. The quality of care was not.
Wells Fargo is probably the most cited example of metric destruction in the corporate world. In the early 2000s, management began closely tracking the number of new accounts opened at each branch and tied that number to bonuses, career progression, and job security. The outcome was predictable: managers started opening accounts without customers’ knowledge. By 2017, the bank admitted to 3.5 million fraudulent accounts. Regulatory fines reached $185 million; the total settlement with the US Department of Justice and the SEC came to $3 billion. 5,300 employees were dismissed. None of them had dreamed up the fraud themselves — they were simply responding to a badly designed incentive system (Ethics Unwrapped, University of Texas, 2020).
The critical point in all these cases is the same: no one consciously set out to corrupt the metric. The system corrupted it by elevating the metric to a goal.
The three most widely used CX metrics NPS, CSAT, and CES share one structural characteristic: they are proxy metrics. They don’t directly measure loyalty, satisfaction, or low friction. They measure what customers report about how loyal, satisfied, or unburdened they feel. That distinction matters enormously.
A proxy metric is inherently vulnerable. Once there is financial or career pressure tied to its outcome, people optimise the proxy not the reality it’s supposed to represent.
Fred Reichheld, the creator of NPS and a partner at Bain & Company, has named this dynamic the biggest systemic failure in NPS implementation: “Bad things happen when you just focus on it as a score. When scores are the objective, they no longer help people. It doesn’t inspire them to learn. All they want is a 10.” (Bain & Company, Refocusing NPS for Earned Growth, 2022). He went further still: programmes that link NPS to frontline employee bonuses typically collapse within one to two years.
In the Harvard Business Review article Net Promoter 3.0 (2021), Reichheld and his colleagues at Bain identified linking NPS scores to bonuses for frontline employees as one of the primary forms of misuse that has undermined the credibility of the entire system. In response, they proposed a supplementary metric Earned Growth Rate which measures the proportion of revenue coming from returning customers and those acquired through referrals. Unlike NPS, it’s a hard accounting metric that cannot easily be gamed through social engineering.
The situation with CSAT and CES is much the same. CSAT typically measured as an average rating on a 1–5 or 1–10 scale immediately after an interaction is highly sensitive to timing. A survey sent right after a problem is resolved generates a very different score from one sent a week later. Teams whose bonuses depend on CSAT know this, and they time their surveys accordingly.
CES, designed by the CEB (now Gartner) team in 2010 as a predictor of customer loyalty through low-effort interactions, is perhaps the most robust of the three but it isn’t immune either. Agents can learn to frame survey questions in ways that imply ease (“I hope we made that easy for you?”), or actively coach customers on how to respond.
Goodhart’s Law is not the product of bad people. It’s the product of rational people responding to badly designed systems.
Psychologists refer to this as surrogation the state in which people begin to conflate a goal (better customer experience) with its proxy (a higher score). Research by Michael Harris and Bill Tayler in Harvard Business Review (Don’t Let Metrics Undermine Your Business, 2019) showed that surrogation kicks in almost automatically once metrics are tied to rewards, even among people who originally understood the goal correctly.
The outcome is well documented: the organisation optimises the number instead of the reality. The number goes up. The reality stagnates or deteriorates. Leadership is happy. Customers are not.
This state is also self-reinforcing. Once metrics are being gamed, they lose their diagnostic value. Managers no longer have reliable data to make decisions with. So they make decisions based on numbers that don’t reflect reality and the cycle continues.
The real question, then, is not “How do we protect our metrics from being gamed?” It’s “What are metrics actually for?”
Reichheld and Bain & Company have consistently argued that NPS and the same applies to CSAT and CES was designed as a learning tool, not a reward mechanism. When survey results feed into a team discussion about what’s frustrating customers and what’s working, the metric retains its value. When they feed directly into a bonus calculation, they lose it.
From a reward system design perspective, this leads to several practical conclusions.
Metrics tied to bonuses should reflect behaviour, not survey responses. Rather than using NPS as a KPI linked to incentive pay, it makes more sense to track retention rates, repeat purchase rates, or Earned Growth Rate in other words, what customers do, not what they say.
Every metric needs a shadow metric a counter-measure that reveals whether the problem has simply shifted elsewhere. If you’re driving down average handle time (AHT), you need to monitor repeat contact rates. If you’re optimising CSAT immediately post-interaction, you need to track NPS with a time delay. Without a shadow metric, Goodhart’s Law is virtually inevitable.
Survey data should be separated from managerial accountability for scores. An employee should be evaluated on whether customer feedback led to action not on what number the survey produced. That is a fundamental structural difference.
To close: metrics stop being useful the moment they become the objective. This isn’t an argument against measurement — it’s an argument for being far more deliberate about why and how we measure. Goodhart’s Law cannot be solved by adding more metrics. It can be mitigated by distinguishing between metrics for learning and metrics for accountability, and by making sure you never use one where the other belongs.
]]>I’ll admit I did the same thing for a long time. For years I saw Customer Effort Score (CES) as the younger sibling of NPS something useful, but not “the” main metric. It took me a few years and several research projects to realise that CES and NPS don’t measure the same thing from a different angle. They measure entirely different things. And if you only use one of them, you’re looking at your customer with one eye closed.
Let’s start with the basic difference, because it often gets blurred in practice.
NPS is by now a classic. Fred Reichheld introduced it in 2003 in a piece that has since become almost required reading in the CX world: “The One Number You Need to Grow” in the Harvard Business Review. The principle is simple: you ask the customer how likely they are to recommend your company to a colleague or friend. You sort the answers on a 0–10 scale into three groups (promoters 9–10, passives 7–8, detractors 0–6), and the result is the difference between the percentage of promoters and detractors. Reichheld argued at the time that this single question correlated with company growth better than any other. And because it sounded compelling, people believed it.
CES came along seven years later with a far more provocative ambition. In 2010, Matthew Dixon, Karen Freeman and Nicholas Toman published an article in the Harvard Business Review with the manifesto-like title “Stop Trying to Delight Your Customers.” It was based on a large study by the Corporate Executive Board (now part of Gartner) covering more than 75,000 customers, and its main finding was uncomfortable: trying to “wow” customers has surprisingly little effect on loyalty. What actually predicts loyalty is something far more mundane how easy it is to do business with the company. That finding gave rise to the question CES still uses today: “To what extent do you agree that the company made it easy for you to handle your issue?” Responses are recorded on a 1–7 scale.
Here’s the crucial difference that’s easy to miss: NPS measures the relationship, CES measures the interaction. NPS asks what the customer thinks about the company as a whole. CES asks how smoothly a specific thing they were dealing with went.
It sounds like a small nuance. But in practice it means that if you only measure NPS, you don’t know where exactly your customer journey is grinding. And if you only measure CES, you don’t know whether your customers actually like you.
This is one of the things I really enjoy about CX: sometimes the data tells you something completely different from what you’d expect. Look at the original research that gave rise to CES, and you get a rather alarming picture. Customers who had a “high-effort” interaction with a company became less loyal in 96% of cases. For low-effort interactions, that figure was just 9%. In other words, a bad experience costs you loyalty far faster than a good experience earns it.
This asymmetry has a deep psychological foundation. Daniel Kahneman and Amos Tversky described it back in 1979 in their prospect theory: losses hurt roughly twice as much as equivalent gains feel good. When a customer has to be transferred between four agents to fix a billing error, they don’t walk away thrilled that it was eventually resolved. They walk away annoyed that it took so long.
In its follow-up studies, Gartner went further and compared the predictive power of individual metrics. According to its research, CES predicts future customer behaviour contract renewal, repeat purchases, increased spend about 12 percentage points more accurately than satisfaction, and nearly twice as accurately as NPS. But only in the context of service interactions. That’s an important caveat I’ll come back to in a moment.
If I were to sum up what I’ve learned over the years working with these metrics, it would look something like this:
NPS makes sense when you want to know how strong your relationship with the customer is. It suits relational measurement not after every interaction, but at regular intervals, typically quarterly or every six months. It works well in contexts where the customer has a long-term relationship with the company and where asking about an overall impression actually makes sense: telecoms, banks, B2B services, software as a service. NPS is a useful strategic metric for leadership a single aggregated indicator you can track over time and (with great care) compare with competitors.
CES makes sense when you want to know whether a specific moment in the customer journey is working. It suits transactional measurement immediately after a particular interaction. After a complaint has been resolved. After an order has been placed. After a call with support. After onboarding. CES is an operational metric that tells you exactly where the customer is struggling. And because it’s tied to a specific event, you can act on it straight away: improve a particular step, retrain a particular team, redesign a particular form.
Think of it this way: NPS is like an annual check-up at the doctor that tells you how you’re doing overall. CES is like a blood pressure monitor you reach for whenever you suspect something might be wrong. You need both. But you need them at different moments and for different decisions.
When I talk to people at companies that only measure NPS, I usually hear two arguments. The first is simplicity: “one number that leadership understands.” The second is tradition: “we’ve been measuring it for five years, we have a time series.” I understand both. But both have a catch.
One number that leadership understands is a lovely thing, but NPS on its own doesn’t tell you what to do. If it drops from 42 to 38, that’s bad. But why? Where? In which part of the customer journey? NPS won’t answer that. So companies often add an open-ended “Why?” question alongside NPS and then drown in hundreds of unstructured responses that are hard to turn into an action plan.
And a time series is only useful if the metric measures what you actually need to manage. If you’re an e-commerce company trying to improve cart conversion, NPS won’t help much. CES after a completed (or abandoned) order will.
It’s worth noting that even Bain & Company, the firm behind the creation of NPS, recommends in its own materials a combination of relational and transactional measurement essentially NPS combined with something else, whether CES or CSAT (Customer Satisfaction Score). NPS was never meant to be the only metric. It became one because it’s so catchy.
Here’s the thing I enjoy most about CX: the best results usually don’t come from picking the “right” metric, but from building a sensible system.
In practice, it pays to split measurement into two layers. Relational NPS is measured at regular intervals (quarterly, every six months) across the whole customer base and serves as a strategic indicator for leadership. You watch the trend, segment by customer type, and correlate it with real behaviour contract renewals, churn, growth in spend.
Transactional CES is measured continuously after key interactions the ones you’ve identified as moments of truth in your customer journey. After a support ticket is resolved. After onboarding is finished. After a first purchase. Here CES acts as an operational signal that lets you quickly spot where the customer is struggling and do something about it immediately.
These two layers inform each other. When relational NPS drops, you look at recent transactional CES data to find the source. When transactional CES flags a problem at a specific step in the journey, you watch to see whether it eventually shows up in relational NPS.
Forrester recommends something similar. In one of its reports on the state of CX measurement, it notes that companies combining multiple types of feedback relational, transactional and behavioural are significantly more likely to translate improvements in customer experience into measurable business results. A single metric, however good, won’t give you that connected picture.
If you’re a company that currently only measures NPS, I don’t think it’s a fatal mistake. NPS has its value, and there’s no point in abandoning it. But it’s worth asking whether that one metric is really enough to tell you where in the customer journey you need to act. If you find that you can’t give a concrete answer to “why has our NPS dropped” over and over again, that’s a signal you’re missing the transactional layer of measurement.
And if you’re a company that only measures CES (which is rarer, but it happens, particularly with product-led tech companies), try asking yourself whether you actually know how strong your relationship with the customer is outside of specific interactions. A product being easy to use doesn’t mean the customer will recommend you. Sometimes they’ve just learned how to use it, and the first competitor’s offer will convince them they can get the same ease elsewhere.
The thing that’s been confirmed for me most over years of working with CX metrics is this: metrics are not the goal. They’re lenses through which we look at the customer. And one lens will never show you the whole picture. So the question isn’t “CES or NPS.” The question is: “What system of perspectives on the customer will give me answers I can actually act on?”
That’s the question I find fascinating. And I think it deserves far more attention than the endlessly recycled debates about whether NPS is dead or not.
]]>And yet the CX team already has a metric sitting in its database that dissolves this imbalance. It’s called Customer Lifetime Value, or CLV. In most companies, though, it lives in a spreadsheet on the finance team’s drive or inside a CRM the CX specialists can’t even access. It shows up at board level once a year, when the acquisition budget gets decided. It rarely enters the conversation about retention, onboarding, or redesigning the customer journey.
This is a strategic mistake. CLV isn’t a financial indicator that happens to involve customers. CLV is a direct numerical expression of the quality of the customer experience over time. And the companies that have grasped this hold a disproportionate advantage when arguing for CX investment, regardless of whether they have five employees, five hundred, or fifty thousand.
The simplified formula is well known: average purchase value multiplied by purchase frequency multiplied by the average length of the customer relationship. More advanced versions add margin and a discount rate to calculate the net present value of future revenue.
The problem isn’t the formula. The problem is what goes into it.
Every variable is a direct consequence of the customer experience. Purchase frequency rises when the customer trusts both the product and the process. Average order value goes up when the customer feels confident and identifies with the brand. Relationship length is a function of whether the customer stays or leaves, and that depends primarily on how they feel every time they interact with the company.
The data on this is consistent across industries and decades. Frederick Reichheld of Bain & Company, in a much-cited study, showed that raising customer retention by just 5% can lift profits by 25% to 95%. Harvard Business Review reconfirmed this in 2014, adding that acquiring a new customer costs five to twenty-five times more than keeping an existing one. These figures get repeated endlessly in the CX literature, but very few people draw the logical conclusion: if retention matters that much to profit, and if retention is a product of customer experience, then the CX team holds one of the most powerful financial levers in the company.
And almost never manages to show it.
There are two reasons CX teams and CLV don’t get along. Both are organisational rather than technical.
The first is the time horizon. CX managers are measured on quarterly NPS scores, ticket resolution speed, or CSAT (Customer Satisfaction Score). CLV, by contrast, needs at least a one-year horizon, realistically three years, before it shows meaningful movement. In an environment that rewards quick wins, a long-term metric is always at a disadvantage.
The second reason is cultural. Customer experience as a discipline grew out of empathy, psychology, and service design. Its native language is one of feelings, journeys, and moments of truth. Financial language was long considered the CFO’s tool, not the CX manager’s. The argument “our customer feels better” is meaningful, but it loses against a line item on a balance sheet. CLV is a number the CX team already has and doesn’t use.
The real question, then, isn’t what CLV is. It’s how the CX team can influence it and prove they’ve done so.
Before getting into how to use CLV in practice, its limits need to be acknowledged. Without that, this article turns into a brochure.
Historical CLV assumes the past predicts the future. In industries with rapid shifts in customer behaviour or technological disruption, that assumption breaks. A customer who stayed loyal to a telecoms provider for ten years can walk out overnight when the tariff structure changes, no matter how positive their previous experience was. Predictive CLV, a model built on the probability of future behaviour, is more accurate but requires a data infrastructure that mid-sized companies typically don’t have.
The second limit is segmentation. Average CLV across the entire customer base is a metric that says almost nothing. The gap between the most valuable and least valuable customer tends to be fifteenfold in retail and up to fiftyfold in B2B (business-to-business). When a company reports “average CLV,” it’s reporting the average of two customers who have nothing in common.
The third limit is the most serious: CLV doesn’t measure why a customer stayed or left. It’s an output metric, not a diagnostic one. A CX team that has only CLV and no qualitative data alongside it can track the result but has no understanding of the cause.
Even so, making decisions about CX investment without any financial indicator at all is worse than making them with an imperfect one. CLV is the bridge between the language of customer empathy and the language of the boardroom. At the moment, there’s no other bridge in the field.
Three steps separate the companies that use CLV strategically from the ones where CLV sits untouched in the CRM.
The first step is segmenting customers by CLV and overlaying that with CX data. High-CLV customers who also report low satisfaction or give detractor-level NPS responses are the most urgent group in the entire database. Losing them is the most expensive outcome. In banking and telecoms, McKinsey & Company calls this group the silent at-risk cohort, customers who leave without escalation, complaint, or any visible signal. This cohort should be priority number one for every CX programme. What actually tends to happen is that companies invest in generally improving NPS among the average customer, while the ones who generate 40% of revenue quietly slip away.
The second step is tracking CLV as the outcome metric of CX initiatives. When a company rolls out a new onboarding flow, proactive care, or a redesigned customer journey, the CLV of the cohort that went through those programmes should measurably differ from a control group. This methodology, standard in product management and growth marketing, is rarely applied in CX. And yet it’s exactly the kind of evidence a CFO will accept.
The third step is linking CX metrics to CLV at the individual level, not just in aggregate. Bain & Company has repeatedly documented that promoters, meaning customers with NPS scores of 9 or 10, have statistically significantly higher lifetime value than passive customers. When this correlation exists in a company’s own data and the CX team can’t communicate it internally, one of the strongest arguments for investing in customer experience gets lost.
The principle is the same at every scale. Only the numbers change.
Amazon Prime is a textbook case of CLV thinking. According to Consumer Intelligence Research Partners, Prime members spent an average of $1,170 on Amazon.com in 2024, while non-Prime customers spent $570. That two-to-one ratio has held steady for the past five years. The entire design of Prime, from free shipping to Prime Video to premium customer service, is built as a programme for increasing CLV, not as a loyalty scheme in the traditional sense. Prime isn’t paid for with discounts. It’s paid for with experience.
Netflix follows the same logic by a different route. In Q3 2024, Netflix hit a monthly churn rate of 2.17%, the best result in the entire streaming segment, ahead of Prime Video at 3.7% and well ahead of Paramount+ at nearly 5%. Netflix hasn’t stopped investing in acquisition, but its decisions about content, interface, and pricing are primarily optimised for retention. A well-documented example is the end of password sharing in 2023: a move that triggered a wave of negative reactions and short-term cancellations, but added roughly 50 million new subscribers between the end of 2023 and Q4 2024. The decision made sense in the language of CLV. In the language of quarterly NPS, it probably wouldn’t have passed.
The real question is whether this logic works at a scale that’s actually useful to most readers. The answer shows up in every mid-sized e-shop.
Take a typical B2C (business-to-consumer) e-shop with annual revenue somewhere between 30 and 200 million Czech crowns, whether in cosmetics, fashion, or speciality food. Benchmark data for e-commerce is fairly consistent: the average repeat purchase rate, meaning the share of customers who buy from the company more than once, sits between 20% and 30% according to analyses from Bluecore and other sources. Average CLV in Europe and the US, based on Shopify and aggregated e-commerce data, lands between $100 and $300 per customer. And according to Shopify Commerce Trends, average customer acquisition cost (CAC) has risen 222% over the past eight years.
The maths that follows is simple but brutal. If a mid-sized e-shop acquires a customer for 500 crowns and average CLV is 1,500 crowns, the LTV-to-CAC ratio is 3:1, which Shopify and other platforms consider the healthy threshold. Once CAC rises to 700 or 800 crowns, which has become common in highly competitive segments over the past two years, the company suddenly isn’t making money on each acquired customer. The acquisition model has stopped working.
At that point, the company has two options. Either raise CLV or lower CAC. Lowering CAC in a world of rising ad prices on Meta and Google is almost impossible without scaling back reach. Raising CLV, on the other hand, sits entirely in the hands of the CX team, or, where no formal CX team exists, the operations or e-commerce manager. Every percentage point of improved retention, every extra crown in average order value, every additional order per year from the same customer pushes CLV up and the budget equation back into the black.
The practical levers a mid-sized e-shop has are modest but effective. Segmenting customers by purchase frequency and value, known as RFM analysis (Recency, Frequency, Monetary), is available in Shopify and in Czech tools like Meiro or Samba.ai. Automated win-back campaigns, according to aggregated e-commerce data, bring 10% to 15% of lapsing customers back. Welcome emails, according to ActiveCampaign, have open rates above 90% and generate revenue per message that’s an order of magnitude higher than standard promotional campaigns. None of this requires a billion-crown budget or a data scientist. It requires a CX mindset that starts with the question “what does it cost me to lose this customer?” and continues with “what can I do to keep them?”
The difference between Amazon and a mid-sized e-shop isn’t that one knows CLV and the other doesn’t. It’s that Amazon plans every strategic decision around CLV, while a mid-sized e-shop typically can’t even calculate it. And for a smaller company, this number matters even more, because it can’t afford to offset the loss of a few hundred valuable customers with aggressive acquisition. The budget isn’t there.
CX programmes that can’t demonstrate financial impact lose internal battles for budget, headcount, and leadership attention. In mid-sized companies these battles look different than in corporates, but they play out the same way. The budget that could go into improving the customer experience goes into performance advertising instead, because the marketing manager can express performance advertising in ROI, while the CX specialist can only say NPS went up three points.
The solution isn’t to teach the CX team the language of finance. The solution is to teach the CX team one specific indicator that makes sense to the rest of the company and directly reflects the quality of the customer experience. CLV is that indicator. It isn’t perfect. It’s historical, averaged, and diagnostically thin. But it’s the one metric the CFO or business owner and the CX manager can read together and reach the same conclusion.
And in an environment where the decision gets made about whether customer experience gets a budget or not, that’s probably the most important thing a CX metric can do.
]]>McKinsey describes the problem even more sharply. A typical CX survey, according to their research, captures only about 7% of customers. Only 13% of CX leaders expressed full confidence in the representativeness of their own measurement, and only 4% stated that their system can calculate the ROI (Return on Investment) of CX decisions (McKinsey, Prediction: The future of CX, 2023). This is exactly why so many companies report scores but are unable to manage economic impact.
At the same time, however, it is not true that the link between experience and business outcomes does not exist. A global study by XM Institute on a sample of 28,400 consumers in 26 countries showed a strong correlation between satisfaction and trust, recommendation, and further purchase: the Pearson coefficient was 0.71 for trust, 0.82 for recommendation, and 0.69 for willingness to buy more. Overall, customers after a five-star experience were 2.2 times more willing to buy more than after an unsatisfactory experience (XM Institute / Qualtrics, ROI of CX, 2024). As an external benchmark, this is a very strong signal; as proof of causality in your company, however, it is not sufficient on its own.
Similarly, ACSI (American Customer Satisfaction Index) has long shown that companies with higher and improving customer satisfaction have better capital performance: the portfolio of ACSI leaders achieved a cumulative return of 2,265% from 2006 to January 2025 compared to 605% for the S&P 500 index (ACSI, 2025). This is also an important indication for boards and investors. But at the level of a specific company, it still holds that a CFO will not invest in CX because a correlation exists in the market. They will invest when they see the mechanics of impact in their own P&L (Profit and Loss statement).
The worst possible question is: “By how much did we increase NPS?” The correct question is: “Which customer behaviors create value in our business?” McKinsey recommends starting exactly here: define the behaviors that generate the company’s economics in a given industry, and only then track how customer experience influences them. In telecom, it may be churn, the number of escalated calls, and upsell of additional services. In airlines, a greater share of trips and lower cost-to-serve. In companies that use CX as a growth engine, metrics such as share of wallet, repeat purchase, or net revenue retention – NRR are then tracked (McKinsey, Linking the customer experience to value, 2016).
Here is the key logic that companies often skip:
experience → customer behavior → economic outcome.
Relational metrics such as NPS, CSAT, and CES are therefore more like sensors than financial outcomes. Beneath them must be operational and journey metrics — first-contact resolution, lead time, OTIF (On Time In Full), number of channel transfers, time-to-value, number of repeated contacts. Only these translate into behavior: retention, purchase frequency, cross-sell, share of wallet, or NRR. And only from these do revenue, gross margin, and cost-to-serve arise. McKinsey also points out that end-to-end journeys have a significantly greater impact on economic outcomes than isolated touchpoints.
1. Retention bridge
The cleanest model for most industries is the retention bridge. Its logic is simple: how many customers do not leave thanks to a better experience — and what gross margin they have. The formula looks like this:
Incremental value = number of customers in the affected segment × churn reduction × annual gross margin per customer
It is important that it is not calculated from revenue, but from gross contribution after deducting service costs. McKinsey recommends working at the customer level and linking survey results with two- to three-year histories of retention, revenue, upgrades, and cost-to-serve. Only then will you find out how much a point of improvement is actually worth.
Model example: redesign of the first invoice reduces 90-day churn from 18% to 15% for 40,000 new customers. The annual gross margin per customer is CZK 2,500. The number of retained customers increases by 1,200, which means CZK 3 million in incremental gross margin. If the process, IT, and communication changes cost CZK 1.2 million, the first-year ROI is 150%. This is the language finance understands.
2. Revenue expansion model: share of wallet, cross-sell, and NRR
In subscriptions, telecom, banking, or B2B SaaS (Software as a Service), CX is often stronger through expansion than through retention alone. It is not just about whether the customer stays, but how much additional value they leave in the relationship. In its analysis of more than a hundred B2B SaaS companies, McKinsey shows that companies in the top quartile of valuations achieved NRR of 113%, while companies in the bottom quartile only 98%. An even more interesting detail: advanced value realization and adoption journeys programs were associated with roughly a seven-point advantage in NRR compared to companies with basic practices. For pricing and packaging, the difference was approximately 16 points (McKinsey, The net revenue retention advantage, 2023).
In practice, this means one thing: the CX team must not stop at sentiment. It must be able to demonstrate how onboarding, product adoption, support quality, and proactive care influence activation, use of key features, contract renewal, and upsell. In B2B, this is often the most convincing bridge to revenue. In B2C, the analogue is purchase frequency, basket size, and share of wallet.
3. Journey economics: cost-to-serve and margin
The third model is paradoxically often the fastest in companies, because it returns money sooner than revenue uplift. McKinsey has long shown that successful CX programs across industries typically bring revenue growth of 5 to 10% and cost reductions of 15 to 25% within two to three years. Companies with exceptional CX can also outperform competitors in gross margin by more than 26%. In more recent work, McKinsey states that data-driven “next best experience” can increase revenue by 5 to 8% and reduce cost-to-serve by 20 to 30% (McKinsey, Customer experience: Creating value through transforming customer journeys, 2016; McKinsey, Prediction: The future of CX, 2023).
This is why a good CX business case is often built from the bottom up: fewer repeated contacts, fewer escalations, fewer channel switches, lower error rates, fewer returns, faster digital service. In one example, McKinsey describes a card company that, thanks to journey analytics across 13 priority journeys, reduced interaction and operational costs by 10 to 25%. This is no longer a “soft benefit.” It is a hard item in operational efficiency.
For the connection between CX and finance to work, the company must above all have a unified customer identifier across CRM (Customer Relationship Management), billing, contact center, digital analytics, and survey. Without customer-level data, everything remains at the level of correlation in PowerPoint. McKinsey recommends linking survey responses with two- to three-year histories of monthly data and tracking outputs over time for segments that are important to the business.
The second prerequisite is discipline in causality. It is not enough to show that customers with better CSAT churn less. It is necessary to separate the effect of experience from price, promotions, segment, length of relationship, acquisition source, or seasonality. In practice, this means working with cohorts, holdout groups, before-after comparisons, and, where possible, controlled pilots. The survey remains important, but on its own it is backward-looking, incomplete, and weak in identifying causes. That is why McKinsey recommends combining it with interaction, transactional, and profile data and creating predictive scores that directly estimate revenue, loyalty, and cost-to-serve.
1. The company confuses correlation with causality. External studies show strong links between CX and loyalty, but an internal business case must be based on internal data, control groups, and time tracking (XM Institute / Qualtrics, 2024).
2. It calculates benefits in revenue, but not in margin. Higher turnover without accounting for discounts, service costs, and cost-to-serve is insufficient for a CFO.
3. It measures touchpoints but manages journeys. McKinsey explicitly points out that journey performance is more strongly linked to economic outcomes than individual touchpoints.
4. It ignores time lag. Some impacts appear only after months.
5. It overemphasizes acquisition and underestimates the economics of the existing base. McKinsey calls this the “acquisition trap” (McKinsey, Experience-led growth, 2022).
The best CX teams today do not primarily talk about scores. They talk about churn, share of wallet, NRR, cost-to-serve, and gross margin. McKinsey shows that companies that have made CX a growth engine start from the desired financial outcome and only then determine which experiences should trigger it. And where customer satisfaction is significantly improved, the effect is also visible in cross-sell and share of customer spend; CX leaders also achieved more than double the revenue growth between 2016 and 2021 compared to lagging companies.
That is ultimately the most important shift. A mature company no longer asks whether CX is important. It asks in which journeys, for which segments, and with what financial effect. The moment you translate customer experience into the language of retention, margin, and cost-to-serve, it stops being a “soft discipline.” It becomes the management of the economics of customer relationships.
]]>But this very simplicity is also a weakness.
Customer experience is not a one-dimensional variable. And the average — however statistically correct — often hides more than it reveals.
Let us imagine two companies with the same NPS of 30.
The same number. A completely different reality.
While the first company operates with a relatively stable, albeit unremarkable experience, the second faces a fundamental reputational risk. However, this difference is not shown by the average score.
This is precisely why, in recent years, the importance of the distribution of responses has been increasingly emphasized. For example, analyses by Bain & Company, who stand behind the popularization of NPS, repeatedly point out that the score itself, without the context of segments and the structure of responses, leads to incorrect managerial conclusions.
The distribution reveals:
Without this view, you are working with incomplete information.
1. Polarization of the experience
A wide spread of responses signals inconsistent performance. A narrow one, on the contrary, indicates a standardized, predictable experience.
Variability is not just a statistical property. It is a reflection of management quality.
According to McKinsey research (2021), companies that are able to systematically reduce variability of customer experience across touchpoints achieve up to 15–20% higher customer satisfaction than the competition.
2. Hidden risk of detractors
The average may remain stable, while the share of detractors in a key segment grows. Typically among VIP customers or in critical phases of the customer journey.
This is a fundamental problem — a study by Temkin Group (now part of Qualtrics XM Institute) shows that detractors have up to 2–3× higher probability of churn and negative word-of-mouth than passive customers.
The average “dilutes” this signal.
3. Top-box effect
In CSAT, it has long been shown that the so-called top-box score (e.g. rating 5 out of 5) has a significantly stronger correlation with loyalty than the average.
For example, research by Harvard Business Review (Keiningham et al., 2014) showed that customers with the highest rating have a significantly higher probability of repeat purchase as well as recommendation than those who are “just satisfied.”
The difference between 4 and 5 is not cosmetic. It is behavioral.
4. Shape of the distribution
The skewness of the distribution can reveal systemic problems.
Left skew may indicate structural failures (e.g. logistics, complaints)
Right skew, on the contrary, indicates stable performance with isolated failures
Such patterns cannot be read from the average.
Relying on the average has concrete consequences.
Underinvestment in CX
If the overall number “looks good,” the organization has no tendency to act. Meanwhile, a dissatisfied minority may be growing, gradually eroding brand value.
According to Forrester (CX Index, 2023), companies that ignore segment differences in experience lose up to 10–15% of potential revenue due to unidentified problems.
Distribution only makes sense in connection with segmentation:
Only here does the true quality of experience management become apparent.
Trend of distribution vs. trend of the average
A stable NPS may give a reassuring impression. But if polarization is growing at the same time, it is a warning signal.
The average stagnates. Risk grows.
Companies that take CX seriously today report differently than five years ago. One number in a dashboard is not enough.
The standard should be:
Technologically, this is not a problem today. Platforms such as InsightSofa or Qualtrics make it possible to track not only the metrics themselves, but also their structural dynamics over time.
And it is precisely this layer that determines whether you truly manage CX — or just report it.
The average is convenient. The distribution is truthful.
If we want to manage customer experience systematically, we must understand its variability, distribution, and differences between segments.
Because CX is not about how many points you get.
It is about how consistently you deliver value — and to whom.
]]>This paradox has a name: the behavioral gap – the difference between what customers declare and how they actually behave.
For customer experience (CX) leaders as well as commercial directors, this is not a methodological detail but a strategic problem. If an organization measures only declarations but does not track real behavior, it is not actually managing retention – it is only managing its illusion.
Empirical research has long shown that satisfaction is a necessary, but not sufficient condition for loyalty. A customer can be “satisfied” and still:
Bain & Company has repeatedly demonstrated that real economic value is delivered only by so-called “top-box” customers – those who give the highest possible rating (e.g., 9–10 in NPS). This group shows significantly higher retention rates and revenue growth. By contrast, “average satisfaction” often does not mean loyalty, but rather comfortable indifference (Reichheld, Bain & Company).
1. Response bias and social desirability
Customers respond politely. They avoid conflict, rationalize their decisions, and tend to provide more socially acceptable answers.
As a result, surveys often capture rather an ex-post rationalization of the purchase than the actual emotions that will drive future behavior. This effect is well described in behavioral economics (e.g., Kahneman, Thinking, Fast and Slow).
2. Measuring the moment vs. cumulative experience
Transactional surveys capture a specific interaction – delivery of a shipment, contact with a call center, a purchase. But customer departure is almost always the result of accumulated frustration.
Individual minor problems – delays, unclear communication, a complicated process – do not appear critical on their own. In aggregate, however, they systematically erode the relationship.
Behavioral data often signals risk earlier than surveys:
A McKinsey study (2021) shows that the combination of behavioral and experiential data increases the ability to predict churn by tens of percent.
3. Low switching costs
In many industries – banking, telecommunications, e-commerce, or SaaS – barriers to switching have dropped dramatically. Alternatives are comparable and onboarding takes minutes.
In such an environment, satisfaction ceases to be a competitive advantage. It becomes a hygiene factor.
As PwC research (Future of Customer Experience Survey, 2022) shows, up to 32% of customers leave after a single bad experience – even if they had previously been satisfied.
1. Link experiential and behavioral data
Metrics such as NPS or CSAT alone do not provide a sufficient picture. Real insight arises from linking them with:
Only this combination reveals which “satisfied” customers are in fact about to leave.
2. Track change, not just absolute values
A static score has limited explanatory power. From a predictive perspective, the trend is often more significant.
A one-point drop for a specific customer may be a stronger signal than a score of 8/10 itself. The dynamics of the relationship are more important than its current state.
3. Measure emotions, not only functional quality
Behavioral economics and neuroscience show that decision-making is primarily driven by emotions.
Alongside traditional metrics (speed, price, availability), it is therefore necessary to systematically track:
These factors often explain the difference between declared satisfaction and actual loyalty (Lemon & Verhoef, Journal of Marketing, 2016).
The behavioral gap highlights the limits of the traditional approach to CX. Standalone surveys are not enough.
Effective management today requires:
Modern platforms, such as InsightSofa, make it possible to connect experiential signals with real customer behavior and identify warning patterns before they manifest in churn. The goal is no longer reporting – but predictive intervention.
Customers rarely leave because they are openly dissatisfied. They leave because they do not have a sufficient reason to stay.
The behavioral gap reminds us of a simple but often overlooked truth:
What customers say is important. What they do determines your growth.
Companies that want to truly manage loyalty must start where average satisfaction ends.
]]>Let us call this principle experience elasticity.
It is not a metaphor. It is a practical framework that explains why some changes in customer experience have almost no impact, while others – seemingly small – lead to significant shifts in loyalty, recommendations, or customer departure.
Experience elasticity describes how strongly customers react to a change in a specific element of the experience – and how this reaction translates into their behavior.
The key premise is simple: not all changes have the same weight.
Extending opening hours by 30 minutes may remain without a measurable effect.
Conversely, a one-day delay in delivery may dramatically increase the share of negative ratings.
The difference is not in the change itself, but in its perceived importance in the given context.
1. Low elasticity: zone of tolerance
There are areas where customers accept variability without a fundamental change in evaluation. This phenomenon corresponds to the concept of the “zone of tolerance,” described by Parasuraman, Zeithaml, and Berry within the SERVQUAL model (Parasuraman et al., 1991).
Typically, these include:
From the perspective of CX management, these are areas with limited impact on loyalty. Investments here often bring only marginal returns.
2. High elasticity: critical moments
In these situations, the opposite logic applies: a small change leads to a significant impact.
Typical examples include:
It is precisely here that loyalty is decided. Research by McKinsey shows that customers rate consistent fulfillment of expectations as one of the strongest drivers of both satisfaction and retention (McKinsey, The Three Cs of Customer Satisfaction, 2016).
The problem is that aggregated metrics – such as NPS – often mask this volatility.
3. Asymmetric elasticity
Customer reactions are not symmetrical. Improvements and deteriorations of the same magnitude do not have the same effect.
Behavioral economics describes this phenomenon as loss aversion – losses hurt more than equally large gains please (Kahneman & Tversky, 1979).
In CX, this means:
Bain & Company, which popularized NPS, repeatedly points out that detractors have a disproportionately strong impact on negative word-of-mouth and churn (Reichheld, 2003).
A stable NPS does not mean a stable experience. Highly elastic moments may be hidden within aggregated data.
According to analyses by Qualtrics (2022), an aggregated score can mask significant differences between individual touchpoints that have fundamentally different impacts on loyalty.
Elasticity differs depending on:
Without segmentation, the interpretation of CX data becomes methodologically inaccurate.
For example, a PwC study (Experience is everything, 2018) shows that customers of premium brands have significantly lower tolerance for errors than customers in low-cost segments.
Not all initiatives have the same impact on the business.
Improving areas with low elasticity leads only to limited growth in loyalty.
Conversely, an intervention in critical moments can disproportionately reduce churn. According to Harvard Business Review (Stop Trying to Delight Your Customers, Dixon et al., 2010), eliminating problems in critical interactions is more effective than striving for a “wow effect.”
Identifying elasticity is not a matter of a single method, but a combination of approaches:
Driver analysis – linking touchpoint evaluations with loyalty
Churn analysis – identifying patterns of behavior before customer departure
Text analytics – detecting emotionally strong reactions in open feedback
A/B testing – measuring real behavior, not only declared satisfaction
Modern CX platforms today make it possible to combine quantitative and qualitative data. For example, InsightSofa integrates structured metrics with the analysis of open responses, which makes it possible to reveal asymmetric reactions – that is, the real zones of high elasticity.
The key question for CX management today is not:
“What is our score?”
But:
“What are our customers actually sensitive to – and where are the boundaries of their tolerance?”
Customer experience is not linear. And customer reactions even less so.
Companies that ignore this nonlinearity will be repeatedly surprised by their own data – and will invest in improvements where it has no real impact on loyalty or business results.
]]>