Skip to content

Three scores, one renewal decision, and only one worth a target

Customer Effort Score vs CSAT vs NPS: Which One Predicts Renewal?

Customer effort score vs CSAT vs NPS, judged on renewal: what each of the three was validated against, why CES is absent from B2B SaaS, and the test to run.

By , Co-founder, GainTrace · Updated · 16 min read · For Head of Customer Success, CS Operations

Short answer

Customer effort score vs CSAT vs NPS is a choice between three metrics that were each validated on something other than a B2B SaaS renewal: CES on contact-centre service interactions, NPS on company growth across industries, CSAT on single transactions. Pick the one closest to your renewal decision, then measure its lift on your own closed renewals before you put a target on it.

Somebody has asked you to settle customer effort score vs CSAT vs NPS before the next planning cycle, and the honest answer is that the question is usually framed wrong. All three are single-number summaries of what one person felt at one moment. A renewal is a decision taken by a committee over several weeks, often by people who never answered your survey. The metric you pick matters less than whether you can show it moved before a renewal did.

This page is for the Head of CS or CS Ops lead choosing which of the three to run and what to attach it to. It gives the origin and validated claim of each of the three, the evidence from 4,978 public reviews and 33,600 practitioner posts on which ones teams use in practice, a test that ranks them on your own renewals, and the one question that beats all three.

Key takeaways
  • CES has the strongest published claim on loyalty and almost no presence in B2B SaaS: it appears in 1 of 4,978 public customer success platform reviews we searched, against 217 for NPS.
  • NPS was validated in 2003 as a correlate of company growth across industries, not as an account-level predictor of a renewal, and it is used as the second thing almost everywhere.
  • In a B2B account the person who answers the survey and the person who signs the renewal are usually different people, which is the fault no amount of question wording repairs.
  • Run the lift test before setting any survey target: renewal rate among top-box responders minus renewal rate among everyone else, on your own closed renewals.
  • The question closest to the decision beats all three: ask the renewal decider how likely they are to renew today, 90 days out, and treat anything under 9 as a risk.
Browse this guide

Questions this page answers

  • Should we run CES, CSAT or NPS?
  • Which customer satisfaction metric actually predicts churn in B2B SaaS?
  • What is a good CSAT response rate?
  • Is NPS useless for B2B?
  • How is customer effort score calculated?
  • Our NPS is high but customers are leaving, what are we missing?
  • What survey should we send after onboarding?

Customer effort score vs CSAT vs NPS: which one predicts a renewal?

The decider gap

The decider gap is the difference between the share of your renewal deciders who answered the last survey and the share of end users who did. A wide gap means the score is user sentiment being used to forecast a decision those users do not take. Measure it once and most arguments about customer effort score vs CSAT vs NPS end, because the wording of the question stops being the problem.

None of the three was validated against a B2B SaaS renewal, so treat every published claim as evidence from an adjacent population. The Customer Effort Score arrived in 2010 in "Stop Trying to Delight Your Customers" by Matthew Dixon, Karen Freeman and Nick Toman in Harvard Business Review, built on a study of more than 75,000 people interacting with contact-centre representatives or using self-service channels; the article states that CES is a better predictor of loyalty than customer satisfaction measures or the Net Promoter Score. Net Promoter Score arrived seven years earlier, in Frederick F. Reichheld's December 2003 Harvard Business Review article, from two years of research linking survey answers to purchasing patterns, referrals and company growth rates.

Read the populations and the claims narrow fast. CES was measured on service interactions, where the customer has a problem and wants it gone. NPS was measured at company level against growth, not at account level against a renewal. CSAT has no single founding study at all; it is a family of transactional questions. In B2B SaaS the renewal decision runs through procurement, a budget holder and a champion, and none of those three studies looked at that.

The three scores compared on origin, respondent and validated claim, with the situation each is the right choice for. Ordered by how close the question sits to a renewal decision, closest last.
ScoreQuestion askedWho answersValidated onUse when
NPSHow likely are you to recommend us to a friend or colleague, 0 to 10Whoever opens the email, usually an end userCompany-level growth rates across industries, Reichheld, HBR 2003You need one comparable number across a large base and can tolerate quarterly noise
CSATHow satisfied were you with this interaction, 1 to 5The person who had the interaction, moments after itNo single founding study; a family of transactional measuresYou are measuring support, onboarding or a training session and want to fix a process
CESHow much effort did you personally have to put forth to handle your requestThe person who did the work, at the end of the taskMore than 75,000 contact-centre and self-service interactions, HBR 2010You are removing friction from a repeated task and want a number that moves when you fix it
Renewal likelihoodIf you had to renew today, how likely are you to do it, 0 to 10The named renewal decider, 60 to 90 days outNothing published; practitioner use only, and it names the decision directlyYou want the earliest honest read on a specific renewal and can ask the decider

Practitioners have already voted, and they did not vote for the metric with the best published claim. Across 4,978 public G2 reviews of five customer success platforms, 217 reviews (4.4%) mention NPS, 49 (1.0%) mention CSAT and exactly 1 (0.02%) mentions CES. That single mention is a feature request.

I'd like CES and CSAT survey options in addition to NPS, and I'd love more options to create formulaic fields.
Small-Business reviewer, public G2 review

What does each of the three scores measure, and on whom?

Each of the three scores compresses a different question into one number, and the arithmetic matters because two of the three throw away most of the answers. NPS discards every 7 and 8. CSAT usually reports top-box only, discarding the middle. CES reports a mean, which keeps everything but is the only one of the three with no agreed scale: five-point, seven-point and reversed variants all circulate, so two CES numbers are rarely comparable.

Net Promoter Score

NPS = Promoters (9 to 10) % Detractors (0 to 6) %

Promoters
respondents scoring 9 or 10, as a share of all respondents
Detractors
respondents scoring 0 to 6; the 7s and 8s are counted in the denominator and nowhere else
What good looks like
Userpilot's 2024 Product Metrics Benchmark Report, first-party data from 547 SaaS companies, reported NPS of 34.5 at $1M to $5M revenue, 23.3 at $5M to $10M, 37.5 at $10M to $50M and 39.1 above $50M; the report does not state whether these are means or medians
Customer Satisfaction and Customer Effort

CSAT = Top-box responses ÷ All responses × 100 CES = Sum of effort ratings ÷ Number of responses

Top-box
4 and 5 on a five-point scale, or 5 alone if your scale is strict; state which, because the two definitions differ by 10 to 20 points
Effort rating
agreement that the company made it easy to handle the request, normally on a seven-point scale where higher is better
What good looks like
a number you can compare with your own last quarter; cross-company comparison of CSAT or CES is meaningless without the identical scale and trigger

The respondent is the part nobody controls. A CSAT survey fires at the person who raised the ticket. An NPS campaign reaches whoever is on the contact list. Neither is targeted at the renewal decider, and in an account with 200 seats and one budget holder, the odds that your score reflects the budget holder are poor. Measuring the decider gap takes an afternoon and changes how much weight the score deserves.

Decider coverage

Decider coverage = Named renewal deciders who answered ÷ All named renewal deciders × 100

Named renewal decider
the economic buyer and the champion recorded on the account, not the whole contact list
What good looks like
above 60%. Below 30% the score describes users only, and no survey metric should carry a renewal target at that coverage

Why does an NPS score move without the renewal moving?

An NPS score at account level moves mostly because the response set changed, not because sentiment did. With six responses from a 200-seat account, one detractor swings the account score by 17 points. Quarter to quarter you are comparing different people, and the accounts that stop responding are frequently the ones going quiet for the reason that matters. Low response rates are the complaint reviewers make most.

NPS was totally unreliable with bad response rates.
Head of Customer Success Management, mid-market SaaS, public G2 review
The days of relying on a once-a-year NPS score are over.
SVP Sales and Partnerships, mid-market SaaS, public G2 review

The second reason is cadence. An annual NPS campaign gives you one reading per account per year, and a renewal decision forms over six to eight weeks. A score you sample once cannot show a change, and change is the only part a CSM can act on. Falling response rates are making the sampling problem worse across the profession, not better.

Everyone in CX circles is talking about falling survey response rates. Survey fatigue, customer concerns on how well their feedback will be addressed, cognitive load of thinking what to respond when there is no wow or severe disappointment
r/CustomerSuccess, 2025

Both problems have the same practical answer: stop treating the score as a level and start treating it as a change against the account's own history, and stop reading any account-level score built on fewer than five responses. Is an NPS target a fair KPI for a CSM covers what to do when the number is already on your scorecard.

How do I test which score predicts renewal on my own accounts?

Run the lift test: compare the renewal rate of accounts that gave a top-box answer against the renewal rate of everyone else, using surveys sent 90 or more days before the renewal date. The metric with the widest gap is the one worth a target. Most teams have never run this, which is why the choice between customer effort score vs CSAT vs NPS gets settled by preference.

Renewal lift

Renewal lift = Renewal rate of top-box responders Renewal rate of all other accounts

Top-box responder
an account whose most recent response 90 or more days before the renewal was a promoter, a 5 on CSAT, or a low-effort answer on CES
All other accounts
everyone else, including accounts that did not respond at all; silence is data and excluding it flatters the metric
What good looks like
above 10 percentage points. Under 5 points the score is not carrying renewal information and should not carry a renewal target
  1. Pull every renewal decision from the last four quarters

    Renewed, churned and contracted, with the decision date. Contractions above 20% count as a partial loss and should be kept in a separate column and never dropped.

  2. Attach the last survey response before the 90-day mark

    Use the response, the date and the respondent's role. Accounts with no response form their own group; do not delete them.

  3. Compute the lift for each metric you run

    Renewal rate of top-box responders minus renewal rate of everyone else. Do it separately for NPS, CSAT and CES if you run more than one.

  4. Compute the decider gap

    What share of the responses came from a named renewal decider. If it is under 30%, the lift computed above is a property of your user base, not your buyers.

  5. Repeat with change instead of level

    Recompute using the movement between the last two responses for each account. A promoter who fell from 10 to 7 is a different account from one who has been at 7 for two years.

  6. Keep one metric and one question

    Retire the ones that showed under 5 points of lift. Running three surveys to hedge produces survey fatigue and three weak signals instead of one usable one.

Worked example

180 accounts closed 62 renewals over four quarters: 54 renewed, 5 churned, 3 contracted. Of the 62, only 29 had an NPS response inside the window. Top-box NPS responders renewed at 93%; everyone else renewed at 84%, a lift of 9 points. CSAT, triggered on support tickets, gave 41 accounts a response and a lift of 3 points, because almost every response was a 4 or 5. The renewal-likelihood question, asked of the named decider in the 28 renewals where a decider was recorded, separated by 31 points. Decider coverage across the NPS responses was 21%. These figures are illustrative; run the test on your own closed renewals.

How do you balance getting enough feedback to improve the product without annoying users or causing survey fatigue?
r/CustomerSuccess, 2025

Which score should a small customer success team run, and when?

A team under ten people should run one relationship survey and one transactional survey, and no more. The relationship survey exists to catch a change in how the account feels about you; the transactional one exists to fix a process. Running all three of CES, CSAT and NPS across 200 accounts produces survey fatigue and three numbers nobody trusts, which is the state most teams describe when they arrive at this question.

Which survey to fire at which moment, and what to do with a poor answer. Ordered by where the moment sits in the customer lifecycle.
MomentScore to runSend it toIf the answer is poor
End of onboardingCES on the setup taskThe admin who did the configurationFix the step they named before the next cohort starts onboarding
After a support ticket closesCSATThe requester onlyRoute to support triage; do not move a health score on one ticket
Twice a year, relationship levelNPS or a single satisfaction questionEvery named contact, tracked by roleCompare against the account's own last reading, then book a call on the movement
90 days before renewalRenewal likelihood, 0 to 10The named economic buyer and championOpen a risk, name the blocker, and work it as a save rather than a survey follow-up
This is very helpful to gather CSAT and NPS which are two of our leading indicators to happy successful customers and at risk customers.
Global Head of Customer Success, mid-market SaaS, public G2 review

Where the survey feeds a customer health score, keep it as one input among several and cap its weight. A score with a 40% survey weight inherits the survey's response-rate problem. CSAT as a CSM KPI covers what happens when a survey number becomes a personal target.

What should I ask instead of customer effort score vs CSAT?

Ask the renewal decider about the renewal. The question that names the decision outperforms every proxy in the corpus, because it removes both the respondent problem and the translation step between sentiment and behaviour. One practitioner describes running it 60 to 90 days out as a standing habit.

The question is on a scale of 1 to 10 how likely would you renew if you had to renew right now? If the answer of the customer is anything less than 9, i would consider the customer at risk
r/CustomerSuccess, 2025

Two objections come up, and both are answerable. Asking feels like inviting the customer to reconsider: in practice the decision is already forming, and a 7 gives you 90 days to work the blocker rather than 10 days to discount. It cannot be asked at scale: it can, as a one-question email to two named contacts, which is a smaller ask than the NPS campaign you already send. Keep one number for trend reporting and this question for decisions.

Before you put a survey target on anyone

  • The lift test has been run on four quarters of closed renewals, and the result is written down.
  • Decider coverage is above 30%, and the figure is shown next to the score.
  • Any account-level score built on fewer than five responses is suppressed, never displayed.
  • The score is reported as change against the account's own history, with the level beside it.
  • Only one relationship survey and one transactional survey are running.
  • The survey weight inside any health score is capped and stated.
  • Nobody can edit a response, and nobody chooses who receives the survey.
  • The renewal-likelihood question is asked of the decider at 90 days, separately from the score.
How often each metric appears in the two corpora behind this page, with what the count implies. Ordered by frequency in the review corpus, most common first.
MetricG2 reviews of 4,978ShareReddit posts of 33,600Reading
NPS2174.4%100The default relationship survey, and the one most often described as unreliable
CSAT491.0%95Used transactionally, mostly beside support; rarely load-bearing for renewal
CES10.02%1Almost absent from B2B SaaS customer success despite the strongest published loyalty claim

That last row is the finding worth carrying away. Searching 33,600 posts from r/CustomerSuccess, r/SaaS, r/sales and r/startups published between May 2024 and September 2026, the phrase "effort score" appears zero times, and exactly one post uses CES to mean the customer effort score at all.

Currently, the platform has NPS, CES, and CSAT as native metrics, while other analyses need to be performed using specific account indicators
r/CustomerSuccess, 2025

How does GainTrace read survey scores against renewals?

GainTrace treats a survey response as one signal beside billing, product usage, support and contact activity, and weights it by who answered and how recently. Because renewals and churn flow through the same system, the lift test on this page runs continuously instead of once a year, so you can see whether your NPS or CSAT is carrying renewal information across your accounts. Customer health shows the signals behind each account, and churn prediction ranks the accounts most likely to move. Survey scores stay inputs, never the verdict.

Frequently asked questions

Should we run CES, CSAT or NPS?

Run one relationship survey and one transactional survey. CSAT after support tickets and onboarding tasks, because it points at a process you can fix. A twice-yearly relationship question, NPS or otherwise, for trend. Then ask the named renewal decider how likely they are to renew, 90 days out, and treat that answer as the one with decision value. Three concurrent survey programmes produce fatigue and three weak signals.

Which customer satisfaction metric predicts churn in B2B SaaS?

None of them reliably, on published evidence. CES was validated on contact-centre interactions in 2010, NPS on company growth rates in 2003, and CSAT has no founding study. All three measure one person at one moment, while a B2B renewal is a committee decision. Measure the lift on your own closed renewals rather than trusting a claim from an adjacent population.

How is customer effort score calculated?

Add the effort ratings and divide by the number of responses. The usual question asks how much effort the customer had to put in to handle their request, on a seven-point agreement scale where a higher score means less effort. No scale is standard: five-point, seven-point and reversed versions all circulate, so record which one you use and never compare your number with a published one.

Our NPS is high but customers are leaving, what are we missing?

Almost certainly the respondents. A high NPS built on end users tells you the product is pleasant to use; it says nothing about whether the budget holder can defend the line item. Compute your decider gap: the share of responses from named economic buyers and champions. Under 30%, the score is measuring a population that does not take the renewal decision.

What is a good CSAT response rate?

No credible B2B SaaS benchmark exists for survey response rates, and the vendor figures in circulation have no disclosed method. Judge your own rate against your own history and against coverage: what matters is whether enough named deciders answered to support the conclusion you are drawing, not whether you beat a published average.

Can a survey score be a fair KPI for a CSM?

Only when the CSM controls the thing being measured and the sample is large enough to be stable. A CSAT score on support tickets fails the first test; an account-level NPS built on six responses fails the second. If a survey number sits on a scorecard, publish the response count beside it and the lift test that justifies it, or move the target to the outcome the survey was meant to predict.

How this was researched

We searched 29,027 sentences from 4,978 public G2 reviews of five customer success platforms and 33,600 posts from r/CustomerSuccess, r/SaaS, r/sales and r/startups, published May 2024 to September 2026. Mentions were counted at review and post level: NPS in 217 reviews and 100 posts, CSAT in 49 reviews and 95 posts, CES in 1 review and 1 post, counting only uses that mean the customer effort score: the raw three-letter token also matches a trade show and French text, which we excluded by reading each hit. The phrase "effort score" appears in no post. Origin claims are taken from the published abstracts of the two Harvard Business Review articles named in the sources, and the NPS benchmark band from Userpilot's 2024 first-party report across 547 SaaS companies, which does not state whether its figures are means or medians. The decider gap, the lift test and the survey-by-moment table are our own analysis; the worked example uses illustrative figures.

Next steps

Run the lift test on last year's renewals, then keep one survey and one question instead of three programmes. Start free or book a demo.

See GainTrace first in your Google results

Add as a preferred
source on Google
View markdown