Skip to content

When 90% of accounts are healthy and 9% of revenue still leaves

Why Is Our Customer Health Score Distribution Mostly Green?

A customer health score distribution that is 90% green is a threshold artefact. The three shapes a distribution takes, and the red budget that sets your cut points.

By , Co-founder, GainTrace · Updated · 15 min read · For CS Operations, Head of Customer Success

Short answer

A customer health score distribution that is 90% green is usually a threshold artefact, not a healthy account base. Set the red share from your own churn: if 3% of accounts give notice in any six-month window and half of your red calls are right, the red budget is near 6%. Then check the colour lift, which is the churn rate among red accounts divided by the rate among green ones.

A customer health score distribution that comes back 90% green is the moment most teams stop trusting the score. Someone builds the pie chart for the board pack, sees 412 green, 31 yellow and 14 red, and remembers that the company lost 22 customers last year. The colours and the outcomes are describing different businesses, and nobody in the room can say which one to believe.

This page is for CS Operations and heads of customer success who have a score running and want to know whether its output is shaped correctly. It gives the three shapes a distribution takes and what each one means, the arithmetic that sets the red, yellow and green cut points, the two tests that tell you whether the colours separate, and what to report instead of the pie chart.

Key takeaways
  • A green-heavy distribution is a statement about where the cut points sit, not about the accounts. Median gross revenue retention for private B2B SaaS was 91% in 2025, so a score claiming 95% of accounts are fine is claiming to beat the market.
  • Set the red share from the forward churn rate divided by your target precision, then cap it at what the team can work in a quarter. That number is the red budget.
  • Colour lift is the one-line test: churn among red accounts divided by churn among green accounts. Below 2 the colours are decoration; 5 or more means the score separates.
  • A distribution with no spread cannot be fixed by moving thresholds. If the middle half of your accounts sit within 10 points of each other, the inputs do not discriminate and the inputs are what to change.
  • Report movement between colours, not the static pie. Thirty accounts that slid from green to yellow this month is a work queue; 90% green is a screenshot.
Browse this guide

Questions this page answers

  • 90% of our accounts are green. Is that right?
  • What percentage of accounts should be red in a health score?
  • How do I set health score thresholds for red, yellow and green?
  • Our health score distribution never changes. What does that mean?
  • Should our health score distribution match our churn rate?
  • How do I know if our health score thresholds are calibrated?
  • What should a healthy customer health score distribution look like?

Is a customer health score distribution that is 90% green normal?

The red budget

The red budget is the share of accounts a score is allowed to mark red: the forward churn rate over the score's horizon divided by the precision you expect from a red call, capped by the number of accounts a CSM can work in a quarter. A red share below the forward churn rate guarantees misses; one far above it guarantees the colour gets ignored.

A customer health score distribution that is 90% green is common, and it is usually evidence about the thresholds, not about the customers. Start with the outcome you already know. SaaS Capital's September 2025 retention brief, a survey of more than 1,000 private B2B SaaS companies, puts median gross revenue retention at 91%, and High Alpha's 2025 report puts it between 88% and 92% by ARR band. Nine points of revenue leaves the median company every year. Because gross retention is revenue-weighted, an SMB-heavy account base loses a higher share of logos than that. A score saying 93% of accounts are healthy is asserting an outcome better than the market median, which might be true and should be checked.

My team is struggling with "Green Churn" ... customers look fine on paper but cancel anyway.
r/CustomerSuccess, 2026

Green churn is what a miscalibrated distribution feels like from the inside. Health scores are heavily discussed in our review corpus, with 459 of 4,978 public customer success platform reviews mentioning one (9.2%) and 83 mentioning a scorecard (1.7%), and the colour system comes up constantly while the shape of its output almost never does. Across 33,600 Reddit posts from May 2024 to September 2026, 106 discuss health scores and the recurring complaint is the same: the colours are there, the separation is not.

You can have a "green" account that churn or a "red" account that renews and expands, because the score isn't tied to real outcomes.
r/CustomerSuccess, 2026

Which three shapes does a customer health score distribution take?

Three distribution shapes cover almost every account base we have looked at, and each one has a different cause and a different fix. Find your shape first, because the work that follows is not the same. The shares below are the pattern, not a rule to copy.

The three shapes a customer health score distribution takes, what causes each one, and the first thing to change. Ordered from the most common to the least.
ShapeWhat it looks likeCauseFirst fix
Green-heavy85% to 95% green, a thin yellow band, a handful of red that never changesCut points set to make the dashboard look calm, or inputs everyone passes, or manual sentiment that drifts upwardSet the red budget from your own forward churn rate, then move the cut points to meet it
Forced thirdsRoughly a third in each colour, month after month, regardless of what is happening in the accountsThresholds set on percentiles, so a fixed share is red by construction even in a good quarterSwitch from percentile cut points to absolute ones tied to observed churn, and let the shares move
CalibratedRed near the red budget, yellow as a working queue, shares that move month to monthCut points derived from outcomes and capacity, with inputs that vary across the accountsNothing. Watch the transitions and re-derive the cut points every two quarters

Forced thirds is the shape people mistake for rigour. A percentile-based cut point always produces the same distribution, so the chart cannot tell you the accounts got better or worse, and CSMs learn within a quarter that being red means sitting in the bottom third of a ranking, not being in trouble. The percentile version is defensible only for ranking a work queue, and it should never be shown to a board as a health distribution.

Health scores can become theater if they only label accounts. Green, yellow, red. Useful, but incomplete.
r/CustomerSuccess, 2026

How do I set the red, yellow and green cut points?

Derive the cut points from two numbers you already have: how many accounts leave in a window, and how many accounts a CSM can work in a quarter. The red budget sets the ceiling on red, capacity sets the floor on usefulness, and the score's own score distribution decides where the line lands to produce that share.

The red budget

Red budget = Forward churn rate over the score horizon ÷ Target precision

Forward churn rate
the share of accounts that gave notice within the score's horizon historically, usually two quarters. Count notices, not contract end dates
Target precision
the share of red accounts that turn out to be at risk. Start at 0.5; a score that has never been backtested will be worse
What good looks like
a red share between one and three times the forward churn rate. Below that the score cannot catch the churn that exists; above it, CSMs stop working the colour

Worked example

460 accounts lose 28 a year, so the forward churn rate over two quarters is about 3%. At a target precision of 0.5, the red budget is 6%, or 28 accounts red at any time. Six CSMs can each run a serious intervention on about five accounts a quarter, which is 30, so capacity and budget agree and 6% is the cut. Yellow gets the next 14%, around 64 accounts, which is the watch list a CSM reviews rather than works. The cut point on the score itself is wherever the 6th and 20th percentile of scores fall, recomputed every two quarters. These figures are illustrative; run the arithmetic on your own accounts.

The four inputs that set cut points, where each one comes from, and how often to refresh it.
InputSourceRefresh
Forward churn rateNotice dates from the last eight quarters, counted over the score's horizonEvery two quarters
Target precisionThe last backtest: red accounts that went on to churn or contract, divided by all red accountsEvery backtest
Intervention capacityCSM count multiplied by serious interventions per CSM per quarter, which is usually four to sixWhen headcount or coverage changes
Score spreadThe percentile table of current scores, which decides where a given share of the accounts fallsMonthly

Capacity is the constraint people leave out, and it is the one that decides whether the distribution gets used. A score that marks 120 accounts red for a team of four turns red into background noise within a month. How many accounts per CSM is too many covers the coverage maths behind the capacity number.

Why does our customer health score distribution have no spread?

A distribution with no spread means the inputs do not discriminate, and no threshold can rescue it. Check it in one query: take the score for every account and find the 25th and 75th percentile. If the middle half of the accounts sit within about 10 points on a 100-point scale, most accounts are scoring on inputs that every account passes, and moving the red line picks a different arbitrary set of accounts to worry about.

There needs to be more finite control over thresholds and assigning weighted values to criteria to accurately reflect customer sentiment.
Senior Sales Operations Specialist, enterprise SaaS, public G2 review

Variance collapse has three usual causes. Inputs that are nearly always true, like having an active contract, an admin user or any login in 90 days, contribute the same points to everyone. Weighted averages of many inputs pull every account toward the middle, which is what averaging does. And manual sentiment fields drift upward over time, compressing the top of the range. Is health score coverage a KPI worth chasing covers the manual-drift problem specifically.

Colour lift

Colour lift = Churn rate among red accounts ÷ Churn rate among green accounts

Churn rate among red accounts
of accounts red at a point in time, the share that churned or contracted within the score's horizon
Churn rate among green accounts
the same measure for accounts that were green at that same point in time
What good looks like
5 or more. Between 2 and 5 the score is weakly informative; below 2 the colours carry no information and should not drive a CTA

Colour lift is worth computing before any redesign, because it tells you whether the problem is the thresholds or the inputs. A score with good lift and a bad distribution needs new cut points, which is an afternoon. A score with lift below 2 needs new inputs, which is a project. Why is my customer health score accuracy so poor covers the nine input-level failures that produce low lift.

How do I check the score distribution against real churn?

Join a past distribution to what happened next, then read the churn rate by score decile. The check needs a historical snapshot of scores, a list of churn and contraction events with notice dates, and a spreadsheet. Running it on today's distribution tells you nothing, because today has no outcomes yet.

  1. Take the score distribution as it stood two quarters ago

    One row per account with its score and colour on that date. If the platform cannot show a past date, rebuild scores from stored inputs, and if that is also impossible, start snapshotting monthly from today and come back in two quarters.

  2. Join the outcome that followed

    Churn, contraction above 20%, or neither, using the notice date and not the contract end date. An account that gave notice in month five counts as an event even if its contract ran on.

  3. Bucket by score decile and count

    Ten buckets, each with its account count and its observed event rate. A calibrated score produces rates that fall as the decile rises, without reversals in the middle.

  4. Compute colour lift and the red budget

    Divide the red event rate by the green event rate, then compare the red share you ran with the forward churn rate you observed. Two numbers, and they tell you which of the two fixes you need.

  5. Move the cut points, then leave them alone for two quarters

    Set red at the budget, yellow at the next band the team can review weekly, and hold. Cut points moved every month cannot be evaluated, because each change resets the outcome data.

  6. Publish the calibration table beside the distribution

    Ten rows, event rate by decile, dated. A distribution shown with its calibration gets trusted for what it is. A pie chart on its own gets argued with.

An illustrative calibration table: churn and contraction rate by score decile over the two quarters after the snapshot, lowest decile first.
Score decileAccountsEvent rate in next 2 quarters
1 (lowest scores)4619.6%
24610.9%
3468.7%
4464.3%
5464.3%
6462.2%
7462.2%
8462.2%
9460.0%
10 (highest scores)462.2%

Read that table the way a calibration check is meant to be read. Event rates fall from 19.6% to near zero across the deciles, which is a score that separates. The reversal in the top decile is small enough to be noise at 46 accounts per bucket. The first two deciles hold 14 of the 26 events, 54% of them, in 20% of the accounts. A red budget of 6% reaches about a fifth of the events where a 20% red band would reach 54%, which is the capacity trade-off made explicit. These figures are illustrative; run the join on your own snapshot. The customer success scorecard is a starting point if you have no historical scores to join yet.

I do think the customer health scores can be misleading, showing a customer as green when sometimes the usage data doesn't reflect that.
Senior Director, enterprise SaaS, public G2 review

Should I report the health score distribution or the movement between colours?

Report the movement, and keep the distribution as context underneath it. A static distribution answers how the accounts look, which changes slowly and tells nobody what to do. A transition table answers what changed since last month, which is the work queue and the story a leadership meeting can act on.

An illustrative monthly transition table: accounts by the colour they held last month and the colour they hold now. Rows are last month, columns are now.
Last monthNow greenNow yellowNow redChurned
Green (392)3582941
Yellow (51)182481
Red (26)39122

Three numbers in that table matter more than the whole pie chart. Twenty-nine accounts slid from green to yellow, which is this month's proactive work. Four went green straight to red without passing through yellow, which means the score is reacting to an event instead of a trend, and those four are worth reading individually. One account churned straight out of green, and that single row is the start of the list you review every quarter. How do I backtest customer health score logic covers the fuller version of that review.

For example, I had a green account that suddenly turned red, then back to green the next day.
Customer Success Manager, mid-market SaaS, public G2 review

Before you publish the distribution

  • The red share is inside the red budget and inside what the team can work this quarter.
  • Colour lift from the last calibration is printed next to the chart, with its date.
  • Cut points are absolute and tied to outcomes, not percentiles of the current account base.
  • The calibration table by decile is available on the same page.
  • Accounts under 90 days old are excluded or scored separately, so onboarding does not colour the account base.
  • Month-on-month transitions are shown above the static distribution.
  • Single-day flips are smoothed, so one event cannot move an account two colours and back.
  • The date of the next cut-point review is in the calendar, two quarters out.

A customer health score distribution is a summary of a prediction, and a prediction that nobody scored against outcomes is a colour scheme. The teams who get value from the chart are the ones who can say what happened to the accounts in each band last time, which takes one join and about an hour.

How does GainTrace shape the customer health score distribution?

GainTrace scores accounts on change against their own baselines, so the distribution spreads out instead of piling up at the top, and it keeps the history that makes a calibration table possible without a snapshot project. Cut points are derived from your own churn and your own capacity rather than set by hand. Health signals show the distribution with the movement between bands, and churn prediction shows which accounts moved and what drove each call.

Frequently asked questions

90% of our accounts are green. Is that right?

Probably not, and the outcome data is how you check. Median gross revenue retention for private B2B SaaS was 91% in 2025, and logo churn in an SMB-weighted account base runs higher than the revenue number implies. A score claiming 90% of accounts are fine is claiming to beat that. Compute the churn rate among your green accounts: if it is close to your overall average, the colour carries no information.

What percentage of accounts should be red in a health score?

Between one and three times your forward churn rate over the score's horizon. If 3% of accounts give notice in a typical six-month window and half your red calls are right, the red budget is about 6%. Then cap it at what the team can work: a red list larger than the number of interventions a CSM can run in a quarter gets ignored within a month.

Should health score thresholds be percentiles or absolute values?

Absolute, tied to observed churn. Percentile cut points force the same share into each colour every month, so the chart cannot show the accounts improving and being red means sitting in the bottom of a ranking. Use percentiles only to order a work queue, never to report a distribution to a board, and recompute absolute cut points every two quarters.

How do I know if our health score distribution is calibrated?

Bucket accounts by score decile as they stood two quarters ago, then count churn and contraction events in each bucket. A calibrated score produces event rates that fall steadily as the decile rises, without large reversals. Then compute colour lift: churn among red divided by churn among green. Five or more is separating, below two carries no information.

Our distribution never changes month to month. What does that mean?

Either the cut points are percentiles, which force a fixed share into each colour, or the inputs have no variance. Check the second by finding the 25th and 75th percentile of scores: if the middle half of the accounts sit within 10 points, the score is built on inputs that nearly every account passes and new thresholds will not help.

Should new customers be in the distribution at all?

Score them separately for the first 90 days. A usage-weighted score marks every new account red before it has had a chance to use anything, which drags the distribution and teaches CSMs to ignore red. Run an onboarding score built on milestones reached, then move the account into the main distribution at first value rather than on a calendar date.

How this was researched

We searched a corpus of 4,978 public G2 reviews of five customer success platforms, 29,027 sentences, for mentions of health scores (459 reviews, 9.2%), scorecards (83 reviews, 1.7%), colour banding and thresholds, and read every complaint about the output being uninformative. We then searched 33,600 posts from r/CustomerSuccess, r/SaaS, r/sales and r/startups published between May 2024 and September 2026, of which 106 discuss health scores. Retention figures come from SaaS Capital Research Brief 32 (September 2025, more than 1,000 private B2B SaaS companies, self-selected survey, medians) and High Alpha's 2025 SaaS Benchmarks (more than 800 respondents, self-selected). The red budget, colour lift, the three-shape taxonomy and the calibration procedure are our own analysis; the calibration table, the transition table and the worked example use illustrative figures and are labelled as such.

Next steps

Join last quarter's distribution to what happened next, compute colour lift, and move the cut points once. Start free or book a demo.

See GainTrace first in your Google results

Add as a preferred
source on Google
View markdown