Customer health score accuracy fails for one reason: the score measures level rather than change and was never tested against churn. In 3,628 public reviews of the three most-reviewed customer success platforms, 347 mention a health or churn score, and the complaints repeat: gamed inputs, new customers scoring red, parent and child accounts averaged, data a day stale. Backtest against four quarters of churn before you touch a weight.
Last quarter three accounts you had marked healthy cancelled, two accounts that had sat red for months renewed without a conversation, and customer health score accuracy became the agenda item nobody wanted. Someone senior has now asked how it is possible to lose customers you said were fine. The score has weights, colours, a dashboard and a CTA that fires when it drops, and none of that predicted anything.
This page is for the CS Ops lead or Head of CS whose score is running and not working. It names the nine ways that happens, shows which four cause most of the damage with the evidence from practitioners who hit them, and gives an afternoon's backtest that tells you exactly how wrong your score is before you change a single weight.
- Backtest before you reweight. Pull every account that churned in the last four quarters, look up its score 90 days before, and count how many were green. That number is your score's real accuracy.
- Remove every input a CSM can type. Sentiment fields and manual overrides are where gaming starts, and the corpus has reviewers admitting it.
- Score change against the account's own baseline, not level against the book. A daily user who drops to weekly is red even if the dashboard says active.
- Give new customers their own score for the first 90 days, or exclude them. Any usage-weighted score marks them red before they have had a chance to use anything.
- One score per product and per child account, rolled up, never averaged. An average of one healthy and one dying subsidiary is a yellow that means nothing.
Questions this page answers
- Our health score said green and they churned. What went wrong?
- Why isn't our health score actually predicting churn?
- How do I stop CSMs gaming the health score?
- Why do new customers always show as high churn risk in our health score?
- How do I test whether our customer health score is accurate?
- Should we track health score trend instead of the score itself?
- The account was green and they churned: what happened?
- Which nine failure modes break a health score, and what fixes each?
- Which four failure modes hurt customer health score accuracy most?
- How do I test the score against real churn in an afternoon?
- What should I change first?
- When can a rules-based score no longer be fixed?
- How does GainTrace score health differently?
The account was green and they churned: what happened?
The green list is every account that churned while its health score said green, written down with the inputs it had 90 days out. Teams argue about score weights in the abstract for months. The green list ends the argument because each row names an input that was missing, stale, gamed or measured on the wrong unit.
Most health scores we have seen are a weighted average of whatever data was easy to get, tuned until the colours looked plausible on a dashboard, and never once checked against the accounts that left. A customer health score is a prediction. A prediction that has never been scored against outcomes is a guess with a colour.
“Some customers look "healthy" right up until they leave, while others get flagged even though they renew. Our score is based on things like product usage, support tickets, NPS, and engagement, but the signals don't always line up.”
We read 3,628 public G2 reviews of the three most-reviewed customer success platforms. 347 of them mention a health score or churn score. Reviewers who bought the platform for a score they could trust describe the same handful of failures: someone gamed it, new customers came out red, the parent account averaged its children into a meaningless yellow, the data was a day old, or the score was one number for three products. None of these is a weighting problem. Reweighting a score with a structural fault produces a different wrong answer.
Underneath all nine failure modes are two design mistakes. The score measures level (how much usage, how many tickets) when churn shows up as change (less than before, different from before). And the score has no feedback loop, so nothing about a churn ever changes the score that missed it.
Which nine failure modes break a health score, and what fixes each?
Each row is a pattern that appears in the review corpus, the Reddit threads, or both. The symptom column is how you notice it in your own book. Find your rows first; most scores have three or four.
| Failure mode | Symptom | Fix |
|---|---|---|
| No outcome data | Nobody can say what share of last year's churn was red 90 days before. Weights were set by opinion and never revisited. | Run the backtest below every quarter. Retire any input whose presence does not separate churned from renewed accounts. |
| Level, not trend | Accounts stay green while declining, then drop to red the week they cancel. A low-but-flat account sits red for a year and renews. | Replace each level input with change against the account's own 90-day baseline. Add a trend field: improving, flat, deteriorating. |
| Gamed or subjective inputs | The sentiment field is always positive. Overrides cluster before pipeline reviews. Red accounts turn green with no change in usage. | Remove every input a CSM types. Keep manual sentiment as a separate note that never enters the number. |
| Stale sync | The CSM finds out about a churn, an upgrade or a closed ticket before the score does. Alerts fire a day after the event. | Know the refresh interval of every input and show it on the score. Anything slower than daily is a report, not a signal. |
| New customers score red | Every account under 90 days old is high risk. CSMs learn to ignore red in the first quarter, then ignore it everywhere. | A separate onboarding score built on milestones reached, not usage volume. Switch to the main score at 90 days or first value. |
| One score for many products | A customer using product A heavily and product B not at all scores yellow, so the product B churn is invisible. | One score per product subscription, rolled up by ARR at risk. Never averaged. |
| Parent and child averaging | The parent shows healthy while a child site has gone dark. Or one big healthy child hides three small dying ones. | Score each child. Roll up as ARR-weighted risk and show the worst child on the parent record. |
| Recency dominates | One good call resets a year of decline to green. One angry ticket turns a healthy account red for a week. | Smooth every input over 30 to 90 days. Let single events raise a flag, never move the score alone. |
| Contact-level signals missing | The champion left three months ago and the account is still green because logins from other users held up. | Track the champion and sponsor as their own inputs: role change, reply latency, meeting attendance. |
Which four failure modes hurt customer health score accuracy most?
Four of the nine appear so often in the evidence, and do so much harm, that they deserve their own explanation.
Why do gamed and subjective inputs break the score?
A sentiment field or a manual override is an invitation. The moment a score feeds a pipeline review, a comp plan or a CTA queue, the people who can type into it will, and the corpus has reviewers saying so.
“The way we have [the platform] configured allows some users to "game" the system and represent false health scores. This leads to inaccurate reporting and misrepresentation of the health of the overall portfolio.”
“Each of us has a different understanding of what "healthy" means. We can not go on this way, because many tasks are marked as completed just by adjusting the customer health scores based on personal judgment.”
The fix is blunt: nothing typed enters the number. Keep the CSM's read of the relationship as a note beside the score, where it is useful context and cannot move a colour.
Why do new customers always score red?
A score weighted on usage volume compares a customer in week two against customers in year three and finds them wanting. CSMs quickly learn that red under 90 days means nothing, and that lesson spreads to the rest of the book.
“Newly onboarded customers will always have very high churn scores, as they have not interacted with the platform as much, which can be misleading.”
Give the first 90 days a separate score built on milestones (admin configured, first integration live, first report shared) and hand over to the main score at first value, not on a calendar date. Onboarding complete but the customer churned covers how to define that milestone.
Why do parent and child accounts score wrong?
Any score that averages children into a parent will hide a dying subsidiary behind a healthy one. Six reviews in the corpus describe exactly this, and none of them found a setting that fixed it.
“Setup can be daunting and there are some weird quirks with the parent/child accounts that lead to the inability to trust some of the big picture data. On a granular level it works out, but it would be nice for the parent churn score to be more reflective of active child accounts.”
Score every child. Roll up to the parent as ARR at risk (the sum of child ARR in red or deteriorating), and show the worst child on the parent record. An average is the one aggregation that guarantees the parent never looks as bad as its worst part.
Why does a stale sync make the score lie?
100 of the 3,628 reviews mention sync, and 65 ask for real-time data. The pattern is a score that refreshes once a day from a CRM that itself syncs once a day, so the CSM hears about the event from the customer, then watches the score catch up.
“Wish it was instantaneous, as sometimes by the time [the platform] alerts me, the customer has already upgraded.”
Print the refresh interval of each input next to the score. If billing is live, support is hourly and usage is weekly, the score is weekly and everyone should know it. A score that presents day-old data as current is worse than no score, because people trust it.
How do I test the score against real churn in an afternoon?
Score last year's churned accounts as they looked 90 days before they left, then count how many the score called red. The test needs a churn list, a way to see historical scores (or a rebuild of the score from historical inputs), and a spreadsheet. Almost nobody runs it, and it is the only step that tells you whether the score is worth fixing.
List every account that churned or contracted in the last four quarters
Include downgrades and seat reductions above 20%. Add the same number of accounts that renewed flat, chosen at random, as the control group.
Look up each account's score 90 days before the event
If your platform cannot show a past date, rebuild the score from the inputs as they stood then. If you cannot do either, that is failure mode one and the rest of the test is moot until you can.
Count the misses
How many churned accounts were green at 90 days. That share is your miss rate. How many renewed accounts were red. That is your false alarm rate.
Test each input on its own
For every input in the score, compare its 90-day value across the churned group and the control group. An input that looks the same in both groups predicts nothing and should leave the score, whatever its weight.
Repeat with change instead of level
Recompute each surviving input as change against the account's own prior 90 days and run the comparison again. Expect the change version to separate the groups better; the practitioners in our corpus who ran this comparison report the same thing, with login decay against the customer's own baseline as the usual example.
Write the result on the dashboard
"This score flagged 6 of 14 churned accounts at 90 days last year." A score that shows its own accuracy gets trusted for what it is and ignored for what it is not.
Worked example
A 220-account book lost 14 accounts and had 6 contractions over four quarters, 20 events in total. At 90 days before the event, 11 were green, 5 yellow and 4 red: a miss rate of 55%. Of 20 randomly chosen renewed accounts, 7 were red at the same point: a false alarm rate of 35%. Testing inputs one at a time, login volume was almost identical across both groups. Inbound support volume was lower in the churned group, but only when measured as change: churned accounts averaged 40% fewer tickets than their own prior quarter, renewed accounts were flat. Champion role change appeared in 9 of the 20 churn events and 1 of the 20 renewals. It had not been in the score at all. These figures are illustrative; run the test on your own book.
Miss rate = Churned accounts that were green 90 days out ÷ All churned accounts × 100
- Green 90 days out
- the score the account carried three months before it gave notice, not the score after the bad news arrived
- What good looks like
- under 30%. Above 50% the score is describing the past, and no change to the weights will fix an input problem
What should I change first?
Change the inputs before the weights, in the order below, and rerun the backtest after each round. The backtest leaves you with a list of inputs that predicted something and a list that did not. Changing everything at once tells you nothing about which change moved the miss rate.
- Remove every input a person types. This is the fastest accuracy gain and the one with the most internal resistance.
- Convert each remaining input from level to change against the account's own baseline, and add a trend field.
- Add the champion and sponsor as inputs: role change, reply latency, meeting attendance.
- Split new customers into their own onboarding score.
- Split by product and by child account, and roll up as ARR at risk.
- Smooth single events over 30 days so a call or a ticket cannot move the colour on its own.
Before you trust the score again
- Every input has a named source system and a known refresh interval, printed on the score.
- No input can be edited by hand.
- Every input passed the churned-versus-renewed comparison in the last backtest.
- The score shows its own miss rate from the last four quarters.
- New customers under 90 days are scored separately.
- Parents show ARR at risk and the worst child, never an average.
- A trend field sits beside the colour, and the CTA fires on deterioration, not on colour.
- The backtest is in the calendar for next quarter.
The seventh item matters more than it looks. Practitioners in our Reddit corpus keep arriving at the same conclusion from different directions: an absolute score answers "is this customer healthy", when the question that changes behaviour is "what has changed in this account, and does the change matter".
“A customer can be healthy today but deteriorating quickly. Another can have a low health score but be improving steadily.”
When can a rules-based score no longer be fixed?
Sometimes the backtest comes back bad after two rounds of changes. The decision rule we would use: if the miss rate at 90 days is still above 50% after you have removed typed inputs and converted to change, the problem is the inputs you have, not the weights. Either you are missing a source (usually the champion, the calendar or billing movements), or the book is too varied for one set of rules.
Missing a source is fixable without software. Early warning signs of churn when your data is scattered shows how to pull the earlier signals from the tools you already have. If the book is varied (three products, five segments, SMB and enterprise on one score) the honest answer is several small scores, and at that point you are choosing between an ops person who maintains them and a model that learns the weights from your own churn history. If you have no score at all yet and are starting from a spreadsheet, health score without a platform is the page to begin with, because the change-based design there avoids most of the nine modes from the start.
How does GainTrace score health differently?
GainTrace does not ask you to pick weights. It connects billing, CRM, product usage and support, scores each account on change against its own baseline, and learns from your own renewals and churn which signals mattered, so the backtest above runs continuously rather than once a year. Churn prediction shows the accounts most likely to move, with the signals that drove the call, and product signals show the champion, usage and support changes behind each one, with the refresh interval visible.
Frequently asked questions
Why did our health score say green right before the customer churned?
How do I stop CSMs gaming the health score?
Why do new customers show as high churn risk in our health score?
How do I test whether our customer health score is accurate?
Should we track health score trend instead of the score itself?
How should a health score handle parent and child accounts?
How this was researched
We read 3,628 public G2 reviews of the three most-reviewed customer success platforms, isolated the 347 that mention a health score or churn score, and grouped every complaint about accuracy by cause; we also counted reviews mentioning sync (100), real-time data (65) and parent-child scoring (6). We then read 1,328 threads from r/CustomerSuccess, r/SaaS, r/sales and r/startups, including 49 about health scores, for how practitioners describe scores failing in their own books. The nine-mode taxonomy, the backtest procedure and the priority order are ours; the worked example uses illustrative figures.
- r/CustomerSuccess: Why isn't our health score actually predicting churn?
- r/CustomerSuccess: Could anyone share some customer health score examples that work?
- r/CustomerSuccess: Is customer health actually the wrong thing to optimise for?
- r/CustomerSuccess: How do you predict churn without a data scientist?
- r/CustomerSuccess: What customer health signal has actually predicted risk for you?
Run the backtest this week, then let the score learn from your own churn instead of your weights. Start free or book a demo.
See GainTrace first in your Google results
Add as a preferredsource on Google