What is a customer success scorecard?
A customer success scorecard turns several separate signals about an account into a single number between 0 and 100. You pick the categories that predict renewal in your business, decide what each is worth, score the account on each category, and take the weighted average. The output is a customer health score that a CSM, an account manager, and a CFO can all read the same way without re-litigating what 'yellow' means.
Every scorecard framework published for CS teams is a weighted mean underneath the branding. That is not a criticism · the weighted mean is the right shape for this problem, because it forces you to state what you believe predicts churn and how strongly. The work is not in the formula. It is in choosing the categories, defending the weights, and calibrating the bands.
What a customer health score should include
Five categories cover most B2B SaaS books. Product usage is the only one that moves before the customer decides anything, which is why it usually earns the largest weight. Relationship strength catches what usage cannot see: the account is using the product fine, but the executive who bought it has left. Support health turns a rising escalation pattern into a number. Commercial signals · late invoices, unused seats, contract term remaining · are the least emotional data on the card. Sentiment belongs there too, but it is sparse and lagging, so it earns the smallest weight.
| Category | What it measures | Typical weight | Leading or lagging |
|---|
| Product usage | Logins, active seats, feature depth, usage trend | 30-40% | Leading |
| Relationship | Exec sponsor, meeting cadence, responsiveness, champion stability | 20-30% | Leading |
| Support health | Ticket volume and severity, escalations, resolution time | 10-20% | Leading |
| Commercial | Payment timeliness, seat utilisation, contract term left | 10-20% | Mixed |
| Sentiment | NPS or CSAT responses, survey participation | 5-15% | Lagging |
Weight ranges are this model's starting guidance, not a published industry standard. Calibrate against your own renewal history.
How to weight a customer health score
Weight by how early and how reliably a signal moves. A category that changes months before a renewal conversation is worth more than one that only changes after the customer has already made up their mind. That single rule explains why usage and relationship dominate most defensible cards and why NPS, despite being the metric executives ask about, earns a small share.
Then check concentration. If one category carries 50% or more of the total, the scorecard is a single metric wearing a scorecard's clothes: it will swing sharply whenever that one input moves and stay blind to risk everywhere else. Spreading weight across at least three or four meaningfully-weighted categories is what makes the composite worth more than its parts.
How to calibrate your score bands
There is no published cross-industry benchmark for what counts as a good customer health score, and any number quoted as one is guesswork. The honest calibration method uses data you already have: score last year's churned accounts and last year's renewals with the same card, then set your thresholds where those two populations actually separate. If they do not separate, the problem is the weights, not the thresholds.
This builder's defaults · 75 and above healthy, 60 and above watch, 40 and above at risk · are stated as the model's own convention so they are easy to argue with. That is deliberate. A band you inherited from a tool and never questioned is worse than one you set yourself from your own renewal history.
Why a manual scorecard stops working
A hand-built scorecard is genuinely useful the first time. It forces the team to agree on what predicts churn, and it produces a number everyone can read. The problem is what happens next: it captures one account at one moment, and the moment passes. A score built in January is describing January.
Scale makes it worse. At 40 accounts a CSM can plausibly refresh the card each quarter. At 150 it stops happening, and the scores quietly go stale while everyone keeps treating them as current · which is more dangerous than having no scores at all. That is the point where the model has to start pulling its own inputs and re-scoring continuously instead of waiting for someone to remember.