Back to all posts

Weighted QA Scorecards: What the Weights Actually Predict

A weighted QA scorecard gives each criterion a share of the total score. Greeting might be worth 5 points, tone 10, compliance 20, resolution 30. The weights are usually set once, in a workshop, by the people who run the QA program, and they record what those people believe matters to the customer. That belief is rarely tested.

What a weight claims

Every weight on a scorecard is a prediction. Giving resolution 30 points and tone 10 says that resolution matters three times as much as tone to the outcome the team cares about, whether that is a satisfied customer or a renewed contract. Nobody writes that sentence down, but a conversation scored 85 out of 100 carries every one of those predictions inside the number.

The weights are the only theory of the customer the scorecard contains, and the only place that theory is written down. The rest of the document is a checklist.

The evidence that weights go untested

COPC Inc., which publishes a widely used contact-center operations standard, surveyed more than 900 customer-experience executives between September and December 2021 for its 2022 Global Benchmarking Series on quality assurance. Its own standard requires organizations to demonstrate, attribute by attribute, the relationship between customer-critical accuracy and measured customer experience. In the survey, 22 percent of organizations said they do not analyze the relationship between QA data and customer satisfaction at all, and a further 3 percent did not know. The same report found that organizations report first-contact-resolution rates well above what their customers report for the same interactions. The internal score and the customer's experience are measured separately and rarely reconciled.

SQM Group, a contact-center benchmarking firm, reported in February 2023 that across its quality-assurance case studies, 81 percent of agents' QA scores did not correlate with customer satisfaction. SQM's explanation was that traditional scorecards use the wrong metrics and the wrong weighting, and that they judge the customer's experience without asking the customer. The same research found that only 17 percent of agents believed QA monitoring improved satisfaction at all. These are findings from a firm that sells QA software, so they are a warning rather than a measurement of the whole market. They still describe a failure most QA leads will recognize.

The mechanism has been documented for years. In 2016 COPC published a worked example of a program reporting 90 percent overall quality while scoring 70 percent on the attributes that mattered most to customers. The gap came from criteria such as correct openings and hold procedures, which agents nearly always pass and which inflate the total. COPC's advice then, and the requirement in its standard now, is to report customer-critical, business-critical and compliance-critical accuracy as separate metrics instead of one rolled-up score.

Why the weights never get tested

Testing a weight requires two things the traditional program lacks. The first is enough scored conversations to see a pattern. COPC's 2022 survey found that 88 percent of organizations monitor each agent at least once a month, but a monthly sample of a few conversations per agent, spread across every criterion and every customer, produces almost no signal per criterion. A handful of calls is enough to coach one agent and nowhere near enough to test a weight.

The second is an outcome to compare the score with. A QA score is recorded at the level of the conversation. The outcomes that matter, a renewal, a downgrade, an escalation, a champion who stopped replying, are recorded at the level of the account, in a different system, weeks or months later. Nobody joins the two, so the weights are never confronted with what happened next.

How to test a scorecard against what the account did

The test is a comparison of two groups of accounts. Take every account that churned or downgraded in the last two quarters. Pull the QA scores on that account's conversations from the six months before the event. Then do the same for a matched set of accounts that renewed cleanly. Look at each criterion separately. The total will hide what the criteria show.

Each criterion lands in one of three places. It scored differently between the two groups, in which case its weight is earning its place. It scored the same in both groups, in which case its weight is noise and it is inflating everyone's total. Or it was rarely scored at all, because the sample never reached the conversations where it mattered, in which case the program has no evidence either way.

COPC's worked example and SQM's finding point at the same suspect for the second group: process and compliance items that nearly everyone passes. The criteria that carry information about the customer are the ones that are hardest to score consistently, such as whether the customer's actual question was answered, or whether a commitment made in one thread was kept in the next.

What scoring every conversation changes

Coverage removes the first obstacle. When every conversation is scored against the same scorecard, the sample is the population, and each criterion has enough observations to compare. This is true of automated scoring in general. It is why Resonant IQ scores every conversation as it arrives, with qualitative criteria, timing thresholds and customer feedback normalised to one scale.

The second obstacle goes away once a score sits on the account's timeline next to the outcome. Resonant IQ ties every signal back to the specific conversation and timestamp that produced it, at the account level, so a QA score sits beside the escalation and the month the champion went quiet. Both inputs to the comparison above are then in one place, which is the part that a quarterly spreadsheet exercise never manages.

What this does not settle

A correlation between a criterion and churn does not prove the criterion caused it, and a scorecard tuned to last year's churn will drift as the product and the customer base change. The comparison should be rerun each quarter, and the weights should be treated as a hypothesis with a date on it. Teams that sample a handful of calls a month can still run it, but they should expect the answer to be "not enough evidence" for most criteria. That answer is a reason to change the sampling before changing the weights.

Sources

Stop guessing which accounts are slipping.

Join the founding cohort and lock your rate for 24 months while we build the evidence layer with you.