Operated function · quality assurance
Most quality programmes have never asked, which means the score being used for coaching, for performance conversations and sometimes for pay may be describing the reviewer rather than the work — and everybody on the floor suspects it already.
A small sample is reviewed against a scorecard. The sample is small because reviewing is slow, so the coverage per person per month is a handful of items out of hundreds — and the conclusions drawn from it are treated as though they described the whole.
The scorecard has criteria that sound objective and are not. "Demonstrated empathy", "followed process", "resolved effectively" — each is a judgement, and each reviewer has a private threshold for it that has never been written down or compared.
Calibration sessions happen occasionally and are the right idea. They surface disagreement, everybody adjusts a little, and the effect decays over the following weeks because nothing measures it continuously.
The people being scored know all of this. They know which reviewer is generous, they know a bad month may be a sampling accident, and that knowledge is what makes the programme feel arbitrary rather than developmental — which is the opposite of what it was built for.
And the fail categories that appear most often are rarely aggregated into anything. A criterion that most people fail most of the time is usually a process or training problem rather than a performance one, and treating it as individual feedback repeats the same coaching conversation across the whole floor.
Quality scores are produced by individual reviewers applying private thresholds, and inter-reviewer agreement is not measured — so nothing distinguishes a score that describes the work from one that describes who reviewed it.
This has to be settled before anything else is worth doing, because coverage, coaching and trend analysis all inherit it. Scaling reviews on an uncalibrated scorecard produces more numbers of the same unreliability, and doing it faster makes it worse rather than better.
Measuring agreement is straightforward and rarely done: give the same items to multiple reviewers without telling them, and compare. The result usually shows agreement is high on the mechanical criteria and low on the judgement ones, which immediately tells you which parts of the scorecard can be scaled and which need rewriting.
Rewriting the judgement criteria into observable ones is the second move. "Demonstrated empathy" cannot be scored consistently; "acknowledged the customer’s stated problem before offering a solution" can. That is a scorecard-design exercise your quality function owns, and it is what makes the whole programme defensible.
Then coverage. Reviewing the mechanical criteria across everything rather than sampling changes what the programme can say — the difference between "your sampled score was low" and "this specific thing happened in this many of your interactions" is the difference between a performance conversation people resent and one they can act on.
What is not delegated is the judgement criteria or the conclusion. A score is not a performance assessment, and deciding what a pattern means for a person is a management act.
Inter-reviewer agreement, measured for the first time — measured by agreement rate per criterion on blind duplicate reviews, against a baseline where it was assumed rather than known.
Which criteria are scoreable at all — measured by criteria separated into mechanical and judgement by measured agreement rather than by how they read.
Coverage on the mechanical criteria — measured by share of the population reviewed, against a sample-based baseline.
Whether a low score is real or a sampling accident — measured by confidence attached to an individual’s score, which a small sample cannot support and full coverage can.
Common fail criteria routed as process rather than performance — measured by criteria failed by most of the population, identified and escalated instead of coached individually.
Whether coaching changed anything — measured by the specific criterion measured before and after an intervention, on the same person, at full coverage.
any performance assessment, any conclusion about a person, and any scoring of criteria that were not shown to be consistently scoreable. Nothing here decides that somebody is underperforming. Where a criterion is a judgement, it stays with your reviewers — and a system scoring it anyway would be manufacturing a number that looks like the others and means less.
Reviews attach to the work where it already is — the ticket, the call record, the case, the deliverable — and results land in the quality tooling your team already uses. Nothing migrates, and no second quality record exists, because two scores for one interaction is an employment-relations problem.
The scorecard is yours. Its criteria, their classification, and any rewrite are owned by your quality function, because a scorecard is a statement about what good work is in your organisation and it cannot be supplied.
Agreement measurement runs continuously rather than as an occasional calibration event, so decay is visible while it is happening rather than at the next session.
A quality score attaches to an individual and can affect their pay, their progression and their standing. That makes this the most employment-sensitive operation in Wave 3 and it is scoped accordingly.
Nothing produces a conclusion about a person. Criteria are scored where they were shown to be consistently scoreable, results go to your managers, and what any pattern means is a management judgement made by somebody accountable for it.
Every score records the criterion, the evidence within the work that it was based on, and the scorecard version. A score somebody cannot see the basis of is a score they cannot contest, and an uncontestable score is the thing that makes a quality programme feel arbitrary.
Where your jurisdiction, works council or collective agreement constrains automated assessment of workers, that constraint governs and it is established before anything is scored rather than discovered afterwards.
Operational access is not permission to train. Recordings and work product belonging to your staff and your customers do not become material improving anything serving another organisation.
This one goes to employee representation before it goes anywhere else. Automated assessment of workers is constrained in several jurisdictions and by many collective agreements, and the constraint is established at scoping rather than discovered when the first score is contested.
HR should see the contestability model: what a person can see about why they scored as they did, and how they challenge it. A programme people cannot interrogate is a programme they will not accept.
Your quality function owns the scorecard and the decision about which criteria remain human-only. That decision should be made on the measured agreement data rather than on preference.
A blind duplicate-review exercise on your existing scorecard — your own reviewers, your own work, nothing scaled and no new score issued — measuring agreement per criterion.
The first phase uses your reviewers rather than replacing them. The same items are reviewed independently without the reviewers knowing there is overlap, and the results are compared per criterion.
What comes out is usually a clean split: high agreement on the mechanical criteria, low agreement on the judgement ones. That result is worth having on its own and it will change how your existing programme is run whether or not anything is delegated — a criterion with low agreement should not be driving a performance conversation today either.
If you continue, scaling begins only on the criteria that met your agreement threshold, with the judgement criteria staying entirely with your reviewers and agreement tracked continuously rather than at occasional calibration.
That is the assumption the first phase tests, and it costs a blind duplicate exercise on work you have already reviewed. Where it holds, you have evidence for a programme that people currently doubt, which is worth having. Where it does not, you have learned that scores being used in performance conversations are partly a function of who reviewed — which is worth knowing regardless of whether anything is ever delegated.
In several jurisdictions and under many collective agreements that is a legal position rather than a preference, and it governs — which is why employee representation is engaged before anything is scored. The scope can be narrowed to aggregate process findings with no individual scoring at all, and that version still surfaces the common fail criteria that should be process fixes. If even that is unacceptable, the agreement measurement alone remains useful to your existing programme.
Correct, and the page says so — a criterion that cannot be scored consistently by two humans should not be scored by anything else. Those stay with your reviewers. What the measurement usually shows is that the scorecard contains more mechanical criteria than anybody assumed, and scaling those frees reviewer time for the judgement ones, which is where reviewer attention is actually worth something.
It changes what you can say to them. "Your sampled score was low" invites an argument about the sample and usually deserves one. "This specific thing happened in this many of your interactions" is a different conversation, and it is the one that can be acted on. If your managers already have that level of evidence, coverage is not your constraint.