Operated function · first line
Most of what arrives is known, and the known part is answered from memory by whoever picked it up — so a customer gets one answer on Tuesday and a colleague gets a subtly different one on Thursday, and neither of them was wrong when it was written.
Sort a week of tickets by subject and the shape is always the same: a long head of a dozen or so questions that recur constantly, and a thin tail of genuinely novel problems. The head is most of the volume and almost none of the difficulty; the tail is almost none of the volume and all of the value.
The head is answered from memory. There is a knowledge base, it was accurate when it was written, and it has drifted — because updating it is nobody’s scheduled work and the person who knows it is stale is the person too busy answering the question to fix it. So the real source of truth is what the experienced agents remember, which is unversioned and leaves when they do.
New agents take months to become useful, and the training cost is paid repeatedly because tenure in this role is short. What they learn is largely the head, which is exactly the part that could be written down.
Escalation is uneven. Some agents escalate early because they are unsure, others hold on too long because escalating feels like failing, and the customer’s experience depends on which they got. Nothing measures which escalations should have happened sooner.
And the last three vendors in this category all sold deflection. Deflection counts a customer who stopped asking as a success, which is true when the answer was good and false when they gave up — and the metric cannot tell those apart.
The recurring questions are answered from individual memory rather than from a maintained source, so answers drift by person and by week — and no amount of capacity fixes an inconsistency that is caused by the answer having no owner.
Adding agents scales the drift. Each new person learns from a different colleague and inherits a slightly different version, and the variance widens with headcount. That is why support quality frequently gets worse as a team grows even when every individual is competent.
Fixing it starts with something unglamorous: establishing what the correct answer actually is for each of the recurring questions, in writing, owned by somebody, with a review date. That is a week of work, it produces no software, and it is the highest-return action available in most support operations.
Once the answers exist and are owned, delegating the head is safe, because the delegated thing is applying a known answer rather than inventing one. And the moment an answer does not apply cleanly, the correct behaviour is to escalate rather than to approximate — which is the opposite of how a deflection-optimised system behaves.
The measure has to change with it. Not deflection: resolution the customer agreed with, plus escalation timing, plus how much of your senior capacity moved from the head to the tail. A support operation should be judged on whether the hard cases got better attention, not on how many people stopped asking.
Consistency of the answer to a known question — measured by responses citing the maintained answer, as a share of contacts on that question, against a baseline where consistency was unmeasurable.
How much of the queue is genuinely known — measured by contacts matching a written owned answer versus not — the honest ceiling on what can ever be delegated.
Where your senior people spend the day — measured by experienced-agent hours on head questions versus tail cases, sampled the same way before and after.
Escalation timing — measured by contacts escalated at the first unmatched turn, against contacts escalated after two or more approximating exchanges.
Answers that have drifted — measured by the count flagged stale by use rather than discovered at a periodic review.
Quality coverage beyond sampling — measured by contacts carrying a machine-readable record of which answer was applied, versus the share reviewed by sampling.
deflection, and it is deliberately not offered as a measure. A customer who stopped asking is not evidence of a resolved problem. Nothing here answers an outage, a billing dispute, a safety issue or a complaint, and nothing invents an answer to a question whose answer is not written down and owned.
Contacts stay in your ticketing system and the maintained answer set lives where your team already keeps knowledge. Nothing migrates, and no second knowledge base is created — a second one is how the drift this page is about becomes permanent.
The answer set is yours. It is written by your people, owned by named people, and carries review dates, which means it keeps working if this operation ends. An answer set that only exists inside a vendor is a dependency rather than an asset.
Where a question needs live data — an account state, an order, an entitlement — that is read from your system rather than remembered, so the answer is current rather than plausible.
Nothing is answered that is not traceable to a written answer your organisation owns. Where no owned answer applies, the case escalates — it does not receive a best approximation, because a plausible wrong answer in support is more expensive than no answer.
Every response records which answer was applied and its version, so when an answer turns out to have been wrong, the affected contacts are findable rather than estimated.
Frequency and persistence are bounded. A customer who says the answer did not help is escalated rather than sent a second variation of the same thing.
Operational access is not permission to train. Your customer support data — which contains complaints, account details and occasionally distress — does not become material improving anything serving another organisation.
The escalation list is the artefact legal cares about — the classes that are routed and never answered. Outages, billing disputes, safety and complaints belong on it, and the list is written before anything goes live.
Support leadership owns the measures. The single most important decision here is refusing deflection as the primary metric, because whatever is measured is what the operation will optimise toward, and deflection optimises toward pushing away the cases that needed a human.
Whoever owns the knowledge base owns the answers. Delegation of answering does not transfer ownership of what is true, and an answer set with no internal owner will drift again inside a year.
One week of contacts on one channel — read-only, nothing answered — sorted by subject, with the recurring head identified and the current answer for each one established and owned.
The first phase is the unglamorous one and it produces the asset: what the recurring questions actually are, what the correct answer is for each, and who owns it. It answers no customer and it is worth doing whether or not anything is ever delegated.
It also produces the honest ceiling. If a small share of your queue matches a knowable, ownable answer, then most of your volume is genuinely novel and delegation would be inappropriate — that is a legitimate finding and it should stop the project.
If the head is large, the first delegation is a single question with a single owned answer, escalating on anything that does not match cleanly, measured on resolution the customer accepted rather than on deflection.
That is what optimising for deflection produces, and it is why deflection is explicitly refused as a measure here. The behavioural difference is what happens on a poor match: a deflection-tuned system tries again with a variation, and this escalates. If your previous vendor could not tell you what share of contacts were answered from an owned, versioned answer, it was generating plausible responses, which is the mechanism that produced the frustration.
It would, which is why the first phase writes and owns the answers and delegates nothing. If the current base is stale, the delegation cannot begin until the head questions have correct owned answers — and a drifted answer flagged in use is treated as a defect routed to its owner, not as something to work around. If your organisation cannot commit an owner to maintaining those answers, this operation should not start.
The first phase measures that directly and costs a week of read-only analysis. Some products genuinely have a flat distribution, and where that is true the honest conclusion is that first-line delegation is inappropriate and your constraint is agent expertise rather than repetition. More commonly the head exists and is under-recognised because the people answering it are too busy to notice how often they repeat themselves.
They should — nothing here impersonates a named employee, and a customer who asks gets a straight answer. The comparison that matters is not with your best agent; it is with the answer they would otherwise get from whoever was on shift, which is inconsistent by construction. Where the interaction genuinely requires your own people, that class escalates by rule and is not attempted.