Automated evaluation tells you whether the agent heard, understood, retrieved and executed correctly. It cannot tell you the agent sounded impatient, talked over the caller twice, or completed a task the caller had already given up on. Those failures are invisible in metrics and obvious within ten seconds of listening.
What listening finds that measurement cannot
Automated evaluation tells you whether the agent heard correctly, identified the intent, retrieved the right fact and called the right system. It cannot tell you that the agent sounded impatient, talked over the caller twice, explained something in a way that confused them, or technically succeeded at a task the caller had already given up on.
Those failures are invisible in every metric and obvious within ten seconds of listening. They are also the ones customers describe when asked why they disliked an automated call — rarely 'it got the facts wrong', frequently 'it kept interrupting' or 'it wouldn't let me explain'.
A scorecard turns that listening from an impression into a record. Twenty items, scored consistently across reviewers, produce trend data and a shared vocabulary for what 'better' means. This is the human half of quality assurance; the automated half covers what it cannot.
Opening and identification (items 1-4)
The first twenty seconds set the caller's expectations for everything after.
- 1. Did the greeting make clear this is an automated agent? Callers are forgiving of a labelled assistant and hostile to one pretending otherwise.
- 2. Was the opening brief? A long introduction before the caller can speak is one of the most common complaints and one of the easiest fixes.
- 3. Was the caller identified where possible, rather than asked for information the system already held?
- 4. Was the first question open rather than a menu, and did it let the caller state their reason in their own words?
Understanding and conversation flow (items 5-10)
The core of the call, and where listening is most revealing.
- 5. Did the agent understand the caller's first statement, or require it to be repeated?
- 6. When the caller was unclear, did the agent ask a sensible clarifying question rather than guessing?
- 7. Did the agent acknowledge what the caller said before moving on? A response that ignores the previous answer reads as a form.
- 8. Did the agent handle interruptions correctly — stopping immediately when the caller began speaking?
- 9. Did the agent avoid repeating itself or looping back to a question already answered?
- 10. Was the pacing natural — pauses matching the complexity of the question, with slow lookups acknowledged rather than silent?
Items 8 and 10 correlate strongly with caller satisfaction and are the two most often failed in early deployments. Both are configuration problems rather than model problems, and the latency instrumentation usually explains item 10.
Accuracy and honesty (items 11-14)
The reviewer checks these against sources rather than judging them by ear.
- 11. Were the facts stated correct according to the approved source? This requires the reviewer to check, not to assume.
- 12. When the agent did not know, did it say so clearly and offer a next step, rather than improvising a plausible answer?
- 13. Were numbers, dates, names and reference codes repeated back to the caller for confirmation before being used?
- 14. Did the agent stay within scope — avoiding pricing beyond published ranges, commitments, advice and anything on the prohibited list?
Item 14 should be scored as a pass or fail rather than on a scale. A single out-of-scope commitment is a finding regardless of how well the rest of the call went, and treating it as a deduction on a numeric score buries it.
Task completion and systems (items 15-17)
Whether the call actually did something, verified against the systems rather than from the audio.
- 15. Was the caller's task completed — booking made, information given, request recorded — and does the record in the destination system match what was said on the call?
- 16. Was the caller given a clear confirmation of what would happen next, with any commitment stated specifically rather than as 'someone will be in touch'?
- 17. When a system was slow or unavailable, did the agent handle it honestly — degrading, queueing or escalating rather than claiming success?
Item 15 requires the reviewer to open the destination system. It is the most time-consuming item on the scorecard and it catches the most expensive errors, so it is worth doing on a smaller sample rather than dropping.
Escalation and closing (items 18-20)
The end of the call, and the handover if there was one.
- 18. If escalation was warranted, did it happen promptly — at the second failed attempt rather than the fourth, and immediately on an explicit request?
- 19. Was the transfer handled well: the caller told what was happening, no silence, and the receiving person opening with the context rather than 'how can I help?' — the test set out in handoff design.
- 20. Did the call close cleanly, with the caller's issue either resolved or clearly owned by someone, and without an abrupt or confusing ending?
Making the scorecard produce consistent results
A rubric applied inconsistently produces trend data that reflects who was reviewing rather than how the system performed.
-
1
Score on observable behaviour, not impression
'Agent stopped within one second of the caller speaking' can be scored consistently. 'Agent sounded natural' cannot. Write each item so two reviewers would agree.
-
2
Calibrate regularly
Have all reviewers score the same three calls, compare, and discuss disagreements. Do this at the start and monthly. Without it, scores drift apart within weeks.
-
3
Separate pass/fail items from scored ones
Scope violations, wrong facts stated as certain, and failed escalations are binary findings. Averaging them into a score hides them.
-
4
Sample deliberately, not randomly only
A random sample for trend, plus a deliberate sample of escalated calls, abandoned calls and long calls — which is where the problems concentrate.
-
5
Keep the sample small and the cadence regular
Ten calls a week reviewed consistently beats fifty reviewed once a quarter, because the point is to notice change.
The calibration sessions are also where most of the useful product feedback comes from — reviewers discussing why a call felt wrong tends to surface design problems nobody had articulated.
What to do with the results
A scorecard that produces a number nobody acts on is worse than no scorecard, because it consumes time and creates a false sense of oversight.
- Route each failure type to its owner: content errors to knowledge owners, interruption and pacing to configuration, escalation failures to the rules, scope violations to whoever set the boundary.
- Track the failure distribution over time rather than the average score. The average is stable and uninformative; the distribution shows what is getting better.
- Feed every failure into the automated test suite as a case, so it cannot recur silently.
- Report scope violations and wrong facts separately and immediately, not in the monthly summary.
- Review the rubric itself quarterly — items that always pass are no longer earning their place, and new failure types deserve their own item.
- Share examples, not just scores. A thirty-second clip of a call that went wrong changes more minds than a percentage.
Where a platform provides call recording, transcripts and review tooling natively — as an AI call centre generally does — the mechanics are simpler, but the calibration discipline still has to be built by the team. Our automation services page covers the integration side of getting recordings and system records into one review workflow.
Failure modes of the review process
The QA process itself fails in recognisable ways.
- Reviewing transcripts instead of listening, which loses pacing, tone and interruption handling — half of what the scorecard exists to catch.
- Subjective items that different reviewers score differently.
- No calibration, so trends reflect reviewer changes.
- Averaging binary failures into a numeric score.
- Sampling only successful or only escalated calls.
- A rubric that never changes as the system matures.
- Findings with no owner.
- Using the scorecard to judge the vendor rather than to improve the system, which turns reviews into a negotiation.
- Reviewing only in the primary language, leaving other languages unexamined.
What the scorecard cannot assess
Be clear about the limits so the results are trusted where they are valid.
- Whether the underlying content is correct, as opposed to whether the agent stated it faithfully.
- Whether a call type should be automated at all.
- System behaviour at volumes or conditions outside the sample.
- What callers who hung up in the first five seconds wanted.
- Aggregate performance — a sample of ten calls describes those ten calls and a trend, not a population statistic.
Pair it with the automated layers for coverage and with call volume data for scale. The scorecard's job is depth, not breadth.
Decision framework and next step
Four questions before starting.
-
1
Who will review, and for how long each week?
Ten calls a week is realistic and sufficient. An ambitious target that lapses after a month is worse than a modest one that holds.
-
2
Are your items observable?
Rewrite any item two reviewers could score differently.
-
3
Can reviewers check the destination systems?
Item 15 needs this, and it catches the costliest errors.
-
4
Where do findings go?
Name an owner per failure type before the first review, or the results will accumulate unread.
Start with the twenty items above, calibrate on three shared calls, review ten calls weekly, and report the failure distribution rather than the average. Pair it with the five-layer measurement framework for the parts listening cannot cover. Our AI solutions overview covers how this operating rhythm is usually established.
Frequently asked questions
-
1
What should a voice agent QA scorecard cover?
Twenty items across five groups: opening and identification, understanding and conversation flow, accuracy and honesty, task completion verified against systems, and escalation and closing.
-
2
Why not rely on automated evaluation alone?
Automated measurement cannot detect that the agent interrupted the caller, sounded impatient, paced the conversation badly, or technically completed a task the caller had already abandoned. Those are the failures customers actually describe.
-
3
How is consistency achieved between reviewers?
Write every item as observable behaviour rather than impression, calibrate on the same three calls at the start and monthly, and keep binary findings such as scope violations separate from scored items.
-
4
How many calls should be reviewed?
Around ten a week, consistently, combining a random sample for trend with a deliberate sample of escalated, abandoned and unusually long calls. Regular and small beats large and occasional.
-
5
What should happen with the findings?
Each failure type routed to a named owner, every failure added to the automated test suite, scope violations and factual errors reported immediately rather than monthly, and the rubric itself reviewed quarterly.
The scorecard is the half of quality assurance that needs ears. Used alongside automated layer testing, it covers what the metrics cannot see.