Call Center Quality Assurance: The Complete Checklist (2026)
Last updated: October 2026
Call center quality assurance (QA) is the process of reviewing customer interactions against a written standard, scoring them, and turning the results into coaching and process fixes. A working QA program needs six things: one owner, a short weighted scorecard, a fair sampling plan, regular calibration, a coaching loop, and separate handling for compliance misses. In 2026 it also needs a seventh: a plan for reviewing the conversations your AI agents handle, not just your human ones.
This guide is a checklist you can copy. Each section lists the items to put in place, with the reasoning kept short. If you want to test a scorecard on a real conversation first, paste one into our free support quality scorecard and see how it grades.
QA, quality control, and quality management
The three terms get used interchangeably, but they cover different jobs:
Quality assurance checks whether agents followed the standard on a given interaction. It is forward looking: the goal is better behavior on the next call.
Quality control finds and fixes defects after the fact, such as correcting a wrong refund or calling back a customer who got bad information.
Quality management is the wider system: the policies, knowledge base, tooling, and reporting that QA and quality control feed into.
The distinction matters when a finding comes in. A missed identity check by one agent is a QA and coaching issue. The same miss by ten agents is a quality management issue, because the process or the training is broken.
Checklist 1: Program setup
Before anyone scores a call, settle who owns the program and what the rules are.
Name one QA owner who maintains the scorecard, sets review rules, and reports findings to support leadership.
Write down what a passing score is (for example, 85 out of 100) and which items are auto-fail regardless of the total.
Decide who can change the scorecard, and how changes are announced to agents.
Define the channels in scope: voice, chat, email, SMS, social, and AI-handled conversations.
Set a dispute process so an agent can challenge a score and get an answer within a fixed time, such as five business days.
Agree on where QA results live. Scores should link back to the exact ticket or recording, not sit in a separate spreadsheet.
Decide how QA scores will be used: coaching only, or also performance reviews. Tell agents before the first review cycle, not after.
Checklist 2: Build the QA scorecard
A good call center QA scorecard is short, observable, and weighted toward what actually hurts customers. Eight to twelve criteria is enough. Past that, reviewers start skimming and scores drift.
Here is a sample weighted scorecard you can adapt:
Category | What the reviewer checks | Weight |
|---|---|---|
Accuracy | The answer matched the current knowledge base and policy | 25% |
Resolution | The issue was solved, or the customer got a specific next step and timeline | 20% |
Understanding | The agent identified the real need and confirmed key details before acting | 15% |
Process | Required steps were followed in order (verification, notes, tagging) | 15% |
Communication | Tone was respectful, clear, and suited to the channel | 10% |
Handoff | Any transfer or escalation carried full context, so the customer did not repeat themselves | 10% |
Efficiency | No unnecessary holds, transfers, or dead air | 5% |
Compliance (auto-fail) | Identity verification and required disclosures completed | Pass or fail |
When you write each criterion, use this checklist:
Describe a behavior a reviewer can see or hear. "Show empathy" is hard to score. "Acknowledged the stated problem before giving instructions" is easy.
Score only what the agent controls. Do not dock an agent for a system outage, unless the criterion is about how they communicated it.
Use yes, no, or not applicable for most items. Five-point scales invite disagreement between reviewers.
Keep compliance items as auto-fail flags, not weighted points. A great call with a missed disclosure is still a failed call.
Give each criterion a one-line example of a pass and a fail, written into the rubric itself.
For ready-made rule wording, see how teams write QA rules in plain English that both human reviewers and AI scorers can apply.
Checklist 3: What to evaluate on every interaction
Keep one core definition of quality across channels, then add channel-specific checks only where the work is different.
Core checks for every interaction:
Understanding: Did the agent find the customer's main need and confirm the details?
Accuracy: Did the answer match the current knowledge base and standard operating procedures?
Resolution: Was the issue solved, or did the customer leave with a clear next step?
Communication: Was the tone respectful and right for the channel?
Process: Were identity checks, disclosures, and internal steps completed?
Handoff: If someone else took over, did they get enough context to continue without re-asking?
Channel-specific checks:
Voice: clarity, active listening, interruptions, hold etiquette, dead air, and verbal disclosures. Our guide to call monitoring software covers the tools that capture this evidence.
Chat: did the reply address every part of the question, and did the customer have to repeat information already given?
Email: accuracy, readability, a complete answer in one reply, and no "let me check and get back to you" loops.
SMS and social: brevity, privacy (no account details in public replies), and a clean move to a private channel when needed.
If "accurate resolution" means one thing on calls and another in email, your scores cannot be compared across teams. Keep the core rubric fixed and treat channel checks as additions.
Checklist 4: Sampling and coverage
Manual QA teams typically review 2% to 5% of conversations. That is enough to coach individual agents, but it routinely misses rare and high-risk problems, which are exactly the ones you most need to catch. We cover the cost of that gap in why sampling 5% of conversations falls short.
Sampling checklist:
Review a fixed number of interactions per agent per period (a common starting point is four to five per agent per week) so every agent gets the same attention.
Stratify the sample across channels, shifts, contact reasons, and new versus tenured agents.
Always include targeted reviews on top of the random sample: escalations, low CSAT surveys, complaints, refunds above a threshold, and repeat contacts.
Record why each interaction was picked (random or targeted) so a low score on a targeted complaint is not read as typical behavior.
Use automated scoring to cover the remaining 95% or more, and route anything it flags to a human reviewer.
Checklist 5: Calibration
Calibration is how you make sure two reviewers give the same interaction the same score. Without it, an agent's score depends on who reviewed them.
Hold a calibration session at least monthly, and also whenever the scorecard changes or a new reviewer joins.
Have every reviewer score the same three to five interactions independently before the session.
Compare scores item by item, not just the total. Disagreement usually sits in two or three criteria.
Set a target spread. A practical goal is reviewers landing within 5 points of each other on a 100-point scale.
Rewrite any criterion that keeps causing disagreement. The problem is usually the wording, not the reviewers.
Log every calibration decision ("a transfer with a warm intro counts as a pass") in a shared rubric notes document.
If you use automated scoring, calibrate the AI too: compare its scores to your calibrated human scores on the same set.
Checklist 6: Metrics to track alongside QA
A QA score describes how an interaction was handled. Operational metrics show what happened afterward. Read them together, because any one of them alone can mislead.
Metric | What it tells you | Pair it with |
|---|---|---|
QA score | How closely agents follow the standard | CSAT, to check the standard reflects what customers value |
First contact resolution (FCR) | Whether the issue was solved the first time | Repeat contact rate, to catch "resolved" tickets that come back |
Average handle time (AHT) | How long interactions take, including wrap-up | QA score and FCR, so speed does not come at the cost of quality |
Compliance pass rate | Share of interactions with every required step done | Critical-miss count, tracked separately |
Customer satisfaction (CSAT) | How customers rated the interaction | QA findings, since surveys only capture customers who respond |
Repeat contact rate | Whether customers came back about the same issue | Resolution score, to find answers that sounded complete but were not |
A simple rule: never put a speed metric on a dashboard without a quality metric next to it. AHT drops are easy to buy with rushed calls. For a deeper look at choosing and tracking QA metrics with AI, see our best practices for AI-driven QA in support.
Checklist 7: Coaching and continuous improvement
QA only pays off when a finding turns into a change in behavior or process. After each review cycle:
Share the score with the agent within a few days, along with the exact moment in the transcript or recording that drove it.
Agree on one behavior to practice, not five.
Separate one-off mistakes from patterns across several interactions before deciding on action.
Assign training when the gap is a skill or product knowledge issue.
Fix the knowledge base or SOP when the right answer was hard to find. If several agents miss the same refund rule, the rule is unclear.
Recheck the same criterion in the next cycle to confirm the behavior changed.
Recognize improvement publicly. Agents engage with QA when it feels like coaching, not policing.
Real-time guidance can close the gap between review and behavior. Agent assist surfaces the right answer and next step during the conversation, so the coaching point is applied on the next customer, not the next review cycle.
Checklist 8: Compliance
Keep compliance separate from service quality. Compliance misses carry legal and financial risk, so they need their own process.
List every required step by contact type: identity verification, recording consent, required disclosures, and approved handling of sensitive data.
Mark each one as auto-fail on the scorecard.
Define which misses need same-day escalation to a supervisor or compliance lead.
Review 100% of conversations for compliance where you can. A 3% sample is a weak defense in an audit.
Limit access to recordings and transcripts to people who need it, and follow your retention policy.
Have legal counsel confirm the rules for your industry and the regions you serve, rather than relying on a generic template.
Industry examples: healthcare queues should emphasize privacy and identity steps, financial services should make required disclosures and account safeguards easy to verify, and ecommerce teams should check that refund and return policies were explained accurately.
Checklist 9: QA for AI agents
Most call center QA programs were built for human agents. If AI agents now answer a share of your calls, chats, or tickets, those conversations need the same scrutiny, and often more, because one bad answer can repeat thousands of times before anyone notices.
Score AI-handled conversations against the same core scorecard as human ones, so the two are comparable.
Add AI-specific checks: was the answer grounded in an approved source, did the AI stay within its allowed topics, and did it hand off when it should have?
Check grounding directly. Our guide on how to check whether AI support answers are grounded walks through the method.
Define what counts as a real resolution, so a customer who gave up is not counted as a success. See what counts as a resolution in AI support.
Review every AI conversation that ended in an escalation, a negative sentiment shift, or a repeat contact.
Track AI and human QA scores side by side in the same report.
Manual vs. automated call center QA
Manual QA | Automated QA | |
|---|---|---|
Coverage | 2% to 5% of conversations | Up to 100% of conversations |
Speed to detect a problem | Days to weeks | Minutes to hours |
Consistency | Depends on calibration | Same rules applied every time |
Best at | Nuance, context, coaching conversations | Finding patterns, compliance misses, and outliers at scale |
Weakness | Misses rare problems | Needs human review of disputed and high-impact scores |
The strongest programs use both. Automation scores every interaction and flags what matters. Humans review the flags, handle disputes, run calibration, and coach. Compare tools for this in our roundup of the best AI tools for support QA and coaching.
How IrisAgent AutoQA fits the checklist
IrisAgent AutoQA scores 100% of AI and human conversations against rules you write in plain English, directly on top of the helpdesk you already use. It maps onto this checklist in four places:
Scorecard (Checklist 2): your criteria become AutoQA rules, with compliance items set as auto-fail.
Coverage (Checklist 4): every conversation is scored, not a 5% sample, and flagged conversations go to a human reviewer.
Compliance (Checklist 8): real-time alerts surface missed disclosures or verification steps as they happen.
AI agents (Checklist 9): AI-handled conversations are scored with the same rules as human ones, in the same report.
Want to see how a scorecard grades a real conversation? Try the free support quality scorecard, or book a demo to see AutoQA running on your own tickets.
Frequently Asked Questions
What should a call center quality assurance checklist include?
A call center QA checklist should cover program setup (an owner, passing score, and dispute process), a weighted scorecard, per-interaction checks for accuracy, resolution, communication, process, and handoff, a sampling plan, regular calibration, a coaching loop, separate compliance checks, and a plan for reviewing AI-handled conversations.
How many calls should a call center QA team review?
Manual QA teams typically review 2% to 5% of interactions, often four to five per agent per week. That supports coaching but misses rare, high-risk issues. Add targeted reviews of escalations, complaints, and low CSAT, and use automated scoring to cover the rest, with flagged conversations sent to human reviewers.
What is a good call center QA score?
Most teams set a passing score between 80 and 90 out of 100, with compliance items as auto-fail regardless of the total. The right threshold depends on your scorecard, so set it after a few calibration rounds and track the trend over time rather than a single number.
How often should call center QA calibration happen?
Hold calibration at least monthly, and also whenever the scorecard changes, a new reviewer joins, or scores start to drift. Reviewers score the same interactions independently, compare item by item, and rewrite any criterion that keeps causing disagreement. A practical target is reviewers within 5 points of each other.
Can AI automate call center quality assurance?
Yes. AI can score 100% of conversations against defined rules, flag compliance misses in real time, and surface patterns that small samples miss. Teams should still calibrate the AI against human reviewers, review disputed or high-impact scores manually, and keep humans responsible for coaching.
Loading the form. Prefer email? Reach us at info@irisagent.comor call +1 617 249 3312.