ToolsTemplate Library · Zendesk · Audit

Support QA Scorecard template.

A six-criterion weighted rubric with written definitions for every score level — calibrated so two reviewers land on the same number.

Download the CSV Free · 5 columns · 6 pre-filled rows · updated July 2026

What's inside — the exact file

This is the complete template, not a sample. The rows are worked examples — replace them with your data (and delete them before any platform import).

criterionweight_percentscore_1_definitionscore_3_definitionscore_5_definition
Accuracy25Wrong or misleading answerCorrect but incompleteCorrect, complete, anticipates the next question
Tone15Robotic, curt, or defensiveProfessionalWarm, matches customer energy
Process adherence20Skipped required steps or documentationMinor gapsFollowed fully, documented cleanly
Resolution20Not resolved, no path forwardResolved with excess effortResolved efficiently, confirmed with customer
Empathy and ownership10Deflected blame or ignored frustrationAcknowledged the issueOwned it personally, named next steps and timing
Customer effort10Customer had to repeat, chase, or re-explainSome back-and-forthOne touch or seamless handoff, zero repetition

What this template is for

Most support QA fails the same way: two reviewers score the same ticket differently, agents call it subjective, and the program dies. This rubric fixes the calibration problem — six weighted criteria, each with written definitions for scores of 1, 3, and 5, so two reviewers reading the same conversation land on the same number.

The six criteria (accuracy, completeness, tone, process compliance, ownership, and writing quality) with their weights are a proven starting point — adjust the weights to your values, but never remove the written score definitions. The definitions are the calibration; the numbers are just arithmetic.

How to use it, step by step

  1. Adapt the definitions to your context. Rewrite each score definition with your product's specifics — what 'accurate' means for a billing dispute vs. a bug report. Keep the 1/3/5 anchor structure.
  2. Weight by what you actually value. The weights should sum to 100 and reflect your priorities — accuracy usually deserves the most. If tone matters more than process in your brand, say so in the numbers.
  3. Calibrate before you score anyone. Two reviewers score the same five tickets independently, then compare. Where scores diverge by 2+, the definition is ambiguous — rewrite it. Repeat until divergence is rare.
  4. Sample fairly and consistently. Score 3–5 tickets per agent per week, randomly sampled with maybe one flagged ticket. Cherry-picked reviews destroy trust faster than no reviews.
  5. Review with the agent, not at them. The rubric turns feedback from opinion into shared criteria. Agents should be able to self-score with the same sheet and predict their result — that's when QA starts improving quality instead of measuring it.
  6. Track scores over time and against CSAT. Rubric scores that don't correlate with customer satisfaction over a quarter mean the rubric measures the wrong things — revisit the criteria.

What each column means

criterionThe quality dimension being scored — six pre-filled.
weight_percentIts share of the total score; weights sum to 100.
score_1_definitionWhat failing this criterion concretely looks like.
score_3_definitionWhat acceptable-but-improvable looks like.
score_5_definitionWhat excellent looks like — written so it's recognizable, not aspirational.

Common mistakes to avoid

  • Scoring without written definitions — the moment scores are contestable, the program is politics.
  • Ten criteria instead of six: reviews take too long, so they stop happening. Fewer, weightier criteria survive contact with real workloads.
  • Using QA scores punitively from day one — score anonymously for the first month while calibrating, or you'll never get honest buy-in.
  • Never scoring AI agent conversations — automated resolutions deserve the same rubric; that's how you find where the AI needs better content or tighter guardrails.

Questions, answered

What's a good QA score target?

After calibration, healthy teams average around 4.0–4.5 out of 5 with the definitions written honestly. A team averaging 4.9 has soft definitions, not perfect quality — tighten the score-5 anchors until excellence is distinguishable.

How is this different from CSAT?

CSAT measures how the customer felt; QA measures whether the work was right. They disagree constantly — a wrong answer delivered charmingly scores high CSAT, and a correct 'no' scores low. You need both lenses, and this rubric is the internal one.

Should this rubric score AI conversations too?

Yes, unchanged. Sample your AI agent's automated resolutions weekly with the same six criteria. It's the fastest way to find the gap between 'the bot answered' and 'the bot answered well' — and it makes the human-vs-AI quality comparison honest.

This is a Market Disrupt worksheet, not a vendor file. Import-format templates follow each platform's documented layout as of July 2026 — platforms evolve, so validate against current documentation before a large import.

Rather have it done than downloaded?

Imports, migrations, and workflow buildouts are the day job — we're a Zendesk Premier Partner, and the same team that made this template runs the real thing.

Talk to a Zendesk Premier Partner