How to score, coach, and improve, step by step.
Practical, vendor-neutral guides to scoring conversations, coaching from them, and getting value past the QA team. No gates, no sign-in.
Built to be useful, not to sell you.
Each guide is something you can act on whether or not you ever use QEval. Filter by what you are working on.
Evaluate a vendor What 1 to 2 percent call sampling actually misses The math of manual sampling, and why coverage is a risk decision, not a quality detail.
If you handle 100,000 conversations a month and review 2 percent, you score 2,000 and never hear the other 98,000. A compliance miss, a churn signal, or a coaching moment in those 98,000 stays invisible until it becomes a complaint, a fine, or a customer who does not come back.
Sampling also assumes the 2 percent is representative. It rarely is. Reviewers tend to pull the calls that are easy to find or already flagged, so the quiet failures — the ones nobody escalated — are the least likely to be seen.
Scoring every conversation changes the question from “what did we happen to catch” to “what is actually happening.” That is the difference between a sample and a system.
The 1 to 2 percent figure is a common industry baseline, not a QEval measurement.
Related: Auto QA, scoring every conversationEvaluate a vendor How to read an accuracy claim before you sign Four questions that separate a measured accuracy number from a marketing one.
Every QA vendor will tell you their AI is accurate. The word does almost no work on its own. Before you sign, ask four things.
First, accurate against what? A number only means something if it is measured against a benchmark, and the strongest benchmark is agreement with a trained human reviewer on the same items.
Second, is it published or just spoken? A figure that lives only in a sales conversation cannot be checked. Look for a number stated plainly, on the record.
Third, is it measured once or continuously? Models drift. An accuracy figure from launch day says nothing about month six unless it is recalibrated against people on an ongoing basis.
Fourth, is it in the contract? An accuracy claim you cannot hold a vendor to is an adjective; a service level you can is a commitment. QEval, for one, writes a 94 percent or higher classification accuracy into the master agreement.
Related: the accuracy commitment, and how it is measuredScore How to build an AI Agent QA scorecard The dimensions that matter, and why AI agents belong on the same scorecard as people.
Start with four dimensions that apply to humans and AI agents alike: compliance, empathy, resolution, and brand voice.
Compliance asks whether required disclosures were present and the rules were followed.
Resolution asks whether the customer’s actual problem was solved, not just whether the call ended.
Empathy and brand voice ask whether it felt like your company.
For AI agents, add two checks that human scorecards often skip: groundedness (whether the agent’s claims were true and on policy), and escalation (whether it handed off when it should have).
The most important design choice is parity. Score your AI agents on the same scorecard as your people. An AI agent that resolves most contacts is only good news if those resolutions would pass the same bar you hold a person to.
Related: AI Agent QACoach Turning QA scores into coaching that sticks A score that sits in a report changes nothing. The loop that changes behavior.
The loop has four moves:
Diagnose: Find the specific behavior behind the score, pinned to the moment it happened.
Target: Pick the one change that will move the metric most, not a list of ten.
Reinforce: Deliver it as coaching the agent can act on, then follow up.
Measure: Check whether the behavior actually changed on later conversations.
The last move is the one almost everyone skips. If you cannot tell whether a coaching action changed behavior, you are running coaching on faith. Tie the action to a later change in CSAT or resolution, and coaching becomes a system instead of a ritual.
Related: Coaching and performanceImprove Getting value past the QA team Quality scores are Layer 1. The other five layers are where most of the value is.
Every conversation also carries: - Customer intelligence — why people call - Revenue intelligence — where deals are won or lost - Operational intelligence — where processes break - Training intelligence — what to coach - Strategic intelligence — what leadership should know
QEval® calls these the Six Layers of Intelligence, and roughly 82 percent of the value sits in layers two through six.
The practical move is to stop treating conversation data as a QA artifact and start routing it to the people who can act on it. The same scored conversation that tells a supervisor to coach can tell product why customers are confused and tell sales which objections keep recurring.
Related: the Six Layers of IntelligenceEvaluate a vendor What to ask any AI Agent QA vendor A vendor-neutral checklist for a category full of adjectives.
Use this list with any vendor, including QEval. The goal is to replace adjectives with answers.
On accuracy: what is it measured against, is it published, and is it in the contract?
On coverage: do you score every conversation, or a sample?
On evidence: can every score be traced to the moment that earned it?
On AI agents: do you score them on the same scorecard as humans, and who owns that?
On compliance: do you check every call for required disclosures, and how fast can you adapt when a rule changes?
On data: do you train any third-party model on our conversations, and is personal information redacted before processing?
A vendor that answers these plainly is one you can evaluate. A vendor that answers with adjectives is one you cannot.
Related: why QEval®No guides in that category yet.
See it on your own conversations.
Reading is one thing. Start a pilot and we will score a sample of your real conversations and show you, with evidence, what your current program is missing.