Move cursor | Click to ripple
Blog / Who calibrates the machine? Making AI scoring transparent enough that frontline teams trust it
blog

Who calibrates the machine? Making AI scoring transparent enough that frontline teams trust it

Ask a quality assurance manager why agents roll their eyes at their own quality score, and the honest answer has almost nothing to do with the number. It has to do with not knowing where the number came from. I have watched agents accept a 72 without a word of complaint. What wears them down is a 72 with no explanation attached to it.

For most of my career, that problem had a human name. Two evaluators scored the same call ten points apart and nobody reconciled the gap. Now the score often comes out of a model, and the question on the floor has changed shape. It is no longer only which evaluator did I get. It is did the machine get this right, and who checked.

That second question is the one worth building a program around. An AI scoring engine that nobody audits is a faster version of the same complaint. What makes a score defensible is a loop: people define the standard, the model applies it at volume, people audit the model against reality, and every disagreement between them changes something. Coverage stops being the headline. The loop becomes the headline.

 

Figure 1. Calibration is not a one time model training exercise. It runs continuously, in both directions.

Why frontline teams stop trusting the score

Most quality programs do not fail because leadership stopped caring about quality. They fail because the score feels arbitrary from the agent’s chair. Manual QA covers one to three percent of interactions and holds every agent to conclusions drawn from that slice. Agents know the math does not work. Four calls a month does not represent a shift that included four hundred.

Verint’s State of Agent Experience 2026 survey put 31 percent of frontline agents as likely to leave their role within six months, with unrealistic performance expectations ranking as their top complaint, ahead of pay. That is market research, not a claim about any single program. It still matches what I see on floors. Agents disengage from a metric they cannot explain to themselves, let alone defend to a supervisor.

Automated scoring solves the coverage half and can make the trust half worse. Rules engines run 65 to 70 percent accuracy on novel data. Keyword analytics cannot tell sarcasm from sincerity. If the honest answer to why did I get this is the model flagged it, you have replaced an inconsistent human with an unaccountable machine, and agents will trust the second one even less than the first.

Full coverage is table stakes. Accuracy is the argument.

Scoring every interaction has been possible since 2024, and most platforms will tell you they do it. The collapse happens after ingestion. Sampling reappears quietly inside the analytics layer, and the accuracy of the score itself is almost never published, let alone committed to. Ask a vendor what their classification accuracy is and see whether you get a number or an adjective.

Coverage answers how many conversations you looked at. It says nothing about whether the score was right.

This is why we wrote 94 percent classification accuracy into the QEval™ master agreement rather than into a datasheet, alongside 98 percent compliance accuracy and 95 percent recall on flagged violations. A number in a contract is a number someone has to defend. That is a different posture than a marketing claim, and it is the posture a program needs if agents are going to be held to the output.

Humans calibrate the model

Start with where the ground truth came from. Most QA models learn from a public dataset or a textbook framework. The QEval™ experts were trained inside a working contact center, on the rubrics a live QA team ran in operations, against the compliance edges real auditors flagged. ETSLabs/Etech runs a floor of more than four thousand agents. The product was built on top of it, not adjacent to it. That difference is not academic. It is the gap between a score that looks reasonable in a demo and a score a supervisor will stand behind in a dispute.

That origin story only matters if the calibration continues after deployment, and this is the part most AI QA conversations skip. Your evaluators score an audit sample blind, without seeing what the model returned. Then you compare. Where the human and the model agree, you have evidence the rubric is holding. Where they diverge, you have work. Somebody has to sit down and decide which reading was correct, and why.

Those disagreements are the most valuable thing your QA team produces. Sometimes the model missed context and the expert needs tuning. Sometimes the evaluator applied a criterion the rubric never actually said. Either way the answer changes something concrete: a rewritten criterion, a reweighted item, a retuned expert. A model that never gets corrected by the people who own the standard is not human in the loop. It is human adjacent.

The model calibrates the humans

The direction nobody expects is the more useful one day to day. Once you have a scoring engine that applies the same weighting at two in the morning and two in the afternoon, you finally have a stable reference point to measure your evaluators against. Calibration variance under 2 percent is not interesting because it flatters the model. It is interesting because it makes evaluator drift visible.

Take the example every QA leader has lived through. Two evaluators review the same escalated call. One came up through overnight support and reads a long pause as the agent thinking through a difficult account. The other came up in retail sales and reads that same pause as the agent losing control of the conversation. Neither is wrong about what they heard. They disagree about what it means, and that disagreement used to vanish into one agent’s file because only one of them ever scored the call.

Figure 2. The same fourteen point gap, with and without a reference point to measure it against.

Inside a loop, that gap surfaces instead of disappearing. Both readings sit against a score that does not move. Evaluator scores get compared against the reference on the same recordings, so the outliers show up before the next session rather than after the next attrition spike. Your calibration meeting stops being a random call pulled from last week and starts with the specific evaluators and the specific scorecard items that are actually drifting.

Cadence matters here too. Weekly or biweekly evaluator calibration is common in the strongest programs. Quarterly is where most teams default, and quarterly is not frequent enough to hold a rubric together when the model is scoring everything in between.

Human in the loop is a division of labor, not a disclaimer

The phrase gets used loosely enough that it has stopped meaning much. In practice it should be a specific list of what the model decides and what a person decides, written down, with names attached.

Figure 3. Automating the scoring does not shrink the QA team. It moves them off the sample and onto the judgment.

This is also the answer to the question every QA director asks in the second meeting. Does automating the scoring mean cutting the team. It has not played out that way. QA analyst productivity rises roughly 65 percent, and that time gets redirected to calibration, coaching, and the edge cases a model should not rule on alone. In one Fortune 500 program running QEval™ across 1,200 agents, eleven full time analysts were retasked rather than removed. You get capacity back and your most experienced people stop spending the week on a two percent sample.

Explain the reasoning behind every score, not just the outcome

A loop only works if the model can show its work. General purpose models score compliance, empathy, resolution, and brand voice with the same weights, and the accuracy shows it. QEval™ routes each item on your scorecard to a purpose trained expert instead. A 47-item scorecard becomes 47 routed classifications. Compliance language goes to a compliance expert, empathy to an empathy expert, brand voice to a brand voice expert. A proprietary vocabulary library tuned across more than 35 languages gives those experts the domain fluency a generic model lacks, because contact center language is not general English.

Every result comes back tied to the exact moment in the transcript that produced it. That is what makes a human audit possible in the first place. An evaluator checking the model cannot say the score was wrong unless the model has already said which line it was reading. QEval™ runs 326 million classifications every five minutes across voice, text and digital channels, more than three billion conversations year to date, and that volume is only useful because each classification is traceable to something specific.

Figure 4. A scorecard your team wrote, routed to specialists, cited back to the transcript, and reviewable by a person.

Give agents a real way to challenge a score, then use what they tell you

Build a dispute path an agent can actually use and resolve it inside a couple of business days. Not a couple of weeks. An appeal process nobody uses tells the floor the score was final before they ever saw it. One that gets used, with a clear timeline and a real reviewer on the other end, tells them the program cares more about getting the score right than about defending whatever the first evaluator wrote down.

The part most programs miss is what happens after the dispute is resolved. An overturned score is free labeled data on exactly the case your rubric handles badly. Route it back into the loop. If agents keep winning appeals on the same criterion, the criterion is the problem, not the agents. A dispute log that never changes rubric is a complaints department.

Regulation is moving in the same direction. Transparency provisions under the EU AI Act tied to high-risk AI systems phase in through August 2026, and a scoring system that affects an employee’s pay or standing sits squarely inside that conversation, whether your organization operates in the EU or is preparing to sell into it. ISO 42001 covers AI management systems specifically and asks many of the same questions about oversight and traceability. A program that can already explain, audit and contest every score is not scrambling to build that capability later.

Close the loop with coaching, or the score is just paperwork

A quality score that never turns into a specific coaching conversation is a row in a spreadsheet. Tie every low score to a development plan built around the actual transcript moments involved, not a generic training module assigned because the category happened to match. Have the conversation privately. Then track whether the coaching moved next month’s score, not whether the session happened. Session completion is an easy number to hit, and it tells you nothing.

Score AI agents and human agents against the same rubric

More centers now run a mix of human agents and AI agents on live conversations, and the fairness question extends past the human floor. A rubric that exists only for people, while an AI agent handling the same call type gets no equivalent evaluation, confirms exactly what the human side of the team already suspects. The scrutiny only runs in one direction.

There is a technical reason this gets skipped. AI agents do not grade themselves the way you grade humans. Containment, deflection and resolution metrics come out of the AI vendor’s own platform and measure the vendor’s definition of success, not your scorecard, your brand voice, or your compliance rules. Scoring both sides against the same criteria requires a vendor neutral layer sitting above both, with drift detection on the AI side, because an AI agent’s behavior changes when its model gets updated and nobody on your floor gets a memo.

Calibrating people only, versus calibrating the model and the people

Calibrating people only Calibrating the model and the people
What gets scored A sample, often one to three percent Every interaction, with coverage throttled by program
The reference point Whichever evaluator drew the call A model that scores the same way every time, calibration variance under 2 percent
How drift is found Someone notices, eventually Evaluator scores are compared against the reference and the outliers are flagged before the next session
What a disagreement produces A number in the agent’s file A rewritten criterion and a tuned expert
Explanation A score with no reasoning attached The transcript line that produced the result
Appeal No real path to challenge a score A defined dispute process, resolved by a person, that feeds back into the rubric
AI and human agents Different standards for each The same scorecard applied to both, with drift monitoring on the AI side

Frequently asked questions

What does human in the loop mean for AI quality scoring?

It means people own the parts a model should not: writing the criteria, blind scoring audit samples to check the model against reality, ruling on edge cases, deciding disputes, and running the coaching. The model owns volume and consistency. If a vendor cannot tell you which decisions stay with your team, the phrase is decoration.

How do you know the AI is scoring correctly?

You check it the same way you would check a new evaluator. Your team scores an audit sample blind, then compares against what the model returned. Agreement is evidence the rubric holds. Disagreement gets resolved by a person and fed back into the rubric or the expert that produced the miss. Ask any vendor what their classification accuracy is and whether it is contractual. QEval™ commits to 94 percent, with 98 percent compliance accuracy and 95 percent recall on flagged violations.

Does automated scoring mean we need fewer QA analysts?

It has not played out that way in production. Analyst productivity rises roughly 65 percent, and that time moves to calibration, coaching, and judgment calls. In one Fortune 500 deployment, eleven analysts were retasked rather than removed.

How often should quality teams run calibration sessions?

Weekly or biweekly for evaluators, in both directions. Human against human, and human against the model. Quarterly is where most teams default and it is not frequent enough when the model is scoring everything between sessions.

Should agents be able to dispute a score generated by AI?

Yes, and more so than with a human score, because the agent has no relationship with the reviewer to fall back on. Resolve it within a couple of business days, have a person make the call, and route overturned scores back into the rubric. A pattern of successful appeals on one criterion tells you the criterion is written badly.

Does scoring every interaction automatically fix agent trust?

No. Full coverage removes the sampling complaint. If the score still cannot explain itself, and nobody is auditing the thing producing it, agents will trust it exactly as little as they trusted the old two percent sample.

Where this leaves your QA program

None of this requires a new philosophy about employee engagement. It requires a scoring program built so that every score can defend itself, and a loop that keeps it honest. People write the criteria. The model applies them at volume and cites its reasoning. People audit the model, resolve the gaps, and feed the resolution back. The model shows you which of your evaluators are drifting, so the next calibration session has an agenda.

QEval™ was built around that requirement by an operator running its own floor, with calibration tooling, a defined dispute workflow, expert routing, and scoring reasoning that traces back to the transcript rather than a black box verdict. If you want to see how a specific call from your floor gets scored and why, book a demo at qeval.ai.

Jim Iyoob

Jim Iyoob

Author

Jim Iyoob is the Chief Revenue Officer for Etech Global Services and President of ETS Labs. He has responsibility for Etech’s Strategy, Marketing, Business Development, Operational Excellence, and SaaS Product Development across all Etech’s existing lines of business – Etech, Etech Insights, ETS Labs & Etech Social Media Solutions. He is passionate, driven, and an energetic business leader with a strong desire to remain ahead of the curve in outsourcing solutions and service delivery.

Score every conversation

Move from keywords to context.

Bring your scorecards, your AI agents, your CCaaS. We will score a real call in 30 minutes and show you what keyword rules missed last week.