Inside the QEval Core.
QEval does not call out to OpenAI, Anthropic, or any third-party foundation model. ETS Labs built a proprietary, closed-source Mixture-of-Experts model, trains it on real contact-center interactions, and runs it inside a live contact center. That is why the accuracy is contractual and your data never leaves.
Hundreds of experts. One deterministic score.
Most QA tools run one general-purpose model over every task. QEval runs hundreds of purpose-trained experts. The routing is deterministic, the experts are machine-learning models, and an ML calibration layer keeps the same conversation scoring the same way every time.
Same conversation in, same score out. Reproducible and audit-defensible, even though the experts are machine-learning models.
Why one general-purpose model is not enough.
Most QA tools, and the CCaaS features built on them, run one general-purpose model over every task. It scores compliance, empathy, resolution, and brand voice with the same generalist weights. It is fluent, but a specialist in none of them. In QA, where a missed disclosure is a fine and a misread tone is a bad coaching call, generalist is not good enough.
Same weights for every judgment
Compliance, empathy, resolution, and brand voice all pass through one set of generalist weights. Industry-average classification accuracy sits at 65 to 70% (industry market context, not a QEval result), and you cannot see why any single score was given.
A purpose-trained expert per dimension
Compliance language routes to a compliance expert. Tone routes to an empathy expert. Each expert is trained on contact-center data for its one job, which is how 94%+ accuracy holds across every dimension at once.
See how a conversation is scored.
Pick a conversation. The classification engine sends each judgment to the expert built for it. Every expert pins its score to the exact line that earned it, and they compose into one grade. Turn an expert off to see how much it moved the score.
The three parts of the Core.
The model is the intelligence layer. Two more components make it operational at the speed and accuracy QEval commits to.
The MoE model
A proprietary, closed-source Mixture-of-Experts language model. Hundreds of purpose-trained expert sub-models, each specialized for one task in one vertical, not one generalist. A deterministic engine routes each scorecard item to the right expert, and an ML calibration layer holds the score reproducible. ETS Labs owns it, trains it, and operates it end to end.
The Classification Engine
Maps each item on your scorecard to the right expert pathway. Define a 47-item scorecard with compliance gates, empathy markers, and brand rules, and the engine routes each item to the expert built for it. This is what enables 326 million classifications every 5 minutes.
The Vocabulary Library
A proprietary lexicon tuned to contact-center language across 35+ languages and 80+ platform integrations. Hold procedures, transfer protocols, disclosure requirements, de-escalation patterns: the domain fluency general models frequently misread.
Ask a question across every conversation.
Ask QEval is an agent, not a search box. Give it a question and it plans an approach, calls its tools to search across millions of scored conversations, reads the calls that matter, synthesizes an answer with citations, and takes the next step on its own. Below is a full run, start to finish.
Illustrative. In production, answers are generated across your scored conversations and cite the calls behind every claim. Your data never leaves QEval.
Hundreds of expert agents, each specialized.
Each is a Mixture-of-Experts agent trained for one task in one vertical, and the classification engine routes every scorecard item to the exact expert for that judgment. A compliance expert for healthcare is not the same model as a compliance expert for collections. The tasks below are a sample. The hundreds come from specializing each task for each vertical.
The circled node is a single expert: the compliance model tuned for healthcare. Every task is specialized for every vertical, and that is where the hundreds of expert agents come from.
Disclosures, consent, and regulated language, by jurisdiction.
Whether the agent named and validated the customer emotion.
Issue solved on this contact, not deferred or bounced.
Greeting, tone, and close measured against your style guide.
The emotional trajectory across the whole conversation.
Why the customer reached out, mapped to your taxonomy.
Churn intent, legal language, and cease-contact demands.
Repeated phrasing and off-script loops in AI agents.
PCI, PHI, and PII removed before any scoring begins.
The summary, action items, and outcome for every contact.
These are some of the tasks. Hundreds of expert agents run across tasks and verticals. A new task or vertical gets its own expert.
How a conversation becomes a score.
This is the real data flow, drawn from our architecture. Note where redaction sits: sensitive data is stripped before any model ever sees the text, not after.
Every channel in
Voice, chat, email, and SMS land through the call-ingestion pipeline and your CCaaS connectors. No sampling.
PII stripped at the gate
Named Entity Recognition removes PCI, PHI, and PII before scoring. Certified PCI DSS Level 1.
Classification engine
Each scorecard item is routed to the expert sub-model built for that judgment.
Held to your reviewers
Scores are calibrated against your human reviewers and monitored for drift, within 2%.
Grade, evidence, outputs
A grade with the transcript span behind it, plus coaching, intent, and the Six Layers of intelligence.
Source: QEval Integration Data Flow Architecture. Redaction runs before the model, by design.
The same scorecard, scored two ways.
This compares the QEval Mixture-of-Experts approach with a single general-purpose model, the pattern behind most QA features. No vendor is named; the difference is architectural. The 65 to 70% industry-average figure is market context observed across vendor materials, not a QEval result.
Measured against your own reviewers.
The number is not a single-task lab benchmark. It is held up against your own reviewers, and we are open about where automated scoring has limits.
The calibration loop
What we do not claim
Tuned to contact-center language.
The same sentence means something different on a contact-center call than it does on the open internet. The vocabulary library is why the experts read it the way a supervisor would.
How accuracy holds over time.
Training the experts is the start. Keeping them honest is the job. Every score is held against your reviewers, and when they diverge, the experts are recalibrated and the version is rolled forward.
The calibration flywheel
What keeps the number honest
Where your data goes.
Accuracy, in the contract.
94%+ classification accuracy is written into the master service agreement as an SLA. 98%+ on compliance. Calibrated to within 2% of your own human reviewers.
Everything on this page, in one document.
The QEval model card lays out the architecture, the data policy, and the accuracy methodology behind every number on this page. It is available under NDA through our Trust Center.
Request the model card- Architecture: the MoE model, classification engine, and vocabulary library
- Data policy: no third-party model calls, no customer data in an outside training loop
- Accuracy methodology: how 94%+ is measured and calibrated to your reviewers
- Redaction: the PCI, PHI, and PII entity list and the PCI DSS Level 1 attestation
- Version history: how the experts are recalibrated and rolled forward
Built and trained inside a live contact center.
Software companies build QA models in a lab and ship them to you. QEval was built inside Etech, a contact center running thousands of agents, and ETS Labs is the first customer of every release. The experts and the vocabulary library are tuned on real scorecards, real compliance edges, and real coaching, which is why they hold up on yours. The model is built and operated by the ETS Labs engineering team, led by Ileshkumar Sisodiya.
One Core, behind every product.
The same engine that scores a call powers every product in the platform. Quality, coaching, compliance, and analytics all read from one model, not five models from five vendors.
Your products
Every score, summary, and signal in the platform comes from the same Core.
MoE router, expert sub-models, vocabulary library
The classification engine routes each item to a purpose-trained expert, fluent in contact-center language across 35+ languages.
Owned infrastructure
ETS Labs-owned compute. No third-party foundation model, and no customer data in an outside training loop.
The Core scores Layer 1, and feeds the other five.
The same model that grades every conversation extracts the intelligence behind it. QA and compliance is Layer 1. The five layers above it, where 82% of the value sits, all draw on the Core.
What buyers ask about the Core.
Is this just GPT or Claude under the hood?
No. QEval runs on a proprietary, closed-source Mixture-of-Experts model that ETS Labs built, trains, and operates. It is not a fine-tuned wrapper around a third-party foundation model, and there are no API calls out to one.
Where does my data go?
Nowhere outside QEval's infrastructure. There are no third-party model calls, and your conversations never enter an outside model's training loop. Redaction runs before scoring, and every decision is traceable to the expert sub-model and the transcript span that produced it.
Is a Mixture-of-Experts model unique to QEval?
No, and we will not pretend it is. At least one peer also runs proprietary AI. Our edge is what the model is trained on: real interaction data from a working contact center, paired with the coaching lifecycle and the four contractual commitments. The architecture is table stakes; the operator data and the accountability are the difference.
What does QEval not claim?
Three things, stated plainly. Automated redaction is never 100%; on poor audio or heavy accents we layer transcription-quality controls and human review rather than pretend otherwise. The 94%+ figure is calibrated against your own reviewers across every dimension at once, not a cherry-picked single-task benchmark. And a Mixture-of-Experts is not unique to QEval; the edge is the operator data the experts train on and the contractual accountability around the number.
How is 94%+ accuracy measured?
Across compliance, empathy, resolution, and brand voice at the same time, calibrated against your own human reviewers to within 2%. It is a contractual SLA in the master agreement, not a single-task lab benchmark. Industry-average classification accuracy sits at 65 to 70% (market context, not a QEval result).
What are the three parts, again?
The MoE model (purpose-trained expert sub-models), the Classification Engine (routes each scorecard item to the right expert), and the Vocabulary Library (contact-center language across 35+ languages and 80+ integrations). Together they deliver 326 million classifications every 5 minutes at 94%+ accuracy.
Can we add our own scoring dimensions?
Yes. The classification engine maps any scorecard item to an expert. For a dimension we do not already cover, we train or extend an expert for it, then route your items to it. Your 47-item scorecard is scored item by item, by the expert built for each one.
What happens when the model and our reviewers disagree?
We recalibrate. Scores are monitored for drift and held to within 2% of your human reviewers. When the model and your reviewers diverge, the experts are recalibrated and the version is rolled forward, so the 94%+ holds over time rather than only at kickoff.
Score a real call in 30 minutes.
Bring a real call and a real scorecard. We will route it through the Core in 30 minutes and show you the score, the expert behind it, and the moment that earned it.