Move cursor | Click to ripple
The accuracy commitment

The accuracy number, and how we prove it.

94%+ classification accuracy, calibrated against human reviewers and written into the master agreement. Here is the number, how it is measured, and the clause that stands behind it.

94%+Classification accuracyService level in the master agreement
98%+Compliance accuracyOn disclosure and script checks
95%+Violation recallOf the violations present, caught
<2%Drift from humansModel score vs a human reviewer
The accuracy question

We publish the number, and stand behind it.

Most platforms answer the accuracy question with coverage and an invitation to go check the work yourself. That is not the same as telling you how accurate the scoring is, and standing behind it.

How accuracy usually gets answered

  • ~"See and verify everything yourself." Coverage stands in for a measured number.
  • ~A relative claim, more accurate than a general model, with no absolute figure you can hold.
  • ~A scale number, proven on billions of interactions, that says nothing about how often it is right.
  • ~No published accuracy, and nothing about accuracy in the contract.

How QEval answers it

  • +A published number: 94%+ classification accuracy, stated as a service level.
  • +Measured against human reviewers, with the drift between them kept under 2%.
  • +Every score pinned to the moment in the conversation that earned it.
  • +The number is written into the master agreement, before you commit.

Left column describes patterns observed across competitor sites in June 2026. Market context, not QEval results.

See how the number is earned

Pick a call. Trace every score to the turn that earned it.

QEval does not hand you a grade and ask for trust. Click any score to see the exact moment it came from, then check the model against a human reviewer on the same call.

Transcript
A-Composite score
across four dimensions
EvidenceClick a dimension to see the turn that earned it.
Illustrative conversations and scores. In a live deployment, QEval scores against your scorecards and your compliance rules, and the calibration sample is drawn from your own reviewers.
How the number is measured

Accuracy is a measurement, run continuously.

The 94%+ figure is not a launch benchmark that ages out. Here is what stands behind it.

Calibrated
Measured against human reviewers
Scores are checked against trained human reviewers on a continuous sample. The gap between the model and the human is kept under 2%. When they diverge, the model is recalibrated, not the reviewer overruled.
Architecture
Experts, not one general model
QEval runs a proprietary mixture of experts: purpose-trained sub-models for each scoring dimension, rather than one general model asked to do everything. Customer conversations never enter a third-party model's training loop.
Evidence-pinned
Every score ties to a moment
A score is only as trustworthy as the evidence under it. Each dimension links to the exact turn that earned or cost the points, so any score can be audited in seconds, by you or by an auditor.
At scale
326 million classifications every 5 minutes
The same accuracy holds at volume: 326 million classifications every five minutes, more than 3 billion conversations scored a year, across human and AI agents on one scorecard.

Industry average automated-QA accuracy sits around 65 to 70 percent. Market context, not QEval results.

What the number means

What 94% means in practice.

A percentage is easy to print and hard to trust. Here is what ours actually refers to, and where the gap goes.

The benchmark is a person
Agreement with a human reviewer
On a reviewed sample, QEval's classification matches a trained human reviewer at least 94 percent of the time on the same items. The bar is a person, not a model grading its own work.
Where the gap goes
The few percent are the hard cases
Most of the disagreement is genuinely ambiguous moments: overlapping rules, borderline tone, a half-finished sentence. When disagreement forms a pattern rather than a one-off, that pattern is retuned.
Compliance is stricter
Held higher than the headline
Disclosure and script checks run at 98 percent or higher accuracy, with 95 percent or higher recall, because a missed disclosure is a regulatory event, not a matter of style.
In the contract

The number is in the agreement.

Accuracy that lives only in a sales deck is an adjective. Ours is a service level in the agreement you sign.

Master agreement, accuracy service level (paraphrased)
QEval commits to a classification accuracy of 94 percent or higher, measured against a human-reviewed sample. Deployment completes within 30 days or the deployment fee is returned. The customer may exit on 60 days notice, without penalty.
Paraphrased for the web. The binding language lives in the master agreement. The 120-day ROI figure is a documented customer-average outcome, not a contractual term.
See how that compares
Where accuracy sits

Accuracy is the floor the other five layers stand on.

L1 Quality and Compliance L2 Customer L3 Revenue L4 Operational L5 Training L6 Strategic

If the Layer 1 score is not accurate, every layer above it inherits the error. That is why the accuracy SLA is the precondition for the 82 percent of value that comes from Layers 2 through 6, not a detail.

Questions buyers ask

Accuracy, in plain terms.

Is the 94% accuracy actually in the contract?

Yes. The 94%+ classification accuracy is written into the master agreement as a service level, alongside the 30-day deployment guarantee and the 60-day exit clause. You see the language before you commit.

Accuracy of what, exactly?

Of the classifications QEval makes when it scores a conversation: whether a disclosure was present, whether the issue was resolved, the sentiment of a turn, and so on. Compliance checks run higher, at 98%+ accuracy with 95%+ recall on the violations that are present.

How do you measure it?

Against people. A continuous sample of conversations is scored by trained human reviewers, and QEval's scores are compared to theirs. The drift between the model and the human is kept under 2%. When the two diverge, the model is recalibrated.

How is this different from a vendor saying "we are 95% accurate"?

A slide claim is not published, not measured against a stated method, not tied to evidence, and not in the contract. The QEval number is all four: published here, calibrated against humans, pinned to the turn that earned each score, and written into the agreement.

What if accuracy slips after we go live?

Accuracy is monitored continuously, not certified once. Scores are recalibrated against the human sample on an ongoing basis, and the service level holds for the life of the agreement, not just at launch.

What counts as a miss?

A miss is when QEval's classification on an item differs from a trained human reviewer's on the same item. It is measured on a continuous sample, not a one-time launch test, so the number reflects how the model performs on your conversations over time.

How often do you recalibrate?

Continuously. The human-reviewed sample runs on an ongoing basis, and when disagreement clusters around a pattern, the model is retuned. That is why the service level holds for the life of the agreement, not just at launch.

Put it to the test

See the number on your own calls.

Start a pilot and we will score your conversations against your scorecards, then show you each score next to a human reviewer. The 94%+ SLA is in the agreement before you commit.

Contractual commitments

Four numbers no peer publishes.

94%+
Accuracy SLA
Written into the master agreement
30 days
Deployment
Money-back guarantee
60 days
Exit clause
Cancel with notice, no penalty
120 days
ROI
Documented customer-average outcome