How the model works, and how we prove it.
Procurement and security teams do not buy adjectives. This is the architecture, the data policy, the accuracy methodology, and the limitations of the QEval Core, the proprietary Mixture-of-Experts model behind every score. The full card is available under NDA.
01 · Model details
02 · How QEval® uses machine learning
QEval® is not a single model. It is a pipeline of machine-learning systems, each doing one job:
- Transcription (ASR): Voice converted to text. Recognition quality bounds everything downstream.
- Redaction (NER): Named Entity Recognition removes PCI, PHI, and PII before scoring. Sensitive data never reaches the classifiers.
- Speech and text analytics: NLP engine extracts key phrases, topics, silence, overtalk, pace, and vocal characteristics.
- Classification (the MoE): Hundreds of purpose-trained classifier sub-models score each scorecard item. The classification engine routes each item to the classifier built for it, by purpose and tuned by industry.
- Sentiment, intent, and summarization: Dedicated models produce the sentiment trajectory, the reason for contact, and the summary and action items.
- Validation: An internal validation step checks summaries and intent for grounding. It runs inside QEval®; no third-party foundation model is called.
Training: Classifiers trained with supervised learning on labeled, real contact-center interactions, then held to each customer’s own reviewers through a human calibration loop.
What we withhold: Architecture family and data categories are disclosed. Model weights, parameter counts, data-mix proportions, and compute are not published.
03 · Intended use
04 · Out-of-scope and prohibited use
- Not a sole basis for adverse employment or disciplinary action without human review.
- Not a legal-compliance adjudicator, and not legal advice.
- Not intended for data outside contact-center conversational data.
- Not a lie detector or a truth-of-emotion detector.
- Not a substitute for a Data Processing Agreement.
05 · Factors
QEval® performance varies across, and is validated across, these factors:
- Channel: voice, chat, email, SMS.
- Agent type: human agent vs AI agent.
- Industry vertical: financial services, healthcare, telecom, insurance, retail, energy and utilities, collections, government, travel and hospitality, BPO.
- Language: 35+ supported languages.
- Audio quality and accent (voice), and conversation difficulty (escalations, compliance edges, repeat contacts).
06 · Metrics and decision thresholds
| Metric | Value | How it is measured |
|---|---|---|
| Classification accuracy | 94%+ | Across compliance, empathy, resolution, and brand voice concurrently, against a human reference. Contractual SLA. Industry average 65 to 70% for context. |
| Compliance accuracy | 98%+ | On compliance-specific classifications. |
| Calibration | Within 2% | Held to within 2% of the customer’s own human reviewers. |
| Redaction | Recall-first | Measured on recall before precision; a missed identifier is the privacy risk. |
Uncertainty management: Human-labeled ground-truth set + human auditors validating a sampled subset, with recalibration on divergence. Thresholds and the exact scorecard are configured per customer.
07 · Training data
08 · Preprocessing
- Transcription of voice to text (ASR) for voice channels.
- Redaction of PCI, PHI, and PII before scoring (Section 10).
- Normalization and channel handling for voice, chat, email, and SMS.
09 · Evaluation and human calibration
- Each deployment establishes a customer-specific ground-truth set, labeled by human experts, mirroring the customer’s scorecard and channels.
- Human auditors validate a sampled subset of model scores on a standing cadence.
- Where the model and the customer’s reviewers diverge, classifiers are recalibrated and the version is rolled forward.
- Segment-level performance (by language, channel, and vertical) is validated during the 30-day deployment against the customer’s own reviewers and documented in the deployment report.
10 · Redaction: PII, PHI, and PCI handling
Redaction runs before scoring, on both the transcript and audio file. Uses Named Entity Recognition models trained on millions of conversations, backed by the vocabulary library and pattern logic.
| PCI entity (required, cannot be disabled) | What it covers |
|---|---|
| BANK_ACCOUNT | Bank account numbers and international equivalents (IBAN) |
| CREDIT_CARD | Card numbers including debit, ATM, prepay, and charge cards; full numbers and last-four mentions |
| CREDIT_CARD_EXPIRATION | Card expiration dates. |
| CVV | 3- and 4-digit verification codes and institution-specific variants. |
| ROUTING_NUMBER | Routing numbers plus sort codes, BSB, IFSC, transit, and SWIFT equivalents. |
| PII entity (redacted by default) | What it covers |
|---|---|
| DOB | Dates of birth. |
| DRIVER_LICENSE | Driver permit numbers. |
| PASSWORD | Passwords, PINs, access keys, and verification answers. |
| SSN | Social Security numbers and international government identifiers. |
Optional entities (configurable per customer, off by default):
- Primary metric is recall:A false negative is the privacy risk; recall is the release criterion; precision is secondary.
- Factors that affect accuracy: audio quality, compression, heavy accents, rapid speech, slang, and background noise. Automated redaction is never 100%.
- Configurability: The customer selects which entities to redact. PCI entities are required and cannot be unchecked.
- Certification: Hosted environment is certified PCI DSS Level 1. In an audited test, no exploitable information was retained after redaction
- Residual risk: Quasi-identifiers may remain; redaction is not a substitute for a Data Processing Agreement.
11 · Data handling, privacy, and security
Signed artifacts available under NDA through the Trust Center.
12 · Bias, risks, and limitations
- Limitation. Transcription errors on heavy accents, noisy audio, or rapid speech propagate downstream. Transcription-quality controls and human review are layered on top.
- Limitation. Lower-resource languages and unusual domain shifts can reduce accuracy until the classifiers are recalibrated.
- Bias. Automated scoring can perform unevenly across accents or dialects. Per-customer calibration and human auditing are the controls.
- Risk. Scores influence coaching and performance, so they must not be the sole basis for adverse decisions. Human review stays in the loop.
- Architecture. A Mixture-of-Experts is not unique to QEval®. The differentiators are the operator data, the calibration loop, and the contractual commitments.
13 · Mitigations and recommendations
- Keep a human in the loop for consequential decisions.
- Validate on your own conversations during the 30-day deployment.
- Maintain the calibration cadence and recalibrate on divergence.
- Redact before any downstream processing.
14 · Ethical considerations
- Sensitive data (PII, PHI, PCI) is redacted before scoring.
- Recording and disclosure obligations remain the customer's; QEval® scores whether the required disclosures were given.
- Employment impact: scores support coaching and review, not automated adverse action.
- Fairness: performance is reviewed against the customer's reviewers across segments.
15 · Governance, versioning, and change management
16 · Environmental
QEval® runs on carbon-aware infrastructure, managed against the green-hosting posture reflected in the trust grid.
17 · Technical and integration
18 · Commitments
19 · Glossary
20 · Document and disclaimer
See the full model card
The complete QEval Core model card is available to verified work emails. Enter yours to unlock it here, and we will also send the signed PDF.