Move cursor | Click to ripple
AI Agent Scorecard

Everyone will sell you AI agents. Who audits them?

QEval® audits AI agents in production under the VERDICT Method, the independent audit standard for the AI workforce. Pick a voice, chat, or email conversation below and read the full audit: what the agent did, whether it was good, and every issue with the fix. Free with a work email.

1M+
voice AI calls scored daily
326M
classifications every 5 minutes
94%+
accuracy SLA, in the contract
3B+
conversations scored YTD 2026
The pattern

Every public AI agent failure was self-graded.

The most documented AI agent incidents share one root cause: the metrics were the vendor's own, and nobody independently audited the conversations until a customer, a court, or a union did.

Case file 01 · Air Canada · Feb 2024

The chatbot that invented a policy

Self-reportedA helpful website chatbot answering policy questions.
Audited realityIt invented a bereavement refund policy. A tribunal rejected the airline's argument that the bot was "a separate legal entity" and held the airline liable for its words.
The missing audit: every policy answer checked against the policy document before a customer relies on it.
Case file 02 · Klarna · 2024 to 2025

The assistant worth 700 agents, until it wasn't

Self-reported"Doing the work of 700 full-time agents."
Audited realityThe CEO said on the record that cost had been "a too predominant evaluation factor" and the result was lower quality. Human agents were hired back.
The missing audit: a quality scorecard tracked alongside the cost line, so decay shows up before customers do the auditing.
Case file 03 · Commonwealth Bank · Aug 2025

The bot that cut 2,000 calls a week, on paper

Self-reportedThe voice bot reduced call volume by 2,000 calls a week, so 45 roles were made redundant.
Audited realityCall volumes were rising, with staff on overtime. Facing the tribunal, the bank formally admitted the assessment was wrong and reinstated the roles.
The missing audit: independent measurement of resolution, not deflection numbers reported by the team that deployed the bot.
Case file 04 · Cursor · Apr 2025

The AI company whose own support bot made up a rule

Self-reportedFront-line AI support, working as designed.
Audited realityThe bot invented a "one device per subscription" policy to explain a login bug. Customers cancelled publicly within hours, and the company apologized.
The missing audit: groundedness scoring, so a confident answer with no source behind it gets flagged before it ships.
40%+
of agentic AI projects forecast to be cancelled by end of 2027, with inadequate risk controls a named cause (Gartner)
40%
of enterprises will demote or decommission AI agents by 2027 over governance gaps found only after production incidents (Gartner)
Public record and analyst forecasts. Market context, not QEval results or QEval customers

The fix in every working deployment is the same: an independent measurement loop that the agent's builder does not control. Ours has a name.

The standard

The VERDICT Method.

QEval®'s independent audit standard for AI agents. Seven commitments, each one auditable, applied to production conversations on the same scorecard your human team is held to. Vendor benchmarks test agents before launch. VERDICT audits them after, conversation by conversation.

QEval audit standard · VERDICT · applies to voice, chat, and email agents
V

Verify the telemetry

Start from what the agent actually did: latency, turns, tool calls, escalations, containment. Telemetry is the input to the audit, never the grade. A green trace can sit on top of a conversation that broke a disclosure rule.

Observability shows the trace. The audit judges the conversation.
E

Evidence-pin every score

Every dimension score traces to the exact transcript turn that earned it, so a supervisor or compliance officer can see the moment instead of trusting a summary. No black-box verdicts.

Turn-level evidence is the scarcest and most diagnostic form of evaluation.
R

Resolution over containment

A contact closed without a human is not a contact resolved. The audit scores whether the issue was actually solved, correctly and compliantly, and treats repeat contact as the truth serum. Containment is coverage, not quality.

Industry studies: a bot can report 90% deflection at roughly 40% real resolution.
D

Disclosure and compliance gates

AI identity, recording consent, payment disclosures, and channel law are checked as hard gates, not style points. The rules are now statute: the TCPA covers AI voices, state laws carry private rights of action, and the EU AI Act requires disclosure from August 2026.

A gate either passed or it did not. There is no 7 out of 10 on consent.
I

Issue taxonomy, not vibes

Every failure maps to a named failure family, from disclosure miss to hallucination to context loss, with a confidence score and a recurrence count. One bad conversation finds every conversation like it.

A named failure can be counted, trended, and fixed. An anecdote cannot.
C

Calibrate against humans

Deterministic and classifier checks run on 100% of interactions. Model-judge scoring is sampled and validated against human supervisors, and recalibrated whenever the two diverge. Accuracy is contractual: 94%+, written into the master agreement.

Machine scoring earns trust by being checked, not by being asserted.
T

Track the fix

Every issue ships with the fix: the knowledge-base or guardrail action to apply, with a projected lift. Every applied fix is re-scored to confirm the lift held. An audit without action is just a report.

The loop closes when the next audit proves the fix worked.

Why it has to be independent. The vendor that built your AI agent controls both the design and the disclosure of its own evaluations. QEval® is built by an operator that runs 4,000+ human agents and scores its own floor on the same standard, and it has no AI agent to sell you. The scorecard below runs the method in front of you.

The audit room

Pick a conversation. Read the verdict.

Six real-pattern conversations across voice, chat, and email, each audited under the VERDICT Method. The seven stages run in order: verified telemetry, evidence-pinned scores, the resolution call, compliance gates, the issue ledger, calibration, and the projected lift if you apply the fixes. Add a work email to lift the glass.

Free tool · behind the glass

Open the audit room. What you can see under the glass is the live scorecard: a graded voice, chat, or email AI conversation with evidence-pinned scores, compliance gates, and every issue with the fix. Add your work email and the glass lifts.

    Enter a valid work email to start.

    Your scoring stays in your browser. We will email your scorecard and the VERDICT audit checklist. No spam, unsubscribe anytime.

    The riskiest channel

    Voice AI took the phone channel. The phone channel has laws.

    Voice is where AI agents are scaling fastest and failing hardest. Production latency, interruptions, accents, and background noise break agents that looked perfect in the demo, and the disclosure rules are now statute, not etiquette. QEval® scores 1M+ voice AI calls every day, so the audit reflects production, not a lab.

    Latency degrades 40 to 120% under production loadBarge-in: stop within 200ms and recover context. Rarely metThe same ASR API: 92% accuracy on clean headsets, 65% on noisy mobile callsAccent bias widens as noise risesCustomers game looping bots just to reach a humanSome bots deny being AI when asked. That is now a legal exposure
    Feb 2024
    The FCC rules AI-generated voices are "artificial" under the TCPA. Prior consent, identification, and opt-out required on AI calls.
    Jan 2026
    California SB 243 takes effect with a private right of action for chatbot disclosure failures. One of 70+ state AI laws passed in 2025.
    Aug 2026
    EU AI Act Article 50: users must be told when they are talking to AI. The transparency deadline survived the Omnibus delays.
    Market context and public regulation, not QEval results
    Observability is not QA

    Your stack sees the trace. QEval scores the call.

    AI observability tools show what the agent did: latency, tokens, tool calls, traces, and the containment rate. None of that says whether the conversation was good for the customer or safe for the business. The VERDICT Method starts where the trace ends.

    Observability sees

    • Latency, cost, and token usage
    • Tool calls, traces, and error rates
    • Task and run success or failure
    • Containment rate, resolved without a human

    QEval scores

    • Resolution quality, not just containment
    • Compliance, disclosure, and PII handling
    • Empathy, predicted CSAT, and churn risk
    • Brand voice, and the fix for every gap

    Containment is not resolution. An AI agent that contains a contact can still leave the customer unresolved, and a green trace can sit on top of a conversation that broke a disclosure rule. The telemetry is the input; the verdict is the judgment.

    How the scoring stays honest

    An auditor you can audit.

    Most automated QA fails on trust: scores that do not match the company's own scorecard, and no way to check the machine. The VERDICT Method is built to be inspected.

    Coverage

    Deterministic on 100%

    Disclosure, consent, and policy gates run as deterministic and classifier checks on every interaction, not a sample.

    Judgment

    Sampled, human-calibrated

    Model-judge scoring on qualitative dimensions is sampled and validated against human supervisors, the way a QA team calibrates its reviewers.

    Correction

    Recalibrated on divergence

    When human and machine judgments drift apart, the expert model is recalibrated before its scores count. Disagreement is a signal, not a footnote.

    Audit trail

    Traceable to the span

    Every classification traces to the expert sub-model that produced it and the transcript span that triggered it. You can always ask the system to show its work.

    The same scorecard as your humans. Scoring runs on a proprietary multi-expert model with a purpose-trained expert for each dimension, at a contractual 94%+ accuracy. AI agents are graded on the rubric your human team is held to, so "good enough for the bot" stops being a separate, lower bar.

    An AI audit is Layer 1. The conversation carries five more.

    Quality and compliance is the first of QEval®'s Six Layers of Intelligence. The same audited conversation also feeds customer, operational, and training intelligence, which is where most of the value is.

    L1 QualityL2 CustomerL4 OperationalL5 Training See the Six Layers
    Take these to the meeting

    Five governance questions for your next vendor meeting.

    Whether you evaluate QEval® or anyone else, these questions separate production platforms from demonstrations.

    1. Where does customer data go when AI processes it? Is there zero external data transfer to third-party LLMs?
    2. Is the model accurate enough for compliance-grade decisions, and will the vendor commit to accuracy in the contract?
    3. Can you audit and explain how the AI reached a specific classification?
    4. Is the audit independent of the team or vendor that built the AI agents?
    5. Does your auditor publish a named methodology and a contractual accuracy number, or adjectives?
    SOC 2 Type IIISO 27001ISO 42001 (AI Management)PCI DSS Level 1HIPAAGDPRCCPAPII redaction at ingest
    Common questions

    The VERDICT Method, answered.

    What is the VERDICT Method?

    The VERDICT Method is QEval's independent audit standard for AI agents. Seven commitments: Verify the telemetry, Evidence-pin every score to the exact transcript turn, score Resolution over containment, check Disclosure and compliance as hard gates, map every failure to a named Issue taxonomy, Calibrate machine scoring against human supervisors, and Track every fix by re-scoring to confirm the lift. It runs on production conversations, on the same scorecard as human agents, backed by a contractual 94%+ accuracy SLA.

    What is an AI agent scorecard?

    An AI agent scorecard is a structured quality report that audits a voice, chat, or email bot conversation. It separates observability telemetry (what the agent did) from a quality score across resolution, containment quality, groundedness, compliance, empathy and CSAT, brand voice, and action accuracy, then lists each issue with what went wrong, why, where in the transcript, and the recommended fix.

    Why is containment not the same as resolution?

    Containment measures the absence of a human handoff, not whether the customer's issue was solved. A bot can report high containment while the customer abandoned, got a wrong answer, or contacts again the next day. The VERDICT Method scores resolution and repeat contact, and treats containment as a coverage metric, not a quality metric.

    How does auditing a voice bot differ from a chat or email bot?

    Voice adds timing and speech metrics (word error rate, turn latency, barge-in recovery) and hard legal gates at connection: recording consent and AI-identity disclosure under the TCPA and state law. Chat centers on retrieval groundedness, citation accuracy, and the deflection versus resolution gap. Email centers on full-thread comprehension, answer completeness, routing accuracy, and auto-reply-loop safety under RFC 3834.

    How does QEval keep AI scoring trustworthy?

    Deterministic and classifier checks run on 100% of interactions, while model-judge scoring is sampled and validated against human reviewers, with recalibration when human and model judgments diverge. Scoring runs on a proprietary multi-expert model backed by a 94%+ accuracy SLA, and every classification traces to the expert and the transcript span that triggered it.

    The VERDICT Method

    Your AI agents have metrics. Get them a verdict.

    Bring your scorecards, your AI agents, your CCaaS. We will audit a real conversation in 30 minutes under the VERDICT Method and show you what your current program missed last week.

    Contractual commitments

    Four numbers no peer publishes.

    94%+
    Accuracy SLA
    Written into the master agreement
    30 days
    Deployment
    Money-back guarantee
    60 days
    Exit clause
    Cancel with notice, no penalty
    120 days
    ROI
    Documented customer-average outcome